Audio noise reduction system with multiple mic arrays
Through the synchronous calibration, source separation, weight configuration and noise discrimination modules of the multi-mic array audio noise reduction system, the signal separation and noise suppression problems under the conditions of changing sound source positions and unstable environmental noise in traditional technologies are solved, thereby improving signal separation accuracy and voice quality.
Patent Information
- Application Number
- CN202510921107.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional multi-microphone array audio noise reduction technology has difficulty achieving accurate separation and suppression when the sound source position frequently changes and the ambient noise is unstable, resulting in synchronous drift of the picked-up data, distortion of the target voice, and residual interference noise, affecting voice quality.
The synchronous calibration module analyzes the energy change curves of multiple channels and adjusts the channel time starting point; the source separation module constructs the time series vector and filters the consistent channels; the weight configuration module adjusts the frequency weight; the noise discrimination module identifies the noise characteristics and matches the processing parameters; the segment scoring module evaluates the speech processing configuration and optimizes the channel response and noise suppression effect.
It achieves improvements in signal separation accuracy and voice processing quality in complex environments, enhances the adaptation performance in changing environments and far-field pickup scenarios, and optimizes the noise suppression effect and voice enhancement experience.
Smart Images

Figure CN120808807A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of multi-microphone noise reduction technology, in particular to a multi-mic array audio noise reduction system. BACKGROUND
[0002] The multi-microphone noise reduction technology field includes the key means of cooperative collection and processing of audio signals based on multiple acoustic sensors to suppress environmental noise and enhance target speech signals, including sound source positioning, beamforming, spatial filtering, covariance matrix estimation, speech signal enhancement, for speech communication, speech recognition, intelligent terminals, vehicle systems, far-field pickup, etc. Multi-microphone noise reduction utilizes multiple pickup units in the array structure to obtain the correlation between signals through time delay estimation and spatial signal analysis technology, further constructs filters suitable for different sound source directions to separate target signals and interference signals. Combined with acoustic modeling, array geometry design and low-delay implementation architecture, a noise suppression method with spatial selectivity is formed. A multi-mic array audio noise reduction system refers to constructing a spatial array structure composed of multiple microphones to obtain audio signals at different positions, and based on the preset array topology and pickup configuration, using delay alignment and weighted synthesis to process the signals, covering the design of spatial signal capture method, the construction of array geometry, the configuration of microphone position parameters, the setting of delay calculation method, the calculation method of weighting factor, the estimation process of covariance matrix, the selection strategy of noise reference path, etc. By constructing a linear or ring array, the beamforming is completed by combining the time delay summation mechanism based on fixed direction or the generalized sidelobe cancellation algorithm. The spectral subtraction method with fixed window length, the statistical quantity estimation method based on the reference channel, and the sub-band filtering structure are used to weight and fuse the frequency domain signals to realize spatial noise suppression processing of the microphone received signals.
[0003] The traditional multi-microphone array audio noise reduction technology relies on fixed rules and parameter inference for channel time alignment, energy trend classification, frequency band weight distribution and noise discrimination, lacks periodic adjustment and behavior feedback based on real-time data, and in practical applications, it is difficult to achieve accurate separation and suppression in cases where the sound source position changes frequently, the environmental noise is unstable or the frequency band response is uneven, resulting in synchronous drift of pickup data, distortion of target speech, residual interference noise, and large fluctuations in voice quality and discontinuous noise reduction effect in far-field communication and vehicle systems, affecting the overall voice enhancement experience. SUMMARY
[0004] To solve the technical problems existing in the prior art, the present application embodiment provides a multi-mic array audio noise reduction system. The technical solution is as follows: In one aspect, an audio noise reduction system of a multi-mic array is provided, comprising a multi-mic array device body, a mainboard, a memory and a processor, a microphone array, the memory storing a computer program, the processor executing the computer program to realize the audio noise reduction system of the multi-mic array, the system comprising: The synchronization calibration module uses the audio sampling information to analyze the energy change curves of multiple channels, judges the response offset through cross-point distribution and slope matching, calculates the offset period and adjusts the channel time starting point, and establishes the offset adjustment result. The source separation module calls the offset adjustment result, analyzes the current frame channel energy change, constructs a time sequence vector, compares the energy distribution trend, selects consistent channels, and generates a channel decoupling structure. The weight configuration module uses the channel decoupling structure to compare the channel frequency energy distribution and response change, judges the frequency band directivity, analyzes the frequency response and time compensation trend, adjusts the frequency weight, and constructs a speech synthesis configuration. The noise discrimination module uses the speech synthesis configuration to analyze the extreme value distribution and change trend of the speech energy curve, judges the interval stability, identifies the noise characteristics, and matches the noise processing parameters.
[0005] As a further scheme of the present application, the offset adjustment result includes the channel starting point adjustment position, the cross-point sequence distribution and the time offset period, the channel decoupling structure includes the sound source subset division result, the channel cooperative distribution matrix and the signal difference identification parameter, the speech synthesis configuration is specifically the channel frequency band combination selection, the frequency response trend judgment and the weight allocation parameter, and the noise processing parameter includes the extreme point change direction, the fluctuation interval stability feature and the noise path matching identification.
[0006] As a further scheme of the present application, the synchronization calibration module comprises: The curve distribution detection submodule obtains audio sampling information, analyzes the energy change curves of multiple microphone channels, compares the time sequence arrangement of the cross-point of each channel curve according to the distribution position of the curve on the time axis, calculates the period change of the cross-point on the time axis, and establishes the cross-period analysis quantity. The trend sequence screening submodule calls the cross-period analysis quantity, analyzes the change relationship between the cross-point time position and the adjacent curve slope direction, screens the cross-point sequence with consistent curve change trend according to the continuous matching sequence, and establishes the trend consistent sequence data. The time axis adjustment submodule calculates the time offset period of each channel corresponding to the reference channel based on the trend consistent sequence data, adjusts the starting point position of the time axis of each channel by compressing the difference amplitude in the processing period, and obtains the offset adjustment result.
[0007] As a further scheme of the present application, the source separation module comprises: The energy sequence construction submodule calls the offset adjustment result, analyzes the signal energy changes of multiple channels in the current frame segment, monitors the energy data of each channel, constructs a time sequence vector, establishes an energy change vector group according to the numerical distribution of the time sequence vector; The coordination comparison submodule compares the energy distribution trends between each channel based on the energy change vector group, judges the coordination characteristics of the energy changes between multiple channels, generates a channel coordination comparison coefficient by comparing the coordination relationship; The sound source classification submodule classifies channels with consistent energy change trends as the same sound source subset according to the channel coordination comparison coefficient, constructs an energy distribution matrix between sound source subsets, and establishes a channel decoupling structure.
[0008] As a further scheme of the present application, the weight configuration module comprises: The frequency distribution comparison submodule calls the channel decoupling structure, compares the energy distribution characteristics and time response change amplitudes of multiple channels in each frequency range, judges the consistency of channel energy distribution, and establishes a frequency band consistency coefficient; The main direction screening submodule analyzes the coverage of frequency band energy in the main direction of the beam based on the frequency band consistency coefficient, calculates the angle between the frequency response change trend of each channel in the main direction and the main direction, screens the direction consistent channels by judging the angle range, and obtains a main direction screening result; The synthesis configuration adjustment submodule analyzes the cooperative trend of the frequency response and time compensation deviation of each channel according to the main direction screening result, adjusts the weight distribution in each frequency range, optimizes the channel frequency band combination, and establishes a speech synthesis configuration.
[0009] As a further scheme of the present application, the specific formula for comparing the energy distribution characteristics and time response change amplitudes of multiple channels in each frequency range is: ; Calculate the frequency band consistency coefficient; Wherein, is the energy normalization value of the channel at the th frequency point, is the energy normalization value of the channel at the th frequency point, is the time response change amplitude normalization value of the channel at the th frequency point, is the time response change amplitude normalization value of the channel at the The normalized value of the time response change amplitude at each frequency point, For channel With channel Band consistency coefficient in the entire frequency dimension, is the first channel number index to which the energy and time response parameters belong, The second channel number index to which the energy and time response parameters belong. is the index number of the frequency point, is the total number of frequency points.
[0010] As a further solution of the present invention, the noise discrimination module includes: The fluctuation monitoring submodule calls the speech synthesis configuration, analyzes the local fluctuation pattern of the energy curve in the continuous speech segment, detects the fluctuation of the energy curve in each time period, counts the fluctuation frequency and amplitude, and obtains the fluctuation feature value; The extreme value sequence analysis submodule compares the density and direction change trend of extreme value points in adjacent time periods based on the fluctuation characteristic quantity, sorts the time interval sequence of extreme value points, determines the fluctuation stability, and establishes the extreme value distribution coefficient; The characteristic matching submodule identifies the noise type by analyzing the repetitive characteristics of energy changes based on the extreme value distribution coefficient, matches the noise processing path according to the noise type, and establishes noise processing parameters.
[0011] As a further solution of the present invention, the specific formula for determining the fluctuation stability is: ; Calculate the extreme value distribution stability parameter, determine the fluctuation stability, and establish the extreme value distribution coefficient; in, For the The normalized time interval between extreme points, is the average value of the normalized time intervals of all extreme points, For the The direction jump state value corresponding to the extreme point, is the average value of the state values of all extreme point direction jumps, is the serial index number of the extreme point, is the total number of extreme points in the current analysis segment, is the stability parameter of the extreme value distribution.
[0012] As a further embodiment of the present invention, the system further comprises: The segment scoring module uses the noise processing parameters to analyze the degree of overlap of adjacent cycles at the frequency start, compare the signal connection smoothness at the cycle splicing position, determine the continuous distribution position of the structural change point, construct a consistency index for the speech cycle segment, and adjust the speech processing configuration based on the integrity of the segment to establish a noise reduction effect feedback record; The noise reduction effect feedback record specifically refers to the splicing smoothness status, cycle consistency index, and voice processing configuration adjustment information.
[0013] As a further embodiment of the present invention, the segment scoring module includes: The frequency overlap judgment submodule calls the noise processing parameters, analyzes the overlap degree of adjacent cycles at the frequency start, calculates the overlap degree of the frequency start distribution of each cycle splicing point, and establishes the overlap distribution value; The connection smoothness comparison submodule analyzes the connection smoothness of the periodic splicing position signal based on the overlap distribution value, determines the signal amplitude jump state of the splicing point, and establishes a jump smoothing coefficient; The consistency adjustment submodule selects non-stationary segments based on the jump state of the splicing point according to the jump smoothing coefficient, constructs a consistency index of the speech cycle segment, adjusts the speech processing configuration, including speech synthesis and noise processing parameters, and establishes a noise reduction effect feedback record based on the segment integrity.
[0014] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least: Through time domain intersection analysis and dynamic slope sequence screening based on energy change curves, the response offset of multiple channels is corrected in real time. The energy trend comparison and cooperative relationship classification mechanism are superimposed to achieve accurate separation of the signal source. The frequency distribution characteristics and time response parameters are combined to coordinately adjust the weights, enhance the channel's ability to distinguish the consistency of different frequency ranges and directions, use extreme value sequence and energy fluctuation characteristic analysis to improve the sensitivity of noise recognition, integrate periodic overlap and splicing smoothness to dynamically score the integrity of the segment, improve the adaptive capture capability of complex sound sources, frequency band energy coverage, delay drift and noise fluctuations, optimize the noise suppression effect, signal separation accuracy and voice processing quality, and enhance the adaptation performance of changing environments and far-field pickup scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Fig. 1 is a system flow chart of the present invention; Fig. 2 A schematic diagram of a system framework of the present application; Fig. 3 A schematic diagram of a back structure of a multi-mic array device body of the present application; Fig. 4 A schematic diagram of a front structure of a multi-mic array device body of the present application.
[0017] Legend: 1, multi-mic array device body; 2, mainboard; 3, memory and processor; 4, microphone array. DETAILED DESCRIPTION
[0018] The technical solutions in the present application will be described below with reference to the drawings.
[0019] In the embodiments of the present application, the words such as "example", "for example" are used to represent as an example, illustration or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.
[0020] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times. It should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent. The words "of", "corresponding" and "corresponding" can be used interchangeably at times. It should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent.
[0021] In the embodiments of the present application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1. When the distinction is not emphasized, the meanings expressed are consistent.
[0022] In order to make the technical problems, technical solutions and advantages to be solved by the present application more clear, the following will be described in detail with reference to the drawings and specific embodiments.
[0023] The embodiments of the present application provide a multi-mic array audio noise reduction system, please refer to Figs. 1 to 4 The present application provides a technical solution, a multi-mic array audio noise reduction system, including a multi-mic array device body 1, a mainboard 2, a memory and a processor 3, and a microphone array 4 inside the multi-mic array device body 1. The memory stores a computer program, and the processor executes the computer program to realize the multi-mic array audio noise reduction system. The system includes: The synchronization calibration module utilizes audio sampling information to analyze energy change curves of multiple channels, judges response offset through cross point distribution and slope matching, calculates offset period and adjusts channel time starting point, and establishes offset adjustment result; The source separation module calls offset adjustment result, analyzes current frame channel energy change, constructs time sequence vector, compares energy distribution trend, selects consistent channels, and generates channel decoupling structure; The weight configuration module adopts channel decoupling structure, compares channel frequency energy distribution and response change, judges frequency band directionality, analyzes frequency response and time compensation trend, adjusts frequency weight, and constructs speech synthesis configuration; The noise discrimination module utilizes speech synthesis configuration to analyze extreme value distribution and change trend of speech energy curve, judges interval stability, identifies noise characteristics, and matches noise processing parameters; The speech segment scoring module utilizes noise processing parameters to analyze coincidence degree of adjacent periods at frequency starting points, compares signal connection smoothness of period splicing positions, judges continuous distribution position of structure change points, constructs consistency index of speech period segment, adjusts speech processing configuration in combination with speech segment integrity, and establishes noise reduction effect feedback record.
[0024] The offset adjustment result includes channel starting point adjustment position, cross point sequence distribution, and time offset period, the channel decoupling structure includes sound source subset division result, channel cooperative distribution matrix, and signal difference identification parameter, the speech synthesis configuration specifically is channel frequency band combination selection, frequency response trend judgment, and weight distribution parameter, the noise processing parameter includes extreme point change direction, fluctuation interval stability feature, and noise path matching identification, and the noise reduction effect feedback record specifically refers to splicing smooth state, period consistency index, and speech processing configuration adjustment information.
[0025] The synchronization calibration module includes: The curve distribution detection submodule obtains audio sampling information, analyzes energy change curves of multiple microphone channels, compares time sequence arrangement of cross points of each channel curve according to distribution positions of the curves on a time axis, calculates period change of the cross points on the time axis, and establishes cross period analysis quantity; The curve distribution detection submodule obtains audio sampling information, the audio sampling information collected is the continuous collection and recording of multiple microphone channels on a unified time window on a voice signal, the sampling frequency is set to 16000Hz, the time window is set to 20ms, the original audio signal collected for each channel is subjected to short-time energy calculation, forming the corresponding energy change curve, for the energy curves of channel 1 and channel 2, in the scene where the voice signal input is the standard Mandarin sentence "today the weather is fine" of the far-field speaker, the curve intersection points are observed at the time positions of 160ms, 320ms, 480ms and 640ms, further, the energy curves formed by channel 3 and channel 1 form intersections at 200ms, 360ms and 540ms, the intersection time points of each two channels are extracted and uniformly summarized, and a cross-time sequence data table is established, under the premise that channel 1 is the reference channel, the cross-point time interval distribution between channel 1 and other channels is counted, the average interval of channel 1 and channel 2 is 160ms, the average interval of channel 1 and channel 3 is 180ms, and the average interval of channel 1 and channel 4 is 170ms, the standard deviation value between the average cross-interval data of each channel is compared, which is 30ms, if it is lower than the set periodic stability judgment reference value 40ms, it constitutes the cross-period judgment condition, the cross-point time distribution between the channel combinations is summarized according to the judgment result, and the cross-interval interval of each pair of channels is recorded into a structure index table, and finally the cross-period analysis quantity is obtained.
[0026] The trend sequence screening submodule calls the cross-period analysis quantity, analyzes the change relationship between the cross-point time position and the adjacent curve slope direction, screens the cross-point sequence with consistent curve change trend according to the continuous matching sequence, and establishes a trend consistent sequence data; The trend sequence screening submodule calls the cross-period analysis quantity, selects the cross-point time coordinates of each group of channels, obtains the energy change curve slope trend in the adjacent region on the time axis, for example, when the cross-point time of channel 1 and channel 2 is 200ms, 360ms and 520ms, the energy change direction within 10ms before and after these time points is observed, if the slope is all positive upward or all negative downward, it is marked as a trend consistent point, and the cross-point of channel 3 and channel 1 at 180ms and 340ms is excluded because the slope directions are inconsistent, the matching statistics of the consistency of the trend direction of not less than three cross-points are performed for each group of channels, if the consistent point proportion exceeds 70%, the trend sequence of this group is retained, and the retained trend sequence is recorded as a time index vector, wherein the elements are each matched cross-point, in the above manner, finally, 3 trend consistent points of the cross-points of channel 1 and channel 2 are screened out, 2 effective points of channel 1 and channel 4 are screened out, which do not meet the judgment condition, the record of this group is excluded, a trend consistent time sequence set with channel 1 as the reference is constructed, which is input information for subsequent time axis adjustment, and finally the trend consistent sequence data is obtained.
[0027] The time axis adjustment submodule calculates the time offset period of each channel corresponding to the reference channel based on the trend consistent sequence data, adjusts the starting point position of each channel time axis by compressing the difference amplitude within the processing period, and obtains an offset adjustment result; The time axis adjustment submodule extracts the time position offset of the trend consistent intersection point in each channel relative to channel 1 based on the trend consistent sequence data. At the intersection points 200ms, 360ms, and 520ms, the time coordinates of channel 2 are 205ms, 365ms, and 525ms, respectively, with a +5ms offset relative to channel 1. The time coordinates of channel 3 are 192ms, 348ms, and 510ms, respectively, with a -8ms, -12ms, and -10ms offset relative to channel 1. After counting these offset amounts, the average value of the offset amount of each channel is selected as the offset reference of the overall time axis of the channel. The average offset amount of channel 2 is +5ms, and the average offset amount of channel 3 is -10ms. When the offset amount change amplitude is less than the system stability reference value of 20ms, compression adjustment is performed, and the compression ratio is set to 50%, i.e., the offset adjustment value is half of the original value. The time starting point of channel 2 is adjusted to -2.5ms, and the time starting point of channel 3 is adjusted to +5ms. The adjustment value is applied to the original sampling data starting time of each channel to form a new time reference point after adjustment. The adjustment parameters of all channels are summarized and arranged into a structured data table for time synchronization control, and the final offset adjustment result is obtained.
[0028] The source separation module comprises: The energy sequence construction submodule calls the offset adjustment result, analyzes the signal energy changes of multiple channels in the current frame segment, monitors the energy data of each channel, constructs a time sequence vector, and establishes an energy change vector group according to the numerical distribution of the time sequence vector; The energy sequence construction submodule calls the offset adjustment result, which is the time reference point data of each channel after adjustment. After obtaining, each channel of the audio signal frame is divided according to the unified time reference, with 20ms as a frame length. The energy value is calculated within the frame length to obtain the frame-level energy sequence of each channel. Each energy curve represents the energy value change within a group of consecutive time frames. For example, the energy values of channel 1 within the first 10 frames are 102, 98, 105, 110, 108, 103, 95, 99, 101, and 106. The energy values of channel 2 are 100, 96, 104, 112, 109, 105, 97, 98, 100, and 107. The energy data of each channel is arranged into a time sequence vector, named wherein is the channel number, For frame number, the vectors of each channel are normalized so that all channel energy values are between 0 and 1, avoiding trend misjudgment due to absolute energy intensity difference. After normalization, the channel 1 energy vector is 0.83, 0.80, 0.85, 0.89, 0.87, 0.84, 0.77, 0.79, 0.82, 0.86, and the channel 2 energy vector is 0.81, 0.78, 0.84, 0.91, 0.89, 0.86, 0.79, 0.80, 0.81, 0.87. The corresponding time series energy vector group of each channel is constructed, and is uniformly arranged into a matrix structure, each row corresponding to a channel and each column corresponding to a time frame. The energy change vector group is established based on the matrix structure, and the energy change vector group is obtained.
[0029] The coordination comparison submodule compares the energy distribution trend between each channel based on the energy change vector group, judges the coordination characteristics of energy change between multiple channels, and generates a channel coordination comparison coefficient by comparing the coordination relationship. The coordination comparison submodule performs frame-by-frame comparison on the energy distribution trend between channels based on the energy change vector group. Assuming that the comparison period is 10 frames, the energy difference of each pair of channels in each frame is calculated. If the absolute value of the energy difference between the two channels is less than 0.05, the frame is marked as a consistent point. The proportion of consistent points in the total number of frames is calculated to obtain a channel coordination consistency coefficient. wherein , is the channel number. If 9 frames in channel 1 and channel 2 satisfy the consistent condition, the coordination consistency coefficient is 0.9. The coordination judgment reference is 0.8, which is set according to experience and is derived from the energy synchronous change frequency between multiple microphone channels in the speech collection process. When the coordination consistency coefficient is greater than the reference, it is determined that the channel pair has strong coordination. The coordination consistency coefficients of all channel pairs are arranged into a matrix structure, and the elements in the matrix are the coordination coefficient values between channels. For example, the coordination coefficient between channel 1 and channel 2 is 0.9, the coordination coefficient between channel 1 and channel 3 is 0.65, and the coordination coefficient between channel 2 and channel 3 is 0.67. According to the matrix data, the differences in coordination degree between channels are compared item by item, and the channel pair with significantly higher coordination than the average value is selected as the reference for subsequent classification. The coordination comparison results between all channel pairs are recorded in a structure table to obtain the channel coordination comparison coefficient.
[0030] The sound source classification submodule filters channels with consistent energy change trends based on the channel coordination comparison coefficient and classifies them into the same sound source subset. An energy distribution matrix between sound source subsets is constructed, and a channel decoupling structure is established. The sound source classification submodule selects channel pairs with a co-motion coefficient greater than 0.8 from the co-motion coefficient matrix according to the channel co-motion comparison coefficient, and judges them as channel pairs with consistent trends. If channel 1 and channel 2 meet the conditions, channel 2 and channel 4 meet the conditions, and channel 3 and any other channel do not meet the conditions, then the sound source subsets are constructed as {channel 1, channel 2, channel 4} and {channel 3}, and a channel group index table is established. For the channels of each subset, the normalized energy value is extracted at each frame time point, and the average of the energy values of all channels in the subset is taken as the frame-level energy value of the subset. The subset energy curve is accumulated frame by frame to construct the energy change trend data of each subset in the entire time period, which is uniformly represented in matrix form, where the rows represent the subset number and the columns represent the time frame number. This matrix is the energy distribution matrix between the sound source subsets. All subsequent speech separation operations are based on this structure for signal decoupling configuration and finally a channel decoupling structure is established.
[0031] The weight configuration module includes: The frequency distribution comparison submodule calls the channel decoupling structure to compare the energy distribution characteristics and time response variation amplitude of multiple channels in each frequency range, determine the consistency of channel energy distribution, and establish the frequency band consistency coefficient; The specific formula for comparing the energy distribution characteristics and time response variation of multiple channels in each frequency range is: ; Calculate the frequency band consistency coefficient; in, For channel In the The normalized energy value at the frequency point, For channel In the The normalized energy value at the frequency point, For channel In the The normalized value of the time response change amplitude at each frequency point, For channel In the The normalized value of the time response change amplitude at each frequency point, For channel With channel Band consistency coefficient in the entire frequency dimension, is the first channel number index to which the energy and time response parameters belong, The second channel number index to which the energy and time response parameters belong. is the index number of the frequency point, is the total number of frequency points.
[0032] The formula used in this paragraph is: The meaning and operation of each parameter in the formula are as follows: Indicates channel In the The energy normalized value of the frequency point, For channel In the The energy normalization value of each frequency point is based on the historical maximum energy of the frequency band of each channel. obtained in the following way, For the original energy, For channel Full belt maximum energy, and Channel 、 In the The normalized value of the time response change amplitude of each frequency point is defined as the ratio of the maximum change of the energy envelope within the frame to the standard maximum change. is the total number of frequency points selected for operation, is the frequency point number, is the frequency band consistency coefficient. In the numerator, the absolute value operation is used to evaluate the inter-channel difference in energy distribution at each frequency point, multiplied by the sum of the normalized values of the time responses of the two channels, to enhance the weighted influence of the difference in the dynamic response segment. The square root of the energy term in the denominator is taken for global normalization. The absolute value sum of the time response difference serves to further normalize the difference. The overall logic shows that the consistency of the energy response is modulated by the dynamic response, and abnormal deviations are suppressed through the full-band amplitude and dynamic distribution. The formula couples the energy difference with the dynamic time response through a dual normalization and weighting mechanism, reflecting the static energy distribution of the channel and the synergy of the dynamic changes of the response at different frequencies. It helps to screen microphone channels in complex environments, improves the accuracy of frequency band consistency judgment, and provides a technical foundation for subsequent beamforming or channel weight allocation.
[0033] The actual parameter acquisition process is as follows: Energy normalization value The original power spectrum data of each microphone channel is directly collected through a spectrum analyzer, and then the maximum value of all frames in the same frequency band is normalized. The frequency point distribution is divided into 5 frequency bands according to the 8kHz sampling system, namely 0–500Hz, 500–1000Hz, 1000–2000Hz, 2000–4000Hz, and 4000–8000Hz. Taking actual data collection as an example, assuming that the energy value of channel 1 in the first frequency band is 0.12 and the maximum is 0.16, then ; The energy of channel 2 in the same frequency band is 0.14, and the maximum is 0.19, then Time response normalized value The ratio of the maximum value of the energy difference between two adjacent frames in the same frequency band and the maximum difference of the whole segment is obtained. The table lists the typical acquisition parameters of the experimental recording period (such as the voice communication conference scene) as follows: Table 1 Normalized energy and time response experimental data table of frequency points As shown in Table 1, , and , are the normalized calculation results after acquisition. The above data is brought into the formula: ; ; ; The result shows that the frequency band consistency coefficient of channel 1 and channel 2 is 0.0425. If the preset frequency band consistency screening reference is 0.08, the value is lower than the reference, indicating that the two channels are highly consistent in the energy and response distribution of each frequency band, and can be preferentially retained in the subsequent weight configuration and speech synthesis. The benefit of the formula is that by coupling the normalized energy difference and the dynamic change of the time response, the screening process has stronger robustness to static energy distribution and dynamic response, and the criterion accuracy is improved when fusing multiple channel signals.
[0034] The main direction screening submodule analyzes the coverage of the frequency band energy in the main direction of the beam based on the frequency band consistency coefficient, calculates the frequency response trend of each channel in the main direction and the angle between the main direction, and selects the direction consistent channels by judging the angle range, to obtain the main direction screening result; The main direction screening submodule based on the frequency band consistency coefficient calls the score matrix of the channel pair in five frequency bands, scans and analyzes the channels one by one, selects the sound source direction as the reference direction, extracts the channels with an angle within 20 degrees from the direction as the direction candidate channels, further calls the energy response distribution values of the candidate channels in each frequency band, constructs the direction frequency response vector, calculates the angle value between the main frequency response trend of the channel and the reference direction vector under the three-dimensional coordinate direction vector, and if the angle is less than 25 degrees, it is judged as direction consistent. In combination with the aforementioned frequency band consistency score, the direction consistent channels are further screened twice. When a channel simultaneously satisfies the frequency band consistency score greater than 0.8 and the direction angle less than 25 degrees, it is added to the direction consistent channel set. For example, the angle between channel 1 and channel 2 is 18 degrees, and the score is 0.85, which meets the conditions; the angle of channel 3 is 28 degrees, and the score is 0.82, which does not meet the angle condition and is excluded; the screening result is recorded in the direction channel marker table, and all channel numbers in the direction consistent channel set are registered to the main direction channel set index, forming an input basis for subsequent synthesis path construction, to obtain the main direction screening result.
[0035] The synthesis configuration adjustment submodule analyzes the synergistic trend of each channel's frequency response and time compensation deviation based on the main direction screening results, adjusts the weight distribution within each frequency range, optimizes the channel frequency band combination, and establishes the speech synthesis configuration; The synthesis configuration adjustment submodule selects all channels that meet the main direction screening results based on the main direction. It extracts their energy response value sequences and time synchronization compensation deviation values within five frequency ranges, constructs channel frequency response vectors and compensation deviation vectors, compares the response synchronization trends between channels on a frequency band-by-frequency basis, and determines whether the energy synchronization deviation between channels persists. If a channel has synchronization differences greater than 0.04 in two or more frequency bands, the frequency band is removed from the synthesis weight configuration. The remaining frequency bands are weighted according to the inverse of the compensation deviation amplitude, with a weight range of 0.1 to 0.3. The synthesis weight matrix is updated based on the weighted cumulative values within the remaining frequency bands of each channel. For example, if channel 1 is valid in three frequency bands, the weighted sum is 0.75, and channel 2 is 0.68. The synthesis contribution values are adjusted proportionally, and the synthesis paths of each channel within the valid frequency band are recombined. The updated channel-band combination parameters, weight allocation coefficients, frequency band shielding indicators, and other information are packaged into a synthesis configuration structure table to ultimately establish the speech synthesis configuration.
[0036] The noise discrimination module includes: The fluctuation monitoring submodule calls the speech synthesis configuration, analyzes the local fluctuation pattern of the energy curve in the continuous speech segment, detects the fluctuation of the energy curve in each time period, calculates the fluctuation frequency and amplitude, and obtains the fluctuation feature value; The fluctuation monitoring submodule calls the speech synthesis configuration to detect the energy curve of the weighted synthesis of multiple channels in the continuous speech segment. It obtains a sequence of energy values in a frame window of 20 milliseconds in length, and advances frame by frame with a sliding step of 10 milliseconds. The maximum and minimum energy points in each frame are extracted and their frame sequence positions are recorded. The amplitude difference and occurrence interval between adjacent extreme points are calculated to determine a local fluctuation event. The total number of fluctuation events occurring within 1 second of speech is counted as the fluctuation frequency, and the average of all fluctuation amplitudes is recorded as the fluctuation intensity indicator. For example, in a continuous speech signal, 37 fluctuation events were detected within 1000 milliseconds, corresponding to an average amplitude change of 0.17. The same operation is performed on the synthesis results output by different microphone combinations, selecting combinations with an average fluctuation frequency between 25 and 45 and an average amplitude between 0.12 and 0.20 as the valid input range. The maximum and minimum values, change difference, and fluctuation duration of each frame within this valid range are compiled to generate a fluctuation trajectory for each combination, convert it into a standard vector form, and finally summarize it into a complete fluctuation feature.
[0037] The extreme value sequence analysis submodule compares the density and direction of extreme value points in adjacent time periods based on the fluctuation characteristics, sorts the time interval sequence of extreme value points, determines the fluctuation stability, and establishes the extreme value distribution coefficient; The specific formula for judging fluctuation stability is: ; Calculate the extreme value distribution stability parameter, determine the fluctuation stability, and establish the extreme value distribution coefficient; in, For the The normalized time interval between extreme points, is the average value of the normalized time intervals of all extreme points, For the The direction jump state value corresponding to the extreme point, is the average value of the state values of all extreme point direction jumps, is the serial index number of the extreme point, is the total number of extreme points in the current analysis segment, is the stability parameter of the extreme value distribution.
[0038] The formula used in this paragraph is: , Parameters and operation logic description: Among them, For the The normalized value of the time interval between an extreme point and the previous extreme point, the normalization method is , Obtained from the absolute time difference of the extreme points (seconds); is the average normalized value of the time intervals of all extreme points, For the The trend jump state quantity of the extreme point is determined by the difference in the value of the signal before and after the extreme point (positive value is recorded as 1, negative value is recorded as -1, and no obvious change is recorded as 0), and normalized ( The value range is [-1,1]). is the mean of the trend values of all extreme points, The number of extreme points in the current frame segment, parameter acquisition and quantification method example: In the actual scene, firstly, energy analysis is performed on the speech segment with a duration of 2 seconds, and the energy extreme points, collect the time when the original extreme points appear, and obtain the time interval array seconds, the maximum time interval is Seconds, normalized The trend jump variable is based on the frame difference method, assuming that the collected data , are all normalized and unitless. , .
[0039] Table 2 extreme point data table As shown in Table 2, By the sampling time sequence of extreme points directly calculated, By normalization processing, For point-by-point collection and then normalized according to the rules.
[0040] Parameter into formula operation process: ; ; ; Regarding the threshold or reference value setting of the extreme value distribution stability parameter, according to the existing literature and standard test data, in the field of multi-channel audio noise reduction, Below 0.2 can be judged as relatively stable fluctuation, more than 0.4 is judged as strong unstable area, so the score of this example Fall in the stable interval, indicating that the fluctuation between extreme values and the trend change are relatively stable as a whole. The formula passes And the cooperative deviation amount of Coupling can accurately reflect the synchronization dispersion of time sequence fluctuation and trend jump, avoid the blind area of traditional only looking at mean or single variation, greatly enhance the recognition accuracy of complex speech signal fluctuation sequence, directly improve the effectiveness of subsequent noise type discrimination and mode adaptive adjustment. The results show that the extreme value distribution sequence of the current segment of speech is in a stable state, and the subsequent extreme value distribution coefficient can be used as a key parameter for sound source separation, noise adaptive processing, etc.
[0041] The characteristic matching sub-module identifies the noise type by analyzing the repetitive characteristics of energy changes according to the extreme value distribution coefficient, and matches the noise processing path according to the noise type, establishes the noise processing parameters; The characteristic matching sub-module extracts the characteristic pair composed of the density classification value and the trend identification value according to the extreme value distribution coefficient, traverses the preset noise distribution characteristic modes of various types of noise in the noise type library, establishes a characteristic vector cosine distance function, respectively calculates the cosine distance between the input extreme value characteristic pair and the noise template, screens all noise template labels with a distance less than 0.25 as matching candidates, and then counts the noise type with the highest frequency in each matching template as the current speech segment noise type judgment result. For example, the current characteristic pair is density level 3 and trend identification -1. Among the traffic noise, keyboard clicking sound and air conditioner low frequency noise templates, the matching frequency with the traffic noise characteristic reaches 70%, and the traffic noise is identified. Then, the path setting bound with the traffic noise is extracted from the noise reduction path library, including the filter passband start and end frequency, frequency band gain adjustment value, channel weighting ratio and other parameter sets. The structured configuration table is arranged and packaged for downstream module calling to establish complete noise processing parameters.
[0042] The speech segment scoring module includes: The frequency coincidence judgment sub-module calls the noise processing parameters, analyzes the coincidence degree of the frequency start of adjacent cycles, calculates the coincidence degree of the frequency start distribution of each cycle splicing point, and establishes the coincidence distribution value. The frequency coincidence judgment sub-module calls the noise processing parameters, extracts the start frequency value of each speech cycle in the frequency dimension, constructs a frequency vector sequence in units of the first 20 frames of each cycle, pairs the frequency points of the first frame of the current cycle with the frequency points of the last frame of the previous cycle one by one, takes the absolute value of the frequency difference value of each pair, and then accumulates and averages the value as the frequency coincidence degree index between two cycles. For example, the start frequency of cycle one is 220Hz and the start frequency of cycle two is 225Hz. The difference is 5Hz. If the average difference is less than 10Hz and meets the condition for more than 5 times continuously, it is judged as a high coincidence degree cycle pair. The frequency coincidence threshold range is set to 0 to 15Hz without weight priority order. By counting the number of cycles that meet the coincidence condition, a coincidence degree trend table is generated. The frequency offset characteristic is additionally labeled for cycle groups with more than two trend changes, and the frequency jump trend of each cycle splicing point is recorded. Finally, the frequency difference matrix and the coincidence statistical quantity of all cycle pairs are integrated to establish a unified coincidence distribution value.
[0043] The connection smoothing comparison sub-module analyzes the connection smoothing degree of the cycle splicing position signal based on the coincidence distribution value, judges the signal amplitude jump state of the splicing point, and establishes a jump smoothing coefficient. The connection smooth alignment sub-module intercepts the signal amplitude values of 5 frames before and after each pair of periodic splicing points based on the coincidence distribution values, compares the signal envelope line change slopes before and after the splicing points, extracts the difference between adjacent frames, constructs an amplitude difference sequence, performs linear fitting on the difference sequence on both sides of the splicing points, obtains the average jump slope, and if the average difference of the front and rear frames is greater than 0.1, it is marked as a jump signal, and if the proportion of such jump events in all splicing points exceeds 40%, it is classified as a low smooth section. For example: among 100 periodic splicing points, 42 jump slope events exceeding 0.1 are detected, the slope values of these events are normalized to construct a jump trend array, and the average jump amplitude fluctuation in the maximum slope range is calculated. Finally, the signal boundaries between the smooth section and the non-smooth section are screened, the proportion of the jump section is calculated and mapped into the splicing position density map, the smooth level score of each splicing point is established, and the overall jump smooth coefficient is finally generated from all scores.
[0044] The consistency adjustment submodule adjusts the speech processing configuration, including speech synthesis and noise processing parameters, according to the jump smooth coefficient, and establishes a noise reduction effect feedback record. The consistency adjustment submodule extracts periodic splicing points with scores lower than the average level according to the jump smooth coefficient, judges the distribution density of the continuous low-score splicing points in the whole speech period by recording the positions of the continuous low-score splicing points, and if there are more than 3 times of score continuous decline in 1000 milliseconds, it is marked as a non-stationary speech segment, classified as a section interval without period stability, and a speech stability index table is established. The frequency and amplitude integrity of the section interval are jointly determined, the frequency fluctuation range and energy difference range of all frames in the section interval are statistically distributed, the amplitude fluctuation upper limit is set to 0.3, and the frequency offset range upper limit is set to 12Hz. If any of the conditions is met, it is judged as a poor consistency section. On this basis, all poor consistency section markers are summarized, the consistency distribution vector of the whole speech period is calculated, and the original synthesis configuration and noise parameter configuration are updated in linkage. The frequency weight is redistributed to the stable section channel, and the filter attenuation ratio is applied to the non-stable section channel. Finally, a new speech synthesis ratio table and a noise reduction factor structure are formed, and a noise reduction effect feedback record is established.
[0045] The above-described embodiments can be implemented in whole or in part by software, hardware (such as a circuit), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0046] It should be understood that the term "and / or" herein merely describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents that the associated objects before and after it are in an "or" relationship, but it can also represent an "and / or" relationship, which can be understood according to the context before and after it.
[0047] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0048] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0049] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0050] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0051] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0052] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0053] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0054] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0055] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An audio noise reduction system with a multi-mic array, characterized in that: The invention comprises a multi-mic array device body (1), wherein the multi-mic array device body (1) is internally equipped with a main board (2), a memory and a processor (3), and a microphone array (4), wherein a computer program is stored in the memory, and the processor executes the computer program to implement an audio noise reduction system of the multi-mic array, wherein the system comprises: The synchronous calibration module uses the audio sampling information to analyze the energy change curves of multiple channels, determines the response offset by matching the cross point distribution with the slope, calculates the offset period and adjusts the channel time starting point to establish the offset adjustment result; The source separation module calls the offset adjustment result, analyzes the energy change of the current frame channel, constructs a time series vector, compares the energy distribution trend, screens consistent channels, and generates a channel decoupling structure; The weight configuration module uses the channel decoupling structure to compare the channel frequency energy distribution and response changes, determine the frequency band directionality, analyze the frequency response and time compensation trends, adjust the frequency weights, and build the speech synthesis configuration; The noise discrimination module uses the speech synthesis configuration to analyze the extreme value distribution and change trend of the speech energy curve, judge the interval stability, identify the noise characteristics, and match the noise processing parameters.
2. The multi-mic array audio noise reduction system according to claim 1, characterized in that: The offset adjustment result includes the channel starting point adjustment position, intersection point sequence distribution, and time offset period; the channel decoupling structure includes the sound source subset division result, the channel coordination distribution matrix, and the signal difference identification parameter; the speech synthesis configuration specifically includes the channel frequency band combination selection, frequency response trend determination, and weight distribution parameters; the noise processing parameters include the extreme point change direction, fluctuation interval stability characteristics, and noise path matching identification.
3. The multi-mic array audio noise reduction system according to claim 1, characterized in that: The synchronous calibration module includes: The curve distribution detection submodule obtains audio sampling information, analyzes the energy change curves of multiple microphone channels, compares the timing arrangement of the intersection points of each channel curve based on the distribution position of the curves on the time axis, calculates the periodic change of the intersection points on the time axis, and establishes the cross-periodic analysis quantity; The trend sequence screening submodule calls the cross-cycle analysis quantity, analyzes the changing relationship between the time position of the intersection point and the slope direction of the adjacent curve, and screens the intersection point sequence with the same curve change trend based on the continuous matching sequence to establish the trend consistent sequence data; The time axis adjustment submodule calculates the time offset period of each channel corresponding to the reference channel based on the trend consistent sequence data, adjusts the starting point position of each channel time axis by compressing the difference amplitude within the processing period, and obtains the offset adjustment result.
4. The multi-mic array audio noise reduction system according to claim 3, characterized in that: The information source separation module includes: The energy sequence construction submodule calls the offset adjustment result, analyzes the signal energy changes of multiple channels in the current frame segment, monitors the energy data of each channel, constructs a time series vector, and establishes an energy change vector group based on the numerical distribution of the time series vector; The collaborative comparison submodule compares the energy distribution trends between each channel based on the energy change vector group, determines the collaborative characteristics of energy changes between multiple channels, and generates channel collaborative comparison coefficients by comparing the collaborative relationships; The sound source classification submodule selects channels with consistent energy change trends according to the channel co-motion comparison coefficient, classifies them into the same sound source subsets, constructs an energy distribution matrix between the sound source subsets, and establishes a channel decoupling structure.
5. The multi-mic array audio noise reduction system according to claim 4, characterized in that: The weight configuration module includes: The frequency distribution comparison submodule calls the channel decoupling structure to compare the energy distribution characteristics and time response variation amplitudes of multiple channels in each frequency range, determines the consistency of channel energy distribution, and establishes a frequency band consistency coefficient; The main direction screening submodule analyzes the coverage of the frequency band energy in the main direction of the beam based on the frequency band consistency coefficient, calculates the frequency response change trend of each channel in the main direction and the main direction angle, and screens the channels with consistent directions by judging the angle range to obtain the main direction screening result; The synthesis configuration adjustment submodule analyzes the synergistic trend of each channel frequency response and time compensation deviation based on the main direction screening results, adjusts the weight distribution within each frequency range, optimizes the channel frequency band combination, and establishes the speech synthesis configuration.
6. The multi-mic array audio noise reduction system according to claim 5, characterized in that: The specific formula for comparing the energy distribution characteristics and time response variation amplitudes of multiple channels within each frequency range is: ; Calculate the frequency band consistency coefficient; in, For channel In the The normalized energy value at the frequency point, For channel In the The normalized energy value at the frequency point, For channel In the The normalized value of the time response change amplitude at each frequency point, For channel In the The normalized value of the time response change amplitude at each frequency point, For channel With channel Band consistency coefficient in the entire frequency dimension, is the first channel number index to which the energy and time response parameters belong, The second channel number index to which the energy and time response parameters belong. is the index number of the frequency point, is the total number of frequency points.
7. The multi-mic array audio noise reduction system according to claim 5, characterized in that: The noise discrimination module includes: The fluctuation monitoring submodule calls the speech synthesis configuration, analyzes the local fluctuation pattern of the energy curve in the continuous speech segment, detects the fluctuation of the energy curve in each time period, counts the fluctuation frequency and amplitude, and obtains the fluctuation feature value; The extreme value sequence analysis submodule compares the density and direction change trend of extreme value points in adjacent time periods based on the fluctuation characteristic quantity, sorts the time interval sequence of extreme value points, determines the fluctuation stability, and establishes the extreme value distribution coefficient; The characteristic matching submodule identifies the noise type by analyzing the repetitive characteristics of energy changes based on the extreme value distribution coefficient, matches the noise processing path according to the noise type, and establishes noise processing parameters.
8. The multi-mic array audio noise reduction system according to claim 7, characterized in that: The specific formula for judging the fluctuation stability is: ; Calculate the extreme value distribution stability parameter, determine the fluctuation stability, and establish the extreme value distribution coefficient; in, For the The normalized time interval between extreme points, is the average value of the normalized time intervals of all extreme points, For the The direction jump state value corresponding to the extreme point, is the average value of the state values of all extreme point direction jumps, is the serial index number of the extreme point, is the total number of extreme points in the current analysis segment, is the stability parameter of the extreme value distribution.
9. The multi-mic array audio noise reduction system according to claim 1, characterized in that: The system further comprises: The segment scoring module uses the noise processing parameters to analyze the degree of overlap of adjacent cycles at the frequency start, compare the signal connection smoothness at the cycle splicing position, determine the continuous distribution position of the structural change point, construct a consistency index for the speech cycle segment, and adjust the speech processing configuration based on the integrity of the segment to establish a noise reduction effect feedback record; The noise reduction effect feedback record specifically refers to the splicing smoothness status, cycle consistency index, and voice processing configuration adjustment information.
10. The multi-mic array audio noise reduction system according to claim 9, characterized in that: The segment scoring module includes: The frequency overlap judgment submodule calls the noise processing parameters, analyzes the overlap degree of adjacent cycles at the frequency start, calculates the overlap degree of the frequency start distribution of each cycle splicing point, and establishes the overlap distribution value; The connection smoothness comparison submodule analyzes the connection smoothness of the periodic splicing position signal based on the overlap distribution value, determines the signal amplitude jump state of the splicing point, and establishes a jump smoothing coefficient; The consistency adjustment submodule selects non-stationary segments based on the jump state of the splicing point according to the jump smoothing coefficient, constructs a consistency index of the speech cycle segment, adjusts the speech processing configuration, including speech synthesis and noise processing parameters, and establishes a noise reduction effect feedback record based on the segment integrity.