Dual-group human voice synchronous acquisition frequency domain noise reduction and restoration method and system
By constructing a dual-source frequency domain coordinate system through simultaneous acquisition of two sets of human voices, and combining four-quadrant partitioning recognition and frequency domain reconstruction, the problem of distinguishing between physiological distortion and environmental noise under extreme emotions is solved, achieving accurate noise reduction and restoration of voice signals, and ensuring the reliability of emergency communication.
Patent Information
- Application Number
- CN202511141920.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing technologies cannot effectively distinguish between physiological distortion and environmental noise under extreme emotional states, leading to excessive noise reduction or loss of key information in voice signals, which affects the accuracy and efficiency of emergency communications.
By simultaneously acquiring two sets of human voices, a dual-source frequency domain coordinate system is constructed. Combined with four-quadrant partitioning to identify the type of emotional distortion, targeted frequency domain reconstruction and correction are performed to distinguish between physiological distortion and environmental noise, and to retain frequency components related to key information.
It significantly improves the noise reduction and restoration accuracy of voice signals under extreme emotional states, ensuring the accurate transmission of key information in emergency communications and avoiding information loss caused by excessive or insufficient noise reduction.
Smart Images

Figure CN120823841B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice communication technology, and more specifically, to a method and system for frequency domain noise reduction and restoration of simultaneous acquisition of dual human voices. Background Technology
[0002] In the field of voice communication, especially in critical scenarios such as emergency rescue and medical emergency services, the quality of the caller's voice signal directly affects the accuracy of information transmission and response efficiency. Under extreme emotional states, voice signals are often accompanied by complex physiological distortions. How to effectively reduce noise while retaining key information has become a core challenge in improving the reliability of voice communication.
[0003] In the prior art, Chinese patent application CN114550752A discloses a method and device for automatic noise reduction and waveform node emotion analysis of two-way speech signals. This method classifies speech emotions by segmenting the sound waveform signal into time segments and comparing them with a pre-generated emotion training set. Its core lies in associating emotional states through waveform features, while integrating automatic noise reduction and waveform modeling functions. Chinese patent application CN119964537A discloses a mixing noise reduction method and system. This system determines whether passengers in a vehicle are in a negative emotional state through emotion recognition. If the emotion is anxiety, it collects sound signals, extracts noise feature vectors, and generates mixed sound bands to neutralize the noise and reduce its impact on emotions.
[0004] However, the aforementioned existing technologies still have specific shortcomings: the emotion analysis in the Chinese patent application with publication number CN114550752A relies on waveform comparison, which can only achieve emotion classification and basic noise reduction, but cannot identify physiological composite distortions under extreme emotions, such as frequency domain compression caused by vocal cord tension and periodic frequency shifts caused by emotional tremors. It is easy to misjudge the compressed high-frequency enhanced components as noise; although the Chinese patent application with publication number CN119964537A incorporates emotion-adjusting noise reduction strategies, it focuses on the interaction between environmental noise and emotions, and does not distinguish the essential difference between physiological distortion and real noise. It may misjudge the frequency fluctuations of emotional tremors as environmental interference and suppress them. More importantly, the composite distortion frequencies under extreme emotions often overlap with the fundamental frequencies of key information, such as location and numbers. The indiscriminate processing of existing technologies will lead to "over-noise reduction," causing the loss of important information. In emergency communication scenarios, this may delay rescue decisions and even lead to serious consequences. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of existing technologies, this invention provides a frequency domain noise reduction and restoration method and system for simultaneous acquisition of dual-group human voices. By constructing a dual-source frequency domain coordinate system through dual-group acquisition and combining four-quadrant partitioning to identify the type of emotional distortion, targeted frequency domain reconstruction and correction are performed to effectively distinguish between physiological distortion and environmental noise. This significantly improves the noise reduction and restoration accuracy of extreme emotional speech, providing reliable technical support for the accurate transmission of key information in emergency communications.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] Frequency domain noise reduction and restoration methods for simultaneous dual-group human voice acquisition include:
[0008] Two sets of synchronized human voice signals are acquired through two sets of acquisition points, and a dual-source frequency domain coordinate system is constructed. The two sets of human voice signals are then mapped onto the constructed dual-source frequency domain coordinate system to form two sets of energy distribution points.
[0009] In the dual-source frequency domain coordinate system, the spatial relationship between the two sets of energy distribution points is calculated to obtain an index that measures the degree of distortion of the dual-set human voice synchronization signals and to identify whether the current speech state of the dual-set human voice synchronization signals is in an emotional distortion state.
[0010] In the dual-source frequency domain coordinate system, the four-quadrant distortion recognition region is divided, the quadrant affiliation of the two sets of energy distribution points corresponding to the two sets of human voice synchronization signals identified as emotional distortion state is determined, and the corresponding distortion type label is generated.
[0011] Based on the generated distortion type label, the first group of frequency domain signals FreA and the second group of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals. For the different types of frequency domain sub-signals, corresponding frequency domain reconstruction processing and correction are performed to obtain the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B.
[0012] The modified first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B are divided into frequency bands, and the frequency bands with significantly reduced energy are identified and enhanced.
[0013] Furthermore, the dual-set human voice synchronization signal includes a first set of human voice signals and a second set of human voice signals;
[0014] The method for forming two sets of energy distribution points includes: performing Fast Fourier Transform on the first set of human voice signals and the second set of human voice signals respectively to obtain the first set of frequency domain signals and the second set of frequency domain signals; using the amplitude of each frequency component in the first set of frequency domain signals as the Y coordinate and the frequency value as the X coordinate to form the first set of energy distribution points; and using the amplitude of each frequency component in the second set of frequency domain signals as the Y coordinate and the frequency value as the X coordinate to form the second set of energy distribution points.
[0015] Furthermore, the method for calculating the spatial relationship between the two sets of energy distribution points to obtain an index measuring the degree of distortion of the two sets of voice synchronization signals includes:
[0016] Based on statistical analysis of historical normal speech data, a standard elliptical region is defined in the dual-source frequency domain coordinate system. The shortest distances from two sets of energy distribution points to the boundary of the standard elliptical region are calculated and denoted as the first deviation value and the second deviation value. The difference between the first deviation value and the second deviation value, Diff, is then calculated. AB and average Mean AB Difference Diff AB and average Mean AB As an indicator to measure the degree of distortion in two sets of synchronized human voice signals.
[0017] Furthermore, the method for identifying whether the speech state of the current dual-group synchronized voice signals is in an emotionally distorted state includes:
[0018] Set the number of consecutive time windows Cou C Difference threshold Thr D and average threshold Thr M When Cou is continuous C Difference within a time window (Diff) AB All are greater than the difference threshold Thr D And the average Mean AB All are greater than the average threshold Thr M At that time, the speech state of the dual-group human voice synchronization signal is marked as emotional distortion state.
[0019] Furthermore, the method for generating the distortion type marker includes:
[0020] In the dual-source frequency domain coordinate system, four quadrant regions are divided with the origin O as the center. The first quadrant is defined as the frequency domain compression distortion region, the second quadrant is defined as the mixed distortion region, the third quadrant is defined as the jitter distortion region, and the fourth quadrant is defined as the normal region.
[0021] The distribution ratio of the first and second energy distribution points of the dual sets of human voice synchronization signals identified as emotionally distorted in the four quadrant regions was statistically analyzed.
[0022] Based on the distribution ratio of the first group of energy distribution points and the second group of energy distribution points in the four quadrant regions, the corresponding distortion type label is automatically generated.
[0023] Furthermore, the method for automatically generating corresponding distortion type labels based on the distribution ratio of the first group of energy distribution points and the second group of energy distribution points in the four quadrant regions includes:
[0024] Determine the key distribution quadrants of the first and second sets of energy distribution points; based on the key distribution quadrants of the first and second sets of energy distribution points, generate distortion type labels, which include frequency domain compression distortion labels, tremor distortion labels, mixed distortion labels, and normal speech labels.
[0025] Furthermore, the method for determining the key distribution quadrants of the first group of energy distribution points and the second group of energy distribution points includes: setting a distribution threshold T. ratio If the proportion of the first group of energy distribution points falling into the i-th quadrant region exceeds the distribution threshold T, ratio If the i-th quadrant is the key distribution quadrant for the first group of energy distribution points, then it indicates that the i-th quadrant is the key distribution quadrant for the first group of energy distribution points; otherwise, it indicates that the first group of energy distribution points has no key distribution quadrant. If the proportion of the second group of energy distribution points falling into the j-th quadrant exceeds the distribution threshold T, then the i-th quadrant is the key distribution quadrant for the first group of energy distribution points. ratio If the value is true, it means that the j-th quadrant is the key distribution quadrant of the second group of energy distribution points; otherwise, it means that the second group of energy distribution points has no key distribution quadrant; where i and j represent quadrant numbers, i = 1, 2, 3, 4, j = 1, 2, 3, 4.
[0026] Furthermore, the step of segmenting the first group of frequency domain signals FreA and the second group of frequency domain signals FreB according to the generated distortion type marker to generate different types of frequency domain sub-signals includes: if the distortion type marker is a frequency domain compression distortion marker, then a compression distortion sub-signal is generated; if the distortion type marker is a jitter distortion marker, then a jitter distortion sub-signal is generated; if the distortion type marker is a mixed distortion marker, then a mixed distortion sub-signal is generated; if the distortion type marker is a normal speech marker, then a normal sub-signal is generated.
[0027] Furthermore, the method for obtaining the corrected first set of frequency domain signals Fre”A and the second set of frequency domain signals Fre”B includes:
[0028] For different types of frequency domain sub-signals, corresponding frequency domain reconstruction processing is performed to obtain the first set of frequency domain signals Fre'A and the second set of frequency domain signals Fre'B after processing; the first set of frequency domain signals Fre'A and the second set of frequency domain signals Fre'B after processing are corrected to obtain the first set of frequency domain signals Fre”A and the second set of frequency domain signals Fre”B after correction.
[0029] The method for correcting the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B includes: marking the processing error points in the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B, and adjusting the relevant parameters of frequency domain reconstruction for the frequency points marked as processing error points.
[0030] A frequency domain noise reduction and restoration system for simultaneous acquisition of dual-group human voices, used to implement the aforementioned frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices, the system comprising:
[0031] Mapping module: Acquires two sets of synchronized human voice signals through two sets of acquisition points, constructs a dual-source frequency domain coordinate system, and maps the two sets of human voice signals onto the constructed dual-source frequency domain coordinate system to form two sets of energy distribution points;
[0032] Distortion recognition module: used to calculate the spatial relationship between two sets of energy distribution points in the dual-source frequency domain coordinate system, obtain an index to measure the degree of distortion of the two sets of human voice synchronization signals, and identify whether the current speech state of the two sets of human voice synchronization signals is in an emotional distortion state.
[0033] Quadrant marking module: used to divide the four-quadrant distortion recognition region in the dual-source frequency domain coordinate system, determine the quadrant affiliation of the two sets of energy distribution points corresponding to the two sets of human voice synchronization signals identified as emotional distortion state, and generate corresponding distortion type markings;
[0034] Frequency domain correction module: Based on the generated distortion type label, the first group of frequency domain signals FreA and the second group of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals. For different types of frequency domain sub-signals, corresponding frequency domain reconstruction processing and correction are performed to obtain the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B.
[0035] Frequency band enhancement module: used to divide the modified first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B into frequency bands, identify frequency bands with significantly reduced energy, and perform enhancement processing on the frequency bands with significantly reduced energy.
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] This invention utilizes a frequency-space fusion analysis framework constructed through dual-group synchronous acquisition to achieve accurate identification and targeted processing of speech composite distortion under extreme emotional states. By leveraging the spatial relationship analysis of energy distribution points in the dual-source frequency domain coordinate system, it effectively distinguishes between physiological composite distortion and environmental noise, avoiding the misjudgment and suppression of high-frequency enhancement components caused by frequency domain compression and periodic frequency shifts caused by emotional tremors as noise. Through distortion type labeling generated by four-quadrant partitioning, it can perform adaptive frequency domain reconstruction and correction for different composite distortion features. Combined with frequency band enhancement processing, it can effectively preserve and restore frequency components related to key information, preventing the loss or blurring of key information due to excessive or insufficient noise reduction. This significantly improves the noise reduction and restoration accuracy of speech signals under extreme emotional states, providing stable and reliable technical support for scenarios such as emergency communications that rely on clear speech to convey key information. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart illustrating the principle of the frequency domain noise reduction and restoration method for simultaneous acquisition of dual human voices in this invention.
[0040] Figure 2 This is a schematic diagram of the quadrant distortion recognition region division according to the present invention;
[0041] Figure 3 This is a flowchart illustrating the frequency domain reconstruction processing and correction principle of the present invention;
[0042] Figure 4 This is a functional block diagram of the frequency domain noise reduction and restoration system for simultaneous acquisition of dual human voices in this invention. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Example 1
[0045] Please see Figure 1 As shown, this embodiment provides a frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices, including:
[0046] Step S10: Acquire dual sets of human voice synchronization signals through dual sets of acquisition points, construct a dual-source frequency domain coordinate system, and map the dual sets of human voice signals onto the constructed dual-source frequency domain coordinate system to form two sets of energy distribution points.
[0047] Further, step S10 includes:
[0048] Step S11: Acquire dual sets of human voice synchronization signals through dual sets of acquisition points, the dual sets of acquisition points including the first set of acquisition points Pos A The second set of sampling points Pos B The dual-set voice synchronization signal includes the first set of voice signals Sig acquired by the first set of acquisition points. A The second set of human voice signals Sig acquired from the second set of acquisition points B ;
[0049] Step S12, using the first set of sampling points Pos A Establish a dual-source frequency domain coordinate system with the origin O, and start from the first set of sampling points Pos A Pointing to the second set of sampling points Pos B The direction of X is set as the positive direction of the X-axis, and the direction that passes through the origin O and is perpendicular to the X-axis is set as the Y-axis;
[0050] Step S13: Perform frequency domain transformation on the two sets of human voice synchronization signals and map them into the constructed dual-source frequency domain coordinate system to form two sets of energy distribution points.
[0051] Specifically, step S10 aims to provide a basic framework for the identification and processing of complex speech distortion under extreme emotions by combining spatial information and frequency domain features, solving the problem that traditional single-group acquisition systems cannot distinguish between distortion and noise due to the lack of spatial reference. Dual-group synchronous human voice signals are acquired through dual-group acquisition devices. Traditional single-group acquisition systems, lacking a reference from the other signal, cannot distinguish physiological distortion through spatial differences. Dual-group synchronous acquisition provides a basis for differentiation through the spatial correlation of the two signals. Without step S11, the subsequently constructed dual-source frequency domain coordinate system cannot be established due to the lack of effective input signals, and the basis for utilizing spatial information in the entire method will not exist, resulting in the inability to identify distortion features from a spatial dimension. The establishment of the dual-source frequency domain coordinate system is based on the requirement of spatial position quantization, with the origin selected as Pos. A This is to use one set of signals as a reference to facilitate the calculation of the relative spatial relationship between the two sets of signals. The setting of the X-axis direction is directly related to the physical location line connecting the two sets of acquisition points, so that the X-axis coordinate can reflect the frequency value, and in the subsequent mapping, the X-axis represents the frequency, and also implies spatial location information. The Y-axis is set to be perpendicular to the X-axis and is used to characterize the amplitude of the signal, i.e., the energy intensity. This setting organically integrates the frequency domain characteristics, i.e., frequency and amplitude, with the spatial location characteristics, i.e., relative distance, solving the defect of traditional frequency domain analysis that only focuses on a single dimension, i.e., frequency-amplitude, and ignores spatial information.
[0052] The method for forming two sets of energy distribution points specifically includes: processing the first set of human voice signals Sig... A The second group of human voice signals Sig BPerform Fast Fourier Transform (FFT) to obtain the first set of frequency domain signals FreA and the second set of frequency domain signals FreB. Use the amplitude of each frequency component in the first set of frequency domain signals FreA as the Y coordinate and the frequency value as the X coordinate to form the first set of energy distribution points. Similarly, use the amplitude of each frequency component in the second set of frequency domain signals FreB as the Y coordinate and the frequency value as the X coordinate to form the second set of energy distribution points. In the formation of energy distribution points, each frequency component corresponds to a point. For example, FreA contains a frequency f1, such as 500Hz, and an amplitude A1, such as 0.8V. This component corresponds to the point (500, 0.8) in the coordinate system. All such points form the first set of energy distribution points. This mapping transforms abstract frequency domain data into a geometric point set that can be intuitively analyzed. This allows the frequency-energy distribution difference between the two sets of signals to be reflected through the spatial distribution pattern of the point set. For example, when emotions tremble, the periodic fluctuations of the two sets of point sets in the X-axis direction are consistent, while the fluctuations of environmental noise do not have this consistency. This solves the problem that signal features are difficult to visualize and spatially correlated in traditional frequency domain analysis. If step S13 is missing, the frequency domain features of the two sets of signals will still exist in the form of abstract data, and subsequent deviation calculations cannot be performed through spatial geometric relationships. For example, the distance to the standard elliptical region in step S20 leads to the loss of quantitative basis for distortion identification.
[0053] Step S10 achieves the transformation of speech signals from the time-space domain to the frequency-space fusion domain through the synergy of dual-group synchronous acquisition, coordinate system construction, and frequency domain mapping. Dual-group synchronous acquisition provides a spatial reference basis, dual-source frequency domain coordinate system provides a fusion analysis framework, and frequency domain mapping transforms signal features into geometric features. The combination of these three features enables speech composite distortion under extreme emotions, i.e., frequency domain compression and periodic frequency shift, to be identified through the spatial pattern of energy distribution points. Traditional methods cannot achieve this due to the lack of spatial information and fusion framework.
[0054] Step S20: In the dual-source frequency domain coordinate system, calculate the spatial relationship between the two sets of energy distribution points to obtain an index that measures the degree of distortion of the dual-set human voice synchronization signals and identify whether the current speech state of the dual-set human voice synchronization signals is an emotional distortion state.
[0055] Further, step S20 includes:
[0056] Step S21: Based on the statistical analysis of historical normal speech data, a standard elliptical region is set in the dual-source frequency domain coordinate system;
[0057] Step S22: Calculate the shortest distance from the two sets of energy distribution points to the boundary of the standard elliptical region, and denot it as the first deviation value Dev. A Second deviation value Dev B ;
[0058] Step S23, calculate the first deviation value Dev A With the second deviation value Dev B The difference Diff AB and average Mean AB Difference Diff AB and average Mean AB As an indicator to measure the degree of distortion in dual-group synchronized human voice signals;
[0059] Step S24, set the number of consecutive time windows Cou C Difference threshold Thr D and average threshold Thr M When Cou is continuous C Difference within a time window (Diff) AB All are greater than the difference threshold Thr D And the average Mean AB All are greater than the average threshold Thr M At that time, the speech state of the dual-group human voice synchronization signal is marked as emotional distortion state.
[0060] Specifically, step S20 aims to establish an objective distortion recognition standard by quantitatively analyzing the deviation between energy distribution points and normal speech features, thus solving the problem of misjudgment caused by traditional algorithms relying on a single feature threshold. Step S21 establishes a standard elliptical region in a dual-source frequency domain coordinate system based on statistical analysis of historical normal speech data. The specific implementation process is as follows: collect normal speech samples of different genders, ages, and speech rates, with a sample size of no less than 1000 hours, covering non-extreme emotional states such as calmness and slight agitation. These samples are converted into two sets of energy distribution points using the method in step S10. An ellipse fitting algorithm is then used to spatially cluster the energy distribution points of all normal samples, and the resulting ellipse is the standard elliptical region. The major axis of the ellipse is distributed along the X-axis, which is the frequency axis. Its length is determined by statistical analysis of the frequency range of normal speech, and the 95th percentile of the difference between the maximum and minimum frequencies of all samples is taken as the length of the major axis. The minor axis is distributed along the Y-axis, which is the amplitude axis. Its length is determined by statistical analysis of the amplitude fluctuation range of normal speech, and three times the standard deviation of the amplitude of all samples is taken as the length of the minor axis. The center of the ellipse is the intersection of the mean frequency and mean amplitude of the energy distribution points of all normal samples. The ellipse is chosen as the standard region because the frequency and amplitude distribution of normal speech exhibits a continuous distribution characteristic that approximates an ellipse. The frequency range is relatively wide, and the amplitude changes smoothly with frequency. The energy distribution points of distorted signals under extreme emotions will deviate from this region. The geometric characteristics of the ellipse can better fit the natural distribution law of normal speech. If a rectangular or circular region is used, it will reduce the subsequent recognition accuracy due to excessive inclusion of unnaturally distributed edge points or omission of normal fluctuation range.
[0061] Step S22: Calculate the shortest distance from the two sets of energy distribution points to the boundary of the standard elliptical region, and denot it as the first deviation value Dev. A Second deviation value Dev B The calculation method is as follows: For each point in the first group of energy distribution points, calculate the shortest distance from each point to the boundary of the ellipse using the standard equation of the ellipse; a positive distance indicates the point is outside the ellipse, a negative distance indicates it is inside, and a larger absolute value indicates a greater degree of deviation. The arithmetic mean of the distances to all energy distribution points in the first group is taken as Dev. A Similarly, the average value of the second group of energy distribution points is calculated as Dev. B This process transforms abstract signal characteristics into comparable numerical indicators by quantifying the degree of deviation of each energy point from the normal region. This solves the problem of traditional methods judging distortion based solely on subjective experience. Without calculating the deviation value, it is impossible to objectively measure the degree of signal distortion, resulting in a lack of quantitative basis for subsequent indicators.
[0062] In step S23, Diff AB For Dev A Subtract Dev B The absolute value of Mean AB For (Dev) A +Dev B ) / 2. Diff AB Reflecting the consistency of the distortion patterns of the two sets of signals, under normal speech conditions, the deviation between the two sets of signals due to spatial differences is small and the Diff... AB Stable within a smaller range, while physiological distortions under extreme emotions can cause both sets of signals to exhibit similar deviation patterns, Diff AB Keeping the value relatively small, environmental noise may cause a significant increase in the deviation of a certain set of signals, thus increasing the Diff value. AB Enlarge; Mean AB This reflects the overall deviation between the two sets of signals; a greater deviation indicates that the signal is more likely to be distorted. Combining both can distinguish between physiological distortion and environmental noise. For example, frequency fluctuations caused by emotional tremors can cause both sets of signals to deviate from the elliptical region simultaneously. AB Increase and Diff AB The smaller volume, and the wind noise near the microphone on one side, can cause a significant deviation in one set of signals, resulting in a diff. AB If only a single indicator is used, it will be impossible to distinguish between the two situations, leading to misjudgment.
[0063] In step S24, the time window length is determined based on the average duration of the human voice syllables, ensuring that each window contains the complete fundamental frequency period of the speech, typically between 10 and 30 milliseconds; Cou C The value ranges from 2 to 5, and is determined based on the stability requirements of the speech signal; ThrD and Thr M Based on historical data statistics, at least 300 hours of audio samples containing extreme emotional distortion were collected, and the Diff of these samples under emotional distortion conditions was calculated. AB and Mean AB Distribution, take Diff AB The 20th to 30th percentile of the distribution is used as Thr D To ensure that over 90% of physiologically distorted samples are Diff AB Less than this value; take Mean AB The 40th to 60th percentile values of the distribution are used as Thr M This ensures that over 90% of normal speech samples have a MeanAB value less than this threshold. The emotion distortion determination mechanism filters out accidental deviations caused by transient interference through the requirement of temporal continuity, improving the stability of recognition. Without continuous window judgment, brief signal fluctuations would be misjudged as emotion distortion, affecting the accuracy of subsequent processing.
[0064] Step S20 achieves objective quantitative identification of emotional distortion states through the coordinated efforts of setting a standard elliptical region, calculating deviation values, analyzing comprehensive indicators, and verifying temporal continuity. The standard elliptical region provides a benchmark framework for distortion judgment based on a large amount of historical data; deviation value calculation transforms spatial distribution characteristics into numerical indicators; the combined analysis of difference and average values distinguishes between physiological distortion and environmental noise; and continuous time window verification filters out instantaneous interference. The quantitative indicators in step S20 transform the spatial features constructed in step S10 into identifiable state labels. The combination of steps S10 and S20 upgrades distortion identification from qualitative description to quantitative analysis. Without step S20, the energy distribution points generated in step S10 cannot be transformed into clear distortion state labels, causing the quadrant assignment analysis in step S30 to lose its premise. The entire processing flow would be unable to determine whether distortion restoration is necessary, ultimately leaving the noise reduction processing of extreme emotional speech at the level of traditional indiscriminate suppression.
[0065] Step S30: Divide the four-quadrant distortion recognition region in the dual-source frequency domain coordinate system, determine the quadrant affiliation of the two sets of energy distribution points corresponding to the two sets of human voice synchronization signals identified as emotional distortion state, and generate the corresponding distortion type label.
[0066] Further, step S30 includes:
[0067] Step S31, please refer to Figure 2 As shown, in the dual-source frequency domain coordinate system, four quadrant regions are divided with the origin O as the center. The first quadrant is defined as the frequency domain compression distortion region, the second quadrant is defined as the mixed distortion region, the third quadrant is defined as the jitter distortion region, and the fourth quadrant is defined as the normal region.
[0068] Step S32: Statistically analyze the distribution ratio of the first group of energy distribution points and the second group of energy distribution points in the four quadrant regions corresponding to the dual groups of human voice synchronization signals identified as emotionally distorted.
[0069] Step S33: Based on the distribution ratio of the first group of energy distribution points and the second group of energy distribution points in the four quadrant regions, automatically generate the corresponding distortion type label;
[0070] Specifically, step S30 aims to achieve accurate classification of different types of composite distortion through geometric partitioning of spatial quadrants, solving the problem of insufficient targeted processing caused by confusion of composite distortion features in traditional algorithms. Step S31 divides the dual-source frequency domain coordinate system into four quadrant regions centered at the origin O: the first quadrant is defined as the frequency domain compression distortion region, the second quadrant as the mixed distortion region, the third quadrant as the jitter distortion region, and the fourth quadrant as the normal region. The specific criteria for quadrant division are as follows: the positive X-axis corresponds to the frequency range above the center of the standard ellipse, and the negative X-axis corresponds to the frequency range below the center of the standard ellipse; the positive Y-axis corresponds to the amplitude range above the center of the standard ellipse, and the negative Y-axis corresponds to the amplitude range below the center of the standard ellipse. Frequency domain compression distortion manifests as a narrowing of the frequency range and abnormally enhanced amplitude in the high-frequency band, with energy distribution points concentrated in the first quadrant bounded by the positive X-axis and positive Y-axis. Periodic frequency shifts caused by emotional tremor manifest as frequency fluctuations in a lower range and unstable amplitude, with energy distribution points concentrated in the third quadrant bounded by the negative X-axis and negative Y-axis. Mixed distortion, containing both compression and tremor features, has its point set distributed across multiple quadrants, with at least one quadrant being the second quadrant (negative X-axis and positive Y-axis), corresponding to a mixed feature of enhanced amplitude in the low-frequency band. Normal speech energy distribution points are mainly concentrated in the fourth quadrant (positive X-axis and negative Y-axis), corresponding to frequencies and amplitudes within the normal fluctuation range. The four-quadrant division, rather than other partitioning methods, was chosen because the three typical states under extreme emotions and the normal state exhibit orthogonal characteristics in the direction of frequency and amplitude deviation. These three typical states are compression, tremor, and mixed distortion. The orthogonality of the quadrants maximizes the differentiation of spatial distribution differences between different distortion types. If a sector or multi-region division were used, overlapping boundaries would lead to classification ambiguity.
[0071] Step S32 involves statistically analyzing the distribution ratio of the first and second energy distribution points of the dual-group synchronized voice signals identified as being in an emotionally distorted state across the four quadrants. Specifically, for each dual-group synchronized voice signal marked as being in an emotionally distorted state, the corresponding first and second energy distribution points are extracted. The quadrant of the coordinates (f, a) of each energy distribution point is determined, where f is the abscissa and a is the ordinate. If f is greater than the ellipse's center frequency and a is greater than the ellipse's center amplitude, it is classified as the first quadrant; if f is less than the ellipse's center frequency and a is greater than the ellipse's center amplitude, it is classified as the second quadrant; if f is less than the ellipse's center frequency and a is less than the ellipse's center amplitude, it is classified as the third quadrant; and if f is greater than the ellipse's center frequency and a is less than the ellipse's center amplitude, it is classified as the fourth quadrant. The number of energy distribution points in each quadrant is counted, and the number of energy distribution points in each quadrant is divided by the total number of energy distribution points in that group to obtain the distribution ratio across the four quadrants. For example, if the first group of energy distribution points has 100 points, with 60 in the first quadrant, 20 in the second quadrant, 10 in the third quadrant, and 10 in the fourth quadrant, then the distribution ratios in the first to fourth quadrants are 60%, 20%, 10%, and 10%, respectively. The reason for using statistical ratios rather than single points is that speech signals under extreme emotions may exhibit brief, normal fluctuations. Ratios reflect the overall distribution trend and avoid misclassification due to individual outliers. If only a single point is assigned, the classification stability would be reduced due to instantaneous fluctuations.
[0072] The method for generating the corresponding distortion type label in step S33 involves determining the key distribution quadrants of the first and second sets of energy distribution points; and generating distortion type labels based on the key distribution quadrants of the first and second sets of energy distribution points. The method for determining the key distribution quadrants of the first and second sets of energy distribution points includes setting a distribution threshold T. ratio The distribution threshold was determined through statistical analysis of historical distortion samples. 300 hours of speech samples were collected, each containing pure frequency domain compression, pure tremor, and mixed distortion. The proportion of energy distribution points in the corresponding feature quadrants for each type of sample was calculated, and the minimum proportion of feature quadrants across all samples was taken as T. ratio The percentage is typically between 50% and 70%, ensuring that over 95% of purely distorted samples can be correctly identified. If the proportion of points in the first group of energy distribution points falling into the i-th quadrant exceeds the distribution threshold T, then... ratio If the i-th quadrant is the key distribution quadrant for the first group of energy distribution points, then it indicates that the i-th quadrant is the key distribution quadrant for the first group of energy distribution points; otherwise, it indicates that the first group of energy distribution points has no key distribution quadrant. If the proportion of the second group of energy distribution points falling into the j-th quadrant exceeds the distribution threshold T, then the i-th quadrant is the key distribution quadrant for the first group of energy distribution points. ratioIf the value is true, it means that the j-th quadrant is the key distribution quadrant of the second group of energy distribution points; otherwise, it means that the second group of energy distribution points has no key distribution quadrant; where i and j represent quadrant numbers, i = 1, 2, 3, 4, j = 1, 2, 3, 4.
[0073] The method for generating distortion type labels based on the key distribution quadrants of the first and second sets of energy distribution points includes: if the key distribution quadrants of both sets are in the first quadrant, a frequency domain compression distortion label is generated; if the key distribution quadrants of both sets are in the third quadrant, a jitter distortion label is generated; if the key distribution quadrants of the first and second sets are different, and at least one key distribution quadrant is in the second quadrant, a mixed distortion label is generated; if neither set has a key distribution quadrant, or if both sets have a key distribution quadrant, a normal speech label is generated. If only one set has a key distribution quadrant and the other set does not, the difference between the proportion of characteristic quadrants in the set with the key distribution quadrant and the highest proportion of all quadrants in the other set is calculated. If the difference exceeds 20%, the type of the set with the key distribution quadrant is used; otherwise, it is labeled as mixed distortion.
[0074] For example, in a certain emotional distortion signal, if the proportion of the first group of energy distribution points in the first quadrant is 65%, and the second group is 70%, with a Tratio set to 60%, then the emphasis of both distribution points is in the first quadrant, generating a frequency domain compression distortion label. In another signal, if the proportion of the first group in the third quadrant is 62%, and the second group is 58%, with a Tratio set to 55%, then a tremor distortion label is generated. If the first group emphasizes the first quadrant and the second group emphasizes the second quadrant, then a mixed distortion label is generated. This judgment mechanism verifies the consistency of the distribution of two groups of signals, reducing misjudgments caused by noise interference from a single signal. If only the distribution of a single group of signals is considered, the type may be misjudged due to local distortion of that group of signals.
[0075] Step S30 achieves precise classification from emotional distortion state to specific distortion type through the synergy of four-quadrant partitioning, distribution ratio statistics, and multi-condition judgment. Four-quadrant partitioning provides spatial feature templates for different distortion types, distribution ratio statistics quantifies the overall distribution trend, and multi-condition judgment, combined with the consistency of the two sets of signals, improves classification reliability. Step S30, along with the energy distribution points generated in step S10 and the emotional distortion state identified in step S20, works in synergy. The spatial mapping in step S10 provides specific analysis objects for quadrant partitioning, and the state recognition in step S20 limits the classification range of the emotional distortion signal. These three elements combine to form a progressive analysis chain of "feature mapping—state recognition—type subdivision." Without step S30, step S40 cannot determine the specific distortion type and can only use a general processing method, leading to a decreased effectiveness in processing mixed distortions. For example, when compression and trembling coexist, a single processing method will be ineffective in addressing both types of distortion and cannot simultaneously meet the needs of repairing both types of distortion. Step S30 provides precise type guidance for subsequent processing, enabling dynamic adaptation of processing strategies based on different distortion characteristics and avoiding the limitations of "one-size-fits-all" processing.
[0076] Step S40: According to the generated distortion type label, the first group of frequency domain signals FreA and the second group of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals. For the different types of frequency domain sub-signals, the corresponding frequency domain reconstruction processing and correction are performed to obtain the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B.
[0077] Step S40 aims to achieve accurate restoration of speech signals with extreme emotions by specifically processing the frequency domain features of different distortion types, thus solving the problem of loss of key information or aggravated distortion caused by the traditional noise reduction algorithm using a single processing strategy for all signals.
[0078] Please see Figure 3 As shown, step S40 further includes:
[0079] Step S41: According to the generated distortion type label, the first group of frequency domain signals FreA and the second group of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals. For the different types of frequency domain sub-signals, the corresponding frequency domain reconstruction processing is performed to obtain the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B.
[0080] Specifically, based on the generated distortion type markers, the first set of frequency domain signals FreA and the second set of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals: if the distortion type marker is a frequency domain compression distortion marker, a compression distortion sub-signal is generated; if the distortion type marker is a jitter distortion marker, a jitter distortion sub-signal is generated; if the distortion type marker is a mixed distortion marker, a mixed distortion sub-signal is generated; and if the distortion type marker is a normal speech marker, a normal sub-signal is generated. This targeted segmentation and processing approach stems from the fundamental differences in frequency domain characteristics among different distortion types: frequency domain compression distortion is characterized by energy concentration in specific high-frequency bands, jitter distortion is characterized by periodic frequency fluctuations in low-frequency bands, mixed distortion contains both of these characteristics, and normal speech maintains a stable distribution across the entire frequency band. Traditional algorithms employ uniform filtering or enhancement strategies for all frequency bands, which can lead to excessive suppression of key frequency bands in compression distortion and misjudgment of jitter distortion fluctuations as noise. However, sub-signal segmentation guided by type markers ensures that processing only affects the distortion component, avoiding interference with the normal signal.
[0081] The method for generating compressed distortion sub-signals includes: statistically analyzing the frequency range FA1 = [FA1,min,FA1,max] of the first group of energy distribution points in the first quadrant, and the frequency range FB1 = [FB1,min,FB1,max] of the second group of energy distribution points in the first quadrant; and taking the frequency range of the intersection of FA1 and FB1 as the compressed distortion frequency range F. comp Extract the frequency range F from the first group of frequency domain signals FreA. comp The corresponding portion serves as the first group of compressed distortion sub-signals; the frequency range F is extracted from the second group of frequency domain signals FreB. comp The corresponding portion serves as the second set of compressed distortion sub-signals. The method for generating the jitter distortion sub-signals is as follows: Calculate the frequency range FA3 of the first set of energy distribution points in the third quadrant (i.e., the interval formed by the minimum and maximum frequencies of all points in this quadrant), and the frequency range FB3 of the second set of energy distribution points in the third quadrant. Take the intersection of these two as the jitter distortion frequency range F. trem In practice, the frequency values of each point in the third quadrant are sorted, with the minimum value taken as the lower limit of the interval and the maximum value as the upper limit. If a certain group of signals has no points distributed in the third quadrant, the range of another group is used as the reference. The frequency range F is extracted from the FreA of the first group of frequency domain signals. trem The corresponding portion is used as the first group of jitter distortion sub-signals, and the corresponding portion is extracted from the second group of frequency domain signals (FreB) as the second group of jitter distortion sub-signals. For example, the first group has a frequency range of 100Hz to 300Hz in the third quadrant, and the second group has a frequency range of 150Hz to 350Hz. The intersection of 150Hz to 300Hz is taken as F. tremThe signal components within this range are extracted as tremor distortion sub-signals. This quadrant-based range determination is directly related to the correspondence between the third quadrant and tremor distortion in step S30, ensuring that the sub-signals only contain fluctuation components caused by emotional tremors.
[0082] The method for generating the hybrid distortion sub-signal is as follows: statistically analyze the frequency range FA1 in the first quadrant and the frequency range FA3 in the third quadrant of the first group of energy distribution points, and take the union of the two as the first group of hybrid frequency ranges FA. mix The frequency range FB1 in the first quadrant and the frequency range FB3 in the third quadrant of the second group of energy distribution points are statistically analyzed, and the union of the two is taken as the second group of mixed frequency ranges FB. mix ; then take FA mix With Facebook mix The intersection of these frequencies is used as the mixed distortion frequency range F. mix The portion within this range is extracted from each of the two sets of frequency domain signals as a hybrid distortion sub-signal. For example, the first set has a first quadrant range of 500Hz to 1000Hz and a third quadrant range of 100Hz to 300Hz. FA mix The first quadrant of the second group is 100Hz to 1000Hz; the third quadrant is 600Hz to 1100Hz, and the fourth quadrant is 150Hz to 250Hz. FB mix The frequency range is 150Hz to 1100Hz; the intersection F mix The frequency range is 150Hz to 1000Hz, which is the frequency range of the mixed distortion sub-signal. This method ensures that both components of the mixed distortion are included in the processing range by covering the frequency bands corresponding to compression and dithering.
[0083] The method for generating normal sub-signals is as follows: Calculate the frequency range FA4 of the first group of energy distribution points in the fourth quadrant and the frequency range FB4 of the second group of energy distribution points in the fourth quadrant, and take the union of the two as the normal frequency range F. norm The normal sub-signal is extracted from the range of each of the two sets of frequency domain signals. If one set of signals has no points in the fourth quadrant, the range of the other set is used as the reference; if neither set has any points in the fourth quadrant, the normal sub-signal is considered empty. For example, the fourth quadrant range of the first set is 300Hz to 500Hz, and the second set is 250Hz to 600Hz. norm The frequency range is 250Hz to 600Hz, and signals within this range are extracted as normal sub-signals. The fourth quadrant was chosen because the signal components in this region, whose frequency and amplitude are both within the normal fluctuation range, do not require complex processing.
[0084] For different types of frequency domain sub-signals, corresponding frequency domain reconstruction processing is performed, including frequency expansion mapping for compressed distortion sub-signals, center-locked stabilization for jitter distortion sub-signals, hierarchical collaborative processing for mixed distortion sub-signals, and amplitude fine-tuning for normal sub-signals. The frequency expansion mapping method for compressed distortion sub-signals specifically involves calculating the center frequency (Center) of the compressed distortion sub-signal. Freq Center Freq To compress the distortion frequency range F comp The midpoint, i.e. (F) comp minimum value + F comp (Maximum value) divided by 2; preset expansion ratio Ratio R Ratio R Experiments using historical compressed distortion samples determined that 500 hours of compressed distortion speech in the frequency domain were collected. The ratio of the frequency range of the corresponding speech under normal conditions to the frequency range after compression was calculated, and the average of all ratios was taken as the Ratio. R The typical range is 1.5 to 2.5, ensuring that the expanded frequency range closely approximates the natural distribution of normal speech; with Center Freq Based on this benchmark, it will be lower than Center Freq The frequency components expand towards lower frequencies, with an expansion magnitude of Center. Freq With F comp The difference between the minimum values multiplied by (Ratio) R -1), will be higher than Center Freq The frequency components expand towards higher frequencies, with an expansion amplitude of F. comp Maximum value and Center Freq The difference multiplied by (Ratio) R -1), which redistributes the compressed frequency components across a wider frequency band. This method utilizes the reversibility of frequency domain compression distortion—vocal cord tension under extreme emotions only leads to a narrowing of the frequency range rather than the disappearance of components, and its original distribution can be restored through symmetrical expansion. Compared with the traditional amplitude enhancement method, it can restore frequency richness without introducing noise.
[0085] The center-locked stabilization method is used for the jitter distortion sub-signal. Specifically, the frequency variation pattern of the jitter distortion sub-signal on the time axis is analyzed. The average frequency within each window is calculated by sliding window. The window length involved here is consistent with the time window in step S24. The average frequencies of all windows are linearly fitted to obtain a trend line of frequency variation. The midpoint of the trend line is taken as the center frequency of the fluctuation, i.e., the anchor point. Pos For each frequency component within a time window, calculate its relationship with the Anchor. PosThe deviation value is subtracted from the original frequency, causing all frequency components to be realigned to the Anchor. Pos To eliminate periodic fluctuations, linear fitting is chosen over arithmetic mean to calculate the center frequency because the frequency fluctuations of emotional tremors often show a gradual increasing or decreasing trend. Linear fitting can more accurately reflect the overall fluctuation center. If a simple average is used, the anchor point will shift due to extreme fluctuation values, and the instability cannot be completely eliminated.
[0086] A hierarchical collaborative processing method is used for the mixed distortion sub-signal, specifically by splitting the mixed distortion sub-signal into high-frequency and low-frequency bands according to frequency range. The high-frequency band consists of frequencies higher than the Anchor frequency. Pos The mixed distortion sub-signal, with the low-frequency band consisting of frequencies lower than or equal to the Anchor frequency. Pos The mixed-distortion sub-signals are processed using frequency expansion mapping for the high-frequency band (parameters identical to those used in compression distortion processing) and center-locked stabilization for the low-frequency band (parameters identical to those used in jitter distortion processing). After processing, the two signal segments are anchored... Pos Smooth splicing is performed at the splice points, with the amplitude at the splice points ensuring no abrupt changes through a linear transition. This layered processing is based on the physical characteristics of mixed distortion—speech under extreme emotions often exhibits both high-frequency compression and low-frequency trembling. Layered processing can optimize for both features separately. If a single method is used, the mutual interference between high-frequency compression and low-frequency fluctuations will lead to a decrease in processing effectiveness. For normal sub-signals, an amplitude fine-tuning method is used. Specifically, the ratio of the amplitude of the normal sub-signal to the average amplitude of the corresponding frequency within the standard elliptical region is calculated, and lower and upper limit ratio thresholds are set, along with corresponding adjustment coefficients. The lower and upper limit ratio thresholds are determined based on historical normal speech samples, covering the fluctuation range of most normal samples. For example, the lower limit ratio threshold is set to 0.8, and the upper limit ratio threshold is set to 1.2. If the ratio is lower than the lower limit ratio threshold, the amplitude is multiplied by the lower limit adjustment coefficient, which does not exceed the amplitude corresponding to the upper limit ratio threshold. If the ratio is higher than the upper limit ratio threshold, it is multiplied by the upper limit adjustment coefficient, which does not fall below the amplitude corresponding to the lower limit ratio threshold. The intermediate range remains unchanged. This fine-tuning avoids the over-processing of normal signals by traditional algorithms, compensates for amplitude attenuation caused by slight environmental noise, and ensures the naturalness of the speech.
[0087] The processed first set of frequency domain signals Fre'A and the second set of frequency domain signals Fre'B are generated as follows: Each processed sub-signal is placed back into the original frequency domain signal according to its original frequency range. Blank frequency bands, i.e., parts not covered by any sub-signal, retain their original signal amplitudes, but their amplitudes are multiplied by 0.5 and considered as weak noise or invalid components. For overlapping frequency bands, i.e., parts belonging to multiple sub-signals simultaneously, the average of the processed amplitudes of each sub-signal is taken as the final amplitude, ensuring complete frequency band coverage and no conflicts. For example, in the first set of frequency domain signals, 150Hz to 300Hz are dithered sub-signals, 500Hz to 1000Hz are compressed sub-signals, and 300Hz to 500Hz are normal sub-signals. After processing, these are spliced together in this order to form the complete Fre'A.
[0088] Step S41 achieves precise repair of different distortion types through type-guided sub-signal segmentation, physical characteristic-adapted processing methods, and multi-band collaborative integration. In conjunction with the distortion type labeling in step S30, it ensures a strict match between the processing method and distortion characteristics, avoiding the inconsistencies caused by traditional single-processing. In conjunction with the dual-source frequency domain coordinate system in step S10, the frequency range of each sub-signal is determined based on the quadrant distribution in the coordinate system, ensuring the accuracy of segmentation. The independent processing of each frequency domain sub-signal allows subsequent error correction to adjust parameters for specific distortion types without affecting other processing, significantly improving correction efficiency. Without this step, subsequent error correction would be ineffective due to the lack of targeted processing, resulting in key information in extreme emotional speech still being masked or over-suppressed by noise.
[0089] Step S42: Correct the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B: mark the processing error points in the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B, adjust the relevant parameters of frequency domain reconstruction, and obtain the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B.
[0090] Further, step S42 includes:
[0091] Step S421: Calculate the phase difference and amplitude ratio of the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B at the same frequency point;
[0092] Step S422: Set the normal phase difference range and the normal amplitude ratio range. When the phase difference value exceeds the normal phase difference range or the amplitude ratio value exceeds the normal amplitude ratio range, mark the corresponding frequency point as the processing error point.
[0093] Step S423: Adjust the relevant parameters for frequency domain reconstruction for the frequency points marked as processing error points.
[0094] Specifically, step S42 aims to identify and correct deviations introduced during frequency domain reconstruction by verifying the consistency of the two sets of signals, thus solving the problem of distortion amplification that is difficult to detect in single signal processing. Step S421 calculates the phase difference and amplitude ratio of the processed first set of frequency domain signals Fre'A and the second set of frequency domain signals Fre'B at the same frequency point. The phase difference is the difference between the phase value of Fre'A and the phase value of Fre'B at the same frequency point, and the amplitude ratio is the difference between the amplitude of Fre'A and the amplitude of Fre'B at the same frequency point. Phase and amplitude are key correlation characteristics of the two sets of synchronously acquired signals. For signals acquired from two spatially separated points of the same sound source, the phase difference is determined by the difference in propagation distance and remains constant, while the amplitude ratio is determined by the distance attenuation characteristics and remains stable within a fixed range. If the processed signal deviates from this inherent relationship, it indicates that errors have been introduced during the processing. Step S422 sets the normal phase difference range and the normal amplitude ratio range. The normal phase difference range is determined by statistically analyzing the phase differences of two sets of signals from 1000 hours of normal speech samples. The arithmetic mean of the phase difference at each frequency point is taken, and then the standard deviation of the phase differences at all frequency points is calculated. The average plus or minus three times the standard deviation is taken as the normal range. The normal amplitude ratio range is determined using the same method, statistically analyzing the amplitude ratio distribution of normal samples, and using a 95% confidence interval as the range. When the phase difference value or amplitude ratio value at a certain frequency point exceeds the normal phase difference range, the corresponding frequency point is marked as a processing error point. Three times the standard deviation and a 95% confidence interval are chosen because the phase difference and amplitude ratio of two sets of signals in normal speech follow a normal distribution. This range can cover most normal fluctuations, reducing false markings caused by random noise. If the range is too narrow, it will increase invalid error points; if it is too wide, it will be impossible to identify true errors.
[0095] Step S423 adjusts the relevant parameters for frequency domain reconstruction for the frequency points marked as processing error points. The adjustment follows the principle of "error type - parameter correlation": for processing error points of compression distortion, the error originates from the phase asynchrony of the two sets of signals caused by excessive or insufficient frequency expansion. In this case, the expansion ratio Ratio is decreased or increased. R Each adjustment is made by 10% of the initial value, and the frequency expansion is re-executed to calculate the phase difference and amplitude ratio until the point returns to the normal range. For jitter distortion processing errors, the error originates from the anchor point position. Pos If the amplitude ratio is abnormal due to calculation bias, increase the number of sliding windows from 3 to 5, refit the frequency trend line, and determine a new anchor. PosAfter re-performing frequency alignment, the performance indicators are verified until they meet the standards. For errors in mixed distortion processing, the errors are often caused by improper high- and low-frequency band segmentation leading to abrupt phase changes at the splicing point. In this case, the segmentation threshold is moved 50Hz towards the frequency band where the error point is located, and the layers are re-processed and a smooth transition is performed until the deviation is eliminated. The magnitude and method of parameter adjustment are determined based on historical error repair data to ensure that each adjustment significantly reduces the deviation without introducing new fluctuations.
[0096] Step S42 achieves precise correction of frequency domain reconstruction errors through dual-dimensional verification of phase difference and amplitude ratio, and dynamic parameter adjustment. In conjunction with the targeted processing in step S41, the sub-signal processing in S41 provides a specific target for error identification, while the correction in S42 feeds back into optimizing the processing parameters of S41, forming a closed loop of "processing-verification-optimization." The unified frequency benchmark provided by the dual-source frequency domain coordinate system in step S10 ensures the comparability of phase difference and amplitude ratio calculations, avoiding misjudgments due to inconsistent frequency axes. The distribution characteristics of error points can be used to pinpoint weaknesses in previous processing. For example, if compression distortion error points consistently appear in a certain frequency band, it indicates that the expansion ratio of that frequency band needs targeted optimization, upgrading the processing parameters from fixed values to dynamically adaptable variables for different frequency bands, significantly improving the robustness of handling complex composite distortions. If step S42 is omitted, the processing error in step S41 will accumulate and propagate to subsequent steps, causing the frequency band enhancement in step S50 to mistakenly identify the error component as a valid signal. The final output speech will still contain distortion or residual noise. Simultaneously, the consistency of the two signal sets cannot be guaranteed, and subsequent applications based on the fusion of the two signal sets, such as sound source localization, will fail due to signal deviation. Step S42 provides a quality control mechanism for the entire processing flow, ensuring that the pre-processed signals maintain consistency in physical characteristics, laying the foundation for the accuracy of the final speech reconstruction.
[0097] Step S50: Divide the modified first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B into frequency bands, identify frequency bands with significantly reduced energy, and perform enhancement processing on the frequency bands with significantly reduced energy.
[0098] Further, step S50 includes:
[0099] Step S51: Divide the modified first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B into frequency bands.
[0100] Step S52: For the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B, analyze the energy loss value band by band.
[0101] Calculate the first set of original energy values Ori for each frequency band of the first set of frequency domain signals Fre”A and the second set of frequency domain signals Fre”B before processing. AThe second set of original energy values Ori B And the first set of current energy values Cur after processing A The second group's current energy value Cur B By Ori A Subtract Cur A The energy loss value Loss for each frequency band of the first set of frequency domain signals Fre”A is obtained. A By Ori B Subtract Cur B The energy loss value Loss for each frequency band of the second set of frequency domain signals Fre”B is obtained. B .
[0102] Step S53, Loss A Greater than the first loss threshold Thr A The frequency bands are marked as the first group of frequency bands with significantly reduced energy, and the Loss bands are... B Greater than the second loss threshold Thr B The frequency bands are marked as the second group of frequency bands with significantly reduced energy;
[0103] Step S54: Calculate the energy ratio Ene of the first group of significantly reduced energy frequency bands and the second group of significantly reduced energy frequency bands to the first group of adjacent normal frequency bands. RA The energy ratio of the second group, Ene RB ;
[0104] Step S55, based on the first set of energy ratios Ene RA The energy ratio of the second group, Ene RB Generate a first set of gain compensation coefficients and a second set of gain compensation coefficients. Apply the first set of gain compensation coefficients to the first set of frequency bands with significantly reduced energy, and apply the second set of gain compensation coefficients to the second set of frequency bands with significantly reduced energy.
[0105] Specifically, step S50 aims to compensate for any residual energy attenuation in the previous processing by specifically enhancing the energy of key frequency bands, thus solving the noise amplification problem caused by traditional full-band enhancement. Step S51 involves dividing the two sets of corrected frequency domain signals into frequency bands and setting the frequency resolution for each band. This resolution is determined based on the frequency characteristics of the speech signal, covering the human voice audio range of 20Hz to 20kHz. Values such as 128, 256, or 512 are selected to ensure that each band contains a corresponding number of adjacent frequency points. This range is chosen because key speech information is mainly concentrated between 300Hz and 3400Hz. A higher resolution, such as 256, ensures that the subdivided frequency bands within this range are sufficient to capture subtle energy changes. Lower resolutions, such as 128, are used for low and high frequency bands to reduce computational load and balance processing efficiency and accuracy. The division method involves equally dividing the entire frequency range according to the resolution. For example, with a resolution of 256, the width of each band is the total frequency range divided by the number of bands, ensuring that adjacent bands are continuous and non-overlapping, providing independent units for subsequent energy analysis. Step S52 analyzes the energy loss value band by band, calculating the original energy value before processing and the current energy value after processing for each frequency band of the two sets of signals. The original energy value is the sum of the energy of the frequency domain signal mapped in step S10 in that frequency band, obtained by summing the squares of the amplitudes at all frequency points within the band; the current energy value is the sum of the energy of the corrected frequency domain signal in the same frequency band, calculated in the same way. The energy loss value is the original energy value minus the current energy value; a positive value indicates a decrease in energy after processing, while a negative value or zero indicates no decrease or increase in energy. This calculation method directly quantifies the impact of the processing on the energy of each frequency band, avoiding misjudgments of energy changes due to subjective judgment, and providing an objective basis for subsequent enhancement.
[0106] Step S53 marks frequency bands with significantly reduced energy and sets a first loss threshold and a second loss threshold. The first loss threshold is determined for the corrected first group of frequency domain signals Fre”A, and the second loss threshold is determined for the corrected second group of frequency domain signals Fre”B. Both are based on independent statistical analysis of the energy loss distribution of their respective signals. Due to the spatial differences between the two sets of sampling points, such as different distances from the sound source and different levels of environmental interference, the energy loss patterns of the two sets of signals naturally differ and need to be set separately to ensure the accuracy of the marking. The specific process is as follows: Collect 500 hours of the first group of speech samples and the second group of speech samples after correction processing in step S42. For the first group of samples, calculate the energy loss values of all frequency bands to form the first group of energy loss distribution. Sort the distribution from smallest to largest, and take the value corresponding to the 75th position after sorting as the first loss threshold. That is, the energy loss value above the first loss threshold only accounts for 25% of the total sample, ensuring that only the frequency bands with significant energy loss in the first group of signals are marked as the objects that need to be enhanced. Similarly, set the second loss threshold. When the energy loss value of a certain frequency band is greater than the corresponding threshold, it is marked as a frequency band with significantly reduced energy. The 75th percentile was chosen because most of the energy loss after normal processing is concentrated in the lower range. This threshold can filter out slight attenuation and focus on the frequency bands that have significant loss in speech intelligibility. If the threshold is too low, it will lead to over-enhancement, and if it is too high, it will miss the frequency bands that need to be repaired.
[0107] In step S54, the energy ratio Ene of the first group of significantly reduced energy frequency bands and the second group of significantly reduced energy frequency bands to the first group of adjacent normal frequency bands is calculated respectively. RA The energy ratio of the second group, Ene RB For the first group of frequency bands with significantly reduced energy, the normal frequency bands on both sides of each significantly reduced energy frequency band are selected as reference frequency bands. The ratio of the current energy value of the significantly reduced energy frequency band to the average energy value of the reference frequency band is calculated to obtain the energy ratio Ene. RA Ene RBThe calculation was performed using a similar method. The selection of the reference frequency band is based on the continuity of speech energy distribution. The energy levels of adjacent frequency bands are usually correlated. Using this as a benchmark ensures that the enhancement amplitude is coordinated with the surrounding energy and avoids abrupt changes. For example, if the current energy of a significantly reduced frequency band is 10, the average energy of the reference frequency band on the left is 20, and the average energy on the right is 25, then the energy ratio is 10 / (20+25) / 2≈0.44, reflecting the degree of energy attenuation of this frequency band relative to its surroundings. Step S55 generates a gain compensation coefficient and applies it to the significantly reduced frequency band. The gain compensation coefficient is inversely proportional to the energy ratio; the smaller the ratio, the larger the gain compensation coefficient, ensuring that the frequency band with more severe energy attenuation receives stronger compensation. The maximum value of the gain compensation coefficient is set to the ratio of the average energy value of the reference frequency band to the current energy value of the significantly reduced frequency band, avoiding excessive enhancement that could lead to distortion. At the same time, a progressive gain adjustment is used. In the edge region of the significantly reduced frequency band, the gain compensation coefficient gradually transitions from 1 to the target value, with a transition range of 10% of the band width, ensuring smooth energy changes.
[0108] Step S50 achieves precise energy restoration of key frequency bands through fine-grained frequency band division, energy loss quantification, and targeted gain compensation. The corrected signal from step S42 provides high-quality input for frequency band division, ensuring that the enhancement target is a valid signal with actual attenuation rather than noise. By comparing the energy ratios of the two sets of signals, unilateral energy attenuation caused by differences in acquisition equipment can be identified, and targeted enhancement is performed only on that side of the signal, making the energy distribution of the two sets of signals more balanced and improving the stability of subsequent signal fusion. If step S50 is missing, residual energy attenuation from the earlier processing will lead to blurred key information, such as numbers and directional words in emergency scenarios, affecting communication accuracy; at the same time, unbalanced enhancement will cause the speech to sound disjointed, reducing naturalness. As the final optimization step in the entire processing flow, step S50, through an energy compensation mechanism, ensures that the speech signal remains clear and natural after distortion restoration, providing reliable speech quality assurance for emergency communication scenarios.
[0109] Example 2
[0110] This embodiment, based on Embodiment 1, provides a frequency domain noise reduction and restoration system for simultaneous acquisition of dual-group human voices, such as... Figure 4 As shown, it includes:
[0111] Mapping module: Acquires two sets of synchronized human voice signals through two sets of acquisition points, constructs a dual-source frequency domain coordinate system, and maps the two sets of human voice signals onto the constructed dual-source frequency domain coordinate system to form two sets of energy distribution points.
[0112] Distortion recognition module: Used to calculate the spatial relationship between two sets of energy distribution points in the dual-source frequency domain coordinate system, obtain an index to measure the degree of distortion of the dual sets of human voice synchronization signals, and identify whether the current speech state of the dual sets of human voice synchronization signals is in an emotional distortion state.
[0113] Quadrant labeling module: used to divide the four-quadrant distortion recognition region in the dual-source frequency domain coordinate system, determine the quadrant affiliation of the two sets of energy distribution points corresponding to the two sets of human voice synchronization signals identified as emotional distortion state, and generate the corresponding distortion type label.
[0114] Frequency domain correction module: Based on the generated distortion type label, the first group of frequency domain signals FreA and the second group of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals. For different types of frequency domain sub-signals, corresponding frequency domain reconstruction processing and correction are performed to obtain the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B.
[0115] Frequency band enhancement module: used to divide the modified first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B into frequency bands, identify frequency bands with significantly reduced energy, and perform enhancement processing on the frequency bands with significantly reduced energy.
[0116] In the frequency domain correction module, the method for performing corresponding frequency domain reconstruction processing and correction for different types of frequency domain sub-signals to obtain the corrected first set of frequency domain signals Fre”A and the second set of frequency domain signals Fre”B includes:
[0117] Step S41: According to the generated distortion type label, the first group of frequency domain signals FreA and the second group of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals. For the different types of frequency domain sub-signals, the corresponding frequency domain reconstruction processing is performed to obtain the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B.
[0118] Step S42: Correct the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B: mark the processing error points in the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B, adjust the relevant parameters of frequency domain reconstruction, and obtain the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B.
[0119] Further, step S42 includes:
[0120] Step S421: Calculate the phase difference and amplitude ratio of the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B at the same frequency point;
[0121] Step S422: Set the normal phase difference range and the normal amplitude ratio range. When the phase difference value exceeds the normal phase difference range or the amplitude ratio value exceeds the normal amplitude ratio range, mark the corresponding frequency point as the processing error point.
[0122] Step S423: Adjust the relevant parameters for frequency domain reconstruction for the frequency points marked as processing error points.
[0123] In the frequency band enhancement module, the method for identifying frequency bands with significantly reduced energy and performing enhancement processing on these frequency bands includes:
[0124] Step S51: Divide the modified first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B into frequency bands.
[0125] Step S52: For the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B, analyze the energy loss value band by band.
[0126] Step S53, Loss A Greater than the first loss threshold Thr A The frequency bands are marked as the first group of frequency bands with significantly reduced energy, and the Loss bands are... B Greater than the second loss threshold Thr B The frequency bands are marked as the second group of frequency bands with significantly reduced energy;
[0127] Step S54: Calculate the energy ratio Ene of the first group of significantly reduced energy frequency bands and the second group of significantly reduced energy frequency bands to the first group of adjacent normal frequency bands. RA The energy ratio of the second group, Ene RB ;
[0128] Step S55, based on the first set of energy ratios Ene RA The energy ratio of the second group, Ene RB Generate a first set of gain compensation coefficients and a second set of gain compensation coefficients. Apply the first set of gain compensation coefficients to the first set of frequency bands with significantly reduced energy, and apply the second set of gain compensation coefficients to the second set of frequency bands with significantly reduced energy.
[0129] The methods and systems of this application may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the method is for illustrative purposes only, and the steps of the method of this application are not limited to the order specifically described above, unless otherwise specifically stated.
[0130] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.
[0131] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices, characterized in that, The method includes: Two sets of synchronized human voice signals are acquired through two sets of acquisition points, and a dual-source frequency domain coordinate system is constructed. The two sets of human voice signals are then mapped onto the constructed dual-source frequency domain coordinate system to form two sets of energy distribution points. In the dual-source frequency domain coordinate system, the spatial relationship between the two sets of energy distribution points is calculated to obtain an index that measures the degree of distortion of the dual-set human voice synchronization signals and to identify whether the current speech state of the dual-set human voice synchronization signals is in an emotional distortion state. In the dual-source frequency domain coordinate system, the four-quadrant distortion recognition region is divided, the quadrant affiliation of the two sets of energy distribution points corresponding to the two sets of human voice synchronization signals identified as emotional distortion state is determined, and the corresponding distortion type label is generated. Based on the generated distortion type label, the first group of frequency domain signals FreA and the second group of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals. For the different types of frequency domain sub-signals, corresponding frequency domain reconstruction processing and correction are performed to obtain the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B. The modified first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B are divided into frequency bands, and the frequency bands with significantly reduced energy are identified and enhanced.
2. The frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices according to claim 1, characterized in that, The dual-group voice synchronization signal includes a first group of voice signals and a second group of voice signals; The method for forming two sets of energy distribution points includes: performing Fast Fourier Transform on the first set of human voice signals and the second set of human voice signals respectively to obtain the first set of frequency domain signals and the second set of frequency domain signals; using the amplitude of each frequency component in the first set of frequency domain signals as the Y coordinate and the frequency value as the X coordinate to form the first set of energy distribution points; and using the amplitude of each frequency component in the second set of frequency domain signals as the Y coordinate and the frequency value as the X coordinate to form the second set of energy distribution points.
3. The frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices according to claim 2, characterized in that, The method for calculating the spatial relationship between the two sets of energy distribution points to obtain an index for measuring the degree of distortion of the two sets of human voice synchronization signals includes: Based on statistical analysis of historical normal speech data, a standard elliptical region is defined in the dual-source frequency domain coordinate system. The shortest distances from two sets of energy distribution points to the boundary of the standard elliptical region are calculated and denoted as the first deviation value and the second deviation value. The difference between the first deviation value and the second deviation value, Diff, is then calculated. AB and average Mean AB Difference Diff AB and average Mean AB As an indicator to measure the degree of distortion in two sets of synchronized human voice signals.
4. The frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices according to claim 3, characterized in that, The method for identifying whether the current voice state of the dual-group synchronized human voice signals is in an emotionally distorted state includes: Set the number of consecutive time windows Cou C Difference threshold Thr D and average threshold Thr M When Cou is continuous C Difference within a time window (Diff) AB All are greater than the difference threshold Thr D And the average Mean AB All are greater than the average threshold Thr M At that time, the speech state of the dual-group human voice synchronization signal is marked as emotional distortion state.
5. The frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices according to claim 4, characterized in that, The method for generating the distortion type marker includes: In the dual-source frequency domain coordinate system, four quadrant regions are divided with the origin O as the center. The first quadrant is defined as the frequency domain compression distortion region, the second quadrant is defined as the mixed distortion region, the third quadrant is defined as the jitter distortion region, and the fourth quadrant is defined as the normal region. The distribution ratio of the first and second energy distribution points of the dual sets of human voice synchronization signals identified as emotionally distorted in the four quadrant regions was statistically analyzed. Based on the distribution ratio of the first group of energy distribution points and the second group of energy distribution points in the four quadrant regions, the corresponding distortion type label is automatically generated.
6. The frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices according to claim 5, characterized in that, The method for automatically generating corresponding distortion type labels based on the distribution ratio of the first group of energy distribution points and the second group of energy distribution points in the four quadrant regions includes: Determine the key distribution quadrants of the first and second sets of energy distribution points; based on the key distribution quadrants of the first and second sets of energy distribution points, generate distortion type labels, which include frequency domain compression distortion labels, tremor distortion labels, mixed distortion labels, and normal speech labels.
7. The frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices according to claim 6, characterized in that, The method for determining the key distribution quadrants of the first and second groups of energy distribution points includes: setting a distribution threshold; if the proportion of the first group of energy distribution points falling into the i-th quadrant exceeds the distribution threshold, then the i-th quadrant is considered a key distribution quadrant for the first group of energy distribution points; otherwise, it indicates that the first group of energy distribution points has no key distribution quadrant; if the proportion of the second group of energy distribution points falling into the j-th quadrant exceeds the distribution threshold T... ratio If the value is true, it means that the j-th quadrant is the key distribution quadrant of the second group of energy distribution points; otherwise, it means that the second group of energy distribution points has no key distribution quadrant; where i and j represent quadrant numbers, i = 1, 2, 3, 4, j = 1, 2, 3, 4.
8. The frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices according to claim 7, characterized in that, The step of segmenting the first group of frequency domain signals FreA and the second group of frequency domain signals FreB according to the generated distortion type marker to generate different types of frequency domain sub-signals includes: if the distortion type marker is a frequency domain compression distortion marker, then a compression distortion sub-signal is generated; if the distortion type marker is a jitter distortion marker, then a jitter distortion sub-signal is generated; if the distortion type marker is a mixed distortion marker, then a mixed distortion sub-signal is generated; if the distortion type marker is a normal speech marker, then a normal sub-signal is generated.
9. The frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices according to claim 8, characterized in that, The method for obtaining the corrected first set of frequency domain signals Fre”A and the second set of frequency domain signals Fre”B includes: For different types of frequency domain sub-signals, corresponding frequency domain reconstruction processing is performed to obtain the first set of frequency domain signals Fre'A and the second set of frequency domain signals Fre'B after processing; the first set of frequency domain signals Fre'A and the second set of frequency domain signals Fre'B after processing are corrected to obtain the first set of frequency domain signals Fre”A and the second set of frequency domain signals Fre”B after correction. The method for correcting the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B includes: marking the processing error points in the processed first group of frequency domain signals Fre'A and the second group of frequency domain signals Fre'B, and adjusting the relevant parameters of frequency domain reconstruction for the frequency points marked as processing error points.
10. A frequency domain noise reduction and restoration system for simultaneous acquisition of dual-group human voices, used to implement the frequency domain noise reduction and restoration method for simultaneous acquisition of dual-group human voices as described in any one of claims 1-9, characterized in that, The system includes: Mapping module: Acquires two sets of synchronized human voice signals through two sets of acquisition points, constructs a dual-source frequency domain coordinate system, and maps the two sets of human voice signals onto the constructed dual-source frequency domain coordinate system to form two sets of energy distribution points; Distortion recognition module: used to calculate the spatial relationship between two sets of energy distribution points in the dual-source frequency domain coordinate system, obtain an index to measure the degree of distortion of the two sets of human voice synchronization signals, and identify whether the current speech state of the two sets of human voice synchronization signals is in an emotional distortion state. Quadrant marking module: used to divide the four-quadrant distortion recognition region in the dual-source frequency domain coordinate system, determine the quadrant affiliation of the two sets of energy distribution points corresponding to the two sets of human voice synchronization signals identified as emotional distortion state, and generate corresponding distortion type markings; Frequency domain correction module: Based on the generated distortion type label, the first group of frequency domain signals FreA and the second group of frequency domain signals FreB are segmented to generate different types of frequency domain sub-signals. For different types of frequency domain sub-signals, corresponding frequency domain reconstruction processing and correction are performed to obtain the corrected first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B. Frequency band enhancement module: used to divide the modified first group of frequency domain signals Fre”A and the second group of frequency domain signals Fre”B into frequency bands, identify frequency bands with significantly reduced energy, and perform enhancement processing on the frequency bands with significantly reduced energy.
Citation Information
Patent Citations
Two-way voice signal automatic noise reduction waveform node emotion analysis method and device
CN114550752A
Sound mixing and noise reduction method and system
CN119964537A
System and method for analyzing and displaying emotion
CN104112055A
Voice active noise reduction optimization method and system
CN114743535A