A meeting window talk system based on a voice sensor

Through technical means such as array sensor acquisition and dynamic confidence assessment, the problem of low-energy voice misjudgment in the meeting window call system has been solved, and accurate recognition and fidelity enhancement of low-volume voice have been achieved, thereby improving information integrity and security.

CN120636432BActive Publication Date: 2025-10-17HANGZHOU HUA TING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511117912.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-17
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

The existing meeting window call system has the problem of misjudging low-energy voice states as environmental background noise, resulting in information loss, semantic interruption or misunderstanding, especially in dialogue-sensitive scenarios such as legal visits and psychological counseling.

Method used

It adopts array sensor acquisition, dynamic confidence assessment, phase coupling and harmonic enhancement, multi-dimensional speech recognition and dynamic gain adjustment, combined with full-link integrity verification to ensure accurate recognition and fidelity enhancement of low-volume speech.

Benefits of technology

It achieves complete restoration and accurate reception of low-volume voices, improves the comprehensive capabilities of the call system in terms of privacy, security, and information integrity, and is particularly suitable for application scenarios with high conversation sensitivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636432B_ABST
    Figure CN120636432B_ABST
Patent Text Reader

Abstract

The application discloses a meeting window call system based on a voice sensor and relates to the technical field of communication.The meeting window call system comprises a full-space voice collection module, a dynamic confidence evaluation module, a phase coupling enhancement module, a multi-dimensional speech mode recognition module, a dynamic gain adjustment and sound quality enhancement module and a full-link voice integrity verification module.The full-space voice collection module is arranged in a meeting window space based on audio reflection characteristics, an array type voice sensor unit is arranged, a multi-frequency full-space collection matrix is constructed, voice signals of both parties in the meeting are collected, and basic audio data streams with high signal-to-noise ratio are output.The application realizes accurate identification and faithful enhancement of low-volume voice through array type sensing collection, dynamic confidence evaluation, phase coupling and harmonic enhancement, multi-dimensional speech recognition and dynamic gain adjustment, and ensures that weak signals are not missed and distorted through full-link integrity verification and false positive backtracking, thereby improving the privacy, security and information integrity of the meeting call.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a meeting window communication system based on a voice sensor. Background Art

[0002] A voice sensor-based meeting window communication system is designed for use in isolated locations such as meeting rooms, banks, and service windows. By integrating a highly sensitive voice sensor, a digital signal processing module, and a voice amplifier, it enables clear voice communication between people on both sides of soundproof glass. The system uses a voice sensor to accurately capture the speaker's voice signal. It also incorporates noise suppression, echo cancellation, and voice enhancement algorithms to effectively filter out ambient noise and glass reflections, enhancing the clarity and confidentiality of voice calls. The system supports full-duplex and half-duplex call modes, avoiding voice overlap and crosstalk, ensuring natural and smooth communication between the two parties. It is particularly suitable for window-based communication scenarios where security, privacy, and call quality are paramount.

[0003] The existing technology has the following deficiencies:

[0004] Existing meeting window call systems lack the ability to recognize special voice states. Specifically, when one of the two parties in a meeting is in a depressed mood, psychologically depressed, or uses low-energy voice expressions such as whispering for privacy reasons, the resulting voice signal has characteristics such as low volume, slow speech speed, and small frequency variation. These signals are easily misidentified by the call system as background noise and automatically blocked or ignored. Because these voice signals are not effectively collected, amplified, and forwarded in the audio transmission link, the speech content is completely unperceivable at the receiving end, leading to information loss, semantic interruption, or misunderstanding during the call. This problem can lead to the omission of important information, the failure of communication goals, and even the inaccurate assessment of key factors such as the parties' psychological state and behavioral tendencies in highly sensitive conversations, such as legal visits, psychological counseling, and family comfort. This poses a high risk.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the present application is to provide a meeting window call system based on a voice sensor, which realizes accurate recognition and faithful enhancement of low-volume voice through array sensing collection, dynamic confidence evaluation, phase coupling and harmonic enhancement, multi-dimensional speech recognition and dynamic gain adjustment, combines full-link integrity verification and false-positive backtracking, ensures that weak signals are not missed or distorted, and improves the privacy, security and information integrity of meeting calls, to solve the problems in the background art.

[0007] In order to achieve the above-mentioned purpose, the present application provides the following technical scheme: a meeting window call system based on a voice sensor, comprising a full-space voice collection module, a dynamic confidence evaluation module, a phase coupling enhancement module, a multi-dimensional speech pattern recognition module, a dynamic gain adjustment and sound quality enhancement module, and a full-link voice integrity verification module:

[0008] The full-space voice collection module, based on the audio reflection characteristics, deploys array voice sensing units in the meeting window space, constructs a multi-band full-space collection matrix, collects voice signals of both parties in the meeting, and outputs a high signal-to-noise ratio of basic audio data stream;

[0009] The dynamic confidence evaluation module performs dynamic confidence evaluation on the basic audio data stream, extracts energy distribution and spectral stability features, and generates a feature weight set that distinguishes voice from background noise;

[0010] The phase coupling enhancement module, based on the feature weight set, constructs a phase coupling enhancement channel, and through phase consistency mapping and harmonic synchronous amplification, enhances the structural coherence and energy density of the audio signal;

[0011] The multi-dimensional speech pattern recognition module, based on the enhanced audio signal, performs multi-dimensional speech pattern recognition, and identifies human voice attributes according to pitch continuity, rhythm change and stop-continuous mode;

[0012] The dynamic gain adjustment and sound quality enhancement module dynamically adjusts the gain mapping curve and spectral channel parameters of the audio signal transmission chain according to the identified human voice attributes, and implements sound quality faithful enhancement;

[0013] The full-link voice integrity verification module performs full-link voice integrity verification on the enhanced audio signal, and through attribute feature comparison and false-positive backtracking verification, realizes complete restoration and accurate reception of low-volume voice signals.

[0014] Preferably, the basic audio data stream acquisition step is as follows:

[0015] Based on the material, position and acoustic reflection characteristics of the walls, ceiling, floor and soundproof glass of the meeting window space, sound field simulation and modeling are performed to determine the sound reflection path, attenuation coefficient and standing wave distribution;

[0016] According to the sound field simulation results, a high-sensitivity capacitive or piezoelectric voice sensor is selected, a mixed arrangement of equal intervals and non-equal intervals is adopted, and a sensor with tilt or directional pickup is set to construct a full-space acquisition matrix covering multiple frequency bands;

[0017] The signal of each voice sensor is pre-amplified and digitized by a pre-amplifier circuit and an analog-to-digital conversion unit, enters a multi-channel signal synchronization platform based on time synchronization, and completes phase consistency and time synchronization;

[0018] Based on a multi-band filter bank, the digitized signal is divided into frequency bands and weighted, weighted average and spectral subtraction are performed to reduce background noise, and dynamic amplitude normalization processing is performed through adaptive gain control to output a basic audio data stream with high signal-to-noise ratio.

[0019] Preferably, the feature weight set generation step is as follows:

[0020] Based on the basic audio data stream, the energy distribution feature, spectral stability feature, spectral envelope similarity, spectral entropy, and dominant frequency component persistence of each frame signal are extracted to form a time series multi-dimensional feature vector;

[0021] The multi-dimensional feature vector sequence is input into a deep learning model composed of a convolutional neural network and a bidirectional long short-term memory network in series to output a confidence score of each frame signal;

[0022] The score sequence is smoothed and thresholded to identify valid speech regions with consecutive multiple frames of scores higher than the threshold;

[0023] Based on the confidence score and the features, a feature weight set for signal enhancement and recognition processing is generated.

[0024] Preferably, based on the feature weight set, the specific steps for enhancing the structural coherence and energy density of the audio signal using phase consistency mapping and harmonic synchronous amplification are as follows:

[0025] Based on the energy distribution feature and spectral stability feature of the feature weight set, a phase tracking model of the signal is established, the instantaneous phase information is extracted through Hilbert transform, the phase trajectory is constructed, and the target frequency with continuous phase and stable amplitude is selected;

[0026] Perform phase consistency mapping on the target frequency and adjacent frequencies, perform phase compensation according to phase similarity, correct phase offset, and maintain consistent phase of frequency components;

[0027] Apply dynamic gain control to the dominant frequency and its harmonic components, dynamically adjust the gain according to the feature weight set, enhance the energy density, and maintain the natural harmonic energy distribution ratio;

[0028] Performing time-frequency domain structure coherence review, dynamically optimizing according to phase spectrum difference, spectral envelope change rate and waveform smoothness index, outputting enhanced audio signal with high energy density and strong structure coherence.

[0029] Preferably, based on the enhanced audio signal, the specific steps of identifying whether the signal has human speech attributes through multi-dimensional analysis of pitch continuity, rhythm change and pause-continuity mode are as follows:

[0030] Extracting the fundamental frequency and harmonics of the enhanced audio signal through short-time Fourier transform to form a pitch contour curve and calculate the fundamental frequency smoothness and stability;

[0031] According to the joint analysis of energy envelope curve and fundamental frequency change rate, rhythm density, beat interval and accent position are analyzed, and periodic and aperiodic rhythm components are identified by combining wavelet transform and rhythm synchronization function;

[0032] Apply long-short term memory network to predict the continuation probability and pause probability of speech segments based on time series characteristics, distinguish continuous pronunciation and natural pause, and form pause-continuity mode;

[0033] Input the features of pitch continuity, rhythm change and pause-continuity mode into the support vector machine classification model, output the human speech attribute confidence score of each signal segment, and determine the speech attribute of the audio signal.

[0034] Preferably, the specific steps of dynamically adjusting the gain mapping curve and spectral channel parameters of the audio signal transmission chain to implement audio quality fidelity enhancement are as follows:

[0035] Based on the identified human speech attributes, construct a dynamic gain mapping curve, set a gain boost coefficient according to the main frequency and harmonic distribution, and set a gain suppression for non-speech segments;

[0036] According to the speech attribute weight, dynamically adjust the spectral channel parameters to adjust the center frequency, bandwidth and quality factor of the filter, strengthen the speech segments and suppress the noise segments;

[0037] Map the dynamic gain mapping curve and spectral channel parameters to each time frame and frequency unit through digital signal processing algorithm, apply phase compensation and spectrum reconstruction, and perform frame-by-frame and frequency-by-frequency enhancement of the audio signal;

[0038] Perform real-time audio quality evaluation on the enhanced signal, dynamically optimize the gain mapping curve and spectral channel parameters through feedback mechanism according to the signal-to-noise ratio, spectral smoothness, dynamic range and harmonic distortion rate index, and ensure audio quality fidelity enhancement.

[0039] Preferably, through attribute feature comparison and false positive backtracking verification, the specific steps of realizing complete restoration and accurate reception of low volume speech signal are as follows:

[0040] The five types of attribute features of the audio signal after gain mapping curve and spectrum channel parameter adjustment are extracted, including energy distribution, spectral stability, pitch continuity, rhythm density and stop continuity mode;

[0041] The five types of attribute features of the signals before and after enhancement are compared by using cosine similarity, dynamic time warping algorithm, Euclidean distance, rhythm density comparison and edit distance algorithm to form a global speech similarity score;

[0042] The signal segment with a score lower than the consistency threshold is traced back to the original signal, enhanced by applying different gain and spectrum parameter configurations, and the optimal configuration is selected by Bayesian optimization to replace the original signal segment;

[0043] The five types of features of the corrected signal are extracted again, and the signal-to-noise ratio, harmonic distortion rate and dynamic range are calculated to review the integrity and sound quality indicators, and the final enhanced signal is output after meeting the standards.

[0044] In the above technical solution, the technical effects and advantages provided by the present application are as follows:

[0045] The present application improves the spatial and frequency coverage of signal acquisition through array sensor layout, realizes intelligent differentiation of human voice and background noise through dynamic confidence evaluation and feature weight construction, and effectively restores the energy density and structural coherence of low volume voice through phase coupling and harmonic enhancement. Further through multi-dimensional speech mode recognition and dynamic gain adjustment, the human voice attributes are precisely strengthened, noise amplification is avoided, and under the protection of full-link integrity verification and false positive backtracking mechanism, the weak signal information is not missed and distorted. The system is particularly suitable for dialog-sensitive and information-integrity-demanding application scenarios such as legal visits, psychological counseling and family pacification, significantly improving the comprehensive ability and application value of the call system in privacy, security and complete transmission of voice information. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0047] Figure 1 A module schematic diagram of a meeting window call system based on a voice sensor. DETAILED DESCRIPTION

[0048] Example implementations are now described with reference to the drawings. Example implementations can, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these example implementations are provided so that this disclosure will be thorough and complete, and will fully convey the inventive gist to those skilled in the art.

[0049] The present application provides a meeting window talk system based on a voice sensor as shown in Figure 1 The meeting window talk system based on a voice sensor as shown in

[0050] The full-space voice collection module, based on the audio reflection characteristics, deploys an array type voice sensing unit in the meeting window space, constructs a full-space collection matrix covering multiple frequency bands, collects full-band voice signals generated by the meeting parties, and outputs a high signal-to-noise ratio basic audio data stream.

[0051] The array type voice sensing unit deployed in the meeting window space to achieve high signal-to-noise ratio collection of full-band voice signals specifically includes the following steps:

[0052] Based on the physical space layout of the meeting window, according to the material, position and acoustic reflection characteristics of the walls, ceiling, floor and soundproof glass in the meeting space, sound field simulation and modeling are performed to determine the reflection path, attenuation coefficient and standing wave distribution of sound in the space. Based on the sound field simulation results, high-sensitivity capacitive or piezoelectric voice sensors are selected, and the array type is arranged according to the determined sound reflection hot spots, low-frequency standing wave accumulation areas and high-frequency reflection dissipation areas. The voice sensors in the array are arranged in a mixed manner of equal spacing and non-equal spacing to ensure optimal collection angle and coverage range for sound waves in multiple frequency bands from low frequency to high frequency. At the same time, sensors with tilt or directional pickup are arranged in front, side and diagonal directions of the window to enhance the direct sound wave collection effect of users on both sides of the window and reduce the interference of space multi-path reflection on the original signal.

[0053] To improve the quality and integrity of the speech signal, the pickup signal of each speech sensor is pre-amplified and digitized through a pre-amplification circuit and an analog-to-digital conversion unit. The pre-amplification circuit has the characteristics of low noise, wide frequency band and high dynamic range, which can effectively improve the energy of weak speech signals during the pickup stage, while suppressing the noise of the sensor itself and environmental electromagnetic interference. The sampling rate of the analog-to-digital conversion unit is set to no less than 96kHz, and the quantization precision is 24 bits, to ensure the spectral restoration and detail preservation in the subsequent digital signal processing link. After digitization, each sensor signal enters a multi-channel signal synchronization platform based on time synchronization, which applies accurate timestamps to all channels of the signal to ensure phase consistency and time synchronization in the subsequent signal fusion and enhancement processing.

[0054] For the digitized audio signals obtained from multiple sensors, a preset multi-band filter bank is used to divide and weight the signals by frequency band. The filter bank divides the audio signal into three main frequency bands: low, medium and high, and sets adaptive filter parameters for each frequency band based on its noise characteristics. In this process, a combination of weighted averaging and spectral subtraction is used to effectively reduce background noise, wind noise and electromagnetic noise in the signal, while improving the signal-to-noise ratio of the speech signal while preserving the integrity of the spectrum. In the weighted averaging process, different weights are applied to the signals of sensors at different positions based on the aforementioned sound field simulation data to strengthen the amplitude of the forward speech signals from the two sides of the window, while weakening the energy of the secondary or multiple reflection sound waves from the spatial reflection path, avoiding signal blurring or ghosting effects caused by reflected sound waves.

[0055] After multi-band filtering and weighted fusion, dynamic amplitude normalization based on adaptive gain control is performed to uniformly adjust the signals of different channels and different frequency bands to the optimal dynamic range, preventing information loss or signal distortion in subsequent signal processing due to uneven amplitude or narrow dynamic range. The normalized signals are subjected to real-time signal-to-noise ratio evaluation and data integrity verification to form the final high signal-to-noise ratio basic audio data stream. This data stream has the characteristics of balanced spectral distribution, moderate amplitude dynamic range and consistent phase timing, and can provide high-quality raw audio basis for subsequent dynamic confidence evaluation, phase coupling enhancement, multi-dimensional speech recognition and other deep processing.

[0056] The role of this step is to lay a high-quality and complete signal foundation for subsequent intelligent recognition, enhancement and faithful transmission of audio signals, especially for the special speech state of low volume and low frequency change in the meeting window space, to achieve accurate and comprehensive audio information collection. The meeting window space usually has soundproof glass, hard walls and ceilings and other structures, which can cause multi-path reflection, sound attenuation and standing wave effects of sound, resulting in weak signal energy directly transmitted from the sound source to the receiving end, and easily covered by background noise. By deploying an array type voice sensing unit, combined with the audio reflection characteristics and spatial layout, a multi-band and full-space collection matrix can be constructed in the space to achieve synchronous capture of sound from multiple directions and multiple frequency dimensions, avoiding sound source omission or signal distortion caused by single-point sensor collection. The collection matrix can balance the pickup of direct sound and some valuable reflected sound, and through sound source positioning and phase alignment of the array structure, the spatial resolution and signal-to-noise ratio of the signal are improved. The output basic audio data stream has the characteristics of high fidelity, high integrity and high signal-to-noise ratio, not only retaining the original speech characteristics of the speaker, but also effectively isolating and weakening environmental noise, glass reflection interference and spatial reverberation. Such a signal foundation is the premise of subsequent dynamic confidence evaluation, phase coupling enhancement, multi-dimensional speech mode recognition and other processing links, ensuring that in complex communication scenarios such as low-energy speech, whispering and low mood, accurate transmission and restoration of information without missing and lossless audio quality can be achieved.

[0057] The dynamic confidence evaluation module performs dynamic speech signal confidence evaluation on the basic audio data stream, extracts energy distribution features and spectral stability features in the audio signal, and generates a feature weight set for distinguishing speech signals from background noise.

[0058] The dynamic speech signal confidence evaluation on the basic audio data stream specifically includes the following steps:

[0059] Based on the collected high signal-to-noise ratio basic audio data stream, a multi-level signal analysis model is constructed for the time domain and frequency domain characteristics of the audio signal. This model obtains the spectral distribution and energy density of each frame of audio through short-time Fourier transform of the audio signal, and calculates the energy envelope curve of the signal in the time domain using a sliding window to capture the energy change trend of the audio signal. To ensure the accuracy of the analysis, the window length of the short-time Fourier transform is set to 32 milliseconds and the frame shift is set to 10 milliseconds, so as to ensure the spectral resolution while considering the time resolution. Through this step, the frequency component, energy intensity and transient change of each time segment can be fully captured.

[0060] The energy envelope curve refers to a description of the energy variation trend of the audio signal in the time sequence, which is formed by integrating or averaging the amplitude or power of the audio signal in each short time segment. Its essence is to fit the amplitude variation of the original audio signal with a smooth curve, reflecting the strength fluctuation of the audio signal in different time periods. The energy envelope curve is used in the present application to track and quantify the energy dynamic change in the basic audio data stream in real time, especially for low-energy speech such as low volume, whisper, and low mood. Although the overall volume is low, the fluctuation in time still has certain speech characteristics, which is different from the randomness and stable low-amplitude change of background noise. By analyzing the energy envelope curve, the rhythm and pronunciation fluctuation of the speech in the signal can be effectively identified, providing a signal strength basis in the time domain for subsequent dynamic confidence assessment, and combining with the frequency spectrum characteristics, the ability to distinguish low-energy speech and noise is improved. This has a key discriminant value for avoiding weak signals being misjudged as noise and being shielded.

[0061] Based on the above signal analysis model, the energy distribution features and spectral stability features of the audio signal in each time window are further extracted. The energy distribution features are described by calculating the energy proportion of each frequency sub-band, the center frequency drift amplitude and the energy change gradient, so as to reflect the strength and stability of the signal. The spectral stability features are quantified by calculating the similarity of the spectral envelope between adjacent frames, the spectral entropy and the persistence of the main frequency component, so as to reflect the persistence and structural coherence of the signal in the frequency domain. Especially for low volume and slow speech signals, the spectrum changes relatively stable but have harmonic characteristics of human voice, and through the above feature extraction, the wide frequency and random spectrum characteristics presented in the background noise can be effectively distinguished.

[0062] The similarity between the spectral envelopes of adjacent frames, the spectral entropy, and the persistence of the dominant frequency component can be used to quantify the persistence and structural continuity of the audio signal in the frequency domain from multiple dimensions. Specifically, first, the audio signal is divided into equal-length short-time frames, and the spectral envelope curve is obtained for each frame by Fourier transform. Then, the cosine similarity or correlation coefficient of the spectral envelopes of adjacent two frames is calculated. If the similarity is high, it means that the spectral form is continuous in time, reflecting the stability of the signal frequency component. Second, the spectral entropy of each frame is calculated. The spectral entropy reflects the degree of randomness of the frequency energy distribution. Low spectral entropy indicates that the energy is concentrated in a specific frequency, representing a good structure of the signal, while the spectral entropy of noise is usually high, and the energy distribution is irregular. Finally, the dominant frequency of each frame, i.e., the frequency component with the most concentrated energy, is detected, and its appearance and maintenance in multiple consecutive frames are counted. If the dominant frequency component persists in multiple frames with little fluctuation, it indicates that the signal has a stable harmonic or fundamental frequency, reflecting the structural continuity. Through the joint evaluation of the three, the stability and regularity of the signal in the frequency domain can be accurately quantified, thereby distinguishing between ordered signals such as human voice and disordered signals such as background noise.

[0063] A dynamic confidence evaluation model based on machine learning is used to distinguish the extracted energy distribution features and spectral stability features. This model combines convolutional neural networks and recurrent neural networks. The former is used to capture spatial features in the spectrum, and the latter is used to learn the time series characteristics of the signal. The input of the model is the feature vector sequence extracted in the first two steps, and the output is the confidence score of each frame of signal. The higher the score, the more likely it is that the frame of signal is human speech signal. In order to improve the recognition ability of the model for low volume and whisper speech, the model introduces a large number of low signal-to-noise ratio and low energy speech samples in the training stage, and improves the sensitivity of the model to subtle speech features through contrastive learning. This step ensures that in a dynamically changing noise environment, the confidence of the speech property of each signal can be accurately evaluated in real time.

[0064] The structure combining convolutional neural network and recurrent neural network refers to a deep model architecture formed by sequentially connecting convolutional neural network and recurrent neural network, wherein the convolutional neural network first performs local perception and spatial pattern extraction on the input feature vector sequence, and the recurrent neural network performs time-dependent analysis and dynamic feature modeling on the sequence data after convolution extraction. In the present application, the convolutional neural network performs feature extraction on the local information such as the spectral features, energy distribution and spectral envelope of each frame signal through multiple convolution kernels, which can effectively extract the local frequency patterns and structural features hidden in the speech signal, thereby enhancing the model's ability to distinguish speech and noise in the spectral level. The recurrent neural network, especially the long short-term memory network (LSTM) or bidirectional LSTM, learns the dynamic changes and dependencies between the front and back frames of the signal along the time sequence dimension based on the feature sequence output by the convolutional layer, captures the coherence and rhythm of the speech signal on the time axis, and the random volatility of the noise signal. Through this combination, the model simultaneously discriminates the signal in both spatial and temporal dimensions, and finally maps the output of the recurrent neural network to a confidence score of each frame signal through a fully connected layer and an activation function, with a higher value representing a higher probability that the frame signal is human speech, thereby achieving accurate signal discrimination and dynamic scoring.

[0065] The feature vector sequence extracted in the first two steps is input to output a confidence score of each frame signal, and the specific steps include:

[0066] The energy distribution features, spectral stability features, spectral envelope similarity, spectral entropy and dominant frequency component persistence of each frame audio signal are combined into a multi-dimensional feature vector in a unified format, and a feature vector sequence is formed in the order of time sequence.

[0067] The feature vector sequence is input into a pre-trained deep learning model, which is composed of a convolutional neural network layer and a bidirectional long short-term memory network layer (BiLSTM). The convolutional layer is used to extract local spatial features and spectral change patterns in the feature sequence, and the BiLSTM layer captures the time-dependent relationships and context information in the feature sequence, thereby improving the model's recognition ability for speech continuity and structure.

[0068] The output of the BiLSTM is mapped to a one-dimensional score sequence corresponding to the number of time frames through a fully connected layer, and the score is normalized to between 0 and 1 through a sigmoid activation function, representing the probability that each frame signal is human speech.

[0069] The score sequence is smoothed and thresholded, and a signal segment with consecutive multiple frame scores higher than a set threshold is determined as an effective speech region, thereby achieving accurate quantization of the confidence of each frame signal.

[0070] According to the score result of the dynamic confidence evaluation model, the energy distribution characteristics and the spectral stability characteristics of each time segment are combined to generate a feature weight set for subsequent signal enhancement and identification processing. The feature weight set is a multi-dimensional weight matrix, which includes the weighting coefficients of the energy of each frequency band, the enhancement factor of the spectral continuity, and the normalized result of the confidence score, as an important parameter basis for guiding the subsequent phase coupling enhancement and multi-dimensional speech mode identification. Through this feature weight set, different enhancement efforts can be applied to signals of different frequency bands and different time segments in the subsequent signal enhancement link, realizing targeted optimization of low-volume speech while suppressing the interference of background noise.

[0071] The purpose of this step is to accurately distinguish which segments in the audio signal are valid speech signals and which segments are background noise through dynamic speech signal confidence evaluation of the basic audio data stream, thereby providing scientific and objective basis for subsequent speech signal enhancement, identification and faithful transmission. In the meeting window call scene, there are various complex noise sources such as glass reflection echo, environmental low-frequency noise, non-speech action sound of the two parties, and the signal strength of special speech states such as low volume, whisper, and emotional depression is weak itself and is easily misjudged as noise and filtered. Therefore, it is difficult to achieve accurate signal discrimination relying only on simple energy detection or traditional noise threshold. Through this step, the energy distribution characteristics of the audio signal are first extracted, such as the energy density of each frequency band and the energy change gradient, which quantize the strength and dynamics of the signal; then the spectral stability characteristics are extracted, such as the similarity of the spectral envelope, the spectral entropy, and the persistence of the main frequency component, which reflect the structural continuity and stability in the frequency domain of the signal. After these features are dynamically evaluated by a deep learning model, a confidence score is generated for each time frame, and a feature weight set is further formed to guide signal enhancement and identification. This feature weight set not only dynamically adapts to changes in different environmental noise and signal strength, but also provides fine-grained parameter guidance for subsequent phase coupling enhancement, multi-dimensional speech mode identification, and other steps, ensuring that low-volume and special speech signals are completely preserved and accurately identified in the entire link, and improving the overall perceptual ability and information restoration of the call system.

[0072] The phase coupling enhancement module constructs a phase coupling enhancement channel for low-volume speech signals based on the energy distribution characteristics and the spectral stability characteristics corresponding to the feature weight set, applies phase consistency mapping and harmonic synchronous amplification, and enhances the structural continuity and energy density of the audio signal.

[0073] Constructing a phase coupling enhancement channel for low-volume speech signals based on the energy distribution characteristics and the spectral stability characteristics corresponding to the feature weight set includes the following steps:

[0074] Based on the feature weight set generated in the previous step, the energy distribution feature and the spectral stability feature are taken as inputs to establish a phase tracking model of the signal. The model continuously tracks the phase of the main frequency component and its harmonic component of each frame of audio signal, extracts the instantaneous phase information using Hilbert transform, and constructs a phase trajectory on the time sequence. By matching with the spectral stability factor in the feature weight set, the frequency components with continuous phase change and stable amplitude change are selected as the target frequencies for subsequent enhancement. For the selected target frequencies, a phase difference mapping table of their adjacent frequency components is further established to ensure that the phase relationship between adjacent frequencies is effectively preserved during the subsequent phase coupling enhancement process, preventing spectral distortion or phase distortion after signal enhancement.

[0075] Hilbert transform is a mathematical transform that can convert real number signals into complex analytic signals. By introducing the orthogonal components of the signal, the instantaneous phase and instantaneous amplitude of the signal at each time point can be calculated. In this step, the role of Hilbert transform is to extract the phase of the frequency components of the low volume speech signal. By obtaining the instantaneous phase information of the signal, the phase change trajectory of the signal on the time axis is dynamically depicted to assist in judging whether the frequency components of the signal have continuity and harmonic characteristics. The specific steps include:

[0076] Apply Hilbert transform to the enhanced audio signal to obtain the corresponding complex analytic signal, with the real part being the original signal and the imaginary part being the 90-degree phase-shifted orthogonal signal.

[0077] Calculate the instantaneous phase at each time point from the analytic signal. The instantaneous phase is obtained from the arctangent function of the real and imaginary parts of the complex signal, which accurately describes the phase angle of the signal at that time.

[0078] Unwrap the instantaneous phase values at consecutive time points and eliminate the phase jump of 2π to form a continuous phase sequence through phase unwrapping operation.

[0079] Based on the continuous phase sequence, construct a phase trajectory and evaluate the phase consistency and continuity of each main frequency and its harmonics in the signal based on the phase change speed of different frequency components. If the phase trajectory is smooth and continuous, it indicates that the frequency component belongs to stable human speech signal, otherwise it is determined as background noise or interference frequency. This process ensures dynamic tracking and structural enhancement of the speech signal at the phase level.

[0080] After the phase tracking and mapping table of the target frequency is completed, a phase consistency mapping process is performed. The specific method is to synchronously adjust the phases of the target frequency and its adjacent frequencies, identify the frequency points with abnormal phase shift by calculating the phase similarity of the adjacent frequency components, and correct them through a phase compensation algorithm, so that the phases of all target frequencies and their harmonic components remain highly consistent. This step eliminates the phase disorder in the multi-band signal caused by noise interference, signal attenuation or recording errors through synchronous phase adjustment, thereby improving the overall structural coherence of the signal in the frequency domain. The signal after phase consistency mapping presents a clearer harmonic structure of the human voice in the frequency spectrum distribution, forming a phase spectrum highly consistent with the characteristics of the original speech signal.

[0081] Based on the audio signal after the phase consistency mapping is completed, a harmonic synchronous amplification process is implemented. The specific method is to apply dynamic gain control to the identified main frequency and its harmonic components, and the gain value is dynamically adjusted according to the energy distribution weight and the spectral stability weight in the feature weight set, to ensure that the enhanced amplitude matches the energy characteristics of the signal itself. In order to avoid the imbalance of sound quality caused by the over-enhancement of a single frequency, a cross-frequency energy balance mechanism is introduced to ensure that the amplitude ratio of the enhanced main frequency and its harmonic components is consistent with the harmonic energy distribution of natural human voice. In this way, the naturalness and clarity of the speech signal can be maintained while improving the overall energy density of the signal, especially when processing low-volume, whispering signals, the audible feeling and transmission fidelity can be significantly enhanced.

[0082] After the audio signal after the harmonic synchronous amplification is performed, the structure coherence in the time-frequency domain is reviewed to ensure the continuity and smoothness of the phase spectrum, amplitude spectrum and time-domain waveform of the signal after enhancement. The specific method is to calculate the phase spectrum difference, spectral envelope change rate and waveform smoothness index of the signal before and after enhancement, dynamically adjust the gain curve and phase compensation amount to avoid introducing new phase distortion or spectral artifacts in the enhancement process. After passing the review, the final enhanced audio signal is output, which has the characteristics of high energy density, strong structural coherence and natural spectral distribution, and can provide a high-quality signal basis for subsequent multi-dimensional speech pattern recognition and dynamic gain adjustment, ensuring clear transmission and restoration of low-volume speech in complex noise environments.

[0083] The purpose of this step is to solve the problem that low-volume speech signals are easily misjudged as background noise by the call system due to weak energy, small frequency variation amplitude, and unstable signal structure. By deeply enhancing the phase and spectral structure of the signal, the structural coherence and energy density of the audio signal are comprehensively improved, and the weak speech signal is accurately restored and robustly enhanced. The harmonic distribution of low-volume speech signals in the frequency domain is not obvious, the amplitude of the fundamental frequency and its harmonic components is low, and the phase information on the time axis is easily disturbed by noise, resulting in a decrease in the completeness and recognition of the signal. By using the energy distribution characteristics and spectral stability characteristics in the feature weight set, the important frequency components and corresponding timing characteristics in the signal can be first selected, and then the phases of the frequency components of the signal are synchronized and aligned through phase consistency mapping, eliminating the phase disorder caused by noise or energy attenuation and ensuring the coherence and coherence between frequency components. Subsequently, combined with the harmonic synchronous amplification technology, the natural energy ratio between the fundamental frequency and its harmonics is used to dynamically amplify the human voice harmonic components, not only improving the overall energy density, but also restoring the harmonic structure of the speech signal, making the weak speech signal more identifiable and natural in the frequency spectrum. Through this series of processing, the audio quality and completeness of the low-volume speech signal are significantly improved when it passes through the noise environment and signal transmission link, solving the technical pain points of the traditional call system in capturing and restoring weak speech signals, and being particularly suitable for meeting windows and other communication scenarios with high requirements for information accuracy and privacy.

[0084] The multi-dimensional speech mode recognition module performs multi-dimensional speech mode recognition based on the audio signal enhanced by the phase-coupled enhancement channel, and identifies the corresponding human speech attributes in the audio signal according to the pitch continuity, rhythm variation, and stop-continuous mode.

[0085] The multi-dimensional speech mode recognition based on the audio signal enhanced by the phase-coupled enhancement channel includes the following steps:

[0086] The audio signal processed by the phase-coupled enhancement channel is subjected to pitch feature extraction, mainly by converting the audio signal into spectral data through short-time Fourier transform, and extracting the fundamental frequency and its harmonic distribution in each time frame to form a pitch contour curve. In order to improve the accuracy of low-volume speech pitch extraction, a weighted autocorrelation method combined with a time-frequency domain fusion algorithm is used to ensure the accuracy of fundamental frequency detection while considering the continuity and variation details of the pitch. In addition, by calculating the smoothness and stability of the fundamental frequency in consecutive time frames, the coherence of the pitch is determined. If the pitch curve shows a continuous and gentle change trend on the time axis, and the harmonic components are distributed in proportion to the fundamental frequency, it is preliminarily determined that the signal has the pitch attribute of human speech.

[0087] Based on the extracted pitch features, further analyze the rhythm change pattern. By dynamically detecting the energy envelope curve and the fundamental frequency change rate of the audio signal in the time sequence, the rhythm density, beat interval and stress position of each frame signal are calculated to form a complete rhythm feature sequence. In order to distinguish natural speech from background noise or mechanical interference, a wavelet transform combined with a rhythm synchronous function analysis method is used to identify the periodic and aperiodic components in the rhythm. The periodic component of the rhythm change often corresponds to the speech speed and stress pattern of human language, and the aperiodic component is often associated with noise or non-speech signals. In this way, the rhythm contour of the speech signal can be effectively captured to provide dynamic feature basis for further judgment of speech attributes.

[0088] The energy envelope curve refers to a smooth curve describing the energy intensity change trend of the audio signal in the time sequence. By envelope fitting the short-time energy or amplitude of the signal, the loudness fluctuation and dynamic change of the sound can be reflected. The fundamental frequency change rate refers to the change speed of the fundamental frequency between adjacent time frames, indicating the fluctuation amplitude and change rate of the pitch in time, which can depict the change rhythm of the pitch in speech. In the present application, both are used to reveal the rhythm and stress position of the speech signal, and then extract the rhythm features. The specific steps are as follows:

[0089] The short-time energy of each frame of audio signal is calculated, and envelope fitting is applied to the energy values of all frames to form a smooth energy envelope curve. The peak value of the curve corresponds to the stress or strong pronunciation part in the speech, and the valley value corresponds to the pause or weak pronunciation.

[0090] The fundamental frequency of each frame is extracted, and the difference between the fundamental frequencies of adjacent frames is calculated to obtain the fundamental frequency change rate, which reflects the fast and slow changes of the pitch.

[0091] The energy envelope curve and the fundamental frequency change rate are jointly analyzed to detect the local maximum value in the envelope curve and the corresponding fundamental frequency fluctuation. These points are marked as potential stress positions, and the beat interval is further calculated by statistically analyzing the time interval between the peak values.

[0092] According to the distribution of stress positions and beat intervals, the rhythm density index is formed, the number of stresses in a unit of time is counted, and finally a complete rhythm feature sequence containing rhythm density, beat interval and stress position is generated, which provides accurate data support for subsequent speech attribute recognition and rhythm pattern analysis.

[0093] The analysis method combining wavelet transform and rhythm synchronization function refers to using wavelet transform to perform multi-scale decomposition on the audio signal, extracting the rhythm features of the signal at different time scales, and at the same time combining the rhythm synchronization function to measure the synchronization of the energy and rhythm changes of the decomposed signal, so as to distinguish the periodic components (such as the prosody and rhythm of natural speech) and non-periodic components (such as background noise, mechanical interference and other irregular fluctuations) in the signal. The role of wavelet transform here is to expand the audio signal in the time-frequency domain, which can not only capture the instantaneous changes of the signal, but also observe the rhythm features in different frequency bands, avoiding the defect that the traditional Fourier transform cannot balance the time-frequency locality. The rhythm synchronization function often uses Tempogram or Autocorrelation Function (autocorrelation function) to quantify the synchronization degree between the audio signal at each time point and different rhythm templates.

[0094] The specific steps are:

[0095] Perform continuous wavelet transform on the audio signal to obtain the energy distribution at different time scales, especially focusing on the frequency band related to speech rhythm, such as the modulation frequency in the range of 1Hz to 10Hz, to capture the rhythm pulse related to speech prosody.

[0096] Calculate the rhythm synchronization function on the energy envelope curve of the wavelet coefficients at each scale. The commonly used method is to perform autocorrelation function on the energy sequence at each scale to extract the periodicity of the sequence, and the peak value of the periodicity represents the basic period of the rhythm.

[0097] Combine the synchronization results of multiple scales to construct a time-rhythm frequency synchronization matrix. By detecting the strong synchronization area in the synchronization matrix, the periodic rhythm component of the signal is identified, and the part with low synchronization and no significant period is marked as aperiodic component.

[0098] Through the separation of periodic and aperiodic components, the natural speech rhythm and background noise, mechanical interference are effectively distinguished, and the real rhythm pattern in the low volume speech signal can be accurately extracted and recognized.

[0099] The recognition of the pause-continuation mode is implemented based on the rhythm feature and the pitch feature. The pause-continuation mode reflects the structure of the pauses and the continuous pronunciation in the speech stream. The time interval between the local minimum value and the local maximum value in the energy envelope curve is calculated, and the natural pause position in the speech is recognized in combination with whether the fundamental frequency is discontinuously changed. In particular, for signals such as low volume and whisper, the pauses are not obvious, and therefore a time series prediction model based on a long short-term memory network is introduced in this step to predict the continuation probability and the pause probability of the continuous speech segment, so as to ensure that the pause-continuation relationship can be accurately distinguished even in weak speech. By quantifying the pause duration, the frequency and the relevance to the rhythm and the pitch, a complete description of the pause-continuation mode is formed, so as to further confirm the natural speech property of the signal.

[0100] The time series prediction model based on the long short-term memory network refers to a deep learning method for modeling and predicting the time series features of the audio signal by using a long short-term memory (LSTM, Long Short-Term Memory). The LSTM is a special recurrent neural network that is good at capturing the long-term and short-term dependence relationship in the time series data and can effectively retain and forget the key information in the time series. In the present application, the function of the model is to learn the time series features such as the energy change, the fundamental frequency change rate and the spectral stability of the audio signal for the low volume speech signal after signal enhancement, to predict the continuation probability (i.e. the possibility of being speech in the future) and the pause probability (i.e. the possibility of pause or speech interruption in the future) of the continuous speech segment at the future time step, so as to accurately determine the pause-continuation relationship in the signal, especially in the case of weak speech signal and unobvious pause, to avoid misjudgment.

[0101] The specific steps are as follows:

[0102] The enhanced audio signal is divided into time windows of a fixed length, and the energy envelope, the fundamental frequency, the spectral stability and other features in each window are extracted to form a time series feature matrix.

[0103] The feature sequence is input into the pre-trained LSTM model. The model captures the dynamic change and trend of the speech signal in the time dimension through the mechanism of its memory cell, forget gate, input gate and output gate, and outputs the continuation probability and the pause probability corresponding to each time point.

[0104] By setting a dynamic threshold, if the continuation probability at a time point is greater than the pause probability and continues for more than a certain number of frames, it is determined that the segment is in a continuous speech state; otherwise, if the pause probability is higher than the continuation probability and continues for several frames, it is marked as a pause position.

[0105] The continuous speech output by the model and the marked sequence of pauses are cross-verified with the energy envelope and the fundamental frequency variation to ensure the accuracy and robustness of the marking, so that the continuity and pauses of the speech can be clearly distinguished even in weak and ambiguous speech, and the rhythm and language structure of the speech can be completely restored.

[0106] The recognition results of the three dimensions of pitch continuity, rhythm variation and pause-continuity mode are integrated, and a decision fusion mechanism is used to determine whether the audio signal has human speech properties. Specifically, the feature indicators of the three dimensions are normalized and input into the trained support vector machine classification model, and the multi-dimensional weight parameters of the model are combined to output a confidence score of human speech properties for each signal segment. The signal segment with a score higher than the set threshold is identified as having human speech properties, and the part with a score lower than the threshold is determined as non-speech or noise signal. Through the execution of the series of steps, not only can human speech be accurately distinguished in low volume and weak speech, but also the accuracy and robustness of speech property recognition can be guaranteed in complex noise background, providing a solid recognition foundation for subsequent dynamic gain adjustment and sound quality fidelity enhancement.

[0107] The trained support vector machine classification model refers to a classification decision model trained by a supervised learning method based on a large number of labeled human speech and non-speech signal samples using a support vector machine (SVM) algorithm. The core principle of the model is to construct an optimal hyperplane in a high-dimensional feature space to divide the input audio signal into two categories of "human speech property signal" and "non-speech or noise signal" in the multi-dimensional feature space. In the present application, the support vector machine model takes the vectors of pitch continuity, rhythm variation, pause-continuity mode and other multi-dimensional features after normalization as input, and the model learns the boundary relationship between these features and signal properties in the training stage, which can output a confidence score for each audio segment in actual application, reflecting the probability that the signal belongs to human speech. Support vector machine has the advantage of small sample and nonlinear feature discrimination, and is especially suitable for processing low volume, speech feature not obvious but with certain rules, so it can effectively identify speech segments in complex noise background, providing reliable decision basis for speech property determination and subsequent dynamic gain adjustment.

[0108] The role of this step is to comprehensively identify whether the signal has the essential attributes of human speech through multi-dimensional speech pattern recognition, ensuring accurate distinction between speech and non-speech signals in complex, low signal-to-noise ratio environments. Due to the sparse energy distribution of low-volume speech in the frequency spectrum and the weak structural coherence, traditional methods based on energy threshold or simple frequency discrimination often misclassify it as noise or invalid signals, resulting in information loss. By introducing multi-dimensional recognition of pitch continuity, rhythm variation, and pause-continuation patterns in this step, the natural attributes of speech can be restored from different speech feature levels. Pitch continuity analysis can identify the stability of the fundamental frequency and harmonics, rhythm variation reflects the unique stress and speed patterns of language, and pause-continuation patterns reveal the natural pauses and pronunciation continuity in speech flow. Comprehensive examination of these dimensions avoids the limitations of single feature recognition. Especially in low-volume scenarios, weak pronunciation still retains the pitch pattern, rhythm, and speech flow structure of human speech. This step can accurately determine the human speech attributes in the signal with high confidence through deep mining and pattern matching, providing accurate identification basis for subsequent dynamic gain adjustment and audio quality fidelity enhancement, ensuring that valuable speech information is not missed in dialogue-sensitive scenarios such as meetings, visits, and pacification, while improving the intelligence and reliability of the communication system.

[0109] The dynamic gain adjustment and audio quality enhancement module dynamically adjusts the gain mapping curve and spectral channel parameters of the audio signal transmission chain according to the identified human speech attributes, and implements parameter mapping and audio quality fidelity enhancement on the audio signal;

[0110] The dynamic gain adjustment and audio quality enhancement module dynamically adjusts the gain mapping curve and spectral channel parameters of the audio signal transmission chain according to the identified human speech attributes, and implements parameter mapping and audio quality fidelity enhancement on the audio signal;

[0111] Based on the human speech attribute results identified in the previous steps, the pitch characteristics, rhythm density, pause-continuation patterns, and energy distribution of the signal in the frequency spectrum are analyzed comprehensively to construct a dynamic gain mapping curve that matches the human speech attributes. The design of this gain mapping curve is based on the energy proportion of the signal in different frequency bands and the dynamic range of the signal, using a multi-segment nonlinear gain strategy: for the main frequency and harmonic frequency bands where human speech energy is concentrated, a higher gain boost coefficient is set to enhance the listening clarity of the speech component; for the low-frequency and high-frequency bands unrelated to speech, a gain suppression curve is set to avoid amplification of background noise or meaningless frequency components. The dynamic adjustment process of the gain curve introduces an adaptive threshold to dynamically adjust the gain proportion based on the speech attribute confidence of each segment of the signal, ensuring that the speech signal is enhanced while avoiding audio quality imbalance due to the presence of noise components.

[0112] The dynamic gain mapping curve refers to a curve that dynamically adjusts the variation law of gain (volume amplification multiple) on each frequency or time period according to the human speech attributes identified in the audio signal and the energy distribution and spectral characteristics of the signal in different frequency bands. Unlike fixed gain, the curve can assign different amplification coefficients to different frequency components of the speech signal according to the actual attributes of each audio signal, achieving selective enhancement. In the present application, the dynamic gain mapping curve is used to focus on improving the energy of the main frequency and harmonic components related to human speech and suppressing the frequency bands related to noise without blindly amplifying the whole signal, thereby avoiding the deterioration of sound quality caused by the simultaneous amplification of noise. The reason for constructing the dynamic gain mapping curve is that low-volume speech signals often lack energy and partially overlap with background noise spectrum. If fixed or linear gain amplification is used, noise and distortion may be amplified simultaneously, which may reduce speech clarity and fidelity. Through the dynamic gain curve that adapts to the speech attributes, the effective components of the speech can be selectively enhanced, so that the energy density, naturalness of sound quality, and clarity of listening experience of the enhanced signal are optimized, while ensuring the completeness of information and the accuracy of transmission during the conversation.

[0113] After the dynamic gain mapping curve is established, the spectral channel parameters of the audio signal are carefully adjusted. This adjustment process is based on the dynamic configuration of the filter bank, and adjusts the center frequency, bandwidth, and Q value (quality factor) of the filter in real time according to the speech attribute weight of different frequency bands in the signal, to achieve fine control of the speech signal spectrum. Specifically, for the frequency bands in the spectrum that are highly related to the identified human speech attributes, the bandwidth of the filter is appropriately widened to retain more natural speech details, and the Q value of the filter is increased to enhance the frequency resolution and prevent interference from adjacent noise frequencies; for frequency bands that are not related to speech attributes or have more noise components, the bandwidth is reduced and the Q value is lowered, thereby effectively suppressing noise frequencies. This step ensures both the integrity of the speech signal and the noise suppression in the frequency domain.

[0114] The dynamic gain mapping curve and the adjusted spectral channel parameters are applied synchronously in the audio signal processing chain to perform parameter mapping operations. This operation maps the gain and spectral parameters to each time frame and frequency unit of the audio signal through a digital signal processing algorithm, achieving precise enhancement of the signal frame by frame and frequency by frequency. To ensure the smoothness and naturalness of the mapping, a dynamic smoothing factor and a transition window function are used to buffer the parameter changes, avoiding signal distortion or spectral tearing caused by sudden changes in gain or filter parameters. At the same time, to maintain the fidelity of the sound quality, a spectral reconstruction and phase compensation mechanism is introduced in the mapping process to ensure that the phase consistency and temporal coherence of the original speech are not destroyed during the enhancement process.

[0115] Digital signal processing algorithm refers to a class of algorithms that analyze, transform and enhance discrete digital audio signals through mathematical methods and computational models. Common algorithms include Fast Fourier Transform (FFT), Short-Time Fourier Transform (STFT), adaptive filtering, dynamic range compression, spectral subtraction, etc. In this step, the role of the digital signal processing algorithm is to accurately map and apply the dynamic gain mapping curve and spectral channel parameters obtained after human speech attribute recognition to each time frame and corresponding frequency unit of the audio signal, achieving targeted enhancement and noise suppression of the signal.

[0116] The specific steps are as follows:

[0117] Using Short-Time Fourier Transform (STFT), the audio signal in the time domain is divided into frames and converted into a time-frequency domain representation, obtaining the amplitude and phase information of each time frame at different frequency units.

[0118] According to the gain coefficients of each frequency in the dynamic gain mapping curve and the spectral channel parameters such as the bandwidth and center frequency of the filter, the amplitude of each frequency unit of each time frame is multiplied and amplified or suppressed, and the bandwidth of the frequency distribution is adjusted according to the spectral parameters to strengthen the frequency band related to the speech feature and suppress the noise frequency band.

[0119] Through the phase compensation algorithm, it is ensured that the phase information of the frequency unit remains consistent with the original signal after amplitude adjustment, avoiding phase distortion after enhancement.

[0120] Apply Inverse Short-Time Fourier Transform (ISTFT) to restore the adjusted time-frequency signal back to the time domain and obtain the enhanced audio signal, achieving accurate enhancement and fidelity recovery of each frame and each frequency band. Through this process, the energy density and clarity of human speech components can be targetedly improved, and background noise and invalid frequencies can be suppressed to improve overall sound quality and call experience.

[0121] After the audio signal is parameter-mapped and sound quality-enhanced, real-time sound quality evaluation and adaptive feedback optimization are performed. By calculating the signal-to-noise ratio, spectral smoothness, dynamic range and harmonic distortion rate, etc., the sound quality effect after enhancement is dynamically evaluated. If the evaluation index does not meet the preset sound quality standard, the gain mapping curve and spectral channel parameter settings are automatically adjusted through the feedback mechanism to achieve iterative optimization. This step ensures that the optimal sound quality and fidelity enhancement effect can be achieved in different environments and different speech attributes, ensuring clear, natural and complete transmission of low-volume, whisper and emotional expression of speech signals in the call system.

[0122] The role of this step is to dynamically adjust the gain mapping curve and spectral channel parameters in the audio signal transmission chain based on the performance of the identified human voice attributes in pitch, rhythm, stop-connection mode, and spectral energy distribution, etc. Fine signal enhancement and sound quality fidelity processing are implemented to ensure that low volume and weak voice signals can be clearly, naturally and completely restored during the call. Since low volume or whisper voice signals have low energy density and overlap with environmental noise spectrum, if fixed gain or simple filtering is used, the full frequency band signal will be amplified simultaneously, and noise and distortion will also be enhanced, which will reduce the perceptibility and clarity of the voice. By dynamically adjusting the gain mapping curve, the gain corresponding to the strength of the voice attribute can be applied to each frequency band of the audio signal, so that the voice-related frequency band obtains a higher amplification ratio, while the non-voice attribute-related frequency band is suppressed to avoid meaningless enhancement. At the same time, the dynamic adjustment of the spectral channel parameters, such as the center frequency, bandwidth, and quality factor of the filter, can adapt the optimal frequency channel configuration according to the identified voice characteristics, further eliminating the interference frequency. Through this parameter mapping and sound quality enhancement linkage mechanism, not only the energy density and voice clarity of the signal are improved, but also the natural tone and spectral structure of the voice are preserved, avoiding the audio distortion or discomfort introduced by traditional enhancement methods. Finally, this step ensures that in the meeting window call scenario, even in the communication of low-energy voice, low-emotion expression or private whisper, the voice information can be accurately captured and restored with high quality, effectively ensuring the smoothness and security of communication.

[0123] The full-link voice integrity verification module performs full-link voice integrity verification on the audio signal adjusted by the gain mapping curve and spectral channel parameters, compares the attribute characteristics of the audio signals before and after enhancement, and verifies the false negatives to ensure the complete restoration and accurate reception of low volume voice signals.

[0124] The specific steps of performing full-link voice integrity verification on the audio signal adjusted by the gain mapping curve and spectral channel parameters, and comparing the attribute characteristics of the audio signals before and after enhancement to verify the false negatives to ensure the complete restoration and accurate reception of low volume voice signals are as follows:

[0125] After the audio signal is processed by dynamic gain mapping curve and spectral channel parameter adjustment, a series of detailed multi-dimensional attribute features are extracted from the enhanced audio signal. The specific features extracted include: energy distribution features in the frequency domain, using segmented Fourier transform to divide the 0-8kHz full frequency band signal into 500Hz sub-bands, a total of 16 sub-bands, and calculate the average energy value of each sub-band; spectral stability features, by calculating the cosine similarity of the spectral envelope of each frame of signal and the envelope of the adjacent frame, forming a similarity sequence of consecutive frames; pitch continuity features, using a fundamental frequency extraction algorithm such as the YIN algorithm to extract the fundamental frequency curve over time, and calculating the fundamental frequency variation smoothness through first-order difference; rhythm density features, counting the number of peak values of the energy envelope curve per second, reflecting the frequency of pronunciation rhythm per unit time; pause-continuous mode features, recording the duration of continuous speech segments and silent segments, and forming a time sequence of pauses and continuous pronunciation. Each of the above features is recorded in the form of a vector or sequence to ensure the rigor and integrity of subsequent comparative analysis.

[0126] The above feature sets of the pre-enhanced and post-enhanced audio signals are strictly compared one by one. The energy distribution feature comparison uses the cosine similarity of each frequency sub-band to calculate the similarity percentage between the energy vectors of each sub-band before and after enhancement. The spectral stability is compared by dynamic time warping algorithm (DTW) to pair and compare the spectral envelope similarity sequences before and after enhancement, and calculate the minimum cumulative distance of the two sequences. In the pitch continuity comparison, the Euclidean distance is used to measure the difference between the fundamental frequency curves before and after enhancement, and the smoothness change rate is calculated. The rhythm density is compared by directly comparing the change in the number of peak values per unit time, and the pause-continuous mode is compared by using the edit distance algorithm to quantify the similarity between the pause and pronunciation time sequences. Finally, the differences in the five types of features are quantified as standardized scores, and a global speech similarity score is formed by weighting. If the score is lower than the 90% consistency threshold, it is considered that there is a risk of false negative or feature loss, and enters the false negative backtracking verification stage.

[0127] The dynamic time warping (DTW) algorithm is a method for measuring the similarity between two time series of different lengths or different speeds. The core of the algorithm is to find a non-linear alignment path through dynamic programming, so that the two sequences are best matched in the time dimension, and the minimum cumulative distance between them is calculated. In this invention, the role of DTW is to align and compare the spectral envelope similarity sequences of the pre-enhanced and post-enhanced audio signals point by point, to solve the problem of slight stretching, compression or offset of the signal time sequence, amplitude or spectral profile caused by the enhancement process, and to ensure that even if the signal has a slight change in the time axis during the enhancement process, the consistency of its spectral form can still be accurately evaluated. The specific steps are as follows: First, the spectral envelope similarity sequence extracted from the pre-enhanced audio signal is recorded as the reference sequence A, and the similarity sequence after enhancement is recorded as the target sequence B. Second, using the DTW algorithm, the amplitude difference of each corresponding frame in A and B is calculated by taking the spectral envelope value of each time frame as the coordinate point in a two-dimensional distance matrix, and the local distance is formed by filling the matrix. Third, DTW finds a cumulative minimum cost path from the starting point to the ending point in the matrix through recursive method, which allows the time axis between sequences to be nonlinearly stretched or compressed, so that the overall path has the minimum cost value (i.e. the sum of the square of the amplitude difference). Finally, the minimum cumulative distance is taken as the difference measure of the spectral envelope similarity before and after enhancement. The smaller the distance, the better the preservation of spectral stability during the enhancement process, otherwise it indicates that the signal structure may be distorted or degraded due to enhancement, and the enhancement parameters need to be further adjusted.

[0128] The edit distance algorithm (EditDistance), also known as Levenshtein distance, is an algorithm for measuring the difference between two sequences by calculating the minimum number of editing operations required to transform one sequence into the other, which usually includes insertion, deletion, and replacement. In this invention, the edit distance algorithm is used to accurately quantify the similarity of the pause and pronunciation time series of the pre-enhanced and post-enhanced audio signals, to determine whether the original pause and continuous pronunciation structure has been mistakenly deleted, expanded, or modified during the enhancement process. Specifically, first, the pre-enhanced and post-enhanced audio signals are converted into binary sequences, with "1" representing the pronunciation segment and "0" representing the pause segment, while recording the duration of each segment. Then, the edit distance algorithm is applied to compare the two sequences bit by bit, and the minimum number of insertion, deletion, or replacement operations required to transform the original sequence into the post-enhanced sequence is counted. For example, if there is a pause after continuous pronunciation in the original sequence, but the pause segment disappears after enhancement, it requires one insertion or replacement operation. The fewer the total number of editing operations, the better the consistency of the pause and continuous pattern before and after enhancement, and the more the number, the more likely the enhancement process has caused deviations in the speech rhythm and pause-continuous structure. Through this quantitative indicator, the fidelity of the speech rhythm structure during the enhancement process can be effectively evaluated, ensuring that the continuous and pause relationship of low-volume speech is not destroyed and maintaining the natural fluency of the speech.

[0129] The false negative backtracking verification process is started, and for the audio segment detected to have insufficient similarity score, the original unenhanced signal segment is backtracked to, and the gain mapping and spectral channel adjustment of different parameter combinations are re-applied. Multiple sets of gain curve adjustment schemes are set, such as adjusting the main frequency and harmonic frequency band gain from +6dB in the original scheme to +4dB or +8dB, adjusting the suppression of non-speech frequency band from -10dB to -15dB or -5dB, and adjusting the center frequency ±100Hz fine tuning and bandwidth of the spectral channel filter from 200Hz to 300Hz or from 100Hz. The five types of features mentioned above are extracted from the audio signal enhanced by each parameter combination, and compared with the original signal again. The parameter configuration with the highest similarity score is automatically selected by the Bayesian optimization algorithm to ensure the optimal balance between audio quality improvement and feature fidelity. After completion, the optimal enhancement result of the false negative backtracking is replaced with the original signal segment.

[0130] The enhanced audio signal after backtracking verification and correction is subjected to full-link integrity review. This review not only extracts the five types of features of complete energy distribution, spectral stability, tone continuity, rhythm density, and pause-connection mode again, and forms the final similarity score, but also calculates the signal full-band SNR, THD, and DR, etc. quality indicators, such as SNR quantified in dB, THD controlled below 1%, and DR maintained at least above 60dB. If all indicators meet the set double thresholds of integrity and quality, it is confirmed that the enhancement link processing of the audio signal meets the standard, realizing the complete restoration and high-fidelity reception of low-volume speech, as the final call output signal.

[0131] The purpose of this step is to evaluate whether the signal enhancement process has caused damage or omission to the original attributes and structure of the low-volume speech signal by performing full-link speech integrity verification on the audio signal after gain mapping curve and spectral channel parameter adjustment, ensuring that all human speech components are restored and accurately received. Low-volume speech signals are relatively fragile in terms of energy, spectrum, rhythm, etc. If the enhancement process lacks targeted integrity monitoring, valuable speech information may be mistakenly deleted (i.e. "mis-killed") due to filtering, gain adjustment, spectral compression, etc., or distortion and information loss may be introduced, especially for whispered, emotionally suppressed speech expressions, which are more likely to be incorrectly processed as background noise to be removed. Therefore, this step extracts multi-dimensional attribute features of the audio signal before and after enhancement, such as energy distribution, spectral stability, tone continuity, rhythm density, and pause-connection mode, and uses dynamic time warping, edit distance, etc. algorithms to realize accurate feature comparison and identify the differences between the enhanced and original signals in various attributes. At the same time, a mis-killing backtracking verification mechanism is introduced to backtrack the signal segment with insufficient similarity to the original signal, re-adjust the gain and spectral adjustment parameters, and after multiple scheme attempts, the enhancement effect closest to the original attributes is retained. Through this full-link verification and backtracking, it is ensured that in the enhancement process of low-volume, weak speech, the signal clarity and energy density are improved, while the original pronunciation rhythm, tone features, and pause-connection relationship are not damaged, ensuring that every piece of speech information delivered in the call is complete, accurate, and true, greatly improving the quality of speech interaction and information restoration in the meeting window call scenario.

[0132] The meeting window call system based on the voice sensor can systematically and fully-link solve the problem of insufficient recognition of special voice states such as low volume, low frequency change and low speech speed in the prior call system, and realize the full-process closed-loop processing of high-sensitivity collection, accurate recognition, intelligent enhancement and faithful restoration of weak voice signals. The scheme improves the spatial and frequency coverage of signal collection through array type sensing layout, realizes intelligent differentiation of human voice and background noise through dynamic confidence evaluation and feature weight construction, and effectively restores the energy density and structural coherence of low-volume voice through phase coupling and harmonic enhancement. Further through multi-dimensional speech mode recognition and dynamic gain adjustment, the human voice properties are accurately strengthened, noise amplification is avoided, and under the guarantee of full-link integrity checking and false kill backtracking mechanism, the weak signal information is not missed and not distorted. The system is particularly suitable for application scenarios such as legal visit, psychological counseling, family pacification and other high-sensitivity dialogue, and strictly requires information integrity, significantly improves the comprehensive ability and application value of the call system in privacy, security and voice information complete transmission.

[0133] The above only describes some exemplary embodiments of the present application by way of illustration, and it is needless to say that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present application. Therefore, the above figures and description are illustrative in nature and should not be understood as limiting the scope of the claims of the present application.

[0134] It should be noted that in this paper, if there are relationship terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0135] It should be understood that in various embodiments of the present application, the size of the sequence number of each process described above does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0136] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0137] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0138] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment according to actual needs.

[0139] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit.

[0140] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any skilled person in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0141] The above only describes some exemplary embodiments of the present application by way of illustration, and it is self-evident that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present application. Therefore, the above figures and descriptions are illustrative in nature and should not be understood as limiting the scope of protection of the claims of the present application.

Claims

1. A meeting window communication system based on voice sensor, characterized in that: It includes a full-space speech acquisition module, a dynamic confidence assessment module, a phase coupling enhancement module, a multi-dimensional speech pattern recognition module, a dynamic gain adjustment and sound quality enhancement module, and a full-link speech integrity verification module: The full-space voice acquisition module deploys an array of voice sensor units within the meeting window space based on audio reflection characteristics, constructs a multi-band full-space acquisition matrix, collects the voice signals of both parties in the meeting, and outputs a basic audio data stream with a high signal-to-noise ratio; The dynamic confidence assessment module performs dynamic confidence assessment on the basic audio data stream, extracts energy distribution and spectrum stability features, and generates a set of feature weights to distinguish speech from background noise; The phase coupling enhancement module constructs a phase coupling enhancement channel based on a set of feature weights. It enhances the structural coherence and energy density of the audio signal through phase consistency mapping and harmonic synchronous amplification. A multi-dimensional speech pattern recognition module performs multi-dimensional speech pattern recognition based on the enhanced audio signal, identifying human speech attributes based on pitch continuity, rhythm changes, and pause patterns; Dynamic gain adjustment and sound quality enhancement module, which dynamically adjusts the gain mapping curve and spectrum channel parameters of the audio signal transmission chain based on the recognized human voice attributes to implement sound quality fidelity enhancement; The full-link voice integrity verification module performs full-link voice integrity verification on the enhanced audio signal. Through attribute feature comparison and false positive backtracking verification, it achieves complete restoration and accurate reception of low-volume voice signals.

2. A voice sensor-based meeting window communication system according to claim 1, characterized in that: The steps to obtain the basic audio data stream are as follows: Based on the material, location, and acoustic reflection characteristics of the walls, ceiling, floor, and soundproof glass in the meeting window space, perform sound field simulation and modeling to determine the sound reflection path, attenuation coefficient, and standing wave distribution; Based on the sound field simulation results, highly sensitive capacitive or piezoelectric voice sensors are selected, arranged in a mixed pattern of equal and non-equal spacing, and sensors with tilted or directional pickup are set up to build a full-space acquisition matrix covering multiple frequency bands. The signal of each voice sensor is pre-amplified and digitized through the preamplifier circuit and analog-to-digital conversion unit, and then enters the multi-channel signal synchronization platform based on timing synchronization to achieve phase consistency and time synchronization; The digitized signal is divided into frequency bands and weighted based on a multi-band filter bank, and background noise is reduced by weighted averaging and spectral subtraction. Dynamic amplitude normalization processing with adaptive gain control is then performed to output a basic audio data stream with a high signal-to-noise ratio.

3. The voice sensor-based meeting window communication system according to claim 1, characterized in that: The steps for generating feature weight sets are as follows: Based on the basic audio data stream, the energy distribution characteristics, spectrum stability characteristics, spectrum envelope similarity, spectrum entropy and main frequency component persistence of each frame signal are extracted to form a multi-dimensional feature vector of the time series; The multi-dimensional feature vector sequence is input into a deep learning model composed of a convolutional neural network and a bidirectional long short-term memory network in series, and the confidence score of each frame signal is output; Smoothing and thresholding of the scoring sequence to identify valid speech areas where the scores of consecutive frames are higher than the threshold; Based on the confidence score and each feature, a feature weight set is generated for signal enhancement and recognition processing.

4. The voice sensor-based meeting window communication system according to claim 1, characterized in that: The specific steps for enhancing the structural coherence and energy density of audio signals using phase consistency mapping and harmonic synchronization amplification based on a feature weight set are as follows: Based on the energy distribution characteristics and spectrum stability characteristics of the feature weight set, a phase tracking model of the signal is established. The instantaneous phase information is extracted through Hilbert transform, the phase trajectory is constructed, and the target frequency with continuous phase and stable amplitude is selected. Perform phase consistency mapping on the target frequency and adjacent frequencies, perform phase compensation based on phase similarity, correct phase offset, and maintain phase consistency of frequency components; Apply dynamic gain control to the main frequency and its multiple harmonic components, dynamically adjust the gain according to the characteristic weight set, enhance the energy density and maintain the natural harmonic energy distribution ratio; Perform structural coherence review in the time-frequency domain, dynamically optimize based on phase spectrum difference, spectrum envelope change rate and waveform smoothness indicators, and output enhanced audio signals with high energy density and strong structural coherence.

5. The voice sensor-based meeting window communication system according to claim 1, characterized in that: Based on the enhanced audio signal, the specific steps for identifying whether the signal has human speech attributes are as follows through multi-dimensional analysis of pitch continuity, rhythm changes, and pause patterns: The enhanced audio signal is subjected to short-time Fourier transform to extract the fundamental frequency and harmonics, form a pitch contour curve, and calculate the fundamental frequency smoothness and stability; The energy envelope curve and the fundamental frequency change rate are used to jointly analyze the rhythm density, beat interval and accent position, and the periodic and non-periodic rhythm components are identified by combining wavelet transform and rhythm synchronization function. Long short-term memory networks are used to predict the continuation and pause probabilities of speech segments based on time series features, distinguishing between continuous pronunciation and natural pauses, and forming pause-continuation patterns. The features of pitch continuity, rhythm changes and pause patterns are input into the support vector machine classification model, and the human speech attribute confidence score of each signal segment is output to determine the speech attributes of the audio signal.

6. The voice sensor-based meeting window communication system according to claim 1, characterized in that: The specific steps for dynamically adjusting the gain mapping curve and spectral channel parameters of the audio signal transmission chain to enhance sound quality are as follows: Based on the recognized human speech attributes, a dynamic gain mapping curve is constructed, the gain boost coefficient is set according to the main frequency and harmonic distribution, and gain suppression is set for the non-speech frequency band; Dynamically adjust the spectrum channel parameters according to the speech attribute weight, adjust the center frequency, bandwidth and quality factor of the filter, enhance the voice band and suppress the noise band; The dynamic gain mapping curve and spectrum channel parameters are mapped to each time frame and frequency unit through digital signal processing algorithms, and phase compensation and spectrum reconstruction are applied to perform frame-by-frame and frequency band enhancement of the audio signal. Real-time sound quality evaluation is performed on the enhanced signal. Based on the signal-to-noise ratio, spectral smoothness, dynamic range and harmonic distortion rate indicators, the gain mapping curve and spectral channel parameters are dynamically optimized through a feedback mechanism to ensure the enhanced sound quality.

7. The voice sensor-based meeting window communication system according to claim 1, characterized in that: The specific steps for achieving complete restoration and accurate reception of low-volume speech signals through attribute feature comparison and false positive backtesting are as follows: Extract five attribute features of the audio signal after adjusting the gain mapping curve and spectrum channel parameters: energy distribution, spectrum stability, pitch continuity, rhythm density and pause pattern; The five attribute features of the enhanced and unenhanced signals are compared using cosine similarity, dynamic time warping, Euclidean distance, rhythm density comparison, and edit distance algorithms to form a global speech similarity score. For signal segments with scores below the consistency threshold, we trace back to the original signal, apply different gain and spectrum parameter configurations for enhancement, and select the optimal configuration through Bayesian optimization to replace the original signal segment. The five types of features are extracted again from the corrected signal and the signal-to-noise ratio, harmonic distortion rate and dynamic range are calculated. The integrity and sound quality indicators are reviewed, and the final enhanced signal is output after meeting the standards.

Citation Information

Patent Citations

  • Audio processing method and device, electronic equipment and computer readable storage medium

    CN115835093A

  • Multi-microphone array beamforming signal enhancement method and device

    CN119811408A