An AI-based method and system for enhancing audio-visual conferencing speech.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-09
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本申请目的是提供一种基于人工智能的视听会议语音增强方法和系统,以解决现有技术中语音增强清晰度较低的问题
[0016]本申请所提供的基于人工智能的视听会议语音增强方法,首先通过提取面部发声关联区域的空间复杂度和时间活跃度,并结合梯度对比度与帧间演进频率滤除干扰,获得了纯粹反映生理发声动作的视觉能量序列。同时通过对语音时域信号提取并平滑处理得到过零率序列,进而计算视觉与音频双模态序列的逻辑同步性,从而能够在高频白噪声严重干扰的环境下准确判别出具备清音特征的语音帧,避免了现有技术中将随机噪声误判放大的缺陷。
Smart Images

Figure CN122575386A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence and speech signal processing technology, and in particular relates to an artificial intelligence-based method and system for enhancing audiovisual conferencing speech. Background Technology
[0002] With the widespread adoption of remote work and online collaboration, AI-based audiovisual conferencing speech enhancement methods, by integrating multimodal information such as audio and vision, have demonstrated significant application potential in improving the quality of voice communication in complex environments. This method not only effectively suppresses background noise but also significantly improves the intelligibility of participants' speech, showing broad application prospects in far-field communication and low-end audio acquisition scenarios.
[0003] Currently, existing AI-based audiovisual conferencing speech enhancement methods typically utilize deep learning networks to extract lip movement features and then concatenate these features with the spectrogram of noisy speech to generate a frequency domain mask for speech restoration. These methods primarily rely on direct spectral mapping, suppressing environmental interference and preserving valid speech components by filtering the overall amplitude of the audio signal or by masking and scaling at the feature level.
[0004] However, when high-frequency voiceless consonant details in speech are completely submerged by ambient high-frequency white noise or lost due to hardware limitations, existing methods relying on spectral mapping often fail to reconstruct these lost acoustic features from scratch, easily leading to the recovery of speech lacking consonant components. This approach not only makes the enhanced speech unclear but also easily misinterprets high-frequency random noise as speech, severely reducing speech intelligibility and listening quality. Therefore, existing technologies suffer from the technical problem of low speech enhancement clarity due to the inability to accurately reconstruct and recover consonant details in cases of severe high-frequency noise interference or loss of signal details. Summary of the Invention
[0005] The purpose of this application is to provide an artificial intelligence-based audiovisual conferencing speech enhancement method and system to solve the problem of low speech enhancement clarity in the prior art.
[0006] To address the aforementioned technical problems, in a first aspect, this application provides an artificial intelligence-based audio-visual conferencing speech enhancement method, comprising: Acquire facial image sequences of the target object and speech time-domain signals including high-frequency white noise from the environment, collected by the conference terminal device; By calculating the spatial complexity and temporal activity of pixels in the speech-related region of the facial image sequence, the image entropy matrix at each time step is obtained. By calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components, the visual energy sequence is obtained. The speech time-domain signal is framed using window function weighting, and the number of times the signal waveform in each frame crosses the zero level is calculated to obtain the zero-crossing rate distribution sequence. The zero-crossing rate distribution sequence is then smoothed and weighted using the energy contrast parameter determined based on the speech time-domain signal to obtain the zero-crossing rate sequence. The logical synchronization between the visual energy sequence and the zero-crossing rate sequence is calculated to obtain the classification probability vector at each time step. When the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time step, the gain parameters, center frequency, and bandwidth of the harmonic oscillator are configured based on the statistical amplitude of the image entropy matrix at the current time step and the distribution density of the zero-crossing rate sequence to obtain the harmonic components. Harmonic components are phase-aligned and superimposed with the speech time-domain signal within a frequency range determined by the center frequency and bandwidth to enhance consonant details.
[0007] Optionally, the image entropy matrix at each time step is obtained by calculating the spatial complexity and temporal activity of pixels in the speech-related region of the facial image sequence, including: The space complexity is obtained by calculating the gray-level gradient values of adjacent pixels in the voice-related region of a facial image sequence. The temporal activity is obtained by calculating the grayscale difference of pixels at corresponding coordinate positions in adjacent image frames in a facial image sequence. By weighting and superimposing the spatial complexity and temporal activity, the image entropy matrix at each time step is obtained.
[0008] Optionally, by calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components, a visual energy sequence is obtained, including: Calculate the gradient of the numerical distribution between each element and its neighboring elements in the image entropy matrix to obtain the gradient contrast of each element; The temporal difference values of the corresponding coordinate positions in the image entropy matrix at adjacent time points are calculated to obtain the inter-frame evolution frequency; Based on the inter-frame evolution frequency, the weight coefficients of each pixel region are obtained according to the preset mapping relationship between the inter-frame evolution frequency and the weight coefficients. The gradient contrast is then weighted and aggregated using the weight coefficients to obtain the visual energy sequence.
[0009] Optionally, the speech time-domain signal is framed using a window function weighting method, and the number of times the signal waveform in each frame crosses the zero level is calculated to obtain a zero-crossing rate distribution sequence. The zero-crossing rate distribution sequence is then smoothed and weighted using an energy contrast parameter determined based on the speech time-domain signal to obtain a zero-crossing rate sequence, including: The speech time-domain signal is time-sequentially segmented according to a preset sampling length, and the amplitude of each segmented signal segment is mapped using a preset weighting window to obtain multiple analysis frames. The number of polarity reversals of adjacent signal values within each analysis frame is accumulated to obtain the zero-crossing rate distribution sequence; Calculate the ratio of the mean signal strength of each analysis frame to the preset environmental noise benchmark value to obtain the corresponding energy comparison parameter; The zero-crossing rate distribution sequence is obtained by using the energy contrast parameter to perform amplitude weighting.
[0010] Optionally, the logical synchronicity between the visual energy sequence and the zero-crossing rate sequence is calculated to obtain the classification probability vector at each time step, including: The degree of matching between the changing trends of the visual energy sequence and the zero-crossing rate sequence at each time step is calculated to obtain the logical synchronization. By mapping logical synchronization to a preset classification probability interval, a classification probability vector is obtained to determine whether the speech time-domain signal belongs to unvoiced, voiced, or environmental interference at each moment.
[0011] Optionally, when the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time, the gain parameters, center frequency, and bandwidth of the harmonic oscillator are configured based on the statistical amplitude of the image entropy matrix at the current time and the distribution density of the zero-crossing rate sequence, respectively, to obtain the harmonic components, including: The statistical amplitude is obtained by summing the values of all matrix elements in the image entropy matrix at the current moment. Based on the statistical amplitude, the gain parameter is obtained according to the preset mapping relationship between the statistical amplitude and the gain parameter. Calculate the statistical proportion of values in the zero-crossing rate sequence at the current time that are within a preset threshold range to obtain the distribution density. Based on the distribution density, and according to the preset correspondence between the distribution density and frequency characteristics, obtain the center frequency and bandwidth. Harmonic components are generated by controlling the output strength and oscillation frequency range of the harmonic oscillator through gain parameters, center frequency, and bandwidth.
[0012] Optionally, the harmonic components are phase-aligned and superimposed with the speech time-domain signal within a frequency range determined by the center frequency and bandwidth to enhance consonant details, including: Frequency extraction is performed on the speech time-domain signal based on the center frequency and bandwidth to obtain the reference signal within the target frequency band. Calculate the phase deviation between the reference signal and the harmonic components, and perform offset correction on the initial phase of the harmonic components based on the phase deviation to obtain the aligned harmonic components. Within a frequency range with a defined center frequency and bandwidth, the aligned harmonic components are numerically superimposed on the speech time-domain signal to enhance consonant details.
[0013] Secondly, this application provides an artificial intelligence-based audio-visual conferencing voice enhancement system, comprising: The acquisition module is used to acquire facial image sequences of the target object and speech time-domain signals including environmental high-frequency white noise collected by the conference terminal device; The computation module is used to obtain the image entropy matrix at each time step by calculating the spatial complexity and temporal activity of pixels in the voice-related region of the facial image sequence. The visual energy sequence is obtained by calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components. The generation module is used to frame the speech time-domain signal using a window function weighting and to calculate the number of times the signal waveform in each frame crosses the zero level, thus obtaining a zero-crossing rate distribution sequence. The zero-crossing rate distribution sequence is then smoothed and weighted using the energy comparison parameter determined based on the speech time-domain signal to obtain a zero-crossing rate sequence. The generation module is also used to calculate the logical synchronization between the visual energy sequence and the zero-crossing rate sequence, obtain the classification probability vector at each time step, and when the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time step, the gain parameters, center frequency and bandwidth of the harmonic oscillator are configured based on the statistical amplitude of the image entropy matrix at the current time step and the distribution density of the zero-crossing rate sequence, respectively, to obtain the harmonic components. The processing module is used to perform phase alignment and signal superposition of harmonic components and speech time-domain signals within a frequency range determined by the center frequency and bandwidth, in order to enhance consonant details.
[0014] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the AI-based audiovisual conferencing voice enhancement method as described in the first aspect above.
[0015] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the AI-based audiovisual conferencing voice enhancement method described in the first aspect above.
[0016] The AI-based audiovisual conferencing speech enhancement method provided in this application first extracts the spatial complexity and temporal activity of facial vocalization-related regions, and then filters out interference by combining gradient contrast and inter-frame evolution frequency to obtain a visual energy sequence that purely reflects physiological vocalization actions. Simultaneously, it extracts and smooths the speech time-domain signal to obtain a zero-crossing rate sequence, and then calculates the logical synchronization of the visual and audio bimodal sequences. This enables accurate identification of speech frames with unvoiced characteristics even in environments with severe high-frequency white noise interference, avoiding the defect of amplifying and misjudging random noise in existing technologies.
[0017] Subsequently, when a frame is identified as unvoiced, the statistical amplitude of the visual image entropy matrix and the distribution density of the auditory zero-crossing rate sequence are used to drive cross-modal collaboration. The gain parameters, center frequency, and bandwidth of the harmonic oscillator are adaptively configured to generate the corresponding harmonic components. Within a defined frequency range, the harmonic components are phase-aligned and superimposed with the original speech time-domain signal. This achieves precise parameterized reconstruction and compensation of masked or lost high-frequency acoustic features at the physical waveform level.
[0018] Therefore, this application effectively solves the technical problem in the prior art that the inability to accurately reconstruct and restore consonant details in the case of severe high-frequency noise interference or loss of signal details leads to low speech enhancement clarity, and significantly improves the intelligibility and listening quality of speech in audiovisual conferencing scenarios.
[0019] Furthermore, this application obtains the statistical amplitude by summing the numerical values of the image entropy matrix and maps the gain parameter. At the same time, it calculates the statistical proportion of the zero-crossing rate sequence in the preset threshold interval to obtain the distribution density and maps the center frequency and bandwidth. In this way, the physiological kinetic energy intensity of facial vocalization is quantitatively converted into the control amplitude of harmonic output, and the remaining acoustic distribution law is converted into the high-frequency oscillation range required for reconstruction.
[0020] This cross-modal quantization control method breaks through the limitations of existing technologies that rely excessively on incomplete spectrum for masking and scaling when dealing with scenarios where high-frequency unvoiced consonant details are completely submerged or physically lost by ambient high-frequency white noise. It can utilize the deep mapping between visual motion amplitude and auditory high-frequency oscillation patterns to precisely drive the harmonic oscillator to generate customized harmonic components.
[0021] By precisely fitting and reshaping the high-frequency acoustic characteristics that best match the current sound generation logic at the parameter level, this application eliminates the risk of mis-amplification of high-frequency random noise caused by direct filtering. Therefore, this application further solves the technical problem of low speech enhancement clarity caused by the inability to accurately reconstruct and recover consonant details in the event of severe high-frequency noise interference or loss of signal details. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart illustrating an AI-based audio enhancement method for audiovisual conferencing, provided as an embodiment of this application; Figure 2A flowchart illustrating a method for generating a zero-crossing rate sequence provided in an embodiment of this application; Figure 3 A flowchart illustrating a method for generating harmonic components provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of an AI-based audio-visual conferencing voice enhancement system provided in an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0024] In AI-based audiovisual conferencing speech enhancement technologies, existing algorithms heavily rely on direct spectral mapping and feature masking strategies. However, this processing logic is essentially a passive scaling and filtering of existing audio signals, leading to a critical flaw in complex scenarios: when high-frequency voiceless consonant details in speech are completely submerged by ambient high-frequency white noise or lost due to hardware limitations, existing methods cannot reconstruct these lost acoustic features from scratch, resulting in the recovered speech lacking consonant components and sounding blurry.
[0025] Meanwhile, simple amplification strategies are prone to misclassifying high-frequency random noise as valid speech, severely reducing speech intelligibility. This contradiction stems from the over-reliance of existing technologies on incomplete audio spectra and the indiscriminate scaling of masked frequency bands, necessitating a cross-modal collaborative method capable of proactively reconstructing and restoring lost speech details.
[0026] To address the aforementioned issues, this application proposes an AI-based audiovisual conferencing speech enhancement method. Its core lies in actively reconstructing lost high-frequency speech at the physical waveform level through the synergistic driving of facial visual kinetic energy and auditory zero-crossing rate features. Specifically, this application first extracts the spatial complexity and temporal activity of the facial vocalization-related regions, filters out interference to obtain a visual energy sequence reflecting pure vocalization actions, and simultaneously acquires a smoothed audio zero-crossing rate sequence.
[0027] During processing, the logical synchronization of the audiovisual bimodal sequences is calculated to accurately identify speech frames with unvoiced characteristics. Once a frame is identified as unvoiced, the gain, center frequency, and bandwidth of the harmonic oscillator are adaptively configured using the statistical amplitude of the visual image entropy and the distribution density of the audio zero-crossing rate as parameters to generate customized harmonic components. These components are then phase-aligned and superimposed with the original signal within the corresponding frequency range.
[0028] This method abandons the traditional passive spectral masking mode and utilizes a deep mapping between visual motion patterns and residual acoustic features to directly drive a harmonic oscillator to generate missing high-frequency components. This parametric reconstruction-based strategy avoids the mis-amplification of high-frequency random noise and accurately compensates for consonant details completely masked by noise. It solves the technical problem in existing technologies where the inability to accurately reconstruct and restore consonant details in the face of severe high-frequency noise interference or loss of signal details leads to low speech enhancement clarity.
[0029] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] To address the problems of existing technologies, embodiments of this application provide an artificial intelligence-based audio-visual conferencing voice enhancement method, apparatus, device, computer storage medium, and computer program product. The artificial intelligence-based audio-visual conferencing voice enhancement method provided in this application embodiment will be described first below.
[0031] Figure 1 This illustration shows a flowchart of an AI-based audio enhancement method for audiovisual conferencing according to an embodiment of this application. Figure 1 As shown, the method includes: S101. Acquire the facial image sequence of the target object and the speech time-domain signal including high-frequency white noise in the environment, collected by the conference terminal device.
[0032] The target audience refers to the speaker in the meeting setting. Meeting terminal equipment refers to terminal devices deployed in the meeting environment that possess both visual and audio capture capabilities. The visual capture device is used to capture real-time images of the target audience, while the audio capture device is used to simultaneously record the voice signals emitted by the target audience.
[0033] A facial image sequence refers to a set of image frames of the facial region of a target object that are continuously acquired at a fixed frame rate. Each frame records the pixel distribution of the target object's mouth, jaw, and other speech-related areas at the current moment and reflects the continuous visual changes in speech movements over time.
[0034] Speech time-domain signals refer to the raw waveform data output by audio acquisition devices, with time as the horizontal axis and signal amplitude as the vertical axis. In far-field conferencing or low-end audio acquisition hardware scenarios, this signal typically includes ambient high-frequency white noise. Ambient high-frequency white noise refers to randomly superimposed interference components with a wide frequency distribution and concentrated in the high-frequency band, continuously generated by equipment such as air conditioners and fans. Its frequency range highly overlaps with high-frequency voiceless consonant components in speech and may cause consonant details to be completely submerged or lost.
[0035] During implementation, facial image sequences are first acquired using visual acquisition devices deployed at the conference terminals. Specifically, the target object is continuously photographed at a preset fixed frame rate to capture facial images. Each frame records the complete pixel state of the facial region at that moment, thus forming a set of image frames arranged in chronological order, i.e., the facial image sequence. ,in Indicates the first Frame image and , This represents the total number of frames captured.
[0036] At the same time, the audio signal of the target subject is simultaneously recorded through the audio acquisition device of the conference terminal to obtain the speech time domain signal. ,in Indicates the first The amplitude value of each sampling point and , This represents the total number of sampling points. Due to high-frequency random interference from equipment such as air conditioners and fans in the sampling environment, the signal already contains high-frequency white noise components from the environment.
[0037] S102. By calculating the spatial complexity and temporal activity of pixels in the speech-related region of the facial image sequence, the image entropy matrix at each time step is obtained. By calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components, the visual energy sequence is obtained.
[0038] Optionally, the process of obtaining the image entropy matrix at each time step S102 by calculating the spatial complexity and temporal activity of pixels in the voice-related region of the facial image sequence can specifically include: S1021. Calculate the gray-level gradient values of adjacent pixels in the voice-related region of the facial image sequence to obtain the space complexity.
[0039] Gray-scale gradient value refers to the variation range of gray-scale values between adjacent pixels in the sound-related region. It can be a scalar result obtained by comprehensively calculating the gray-scale differences between any pixel and its adjacent pixels in the horizontal, vertical and diagonal directions, and its value reflects the intensity of local texture change at the location of the pixel.
[0040] Spatial complexity refers to the spatial distribution feature matrix constructed based on the grayscale gradient values of each pixel within the sound-related region. Each element in the matrix Coordinates of the corresponding sound-related area The local grayscale gradient magnitude reflects the overall complexity of pixel texture distribution caused by the sound-emitting action in the current frame image.
[0041] During implementation, the first step is to start with facial image sequences. The speech-related regions are extracted frame by frame. Specifically, for each frame, a method based on 68 facial feature point detection, such as cascaded shape regression or other facial alignment algorithms, is used to automatically locate the mouth and jaw regions of the target object. A rectangular sub-region is extracted as the speech-related region, using the coordinates of the corners of the mouth, upper lip, lower lip, and jaw as boundaries. If facial feature point detection fails, such as when the target object's face is occluded, the coordinates of the successfully detected region from the previous frame are used to replace the coordinates of the region in the current frame.
[0042] Next, for each pixel within this region, the grayscale difference between it and its immediate neighboring pixels in the horizontal, vertical, and diagonal directions is calculated, and the grayscale gradient value of that pixel is obtained by weighted summation. This allows us to arrange the grayscale gradient values of all pixels within the sound-related region of the entire frame according to their coordinates, thus obtaining the spatial complexity matrix corresponding to that frame. As shown below:
[0043] in and These represent the number of rows and columns of the vocalization-related region, respectively. Representing coordinates The grayscale gradient magnitude at that location.
[0044] S1022. Calculate the grayscale difference of pixels at corresponding coordinate positions in adjacent image frames in the facial image sequence to obtain the temporal activity.
[0045] Temporal activity refers to the temporal motion feature matrix obtained by comparing the grayscale changes of pixels at the same coordinate positions in adjacent image frames, denoted as . Each element in the matrix Indicates at time With time Corresponding coordinates The grayscale difference is the absolute difference in grayscale values of a pixel and reflects the degree of motion activity at that location over time. The grayscale difference is the difference in grayscale values of pixels at the same coordinate position in two adjacent frames of an image. It is a non-negative value obtained by subtracting the grayscale values of the two frames and taking the absolute value. Its magnitude directly indicates the motion amplitude of the pixel at that position between frames.
[0046] During implementation, the current frame is first retrieved. With the previous frame The grayscale value of the pixel in the sound-related region. For the first frame of the sequence, i.e. At that time, because there is no previous frame Replace with an all-zero matrix The corresponding grayscale value of the sound-related region, that is, let To ensure that the time activity matrix at the start of the sequence is not empty and Indicates the first The space complexity matrix corresponding to the frame median coordinate The grayscale gradient magnitude at that location. The frame is normally taken as the current frame. With the previous frame The grayscale value of the pixel in the sound-related region.
[0047] Next, for each pair of corresponding coordinate positions in the two frames For each pixel at a given coordinate, calculate the difference between the two grayscale values and take the absolute value to obtain the grayscale difference at that coordinate. Then, the grayscale differences of all coordinate positions within the sound-related area are arranged by coordinates to form the temporal activity matrix at the current moment. As shown below:
[0048] in and These represent the number of rows and columns of the vocalization-related region, respectively. Representing coordinates The difference in grayscale values at each location.
[0049] S1023. By weighting and superimposing the space complexity and temporal activity, the image entropy matrix at each time step is obtained.
[0050] Weighted superposition refers to a fusion operation in which corresponding elements of the spatial complexity matrix and the temporal activity matrix are scaled proportionally according to preset weight coefficients and then summed element by element. The two sets of preset weight coefficients satisfy the normalization constraint and are used to control the contribution ratio of spatial texture information and temporal motion information in the fusion result.
[0051] The image entropy matrix is a comprehensive feature matrix obtained by weighting and superimposing spatial complexity and temporal activity. Each element in the matrix Comprehensive reflection of coordinates The combined energy intensity of the vocalization action in two dimensions—spatial structural abrupt change and temporal motion amplitude—and its overall distribution describe the physiological kinetic energy state of the vocalization behavior at the current moment.
[0052] During implementation, the space complexity matrix obtained at the current time step is taken first. and time activity matrix Then, using preset weighting coefficients... and The two matrices are element-weighted and superimposed, where The contribution ratio of controlling spatial texture complexity, Control the contribution ratio of temporal motion activity, and satisfy... The image entropy matrix is obtained by calculating element by element. As shown below:
[0053] Where, coordinates element at superscript This indicates that the element belongs to the first... The image entropy matrix corresponding to the frame.
[0054] and The optimization objective was determined using the Pearson correlation coefficient between the visual energy sequence and the synchronized audio zero-crossing rate sequence on a bimodal audio-visual annotation dataset covering various speech scenarios, including normal speech rate, fast speech rate, and speech with head-shaking interference. Within the range, in step size Perform a grid search Select the value that maximizes the correlation coefficient. value.
[0055] Typical values for a standard meeting scenario with 30fps capture are: , In other words, the contribution of spatial texture complexity is slightly higher than that of temporal motion activity. When there are significant differences between the acquisition frame rate or lighting conditions and the calibration conditions, the above search process should be re-executed under the new configuration to update. and Finally, the facial image sequence By performing the above operations on each frame sequentially, the image entropy matrix sequence corresponding to each time step can be obtained. .
[0056] This embodiment enables the vocal kinetic energy state at each moment to be quantitatively represented in a unified matrix structure, which not only preserves the fine-grained features of spatial dimensions such as lip contraction, but also captures the temporal evolution information of vocalization actions between frames.
[0057] Optionally, the process in step S102 of calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components and obtain the visual energy sequence may specifically include: S1024. Calculate the numerical distribution gradient of each element in the image entropy matrix with its neighboring elements to obtain the gradient contrast of each element.
[0058] Gradient contrast refers to the numerical distribution gradient between a given element in the image entropy matrix and its immediate neighbors. It is a scalar measure of the intensity of local abrupt changes obtained by comprehensively measuring the numerical differences between the given element and its surrounding elements. And it reflects the image entropy matrix in coordinates The degree of drastic change in the local structure at that point.
[0059] The higher the value, the stronger the spatial contrast of the region after the sound energy is integrated, and the more likely it is to correspond to the feature edge caused by the actual sound production action; The lower the value, the more uniform the energy distribution in the non-sound-producing area or background.
[0060] During implementation, the image entropy matrix obtained at the current time is first retrieved. Then, for each element in the matrix... Calculate the numerical differences between the element and its four adjacent elements in the four directions (right, bottom, bottom right, and bottom left), and sum them using weighted summation to obtain the gradient contrast value at that position. .
[0061] in These are the weights for the horizontal and vertical directions. The weights are in the diagonal direction, satisfying... And the typical value is , That is, the weights in the horizontal and vertical directions are slightly higher than the weights in the diagonal directions, and this can be adjusted before deployment based on the degree of texture anisotropy in the sound-related region. Specifically, boundary pixels only perform calculations for existing adjacent directions, and the weights corresponding to missing directions are normalized to ensure that the sum of the weights for all directions involved in the calculation is still equal to 1. Then, all coordinate positions are... Arranged by coordinates, they form the gradient contrast matrix at the current moment. .
[0062] S1025. Calculate the temporal difference values of the corresponding coordinate positions in the image entropy matrix at adjacent time points to obtain the inter-frame evolution frequency.
[0063] Inter-frame evolution frequency (IF) refers to the temporal dynamic feature matrix obtained by comparing the changes in element values of the image entropy matrix at the same coordinate positions at adjacent time points. Each element in the matrix Representing coordinates The temporal difference between adjacent time points of the image entropy matrix elements reflects the evolution rate of the sound energy at that location in the time dimension.
[0064] The temporal difference value refers to the absolute value of the difference between the values of elements at the same coordinate position in the image entropy matrix at adjacent time points. Its magnitude directly indicates the degree of change in the acoustic kinetic energy at that position over time and can be used as a measure of the temporal difference between the two time points. and The non-negative value obtained by taking the absolute value after subtraction.
[0065] During implementation, the image entropy matrix at the current moment is first obtained. Image entropy matrix from the previous time step Next, for each pair of corresponding coordinate positions in the two matrices... For each element at a given coordinate, calculate the difference between the two values and take the absolute value to obtain the inter-frame evolution frequency value at that coordinate position. Finally, all coordinate positions will be... Arranged by coordinates, they form the inter-frame evolution frequency matrix at the current moment. .
[0066] S1026. Based on the inter-frame evolution frequency, according to the preset mapping relationship between the inter-frame evolution frequency and the weight coefficient, the weight coefficient of each pixel region is obtained. The gradient contrast is weighted and aggregated using the weight coefficient to obtain the visual energy sequence.
[0067] The weighting coefficient is a scalar parameter determined by a preset mapping relationship based on the magnitude of the inter-frame evolution frequency, used to control the contribution of gradient contrast in weighted aggregation. Its value is positively correlated with the inter-frame evolution frequency, meaning that the more drastic the change in sound kinetic energy, the higher the weight is given to the region.
[0068] Visual energy sequence refers to the scalar sequence obtained by performing weighted aggregation on each moment of a facial image sequence and then arranging them in chronological order. Each element It comprehensively reflects the overall energy intensity of the vocalization action at that moment in two dimensions: spatial structural contrast and temporal motion activity.
[0069] The mapping relationship between the preset inter-frame evolution frequency and the weighting coefficients is shown in Table 1 below: Table 1: Mapping Relationship between Inter-Frame Evolution Frequency and Weighting Coefficient
[0070] As shown in Table 1, Table 1 describes the three-stage hierarchical mapping rule between inter-frame evolution frequency and weighting coefficients. and The preset inter-frame evolution frequency segmentation threshold and When the inter-frame evolution frequency at a certain coordinate location falls into a low-activity segment, it indicates that the region is close to quiescent and is assigned a lower weight. To suppress background interference; when it falls into the mid-active segment, it indicates that there is a slight vocalization in that area, and it is given medium weight. When a region falls into the high-activity segment, it indicates that the corresponding vocalization is intense and is therefore assigned the highest weight. To highlight the contribution of authentic vocal characteristics.
[0071] The three weighting coefficients and segmentation thresholds were determined as follows: First, on a labeled facial image dataset covering various speech scenarios including normal speech rate, fast speech rate, and speech with head movement interference, the optimization objective was to achieve the accuracy of speech frame recognition. 、 Perform a grid search within the range with a step size of 1, and select the grid that maximizes the recognition accuracy. 、 combination.
[0072] Then, the above threshold is fixed, with the goal of minimizing the false recognition rate of background interference frames. Under constraints, the search employs three levels of weighting coefficients, with a typical configuration as follows: 、 、 , 、 This corresponds to a standard conference scenario with a 30fps acquisition condition. If the acquisition frame rate is significantly different, such as below 15fps or above 60fps, or if the dynamic range of illumination exceeds the calibration data range, the above grid search process should be re-executed to determine the new parameter configuration.
[0073] During implementation, the first step is to use the inter-frame evolution frequency matrix. Elements The numerical values are used to look up the segmented mapping relationship defined in Table 1 to determine the weight coefficient corresponding to each coordinate position. Thus, the weight coefficient matrix is formed. Next, the weight coefficient matrix... With gradient contrast matrix The visual energy value at the current moment is obtained by multiplying each element at the corresponding coordinate position and summing the results. Finally, the facial image sequence Perform the above operations one by one at each moment in the process, and then perform the operations on all moments. Arranged chronologically, a visual energy sequence is obtained. .
[0074] This embodiment enables regions with strong vocalization to receive a higher contribution weight during the aggregation process, while low-activity interference components such as head shaking and ambient light and shadow are effectively suppressed due to being assigned lower weights. This allows the output visual energy sequence to purely reflect the actual vocal kinetic energy change pattern of the target object on the time axis.
[0075] S103. The speech time-domain signal is framed using window function weighting and the number of times the signal waveform of each frame crosses the zero level is calculated to obtain the zero-crossing rate distribution sequence. The zero-crossing rate distribution sequence is then smoothed and weighted using the energy comparison parameter determined based on the speech time-domain signal to obtain the zero-crossing rate sequence.
[0076] Optionally, step S103 involves using a window function to weight the speech time-domain signal into frames and calculating the number of times the signal waveform in each frame crosses the zero level to obtain a zero-crossing rate distribution sequence. The process of smoothing and weighting the zero-crossing rate distribution sequence using the energy contrast parameter determined based on the speech time-domain signal to obtain the zero-crossing rate sequence can specifically include: Figure 2 A flowchart illustrating a method for generating zero-crossing rate sequences according to an embodiment of this application is shown. Figure 2 As shown, the method includes: S1031. The speech time-domain signal is time-sequentially segmented according to the preset sampling length, and the amplitude of each segmented signal segment is mapped using a preset weighting window to obtain multiple analysis frames.
[0077] The preset sampling length refers to the fixed frame length parameter used when performing time-series segmentation on the speech time-domain signal. Furthermore, its value determines the number of sampling points covered by each signal segment on the time axis. The preset weighting window refers to the window function used when performing amplitude mapping on each segmented signal. Its function is to scale each sampling point within a signal segment according to a specific amplitude weight distribution in order to suppress amplitude abrupt changes at the beginning and end of the signal segment.
[0078] An analysis frame refers to a weighted signal segment obtained after being segmented by a preset sampling length and mapped by a preset weighted window amplitude. Furthermore, all analysis frames are arranged in chronological order to form an analysis frame set. ,in The frame number and , This represents the total number of frames.
[0079] The preset sampling lengths are shown in Table 2 below: Table 2: Preset Sampling Length Comparison Table
[0080] As shown in Table 2, the preset sampling length is divided into three levels based on the hardware conditions and signal-to-noise ratio characteristics of the audio acquisition scenario. Each parameter value is a fixed value pre-calibrated based on the characteristics of conference audio acquisition, and can be configured and selected according to the actual hardware conditions before deployment.
[0081] The preset weighted windows are shown in Table 3 below: Table 3: Preset Weighted Window Comparison Table
[0082] As shown in Table 3, the three types of preset weighted windows each have their own emphasis on amplitude distribution characteristics. and Smoothing of frame amplitude can effectively suppress spectral leakage and has better feature preservation capabilities in far-field conferencing or high-frequency white noise interference scenarios.
[0083] During implementation, the first step is to select the corresponding frame length from the three preset sampling lengths based on the hardware acquisition conditions of the current deployment scenario, referring to Table 2. The value is used to select the corresponding preset weighted window from Table 3. Then follow the selected... Value and preset frame shift and speech time domain signal Perform continuous time-series segmentation, every... A segment of length is extracted from each sampling point. The signal segment, obtained Each signal segment, total number of frames .
[0084] Frame shift It determines the degree of overlap between adjacent analysis frames: when When it is a non-overlapping frame; when When adjacent frames overlap by 50%, signal continuity at the edges of adjacent frames is ensured, resulting in better temporal resolution in high-frequency unclear sound feature detection scenarios. For standard conference scenarios, this is the recommended approach. Then, for each sampling point within each signal segment, a preset weighting window is applied. The defined amplitude weight distribution is multiplied and scaled point by point to complete the amplitude mapping and obtain the corresponding analysis frame. Finally, all analysis frames are arranged in chronological order to obtain the analysis frame set. .
[0085] S1032. Accumulate and calculate the number of polarity reversals of adjacent signal values within each analysis frame to obtain the zero-crossing rate distribution sequence.
[0086] The polarity reversal count refers to the total number of times the signal value of adjacent sampling points changes from positive to negative or vice versa within an analysis frame. It is an integer statistic that measures how frequently the signal waveform in that frame crosses the zero level. The zero-crossing rate distribution sequence refers to the sequence of analysis frames... The integer statistical sequence obtained by counting the number of polarity flips in each frame and arranging them according to the frame number. ,in Indicates the first The frame analysis shows the total number of polarity flips within the frame and reflects the high-frequency oscillation distribution pattern of the speech time-domain signal on the time axis.
[0087] In far-field conferencing or high-frequency white noise interference scenarios, the corresponding high-frequency uncensored frames Typically, the noise level is significantly higher than in voiced or silent frames, while ambient high-frequency white noise frames can also produce higher levels. Therefore, the zero-crossing rate distribution sequence needs to be weighted and corrected by subsequent energy comparison parameters in order to effectively distinguish between the two.
[0088] During implementation, the current analysis frame is first retrieved. The polarity relationship of each adjacent sampling point is detected one by one, and the following rules are used to determine polarity reversal: if the product of the values of two adjacent points is strictly less than zero, i.e., one is positive and the other is negative, then a polarity reversal is determined to have occurred; if one of the two adjacent points is exactly zero, then the sign of the zero point and the next non-zero sampling point is used as the basis for judging the polarity of that position, that is, the zero value is regarded as the instant when the signal crosses the zero level, and if the signs of the non-zero sampling points before and after are opposite, it is counted as a polarity reversal; if both adjacent points are zero, then a polarity reversal is not counted.
[0089] Summing up all polarity flip counts gives the total number of polarity flips for that frame. Next, the analysis frame set... Repeat the above operation for each frame, and then process all frames. Arranged by frame number, the zero-crossing rate distribution sequence is obtained. .
[0090] S1033. Calculate the ratio of the average signal strength of each analysis frame to the preset environmental noise benchmark value to obtain the corresponding energy comparison parameter.
[0091] The signal strength mean refers to the arithmetic mean of the absolute values of the amplitudes of all sampling points within a given analysis frame. This reflects the overall energy level of the signal within the time window and is used to analyze the frame. The absolute values of the amplitudes of all sampling points are summed and then divided by the frame length. The resulting nonnegative scalar.
[0092] The preset environmental noise baseline value refers to the background energy reference level pre-calibrated based on the type of environmental noise in the current meeting data collection scenario. Its value represents the typical signal strength mean of a pure noise frame in this scenario and is used as an energy reference benchmark to distinguish between speech frames and pure noise frames.
[0093] Energy comparison parameters This refers to the average signal strength of each analysis frame. Compared with the preset environmental noise benchmark value The ratio is , A value greater than 1 indicates that the energy of the current frame signal is significantly higher than the level of environmental noise, and the probability of the presence of real speech components is relatively high; A value close to or less than 1 indicates that the energy of the current frame is comparable to the noise background, and it belongs to the noise-dominated frame.
[0094] The preset environmental noise baseline values are shown in Table 4 below: Table 4: Comparison Table of Environmental Noise Reference Values
[0095] As shown in Table 4, the preset environmental noise baseline values are divided into three levels based on the intensity level of environmental noise in the conference data collection scenario. Each value is a fixed parameter pre-calibrated based on the average signal strength of pure noise frames in a typical conference environment. It can be configured and selected according to the actual acquisition environment before deployment.
[0096] During implementation, firstly, based on the noise characteristics of the current meeting scene, the corresponding preset environmental noise benchmark value is selected by referring to Table 4. Next, the analysis frame set... Each frame The sum of the absolute values of the amplitudes of all sampled points within the frame is divided by the frame length. To obtain the average signal strength of this frame Then Divide by the preset environmental noise benchmark value The energy contrast parameters corresponding to this frame are obtained. Finally, for all... Perform the above operations frame by frame, and convert each frame into its own... Arranged by frame number, the energy comparison parameter sequence is obtained. .
[0097] S1034. The zero-crossing rate distribution sequence is obtained by using the energy comparison parameter to perform amplitude weighting.
[0098] Amplitude weighting refers to using energy comparison parameters As a scaling factor for the zero-crossing rate distribution sequence Polarity reversal statistics of the corresponding frame A product scaling operation is performed. The purpose is to proportionally reduce the zero-crossing rate of frames with energy levels close to the noise baseline while preserving the high-frequency oscillation characteristics of the real speech frames, thereby achieving dynamic suppression of spurious high zero-crossing rates caused by ambient high-frequency white noise.
[0099] Zero-crossing rate sequence refers to the weighted statistical sequence obtained by weighting the energy contrast parameter amplitude of each frame in the zero-crossing rate distribution sequence and then arranging them according to the frame number. Each element This comprehensively reflects the saliency of the high-frequency oscillations in the actual speech of that frame.
[0100] During implementation, the already obtained zero-crossing rate distribution sequence is first taken. Comparison parameter sequence with energy Next, for each frame number in the sequence... The polarity inversion statistics of this frame Comparison parameters with corresponding energy Multiply to obtain the weighted zero-crossing rate value of the frame. Finally, all frames... Arranged by frame number, the zero-crossing rate sequence is obtained. .
[0101] This embodiment effectively suppresses the spectral leakage effect at the edge of the signal segment, so that the real speech frame is significantly prominent in the zero-crossing rate sequence due to its high energy contrast parameter, while the frame dominated by high-frequency white noise in the environment is effectively suppressed because the energy contrast parameter is close to or lower than 1. Thus, the output zero-crossing rate sequence can accurately represent the distribution law of acoustic high-frequency features.
[0102] S104. Calculate the logical synchronization between the visual energy sequence and the zero-crossing rate sequence to obtain the classification probability vector at each time step. When the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time step, configure the gain parameters, center frequency, and bandwidth of the harmonic oscillator based on the statistical amplitude of the image entropy matrix at the current time step and the distribution density of the zero-crossing rate sequence to obtain the harmonic components.
[0103] Figure 3 A schematic flowchart of a method for generating harmonic components according to an embodiment of this application is shown. Figure 3 As shown, the method includes: Optionally, the process of calculating the logical synchronization between the visual energy sequence and the zero-crossing rate sequence in step S104 to obtain the classification probability vector at each time step can specifically include: S1041. Calculate the matching degree between the changing trends of the visual energy sequence and the zero-crossing rate sequence at each time step to obtain the logical synchronization.
[0104] The trend of change refers to the direction and magnitude of the numerical changes of each element in a sequence over time. It is the increase or decrease obtained by performing a difference operation on the elements at adjacent time points in the sequence. Its sign indicates an upward or downward trend, and its absolute value indicates the degree of drastic change.
[0105] Matching degree refers to the similarity between the changing trends of the visual energy sequence and the zero-crossing rate sequence at corresponding moments. It is a similarity scalar obtained by calculating the product of the changing trends of the two sequences or the correlation coefficient. The higher the value, the more synchronized the visual kinetic energy enhancement and the audio high-frequency oscillation enhancement are on the time axis and the more likely they are to correspond to real vocal behavior.
[0106] Logical synchronicity refers to the synchronicity measure obtained by calculating the matching degree of the changing trends of the visual energy sequence and the zero-crossing rate sequence at each moment on the time axis and then arranging them in chronological order. ,in Indicates the first The matching strength of the changing trends of the two modal sequences at each time point. This represents the total number of valid time points after the difference operation. This represents the total number of moments in the aligned sequence.
[0107] During implementation, the visual energy sequence first needs to be analyzed. With zero-crossing rate sequence Perform timeline alignment. Since the temporal resolution of the visual energy sequence is determined by the video acquisition frame rate and... The temporal resolution of the zero-crossing rate sequence is determined by the audio analysis frame length and frame shift, and is total. Frames, the temporal granularity of the two sequences is usually different.
[0108] Alignment is performed using the timestamps of the audio analysis frames as the base timeline, i.e., a total of At each moment, the visual energy sequence is obtained through linear interpolation. Resampling to the time scale corresponding to the audio frame yields the aligned visual energy sequence. and aligned zero-crossing rate sequence ,in, For the first Visual energy interpolation results for each audio analysis frame at the corresponding time point , .
[0109] Next, the difference between elements at adjacent time points is calculated for each of the two sequences to obtain the visual energy change trend sequence. and zero-crossing rate change trend sequence ,in , Then for each moment and ,calculate and The weighted similarity of elements at corresponding positions yields the logical synchronicity value at that moment. ,in To prevent division by zero of extremely small positive numbers, finally, all time intervals will be... Arranged in chronological order, a logically synchronous sequence is obtained. .
[0110] S1042. By mapping the logical synchronization to a preset classification probability interval, a classification probability vector is obtained to determine whether the speech time-domain signal belongs to unvoiced, voiced, or environmental interference at each moment.
[0111] The preset classification probability interval refers to several probability assignment intervals pre-divided according to the range of logical synchronicity values. Each interval corresponds to a high probability assignment of a speech type and is used to convert continuous synchronicity measurement values into discrete classification results.
[0112] The classification probability vector is a three-dimensional probability vector obtained by mapping the logical synchronization value at each time step to a preset classification probability interval. ,in This represents the probability that the sound at that moment is a voiceless tone. This indicates the probability that the sound is voiced. This represents the probability of environmental interference, and the sum of the three factors is 1. The preset classification probability intervals are shown in Table 5 below: Table 5: Preset Classification Probability Interval Comparison Table
[0113] As shown in Table 5, the table illustrates the correspondence between preset classification probability intervals and speech types. The preset classification probability intervals are based on logical synchronicity values. The high and low are divided into three intervals, among which and The preset synchronization threshold and Each interval corresponds to a different probability distribution feature of speech type.
[0114] and The optimization objective was determined as follows: On a bimodal audiovisual annotation dataset including labels for unvoiced frames, voiced frames, and environmental interference frames, the sum of the F1 classification scores for the three types of frames was used as the optimization target. , A grid search is performed within the range with a step size of 0.05 to select the threshold combination that maximizes the overall F1 score. In a typical conference scenario, the default value is... , When changes in the visual acquisition frame rate or audio sampling rate cause a change in the alignment accuracy between the two modalities, the above search process should be re-executed under the new configuration to update the threshold.
[0115] when Falling into the high synchronization range indicates that the visual-sound action is highly consistent with the high-frequency oscillation of the audio, giving a higher probability of unvoiced sound; when Falling into the intermediate synchronicity range indicates the presence of some audiovisual coupling, but the high-frequency characteristics are not significant, thus assigning a higher probability of voiced sound; when Falling into the low synchronicity range indicates a mismatch between audiovisual features or the absence of effective speech, thus assigning a higher probability of environmental interference.
[0116] During implementation, the logical synchronization sequence is first taken. Synchronization value at each moment Then according to Based on the numerical value, refer to the classification probability interval division rules defined in Table 5 to determine the three-class probability distribution corresponding to that moment. Then, organize the query results into a three-dimensional classification probability vector. Each probability component is based on The intervals are assigned values according to the high, medium, low, and very low levels shown in Table 5. Specific values can be determined through linear interpolation or table lookup. Finally, the above mapping operation is performed for each time step to obtain the classification probability vector sequence. 。
[0117] In scenarios with severe high-frequency white noise interference, this embodiment can effectively distinguish between real unvoiced frames and high-frequency random noise frames by using the coupling relationship of the audiovisual dual-modal change trends. This avoids the shortcomings of existing technologies that cannot distinguish between the two simply by relying on the audio spectrum, and ensures that parameterized harmonic reconstruction is performed only on real speech frames.
[0118] Optionally, in step S104, when the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time, the process of configuring the gain parameters, center frequency, and bandwidth of the harmonic oscillator based on the statistical amplitude of the image entropy matrix and the distribution density of the zero-crossing rate sequence at the current time to obtain the harmonic components can specifically include: S1043. By summing the values of all matrix elements in the image entropy matrix at the current moment, the statistical amplitude is obtained. Based on the statistical amplitude, the gain parameter is obtained according to the preset mapping relationship between the statistical amplitude and the gain parameter.
[0119] The statistical magnitude is a scalar obtained by summing the values of all elements in the image entropy matrix at a certain moment. It reflects the total amount of sound-generating kinetic energy in the spatial and temporal dimensions of the sound-generating region at that moment. The higher the value, the more intense the sound-generating action, and the stronger the excitation intensity of the corresponding harmonic reconstruction should be.
[0120] The preset mapping relationship between statistical amplitude and gain parameter refers to the mapping of statistical amplitude... This is transformed into a mathematical function relationship for the gain control parameters of the harmonic oscillator, and this mapping relationship ensures that unison frames with higher statistical amplitudes can drive the harmonic oscillator to produce an output signal of corresponding strength. The gain parameter refers to the scalar control quantity used to control the output signal strength of the harmonic oscillator. Furthermore, its value determines the amplitude of the generated harmonic components.
[0121] During implementation, the first step is to determine the current time. Classification probability vector Does it meet the criteria for determining unvoiced sound? Is it the maximum value among the three probabilities and exceeds the preset unvoiced probability threshold? If it is determined to be an unvoiced frame, then the image entropy matrix corresponding to that moment is taken. For all elements in the matrix By summing the results, we can obtain the statistical amplitude. Next, a pre-defined mapping relationship between statistical amplitude and gain parameters is established:
[0122] in, For the first Gain parameters at time 10:00 For the first The statistical magnitude of the image entropy matrix at time step. The preset statistical amplitude normalization benchmark value, This is the preset gain scaling factor.
[0123] and The following method was used to determine the intensity of the audiovisual bimodal annotated corpus, covering three levels of vocal intensity: weak (whisper level), normal vocal, and strong (speech level). First, the median of the statistical amplitude of the image entropy matrix corresponding to all unvoiced frames was calculated, and this median was set as [value missing]. The typical value, i.e., in a standard meeting scenario, is approximately Furthermore, its dimensions are consistent with those of the elements in the image entropy matrix.
[0124] Subsequently, the improvement in the PESQ objective score of the enhanced speech after harmonic overlay compared to the original noisy speech was used as the optimization target. The search is performed within the range with a step size of 0.1, and the value that maximizes the PESQ improvement is found. The value used as the final configuration, i.e., the typical value in a standard meeting scenario, is... When the image acquisition resolution or frame rate of the deployment scenario differs significantly from the calibration conditions, the calibration corpus should be re-acquired under the newly configured acquisition conditions, and the above process should be executed to update the data. and Subsequently, Substituting into the mapping function above, the gain parameter is calculated. .
[0125] S1044. Calculate the statistical proportion of values in the zero-crossing rate sequence at the current time that are within a preset threshold range, obtain the distribution density, and based on the distribution density, obtain the center frequency and bandwidth according to the preset correspondence between the distribution density and frequency characteristics.
[0126] Distribution density refers to the proportion of elements in the zero-crossing rate sequence whose values fall within a preset threshold range in the current time and its neighborhood time window, denoted as . This reflects the concentration of high-frequency oscillation energy in the time domain at that moment. The higher the value, the more significant the high-frequency unison characteristics, and the narrower the high-frequency bandwidth that the corresponding harmonic reconstruction should cover.
[0127] The preset threshold range refers to a pre-defined discrimination range based on the typical zero-crossing rate range of unvoiced sounds, used to filter out zero-crossing rate elements possessing unvoiced sound characteristics. The center frequency refers to the center position of the frequency at which the harmonic oscillator generates harmonic components. It also determines the dominant frequency component of the reconstructed harmonics. Bandwidth refers to the range of harmonic components in the frequency domain. Furthermore, together with the center frequency, it determines the frequency range covered by the harmonic reconstruction.
[0128] The preset distribution density and frequency characteristics correspondence is shown in Table 6 below: Table 6: Correspondence between Distribution Density and Frequency Characteristics
[0129] As shown in Table 6, the correspondence between the preset distribution density and frequency characteristics is based on the distribution density. The high and low are divided into three intervals, among which and Three sets of frequency characteristic parameters are used to define the preset distribution density segmentation threshold. , , These represent the center frequencies of the high-frequency band, mid-high-frequency band, and mid-frequency band, respectively. , , These represent the corresponding narrow bandwidth, medium bandwidth, and wide bandwidth, respectively.
[0130] Each parameter is determined in the following way: , On a speech dataset labeled with the zero-crossing rate distribution of unvoiced frames, a grid search is performed with the objective of achieving accurate classification of voiced and unvoiced sounds. A typical meeting scenario is used to determine the appropriate classification method. , ; , , Based on the statistical distribution of common consonant frequencies in the target language, typical values for Mandarin Chinese are as follows: , , ; , , Take respectively , , The frequency parameters mentioned above are all calibrated based on a 16 kHz sampling rate. When the actual sampling rate is different, they should be scaled proportionally to maintain consistent frequency characteristics.
[0131] when Falling into the high-density range indicates that the zero-crossing rate is highly concentrated in the high-frequency unison, so the center frequency of the high-frequency band should be selected. With narrow bandwidth Accurately reconstruct the consonant frequency band; when When falling within the medium density range, select the center frequency in the mid-to-high frequency band. With medium bandwidth It takes into account the transition characteristics of voiced and voiceless sounds; when When falling into the low-density range, select the center frequency of the mid-frequency band. With wide bandwidth It is adapted to the low-frequency characteristics dominated by voiced sounds.
[0132] During implementation, the first step is to target the moment currently identified as a clear frame. Taking that moment as the center, a time window of a certain length is selected, and the zero-crossing rate sequence within that window is extracted. The subsequence is then analyzed. Next, the number of elements in the subsequence whose values fall within a preset threshold range is counted, and this number is divided by the total number of elements in the subsequence to obtain the distribution density. Then according to Based on the numerical value, refer to the segmentation mapping rules defined in Table 6 to determine the corresponding center frequency. With bandwidth .
[0133] S1045. Harmonic components are generated by controlling the output strength and oscillation frequency range of the harmonic oscillator through the gain parameter, center frequency, and bandwidth.
[0134] A harmonic oscillator is a signal generating device that generates harmonic signals within a specified frequency range in real time based on input gain parameters, center frequency, and bandwidth parameters. Its output signal can be a composite periodic signal obtained by superimposing sine waves or synthesizing parameterized spectra.
[0135] Output strength refers to the amplitude of the signal generated by the harmonic oscillator and is determined by the gain parameter. Direct control. The oscillation frequency range refers to the frequency interval covered by the signal generated by the harmonic oscillator, and is determined by the center frequency. With bandwidth To be determined jointly, specifically as follows: Harmonic components refer to the parameterized reconstructed signal output by a harmonic oscillator. Furthermore, its time-domain waveform characteristics are determined by the combination of the above three parameters.
[0136] During implementation, the current unvoiced frame time is first taken. Determined gain parameters Center frequency and bandwidth Next, these three parameters are input into the harmonic oscillator, and its output strength is configured as follows: Configure its oscillation frequency range as The harmonic oscillator is activated to generate a harmonic signal with a duration consistent with the current frame duration. The harmonic signal has its energy concentrated in the frequency domain at... Nearby, the frequency extension range is affected Constraints, amplitude is determined by control.
[0137] This embodiment breaks through the limitations of existing technologies that rely on passive filtering of incomplete spectra. It realizes a cross-modal parameterized active reconstruction mechanism that controls harmonic amplitude based on visual vocal kinetic energy and harmonic frequency based on auditory high-frequency distribution patterns. This enables the generated harmonic components to accurately match the real vocal logic in terms of amplitude intensity and frequency characteristics. As a result, it effectively compensates for and reconstructs consonant details that are completely covered or lost by high-frequency white noise at the physical waveform level, significantly improving speech enhancement clarity.
[0138] S105. Within the frequency range determined by the center frequency and bandwidth, the harmonic components are phase-aligned and superimposed with the speech time-domain signal to enhance consonant details.
[0139] Optionally, step S105, which involves phase alignment and signal superposition of the harmonic components with the speech time-domain signal within a frequency range determined by the center frequency and bandwidth to enhance consonant details, may specifically include: S1051. Based on the center frequency and bandwidth, extract the frequency of the speech time-domain signal to obtain the reference signal in the target frequency band.
[0140] Frequency extraction refers to performing bandpass filtering on a speech time-domain signal to retain only frequency components falling within the frequency range defined by the center frequency and bandwidth, while filtering out frequency components outside the range. The target frequency band refers to the frequency corresponding to the center frequency of the current unvoiced frame. With bandwidth The determined frequency range Furthermore, this frequency band covers the main frequency range of the consonant details to be reconstructed. The reference signal refers to the time-domain signal of the original speech. The band-limited signal obtained after frequency extraction within the target frequency band. Furthermore, its time-domain waveform only includes the residual speech and noise mixture components within that frequency band.
[0141] During implementation, the current unvoiced frame time is first taken. Determined center frequency With bandwidth Calculate the lower limit frequency of the target frequency band. With upper limit frequency Next, a bandpass filter is constructed, and the passband of this filter is set to... The stopband covers the entire frequency range outside the passband. Then the original speech time-domain signal... The input to this bandpass filter allows only frequency components within the target frequency band to pass through, filtering out frequency components in other frequency bands, and outputting the reference signal. .
[0142] S1052. Calculate the phase deviation between the reference signal and the harmonic component, and perform offset correction on the initial phase of the harmonic component based on the phase deviation to obtain the aligned harmonic component.
[0143] Phase deviation value refers to the reference signal With harmonic components Phase difference in time-domain waveform It is the phase angle value obtained by subtracting the instantaneous phases of two signals after extracting them through cross-correlation analysis or Hilbert transform. Its value reflects the degree to which the harmonic components lead or lag behind the original speech phase reference.
[0144] Offset correction refers to the correction based on the phase deviation value. For harmonic components The initial phase is rotated in the opposite direction to adjust the harmonic components so that the corrected harmonic components are aligned with the reference signal in the phase dimension. The aligned harmonic components refer to the phase-adjusted harmonic signals obtained after offset correction. Its frequency response and amplitude intensity remain unchanged, and only the phase angle is shifted to make it consistent with the reference signal. Phase synchronization is achieved in the time domain to avoid signal cancellation due to phase mismatch after superposition.
[0145] During implementation, a reference signal is first obtained. With harmonic components The Hilbert transform is performed on both signals to extract their respective instantaneous phase sequences. Then, the point-by-point difference between the two phase sequences in the time domain is calculated to obtain the average phase difference value, thus yielding the phase deviation value. Then, the harmonic components... Apply a reverse offset to the phase angle of each sampling point. Phase correction is completed to obtain aligned harmonic components. .
[0146] S1053. Within a frequency range with a defined center frequency and bandwidth, the aligned harmonic components are numerically superimposed on the speech time-domain signal to enhance consonant details.
[0147] Numerical superposition refers to aligning harmonic components. Compared with the original speech time domain signal In the time domain, numerical values are added point by point to the corresponding sampling points. Enhancing consonant details refers to the energy compensation and structural reconstruction of high-frequency voiceless consonant components, such as those completely submerged by high-frequency white noise in the original speech time domain signal or lost due to hardware limitations, by superimposing and aligning harmonic components.
[0148] During implementation, the current unvoiced frame time is first taken. Corresponding aligned harmonic components Compared with the original speech time domain signal The signal segment corresponding to that moment. Next, confirm that the numerical superposition occurs only at the center frequency. and bandwidth The determined frequency range To achieve this frequency band restriction, internal execution can first be performed on... and Perform a Fast Fourier Transform to convert to the frequency domain, add the frequency components of both domains point by point within the target frequency band, and preserve the frequency components outside the frequency band. The frequency domain components remain unchanged, and then an inverse fast Fourier transform is performed to return to the time domain, yielding the enhanced speech signal. Finally, the enhanced signal segment Replace the original speech time domain signal The corresponding signal segment at that moment is used to enhance the consonant details of that unvoiced frame.
[0149] This embodiment avoids energy loss caused by the mutual cancellation of peaks and troughs during superposition by using phase alignment. At the same time, by limiting the superposition by frequency band, it ensures that harmonic compensation only applies to the high-frequency consonant region that is masked by noise and does not affect other frequency bands. Thus, it achieves accurate parametric reconstruction and energy compensation of lost consonant details at the physical waveform level, which significantly improves the clarity and intelligibility of speech in audiovisual conferencing scenarios.
[0150] Figure 4 This application provides a schematic diagram of a specific implementation of an artificial intelligence-based audiovisual conferencing voice enhancement system, with reference to... Figure 4 The system may include: The acquisition module 410 is used to acquire the facial image sequence of the target object and the speech time-domain signal including environmental high-frequency white noise collected by the conference terminal device; The calculation module 420 is used to obtain the image entropy matrix at each time step by calculating the spatial complexity and temporal activity of the pixels in the voice-related region in the facial image sequence, and to obtain the visual energy sequence by calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components. The generation module 430 is used to frame the speech time-domain signal using window function weighting and calculate the number of times the signal waveform of each frame crosses the zero level to obtain the zero-crossing rate distribution sequence. The zero-crossing rate distribution sequence is then smoothed and weighted using the energy comparison parameter determined based on the speech time-domain signal to obtain the zero-crossing rate sequence. The generation module 430 is also used to calculate the logical synchronization between the visual energy sequence and the zero-crossing rate sequence, obtain the classification probability vector at each time step, and when the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time step, the gain parameters, center frequency and bandwidth of the harmonic oscillator are configured according to the statistical amplitude of the image entropy matrix at the current time step and the distribution density of the zero-crossing rate sequence, respectively, to obtain the harmonic components. Processing module 440 is used to perform phase alignment and signal superposition of harmonic components and speech time-domain signals within a frequency range determined by the center frequency and bandwidth, in order to enhance consonant details.
[0151] The AI-based audiovisual conferencing voice enhancement system of this application embodiment is used to implement the aforementioned AI-based audiovisual conferencing voice enhancement method. Therefore, the specific implementation of the AI-based audiovisual conferencing voice enhancement system can be found in the embodiment section of the AI-based audiovisual conferencing voice enhancement method above. The specific implementation can be referred to the description of the corresponding embodiments, and will not be repeated here.
[0152] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application is shown.
[0153] The electronic device may include a processor 510 and a memory 520 storing computer program instructions.
[0154] Specifically, the processor 510 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0155] Memory 520 may include mass storage for data or instructions. For example, and not limitingly, memory 520 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 520 may include removable or non-removable (or fixed) media. Where appropriate, memory 520 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 520 is non-volatile solid-state memory.
[0156] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of this disclosure.
[0157] The processor 510 reads and executes computer program instructions stored in the memory 520 to implement any of the artificial intelligence-based audiovisual conferencing voice enhancement methods in the above embodiments.
[0158] In one example, the electronic device may also include a communication interface 530 and a bus 540. Wherein, such as Figure 5 As shown, the processor 510, memory 520, and communication interface 530 are connected through bus 540 and complete communication with each other.
[0159] The communication interface 530 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0160] Bus 540 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 540 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0161] The electronic device can execute the AI-based audiovisual conferencing voice enhancement method in the embodiments of this application, thereby realizing the AI-based audiovisual conferencing voice enhancement method described in conjunction with the accompanying drawings.
[0162] Furthermore, in conjunction with the AI-based audiovisual conferencing voice enhancement method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the AI-based audiovisual conferencing voice enhancement methods in the above embodiments.
[0163] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0164] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0165] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0166] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0167] The above provides a detailed description of an AI-based audio enhancement method and system for audiovisual conferencing. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A method for enhancing audio-visual conferencing speech based on artificial intelligence, characterized in that, include: Acquire facial image sequences of the target object and speech time-domain signals including high-frequency white noise from the environment, collected by the conference terminal device; By calculating the spatial complexity and temporal activity of pixels in the vocalization-related region of the facial image sequence, the image entropy matrix at each moment is obtained. By calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components, the visual energy sequence is obtained. The speech time-domain signal is framed using a window function weighting method, and the number of times the signal waveform in each frame crosses the zero level is calculated to obtain a zero-crossing rate distribution sequence. The zero-crossing rate distribution sequence is then smoothed and weighted using an energy comparison parameter determined based on the speech time-domain signal to obtain a zero-crossing rate sequence. The logical synchronization between the visual energy sequence and the zero-crossing rate sequence is calculated to obtain the classification probability vector at each time step. When the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time step, the gain parameters, center frequency, and bandwidth of the harmonic oscillator are configured based on the statistical amplitude of the image entropy matrix at the current time step and the distribution density of the zero-crossing rate sequence to obtain the harmonic components. Within a frequency range determined by the center frequency and bandwidth, the harmonic components are phase-aligned and superimposed with the speech time-domain signal to enhance consonant details.
2. The method according to claim 1, characterized in that, The process involves calculating the spatial complexity and temporal activity of pixels in the voice-related region of the facial image sequence to obtain the image entropy matrix at each time step, including: The spatial complexity is obtained by calculating the grayscale gradient values of adjacent pixels in the voice-related region of the facial image sequence. The temporal activity is obtained by calculating the grayscale difference between pixels at corresponding coordinate positions in adjacent image frames of the facial image sequence. The image entropy matrix at each time step is obtained by weighting and superimposing the spatial complexity and the temporal activity.
3. The method according to claim 1, characterized in that, The step of calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components and obtaining the visual energy sequence includes: Calculate the gradient of the numerical distribution between each element and its neighboring elements in the image entropy matrix to obtain the gradient contrast of each element; The temporal difference value of the corresponding coordinate position element in the image entropy matrix at adjacent time points is calculated to obtain the inter-frame evolution frequency; Based on the inter-frame evolution frequency, according to the preset mapping relationship between the inter-frame evolution frequency and the weight coefficient, the weight coefficient of each pixel region is obtained, and the gradient contrast is weighted and aggregated using the weight coefficient to obtain the visual energy sequence.
4. The method according to claim 1, characterized in that, The process involves using a window function to weight the speech time-domain signal, framing it, and calculating the number of times the waveform in each frame crosses zero level to obtain a zero-crossing rate distribution sequence. This sequence is then smoothed and weighted using an energy contrast parameter determined based on the speech time-domain signal to obtain a zero-crossing rate sequence, including: The speech time-domain signal is time-sequentially segmented according to a preset sampling length, and the amplitude of each segmented signal segment is mapped using a preset weighting window to obtain multiple analysis frames. The number of polarity reversals of adjacent signal values within each analysis frame is accumulated to obtain the zero-crossing rate distribution sequence; The ratio of the average signal strength of each analysis frame to a preset environmental noise benchmark value is calculated to obtain the corresponding energy comparison parameter; The zero-crossing rate distribution sequence is obtained by using the energy comparison parameter to perform amplitude weighting.
5. The method according to claim 1, characterized in that, The calculation of the logical synchronization between the visual energy sequence and the zero-crossing rate sequence to obtain the classification probability vector at each time step includes: The degree of matching between the changing trends of the visual energy sequence and the zero-crossing rate sequence at each time point is calculated to obtain the logical synchronization. By mapping the logical synchronization to a preset classification probability interval, a classification probability vector is obtained to determine whether the speech time-domain signal belongs to unvoiced, voiced, or environmental interference at each moment.
6. The method according to claim 1, characterized in that, When the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time, the gain parameters, center frequency, and bandwidth of the harmonic oscillator are configured based on the statistical amplitude of the image entropy matrix at the current time and the distribution density of the zero-crossing rate sequence, respectively, to obtain the harmonic components, including: The statistical amplitude is obtained by summing the values of all matrix elements in the image entropy matrix at the current moment, and the gain parameter is obtained based on the statistical amplitude and according to the preset mapping relationship between the statistical amplitude and the gain parameter. Calculate the statistical proportion of values in the zero-crossing rate sequence that are within a preset threshold range at the current moment to obtain the distribution density, and based on the distribution density, according to the preset correspondence between the distribution density and frequency characteristics, obtain the center frequency and the bandwidth. The harmonic components are generated by controlling the output strength and oscillation frequency range of the harmonic oscillator through the gain parameter, the center frequency, and the bandwidth.
7. The method according to claim 1, characterized in that, The step of aligning the harmonic components with the speech time-domain signal in phase and superimposing them within a frequency range determined by the center frequency and bandwidth to enhance consonant details includes: Frequency extraction is performed on the speech time-domain signal based on the center frequency and the bandwidth to obtain a reference signal within the target frequency band. Calculate the phase deviation between the reference signal and the harmonic component, and perform offset correction on the initial phase of the harmonic component based on the phase deviation to obtain the aligned harmonic component; Within the frequency range defined by the center frequency and the bandwidth, the aligned harmonic components are numerically superimposed on the speech time-domain signal to enhance consonant details.
8. An audiovisual conferencing voice enhancement system based on artificial intelligence, characterized in that, include: The acquisition module is used to acquire facial image sequences of the target object and speech time-domain signals including environmental high-frequency white noise collected by the conference terminal device; The calculation module is used to obtain the image entropy matrix at each time step by calculating the spatial complexity and temporal activity of the pixels in the voice-related region of the facial image sequence, and to obtain the visual energy sequence by calculating the gradient contrast and inter-frame evolution frequency of the image entropy matrix to filter out interference components. The generation module is used to divide the speech time-domain signal into frames using a window function weighting and to calculate the number of times the signal waveform of each frame crosses the zero level to obtain a zero-crossing rate distribution sequence. The zero-crossing rate distribution sequence is then smoothed and weighted using an energy comparison parameter determined based on the speech time-domain signal to obtain a zero-crossing rate sequence. The generation module is also used to calculate the logical synchronization between the visual energy sequence and the zero-crossing rate sequence to obtain the classification probability vector at each time step. When the classification probability vector determines that the speech time-domain signal is a speech frame with unvoiced characteristics at the current time step, the gain parameters, center frequency, and bandwidth of the harmonic oscillator are configured based on the statistical amplitude of the image entropy matrix at the current time step and the distribution density of the zero-crossing rate sequence to obtain the harmonic components. The processing module is used to perform phase alignment and signal superposition of the harmonic components and the speech time-domain signal within a frequency range determined by the center frequency and bandwidth, so as to enhance consonant details.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the AI-based audiovisual conferencing voice enhancement method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the implementation of the AI-based audiovisual conferencing voice enhancement method as described in any one of claims 1 to 7.