A new media time domain audio separation method and system based on visual features

CN122821980APending Publication Date: 2026-09-25宿州学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610977345.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]针对现有技术中频域处理路径被迫复用混合信号相位导致重构波形产生不可消除的相位失真伪影,且视觉信息仅能以全局偏置方式参与分离的不足,本申请提供了一种基于视觉特征的新媒体时域音频分离方法及系统

Benefits of technology

[0023]本申请构建了一条从时域编码到时域解码的全时域处理路径,分离操作直接作用于时域编码表示空间而非频域幅度谱,解码器直接从编码表示重构目标波形,全链路不引入短时傅里叶变换及其逆变换,从而在路径结构上消除了传统频域方法中因相位复用导致的失真伪影。同时,通过跨模态注意力机制,视觉特征序列能够在每个时间步逐帧引导分离网络的掩膜估计过程,使分离网络在每个时刻均可参考对应的唇部运动信息来确定目标成分,实现了视觉先验对音频分离的精细化时序引导,无需依赖排列不变训练即可指定并提取目标语音对象的语音。此外,时域编码器采用可学习的滤波器组替代固定窗函数,突破了传统短时傅里叶变换固有的时频分辨率限制,能够自适应地学习最优分析基函数。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821980A_ABST
    Figure CN122821980A_ABST
Patent Text Reader

Abstract

The application provides a new media time domain audio separation method and system based on visual features. The method comprises: time domain encoding of a mixed audio waveform to obtain a first feature representation; visual feature extraction and time alignment processing of a synchronous video containing a target speech object to obtain a second feature representation; mask estimation of the first feature representation by a cross-modal attention mechanism guided by the visual features to obtain a third feature representation; and time domain decoding of the third feature representation to obtain a separated audio waveform of the target speech object. The application constructs a full time domain processing path, and the separation operation directly acts on a time domain encoding representation space. The full link does not introduce short-time Fourier transform and its inverse transform, thereby eliminating distortion artifacts caused by phase multiplexing in traditional frequency domain methods from the path structure, and realizing frame-by-frame guidance of audio separation by visual priori through a cross-modal attention mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio and video signal processing technology, and in particular to a new media temporal audio separation method and system based on visual features. Background Technology

[0002] In new media content production scenarios, speech separation of the target speech object from mixed audio by multiple speakers is a fundamental requirement. Existing technologies typically employ a frequency domain processing approach, which involves performing a short-time Fourier transform on the mixed audio waveform to obtain a spectral representation, estimating the amplitude spectrum mask corresponding to the target speech object in the frequency domain, and then reconstructing the target waveform through an inverse transform. However, the frequency domain masking method is forced to reuse the original phase information of the mixed signal when performing the inverse transform. When the interference source energy is significant, there is a deviation between the mixed phase and the true phase of the target speech object, resulting in unavoidable phase distortion artifacts in the reconstructed waveform. Furthermore, existing audio-video joint methods fuse visual embedding as a global conditional vector with frequency domain features. Visual information can only affect the separation result in a global bias manner, failing to achieve accurate guidance at each time step. Summary of the Invention

[0003] To address the shortcomings of existing technologies where the forced reuse of mixed signal phases in the frequency domain processing path leads to unavoidable phase distortion artifacts in the reconstructed waveform, and where visual information can only participate in separation via global bias, this application provides a novel media temporal audio separation method and system based on visual features. This method constructs a full temporal domain processing path and utilizes a cross-modal attention mechanism to inject visual features frame-by-frame into the temporal separation network. This eliminates phase distortion from the path structure and enables accurate frame-by-frame guidance of audio separation by visual priors. Specifically, this application provides the following technical solutions:

[0004] Firstly, this application provides a new media temporal audio separation method based on visual features, including:

[0005] The mixed audio waveform is time-domain encoded to obtain the first feature representation;

[0006] Visual features are extracted from the synchronized video containing the target speech object and time-aligned to obtain a second feature representation with the same frame rate as the first feature representation.

[0007] The first feature representation and the second feature representation are input into a separation network. A cross-modal attention mechanism is used to guide the separation network to perform mask estimation on the first feature representation to obtain a third feature representation.

[0008] The third feature representation is then temporally decoded to obtain the separated audio waveform of the target speech object.

[0009] Optionally, the step of temporally encoding the mixed audio waveform to obtain the first feature representation includes: performing frame-by-frame convolution operations on the mixed audio waveform using a one-dimensional convolution filter bank; and performing nonlinear activation processing on the output of the convolution operation to obtain the first feature representation.

[0010] Optionally, the convolution stride of the one-dimensional convolutional filter bank is half the filter length; the step of temporal decoding of the third feature representation to obtain the separated audio waveform of the target speech object includes: decoding the third feature representation frame by frame using a transposed convolutional filter bank symmetrically arranged with the one-dimensional convolutional filter bank; and determining the separated audio waveform based on the adjacent frames after decoding.

[0011] Optionally, the step of extracting visual features from the synchronized video containing the target speech object and performing time alignment processing to obtain a second feature representation with the same frame rate as the first feature representation includes: extracting spatiotemporal features from the image sequence of the lip region of the target speech object through a spatiotemporal convolutional network; encoding the extracted spatiotemporal features into a lip motion embedding sequence through a temporal modeling network; and interpolating and upsampling the lip motion embedding sequence along the time axis so that the number of frames in the upsampled sequence is consistent with the number of frames in the first feature representation to obtain the second feature representation.

[0012] Optionally, the separation network includes multiple stacked dilated temporal convolutional blocks, the dilation factor of which increases exponentially in each repetition cycle.

[0013] Optionally, the cross-modal attention mechanism is set between adjacent repetition cycles of the dilated temporal convolutional block; the cross-modal attention mechanism uses the intermediate features after the first feature representation is processed by the separation network as the query, and the second feature representation as the key and value, to perform multi-head attention operations.

[0014] Optionally, the step of guiding the separation network to perform mask estimation on the first feature representation using the second feature representation through a cross-modal attention mechanism to obtain a third feature representation includes: generating a mask matrix; and determining the third feature representation based on the mask matrix and the first feature representation.

[0015] Secondly, this application also provides a new media temporal audio separation system based on visual features, comprising:

[0016] The time-domain coding module is used to perform time-domain coding on the mixed audio waveform to obtain the first feature representation;

[0017] The visual front-end module is used to extract visual features from the synchronized video containing the target speech object and perform time alignment processing to obtain a second feature representation that is consistent with the frame rate of the first feature representation.

[0018] The cross-modal fusion separation module is used to guide the separation network to perform mask estimation on the first feature representation using the second feature representation through a cross-modal attention mechanism, so as to obtain the third feature representation;

[0019] The time-domain decoding module is used to perform time-domain decoding on the third feature representation to obtain the separated audio waveform of the target speech object.

[0020] Optionally, the cross-modal fusion separation module includes a temporal convolutional subnetwork and a cross-modal cross-attention subnetwork, wherein the cross-modal cross-attention subnetwork is positioned between adjacent repetition cycles of the temporal convolutional subnetwork.

[0021] Optionally, the visual front-end module includes a spatiotemporal convolutional sub-network, a temporal modeling sub-network, and a temporal alignment sub-network; the spatiotemporal convolutional sub-network is used to extract spatiotemporal features from the image sequence of the lip region of the target speech object; the temporal modeling sub-network is used to encode the extracted spatiotemporal features into a lip motion embedding sequence; and the temporal alignment sub-network is used to interpolate and upsample the lip motion embedding sequence along the time axis so that the number of frames in the upsampled sequence is consistent with the number of frames represented by the first feature.

[0022] The beneficial effects of this application are as follows:

[0023] This application constructs a full-time-domain processing path from time-domain encoding to time-domain decoding. The separation operation directly operates on the time-domain encoded representation space rather than the frequency-domain amplitude spectrum, and the decoder directly reconstructs the target waveform from the encoded representation. The entire link does not introduce short-time Fourier transform or its inverse transform, thus eliminating the distortion artifacts caused by phase multiplexing in traditional frequency-domain methods in terms of path structure. Simultaneously, through a cross-modal attention mechanism, the visual feature sequence can guide the mask estimation process of the separation network frame by frame at each time step, enabling the separation network to refer to the corresponding lip movement information at each moment to determine the target component. This achieves refined temporal guidance of audio separation by visual priors, allowing the specification and extraction of the target speech object's speech without relying on permutation-invariant training. Furthermore, the time-domain encoder uses a learnable filter bank instead of a fixed window function, breaking through the inherent time-frequency resolution limitations of traditional short-time Fourier transform and enabling adaptive learning of the optimal analysis basis function. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the overall process of the new media temporal audio separation method based on visual features provided in the embodiments of this application.

[0025] Figure 2 This is a flowchart illustrating the audio time-domain encoding steps provided in an embodiment of this application.

[0026] Figure 3 This is a flowchart illustrating the visual feature extraction and time alignment steps provided in an embodiment of this application.

[0027] Figure 4 This is a flowchart illustrating the cross-modal attention fusion and temporal separation steps provided in an embodiment of this application.

[0028] Figure 5 This is a flowchart illustrating the time-domain waveform decoding and reconstruction steps provided in an embodiment of this application.

[0029] Figure 6 The above is a comparison diagram of the audio waveform separation effect provided in the embodiments of this application.

[0030] Figure 7 A comparison diagram of the spectra before and after separation provided for an embodiment of this application.

[0031] Figure 8 A schematic diagram of the module architecture of a new media temporal audio separation system based on visual features provided in this application embodiment. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0033] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0034] This embodiment provides a new media temporal audio separation method based on visual features. It achieves speech separation of the target speaker (i.e., the target speech object) by constructing a full temporal processing path from temporal encoding to temporal decoding. The mixed audio waveform is mapped to a high-dimensional temporal encoded representation by a temporal encoder. The synchronous video of the target speech object has lip movement features extracted by a visual front-end and aligned to the audio encoding frame rate. A cross-modal attention mechanism injects the aligned visual features frame by frame into the temporal separation network to guide mask estimation. The separated target encoded representation is directly reconstructed into the separated audio waveform of the target speech object by a temporal decoder. The entire link does not introduce short-time Fourier transform or its inverse transform; the separation operation operates on the temporal encoded representation space rather than the frequency domain amplitude spectrum.

[0035] Example 1

[0036] like Figure 1 As shown, this method includes the following steps:

[0037] S100, perform time-domain encoding on the mixed audio waveform to obtain the first feature representation.

[0038] like Figure 2 As shown, the temporal coding performs frame-by-frame convolution operations on the mixed audio waveform using a set of learnable one-dimensional convolutional filter banks and applies nonlinear activation processing to map the one-dimensional waveform into a high-dimensional non-negative temporal coding representation.

[0039] The mixed audio waveform It is a single-channel time-domain signal, where This represents the total number of sampling points. Before inputting into the time-domain encoder, amplitude normalization is performed on the mixed audio waveform to constrain the amplitude value of each sampling point to [value missing]. Within the interval. The time-domain encoder includes A length of One-dimensional convolutional filters, each filter with a step size Perform a sliding convolution operation on the mixed audio waveform. For the first... Frame, frame start position is The time-domain encoder extracts a length of: waveform segment Compare the fragments with... Perform inner product operation on each filter to obtain A dimensional response vector. For the stated... Applying the ReLU nonlinear activation function to the dimensional response vector and setting all negative values ​​to zero yields the dimensional response vector. Frame encoding vector Wherein, the ReLU activation function is Its function is to perform nonlinear activation processing on the convolution output, constraining the output value to be non-negative. In the encoding representation space, the non-negativity constraint gives the masking operation a monotonically decaying characteristic: for any component in the encoding space, a mask value of zero indicates that the component is completely set to zero, a mask value of one indicates that the component is fully transmitted to the decoder, and intermediate values ​​indicate that the component is attenuated proportionally in the encoding space. It should be noted that the mask is in the encoding representation space, not the final waveform space, because the attenuated components in the encoding space are linearly combined by the decoder's synthesis filter (whose coefficients include both positive and negative values) to produce the final waveform sampling points, while the decoder is responsible for converting the proportional modulation in the encoding space into amplitude reconstruction in the waveform space. The two belong to different representation levels.

[0040] After performing the convolution and activation operations sequentially on all valid frames of the mixed audio waveform, the temporal encoder arranges the encoded vectors of each frame in chronological order to obtain the temporal encoded representation of the mixed signal. The number of encoded frames Each column of the time-domain encoding corresponds to a length of [length missing] in the mixed audio waveform. time segments in The dimension represents the encoding in the space, and each row corresponds to the response sequence of a convolutional filter across all time frames.

[0041] All filter coefficients of the one-dimensional convolutional filter bank are learnable parameters, determined by gradient updates of the loss function through end-to-end training, without any pre-defined fixed basis function forms. Compared to short-time Fourier transforms using fixed sine and cosine basis functions for signal analysis, this learnable filter bank can adaptively learn the optimal analysis basis functions suitable for the target separation task, without being constrained by a fixed time-frequency resolution. The encoder's filter weight matrix... These are static parameters during the inference phase and are not updated after being loaded onto the computing device.

[0042] For example, let the audio sampling rate be... Hz, given a 4-second mixed audio waveform, the total number of sampling points is... Let the filter length be... The corresponding time window width is ms, step size The corresponding frame shift is ms. The encoder contains There are 1 filter. The number of encoded frames is 1. Take the 0th frame segment of the mixed audio waveform. Each of the 256 filters is then used to perform an inner product operation with this segment. Let's assume the inner product result of filter number 0 is... Output after ReLU activation The inner product result of filter number 1 is: Output after ReLU activation Thus, 256 filters complete the encoding operation of frame 0 in parallel, producing the encoded vector. After performing the above operations sequentially on all 6399 frames, the first feature representation is obtained. Approximately 40% to 60% of the elements are set to zero due to ReLU activation, exhibiting sparse characteristics. The filter length... The corresponding 1.25 ms time window is much smaller than one cycle of the speech fundamental frequency. A typical fundamental frequency range of 80 Hz to 200 Hz corresponds to a cycle of 12.5 ms to 5 ms, thus ensuring subphone-level temporal resolution. Coding Dimensions A balance must be struck between computational efficiency and feature representation capability. When the expression level is below 128, insufficient expression capacity leads to a decrease in separation quality. Above 512, the computational load increases linearly while marginal returns decrease.

[0043] In an optional embodiment, the temporal encoder is specifically a one-dimensional convolutional encoder, including a learnable one-dimensional convolutional filter bank and a nonlinear activation layer. The one-dimensional convolutional filter bank contains 256 convolutional kernels with a length of 20 sampling points, performing frame-by-frame convolution operations on the mixed audio waveform with a stride of 10. The nonlinear activation layer performs nonlinear activation processing on the output of the one-dimensional convolutional filter bank, using the ReLU function to constrain the output value to non-negativity. The one-dimensional convolutional encoder is equivalent to a set of learnable finite impulse response analysis filters, where each filter extracts a scalar response for a corresponding time segment of the input waveform. Multiple filters operate in parallel to produce the same segment. 3D encoded vector. Filter weight matrix. It contains 5120 learnable parameters, which are initialized using the Kaiming uniform initialization method.

[0044] For example, the single-frame encoding process of the one-dimensional convolutional encoder can be expressed in the following mathematical form. Let the first... The waveform segment of the frame is The encoder weight matrix is Then the first The frame encoding operation is as follows:

[0045]

[0046] in The OK For the first The coefficient vector of the nth convolutional filter, the encoding vector of the nth... Each component is .by , , For example, for The encoder executes a total of [number] sampling points of input waveform. Each sub-framing encoding operation consists of 256 vector inner products of length 20 and 256 ReLU activation operations, with a total computational cost of approximately [missing information]. This is a multiplication and addition operation.

[0047] In an optional embodiment, the convolution stride of the one-dimensional convolutional filter bank is set to half the filter length. Specifically, the filter length... Sampling points, step size Sampling points, there is a 50% overlap between adjacent frames, that is, the first... Frame coverage sampling point range , No. Frame coverage sampling point range Both share an overlapping area of ​​10 sampling points.

[0048] The 50% frame overlap rate provides ample freedom for the time-domain decoder to synthesize waveforms through overlap addition, enabling the analysis-synthesis filter bank pair composed of the encoder and decoder to approximate the perfect reconstruction condition during end-to-end training. In traditional filter bank applications, the mathematical definition of the perfect reconstruction condition is: for any input signal... The signal is reconstructed by encoding through the analysis filter bank and then decoding through the synthesis filter bank. satisfy (in (where is a non-zero constant scaling factor), meaning the reconstructed signal differs from the original signal only by a global scaling factor. In matrix form, let the encoder weight matrix be... The decoder weight matrix is For a step size of A 50% overlap configuration, the perfect reconstruction condition is equivalent to an overlapped additive matrix composition. The superposition result in the overlapping region is a constant multiple of the identity matrix. It should be noted that the encoding dimension described in this scheme... Greater than the filter length , constitutes an overcomplete representation ( This means that the dimension of the encoding space is higher than the dimension of the input segment. An overcomplete representation implies that the reconstruction condition is met. The solution is not unique (the solution space has degrees of freedom), which provides favorable conditions for end-to-end training, namely, a larger set of feasible solutions for the loss function, making it easier for gradient descent to find solutions that satisfy the reconstruction constraints. Although gradient descent does not guarantee convergence to the global optimum, the training objective (maximizing the scale-invariant signal-to-distortion ratio of the separated waveform to the target waveform) implicitly includes reconstruction quality constraints: the upper bound of the quality of the separated waveform is limited by the reconstruction accuracy of the encoder-decoder pair, and the gradient of the loss function naturally drives this. and The encoder and decoder evolve in the direction that satisfies the reconstruction conditions. During training, their weight matrices are constrained by the same loss function but do not share parameters; they are not pseudo-inverses and each converges to a local optimum for the given separation task through gradient descent.

[0049] The filter parameters of the temporal decoder are symmetrically set with those of the temporal encoder; that is, the temporal decoder uses a transposed convolutional filter bank symmetrically set with respect to the one-dimensional convolutional filter bank, and the filter length... Step length The number of input channels is The output channel number is 1, and the decoded adjacent frames are synthesized by overlapping addition to obtain the separated audio waveform.

[0050] S200, visual features are extracted from the synchronized video containing the target speech object and time alignment is performed to obtain a second feature representation with the same frame rate as the first feature representation.

[0051] like Figure 3 As shown, the visual feature extraction and time alignment process includes extracting spatiotemporal features from the image sequence of the lip region of the target speech object, performing temporal modeling on the extracted spatiotemporal features to generate a lip motion embedding sequence, and interpolating and upsampling the lip motion embedding sequence along the time axis to make it consistent with the number of frames represented by the first feature.

[0052] The synchronized video is a video signal strictly aligned with the mixed audio waveform on the time axis and includes a facial image of the target speech object. Before being input into the visual front-end module, the synchronized video undergoes lip region cropping preprocessing to extract a sequence of lip region images of the target speech object. ,in This represents the number of video frames, with each frame being [number]. A grayscale image with pixels normalized to 1000 pixels. The visual front-end module sequentially performs four stages of operations on the lip region image sequence: spatiotemporal feature extraction, global spatial pooling, temporal modeling, and temporal alignment processing, ultimately producing a second feature representation with the same number of frames as the first feature representation. .

[0053] In the spatiotemporal feature extraction stage, the visual front-end module performs spatiotemporal joint convolution operations on the lip region image sequence through a three-dimensional convolutional layer. The convolution kernel of the three-dimensional convolutional layer slides simultaneously in both spatial and temporal dimensions; a single convolution operation can capture the lip movement pattern between adjacent frames, establishing a spatiotemporal joint representation at the bottom layer of feature extraction. The three-dimensional convolutional layer includes... There are n convolutional kernels, each with a time dimension of 1. Spatial dimension is The input sequence is convolved with a time step of 1 and a spatial step of 2. The convolution output is then batch normalized and ReLU activated, followed by a 3D max-pooling layer for further downsampling in the spatial dimension, resulting in a shape... The spatiotemporal feature tensor. For the spatiotemporal feature tensor in spatial dimension... Global average pooling is performed on the channel, calculating the mean of all spatial location features at each time step. This eliminates spatial location information while preserving motion temporal information, resulting in a shape of... The temporal feature matrix.

[0054] In the temporal modeling stage, the visual front-end module encodes the temporal feature matrix into a continuous lip motion embedding sequence through a temporal modeling network. The temporal modeling network includes multiple stacked one-dimensional residual convolutional blocks, each containing two layers of one-dimensional convolution, batch normalization, and a residual connection structure. By stacking residual convolutional blocks layer by layer, the temporal modeling network progressively aggregates local spatiotemporal features into a continuous lip motion representation with long-range semantic context, with an output dimension of [missing information]. Lip motion embedding sequence The embedding vector at each time step encodes the lip motion semantic information of the corresponding video frame and its neighboring frames.

[0055] During the time alignment process, the lip motion embedding sequence is linearly interpolated and upsampled along the time axis, reducing the sequence length from... Frame expansion to The number of frames in the upsampled sequence is equal to the number of frames represented by the first feature. Consistent. For the first sample after upsampling... For each target time step, calculate its floating-point position in the source sequence. Take integer frames on both sides of the floating-point position. and Perform linear weighting:

[0056]

[0057] Where the interpolation coefficients The range of values ​​is The linear interpolation described here is a parameterless operation, introducing no additional learnable parameters. The interpolation operation ensures accurate alignment of boundary frames; that is, the first frame after upsampling is the same as the first frame of the source sequence, and the last frame after upsampling is the same as the last frame of the source sequence. After this linear interpolation upsampling, the second feature representation is obtained. Its time dimension is similar to the first feature representation. Strict alignment ensures a one-to-one correspondence between the two at each time step.

[0058] For example, let the video frame rate be... FPS: Input a synchronized video with a duration of 4 seconds, then the video frame rate. The shape of the lip region image sequence is... After processing by a 3D convolutional layer, the temporal dimension remains unchanged for 100 frames, while the spatial dimension changes. Reduced to stride 2 convolution Then, after max pooling with a step size of 2, it is reduced to Number of output channels Global average pooling pairs Taking the average value at each spatial location, the feature tensor is transformed from... Compress to The temporal modeling network progressively expands the channel dimension from 64 to 256 through three layers of one-dimensional residual convolutional blocks, outputting a lip motion embedding sequence. The linear interpolation upsamples the 100-frame embedded sequence to... Frame. Retrieve target frame. For example, its floating-point position in the source sequence is The left integer frame is frame 49, and the right integer frame is frame 50. Interpolation coefficients... ,but That is, the visual features of this audio encoding time step are obtained by proportionally interpolating the embedding vectors of frames 49 and 50 of the video. Upsampling ratio That is, the lip movement information of each video frame is linearly distributed across 64 audio coded frames. Assuming that the lip movement changes approximately linearly within a 40 ms video frame interval, this linear approximation error is minimal for typical lip movement frequencies of 3 Hz to 8 Hz.

[0059] In one optional implementation, the visual feature extraction includes performing spatiotemporal feature extraction on the lip region image sequence of the target speech object using a spatiotemporal convolutional network, and then generating a continuous lip motion embedding sequence via a temporal modeling network. The three-dimensional convolutional layer configuration of the spatiotemporal convolutional network is as follows: input channel 1 (grayscale image), output channel 64, temporal kernel size 5 (covering a 200 ms time window of 5 frames at 25 fps, corresponding to a complete syllable-level lip opening and closing cycle), and spatial kernel size... The temporal dimension has a step size of 1 (keeping the temporal dimension length constant) and a spatial dimension step size of 2 (spatial downsampling). The temporal modeling network employs a multi-layer stacked structure of one-dimensional residual convolutional blocks, with channel dimensions of 64, 128, 256, and 256 respectively, and a kernel size of 3. The effective receptive field after stacking 3 layers of residual blocks covers 13 frames (520 ms), which is sufficient to model the co-articulation effect of adjacent syllables in continuous speech. The output dimension of the temporal modeling network is... With audio encoding dimension This ensures that the dimensions of the query and the key are aligned in subsequent cross-modal attention operations, eliminating the need for an additional dimension projection layer.

[0060] In an optional embodiment, the time alignment process involves interpolating and upsampling the lip motion embedding sequence along the time axis to the same number of frames as the first feature representation. Specifically, when the video frame rate is 25 fps, the source sequence length is... Frames, audio encoding frame rate The target sequence length at fps is Frame, upsampling factor is The interpolation upsampling employs linear interpolation. For each target time step, the two adjacent source frames are determined based on their floating-point positions in the source sequence, and a linear weighted sum is performed according to their distance. This linear interpolation ensures precise alignment between the first and last frames: the feature value of the 0th frame after upsampling is the same as that of the 0th frame in the source sequence, and the feature value of the 1st frame after upsampling is... The feature values ​​of the frame and the source sequence The frames are identical. The visual features of each frame transition smoothly between adjacent video frames, without producing a stepped discontinuity.

[0061] S300, the first feature representation and the second feature representation are input into the separation network, and the separation network is guided by the second feature representation to perform mask estimation on the first feature representation through a cross-modal attention mechanism to obtain the third feature representation.

[0062] like Figure 4 As shown, the cross-modal attention fusion and temporal separation includes inputting the first feature representation into a temporal separation network for deep temporal modeling, injecting the second feature representation through a cross-modal attention mechanism during the processing of the temporal separation network to guide mask estimation, and applying the estimated target separation mask to the first feature representation to obtain the third feature representation.

[0063] The separation network receives the first feature representation. and the second feature representation As a bimodal input, the time dimension of both is... Strictly equal. The separation network performs layer normalization preprocessing on the first feature representation, for each time step. The mean and standard deviation of the dimensional encoding vectors are independently calculated and normalized to eliminate differences in encoding magnitude between different time steps. The normalized representation is then mapped to the hidden dimension space of the segregated network through pointwise convolutions in the bottleneck layer to obtain the initial hidden state. .

[0064] The core processing unit of the separation network is a temporal convolutional network, which consists of multiple stacked dilated temporal convolutional blocks. The temporal convolutional network includes... Each repetition cycle contains [number] repetition periods, and each repetition cycle contains [number] repetition periods. A dilated convolutional block, wherein the dilation factor of the dilated temporal convolutional block increases exponentially in each repetition period, i.e., the first... layer( The expansion factor of ) is Each dilated convolutional block performs a depthwise separable one-dimensional dilated convolution operation on the input, with the convolution kernel using a dilation factor. Sampling the time axis at defined intervals allows a single convolutional layer to span [time axis]. Feature extraction is performed across time steps, where... The kernel size is [size]. The convolutional output is normalized and parameterized with ReLU activation, then mapped back to the original hidden dimensions via pointwise convolution, and finally added to the layer input through residual connections to obtain the layer's output. With the stacking of exponentially increasing expansion factors, the effective receptive field of a single repetition cycle grows exponentially to One time frame, After repeated iterations, the total effective receptive field coverage The time frame enables the separation network to utilize a sufficiently wide temporal context window to distinguish speech components from different speakers.

[0065] The cross-modal attention mechanism is configured between adjacent repetition cycles of the dilated temporal convolutional block. Specifically, in each repetition cycle... After the dilated convolutional block is processed, its output serves as the input to the cross-modal attention mechanism, and the output after the cross-modal attention operation serves as the input for the next repetition cycle. The cross-modal attention mechanism uses the intermediate features of the first feature representation processed by the separation network as the query, and the second feature representation as the key and value, to perform multi-head attention operations. For each audio time step, the cross-modal attention mechanism calculates the correlation weights between the audio intermediate features of that time step and the features of all visual time steps, injecting visual information into the audio representation in the form of a weighted sum, thereby achieving frame-by-frame fine-grained guidance of the audio separation process by visual priors.

[0066] The multi-head attention operation decomposes the hidden dimension space into... There are 3 parallel attention heads, each independently learning different types of cross-modal association patterns. For the 1st... Each attention head is obtained by querying the projection matrix. Key projection matrix Sum projection matrix Perform a linear transformation on the input to obtain the query matrix. Key matrix Sum matrix Attention weights are calculated using the scaled dot product of the query and the key.

[0067]

[0068] in For each head's query and key dimensions, This is a scaling factor used to prevent the inner product from becoming too large, causing the gradient of the Softmax function to tend to zero. Attention weight matrix. The Line 1 The column represents the audio number. Frame to visual The attention level of each frame is determined by Softmax normalization, which ensures that the sum of the weights in each row is 1. The attention output of each head is a weighted sum of the value matrix based on the attention weights. , The output of each head is spliced ​​together and then projected through the output matrix. The input is mapped back to the original dimension and added to it via a residual connection. Since the audio encoded representation and the visual feature sequence have been temporally aligned in step S200, the attention weights are naturally concentrated near the temporally corresponding positions, allowing the network to capture the co-articulation delay effect between lip movements and acoustic output.

[0069] The separation network is through all Repeating cycle and After the cross-modal attention operation, the final hidden state is mapped to a target separation mask through a mask generation layer. The mask generation layer includes a pointwise convolutional layer and a sigmoid activation function; the pointwise convolutional layer hides the dimension... Mapping back to encoding dimension The sigmoid activation function compresses the output to a value range between 0 and 1, generating a target separation mask. Each element of the target separation mask. Represents the corresponding position in the first feature representation The proportion of encoded energy belonging to the target speech object. A mask value of 0 indicates that the encoded energy at that position belongs to the interference component, and a mask value of 1 indicates that it belongs to the target component.

[0070] The target separation mask is multiplied element-wise with the first feature representation to obtain the third feature representation:

[0071]

[0072] in This represents the Hadamard product, i.e. Since the first feature indicates that the ReLU activation constraint after step S100 is non-negative, and the value range of the target separation mask is between 0 and 1, the result of element-wise multiplication of the two is... It remains a non-negative value, and each element does not exceed the original encoded value at the corresponding position; that is, the masking operation only performs attenuation and not amplification. The third feature represents... The representation of the target speech object in the time-domain coding space preserves the coded components belonging to the target speech object and suppresses interference components.

[0073] For example, let the number of repetitions of the temporal convolutional network be... The number of dilated convolution blocks repeated each time Hidden channel number kernel size The expansion factor is successively as follows in each repetition: In a single repetition cycle, the dilation factor of layer 0 is 1, and the convolution kernel operates on three adjacent time steps; the dilation factor of layer 7 is 128, and the three sampling points of the convolution kernel span a time range of 256 frames. Taking layer 7 of the first repetition cycle as an example, at time step... ,aisle For example, the dilated convolution operation is as follows: That is, it is directly related to the interval. Time step information for a single frame (approximately 160 ms). The effective receptive field for a single repetition is... Frame, corresponding ms; the total effective receptive field after 3 repetitions was Each frame, approximately 478 ms, covers the time span of a typical speech syllable. The dilated convolution employs a non-causal mode, symmetrically padding zeros at both ends of the time axis, maintaining the time dimension length of each layer's output. constant.

[0074] For example, the number of attention heads in the cross-modal attention mechanism Dimensions per head scaling factor Take the 0th head of the first layer of cross-attention and the audio time step. For example, query vector Calculate the scaled dot product with the key vectors at each visual time step. Attention weights are then concentrated after Softmax normalization. to Nearby (current time step and the previous frame), the corresponding weights are approximately 0.24 and 0.29, totaling 0.53, which conforms to the co-articulation rule where lip movements slightly precede acoustic output. The attention output is a weighted sum of all visual time step value vectors, mainly contributed by visual information from the time-corresponding position and its neighborhood, achieving fine guidance of visual prior for audio separation at the current time step. The cross-modal attention mechanism consists of 3 layers (corresponding to 3 repetition cycles), with visual information injected into the audio processing path at each scale level, allowing the separation network to reference lip movement information in its multi-level representations from shallow to deep.

[0075] For example, after all 3 repetition cycles and 3 cross-modal attention operations, the time step is taken. Taking the mask generation process as an example, the final hidden state is mapped by pointwise convolution of the mask generation layer to obtain a 256-dimensional logits vector, which is then compressed by Sigmoid activation. Interval. Assuming the Sigmoid input value for channel 0 is 2.1, then... This indicates that 89.1% of the encoded energy of this channel at the current moment belongs to the target speech object; the input value of channel 1 is... ,but This indicates that only 18.2% of the channel belongs to the target. Applying a mask to the first feature representation, assuming... ,but ; ,but The encoded components of the target speech object are preserved, while the interference components are significantly attenuated.

[0076] In one optional implementation, the separable network comprises multiple stacked dilated temporal convolutional blocks, the dilation factor of which increases exponentially in each repetition cycle. Specifically, the temporal convolutional network includes 3 repetition cycles, each containing 8 dilated convolutional blocks with dilation factors of 1, 2, 4, 8, 16, 32, 64, and 128 respectively. Each dilated convolutional block employs a depthwise separable one-dimensional dilated convolutional structure, first performing dilation convolution (depthwise convolution) independently on each channel, and then achieving inter-channel information interaction through pointwise convolution, reducing the number of parameters to that of standard convolution while maintaining the equivalent receptive field. Each convolutional block contains residual connections that directly add the input to the layer output, ensuring the effective propagation of gradients in deep networks.

[0077] In a further optional implementation, the cross-modal attention mechanism is set between adjacent repetition cycles of the dilated temporal convolutional block, performing multi-head attention computation with audio features as queries and visual features as keys and values. The three repetition cycles of the temporal convolutional network generate three cross-attention insertion positions: one layer of cross-modal cross-attention is set after the output of the first repetition cycle, the second repetition cycle, and the third repetition cycle. The query source for each cross-attention layer is the intermediate audio features processed in the current repetition cycle, and the key and value source is the second feature representation (visual feature sequence, shared across layers). The multi-head attention computation uses four parallel attention heads, each with a dimension of 64 and a scaling factor of [missing value]. The output, after being projected and connected with the residual, serves as the input for the next repetition cycle.

[0078] In one optional implementation, the target separation mask has a value range between 0 and 1, and the third feature representation is obtained by element-wise multiplying the target separation mask with the first feature representation. The target separation mask is generated using a Sigmoid activation function, which is defined as follows: Its output value range is strictly limited to Within the open interval. The element-wise multiplication operation independently scales each position of the first feature representation, retaining the corresponding encoded energy at positions with mask values ​​close to 1 and suppressing the corresponding encoded energy at positions with mask values ​​close to 0. The sigmoid mask is suitable for single-target speech object separation scenarios, locking the identity of the target speech object through visual features, and can specify and extract the speech components of a specific speaker without requiring training with invariant permutations.

[0079] S400, perform time-domain decoding on the third feature representation to obtain the separated audio waveform of the target speech object.

[0080] like Figure 5 As shown, the temporal decoding performs frame-by-frame decoding of the third feature representation by a transposed convolutional filter bank symmetrically set with the temporal encoder, and synthesizes the decoded adjacent frames by overlapping addition to obtain the separated audio waveform of the target speech object.

[0081] The time-domain decoder receives the third feature representation. As input, where For encoding dimensions, The number of encoded frames. The temporal decoder contains... A length of The synthesized filters are used to form a transposed convolutional filter bank, whose filter weight matrix is: For the third feature represented by the first Frame coding vector The time-domain decoder multiplies the vector by the synthesis filter weight matrix, producing an output of length [missing information]. Waveform segment:

[0082]

[0083] in The Line 1 Column elements For the first The synthesis filter at the th ... The coefficient at the position, the th waveform segment Each sampling point is a weighted sum of all channel coded values ​​and their corresponding synthesized filter coefficients:

[0084]

[0085] The transpose convolution operation is equivalent to a learnable synthetic filter bank operation, which will... The representation in the 3D encoding space is inversely mapped to a one-dimensional time-domain waveform segment. A linear combination of several basis functions reconstructs the target waveform within the corresponding time window.

[0086] The time-domain decoder is for all After each frame undergoes the transpose convolution operation described above, the length of the output of each frame is... Waveform segments by step size The segments are arranged at intervals, and adjacent segments are combined into a continuous waveform through overlapping addition. Specifically, the first... The waveform fragment of the frame is placed at a position on the global time axis. When step length At this time, there is a 50% overlap between waveform segments of two adjacent frames, and the contributions of each frame within the overlap area are directly added together. For any sampling point in the output waveform... Its value is the sum of the values ​​of all waveform segments covering that location at the corresponding local location:

[0087]

[0088] Under 50% overlap, except for the boundary regions of the first frame's start segment and the last frame's end segment, each sampling point is contributed by waveform segments from exactly two adjacent frames. After overlapping and additive synthesis, a length of [length missing] is obtained. A continuous waveform, when Compared with the original number of sampling points Output directly when they match; when... When the last redundant sampling points are truncated, Ending zeros are added. The final output is the separated audio waveform of the target speech object. .

[0089] All filter coefficients of the transposed convolutional filter bank are learnable parameters, determined through end-to-end training and updated by the gradient of the loss function. The filter weight matrix of the temporal decoder... With the filter weight matrix of the time-domain encoder The two waveforms have the same number of parameters but do not share weights. During training, they are constrained by the same loss function and converge to a solution satisfying the reconstruction condition through gradient descent. The training objective is to minimize the difference between the separated waveform and the true waveform of the target speech object. This loss function implicitly drives the encoder-decoder to approximate the perfect reconstruction condition without requiring analytical conditions in filter design to enforce it. Although the third feature representation... All elements in the expression are non-negative, but the coefficients of the synthesized filter contain both positive and negative values. The weighted sum of the channel encoded values ​​and the synthesis coefficients can produce any real value, and the alternation of positive and negative values ​​in the decoded output waveform is a natural physical characteristic of sound waves.

[0090] For example, let the coding dimension be... Filter length Step length Number of encoded frames The decoder output waveform length is , compared with the original number of sampling points Consistent, no truncation or zero padding required. Take the first... Taking a frame as an example, the masked encoded value continues from step S300. , , Assuming the synthesis filter is in The coefficient of position is , , Then the portion of the waveform segment at the 0th sampling point is calculated as follows: The summation of the contributions from 256 channels yields the following result. The complete value. For After performing the above calculations one by one, a complete 20-point waveform segment is obtained. .

[0091] For example, the first The waveform fragment of the frame is placed in a global location. , No. Frame placed at Both are in There is an overlap of 10 sampling points in the interval, and the output waveform within this region is the sum of the contributions from the two frames, for example... , And so on. Second half of the frame With the First half of the frame The first 10 sampling points overlap. Therefore, except for the first 10 sampling points of the first frame and the last 10 sampling points of the last frame, each position in the output waveform is composed of the superposition of two frame waveform segments. The final output is a separated audio waveform. The sampling rate is 16000 Hz, the duration is 4 seconds, and the waveform amplitude is constrained to... Within the specified range, audio files can be written directly or played in real time.

[0092] like Figure 6 As shown, the separated audio waveform and the reference clean waveform are compared and verified in the time domain. Figure 6 The horizontal axis represents time in milliseconds, and the vertical axis represents the normalized amplitude, which ranges from -1 to 1. Figure 6The comparison includes three rows of waveforms: the first row, labeled (a) the mixed audio waveform, shows the waveform of the mixed signal after the speech of two speakers is superimposed; the second row, labeled (b) the separation result of this method, shows the waveform of the target speech object separated by this method; and the third row, labeled (c) the reference clean waveform, shows the original clean speech waveform of the target speech object. Comparing the second and third rows, it can be seen that the separated waveform of this method is highly consistent with the reference clean waveform in the time domain structure. The speech components interfering with the speaker are effectively suppressed, and the waveform amplitude and phase information are well recovered, verifying the effectiveness of the full-time-domain processing path in structurally eliminating phase distortion.

[0093] like Figure 7 As shown, a spectral comparison analysis is performed on the signals before and after separation. Among them, Figure 7 The horizontal axis represents time in seconds, and the vertical axis represents frequency in kHz. The shade of gray indicates the signal amplitude at each time and frequency position in dB. The darker the color, the stronger the energy at that time and frequency position. Figure 7 The spectrum consists of three rows: the first row, labeled (a), shows the mixed signal spectrum, where the harmonic structures of multiple speakers are intertwined; the second row, labeled (b), shows the separation result of this method, where the harmonic components of the interfering speaker are significantly suppressed, and the harmonic structure of the target speech object is clearly distinguishable; the third row, labeled (c), shows the reference clean signal spectrum, serving as a reference benchmark. The separation result of this method preserves the fundamental frequency and harmonic components of the target speech object in the frequency dimension, while effectively suppressing the frequency bands where the energy of the interfering speaker is concentrated. It should be noted that this spectrum is only a visual representation of the separation effect and does not limit the actual effect.

[0094] All learnable parameters of the new media temporal audio separation method based on visual features are jointly optimized through end-to-end training. The goal of the training phase is to maximize the scale-invariant signal-to-noise ratio (SNR) between the separated waveform and the real waveform of the target speech object, i.e., using the scale-invariant signal-to-distortion ratio (SNR) as the training loss function. For the separated waveform... and the target's true waveform The scale-invariant signal distortion ratio is defined as:

[0095]

[0096] in The projection of the target signal onto the direction of the separated waveform. This represents the residual noise component. The training objective is to maximize the scale-invariant signal-to-distortion ratio (SINTR), which is equivalent to minimizing the negative SINTR. This loss function is invariant to the global scaling of the signal and focuses only on the degree of waveform shape matching rather than absolute amplitude, allowing the encoder and decoder to effectively optimize separation quality without learning precise energy recovery.

[0097] The training process uses the Adam optimizer for parameter updates. The Adam optimizer adaptively adjusts the learning rate for each parameter by maintaining first-order and second-order moment estimates of the gradient. The initial learning rate is set to... When the loss function value on the validation set fails to improve for three consecutive training epochs, the learning rate is reduced to half its current value. The training batch size is set to contain several fixed-length mixed audio segments and their corresponding synchronized video segments per batch, and all training samples are iterated once per training epoch. The gradient clipping threshold is set to 5 to prevent gradient explosion from causing training instability. During training, amplitude normalization preprocessing is performed on the mixed audio waveforms, and random horizontal flipping is performed on the lip region image as data augmentation.

[0098] In an optional embodiment, the training loss function is a negative scale-invariant signal-to-distortion ratio (SINTR). The SINTR is invariant to global signal scaling and is calculated as follows: first, the separated waveform is projected onto the target waveform direction to obtain the target component; then, the logarithmic ratio of the target component energy to the residual noise energy is calculated. The training process employs an adaptive learning rate optimizer, with the initial learning rate dynamically decaying based on validation set performance during training. All learnable parameters of the temporal encoder, the visual front-end module, the separation network, and the temporal decoder share the same loss function during training, and gradients are calculated and updated synchronously using a backpropagation algorithm, enabling collaborative optimization among modules to maximize end-to-end separation performance.

[0099] In an optional embodiment, the temporal decoder employs a transposed convolutional filter bank symmetrically arranged with the one-dimensional convolutional filter bank to decode the third feature representation frame by frame, and determines the separated audio waveform based on the decoded adjacent frames. Specifically, the filter length of the transposed convolutional filter bank is... Sampling points, step size is Sampling points, number of input channels The output channel number is 1, and it is strictly symmetrical with the one-dimensional convolutional filter bank of the time-domain encoder in terms of filter length and stride. The weight matrix of the transposed convolutional filter bank... It contains 5120 learnable parameters, initialized using the Kaiming uniform initialization method. The overlapping addition is a parameterless operation, automatically executed by the transposed convolution operation when the stride is less than the filter length, without requiring an explicit window function. The learned synthesis filter implicitly encodes the shape of the synthesis window, adaptively converging to the optimal synthesis basis function that satisfies the reconstruction condition during end-to-end training.

[0100] like Figure 8As shown, this embodiment also provides a new media temporal audio separation system based on visual features, including a temporal encoding module, a visual front-end module, a cross-modal fusion and separation module, and a temporal decoding module. Each of these modules can be implemented by a processor executing computer program instructions stored in memory, or by dedicated hardware circuits or programmable logic devices.

[0101] The time-domain encoding module is used to perform time-domain encoding on the mixed audio waveform to obtain a first feature representation. The time-domain encoding module receives a single-channel mixed audio waveform. As input, the mixed audio waveform is subjected to frame-by-frame convolution operations through a learnable one-dimensional convolutional filter bank, and a nonlinear activation process is applied to the convolution output to constrain the output value to non-negativity, resulting in a time-domain coded representation of the output mixed signal. The filter parameters of the time-domain coding module are static weights during the inference phase and are not updated after being loaded onto the computing device.

[0102] The visual front-end module is used to extract visual features from synchronized video containing the target speech object and perform time alignment processing to obtain a second feature representation with the same frame rate as the first feature representation. The visual front-end module receives a sequence of lip region images of the target speech object. As input, spatiotemporal features are extracted sequentially through a 3D convolutional layer, spatial location information is eliminated through global average pooling, a lip motion embedding sequence is generated through a temporal modeling network, and the embedding sequence of the video frame rate is upsampled to the audio coding frame rate through linear interpolation, outputting a visual feature sequence with the same number of frames as the first feature representation. There is no data dependency between the visual front-end module and the temporal encoding module; they can be executed in parallel.

[0103] The cross-modal fusion separation module is used to guide the separation network to perform mask estimation on the first feature representation using the second feature representation through a cross-modal attention mechanism, thereby obtaining a third feature representation. The cross-modal fusion separation module receives the first feature representation output by the temporal coding module and the second feature representation output by the visual front-end module as bimodal inputs. The cross-modal fusion separation module includes a temporal convolutional sub-network and a cross-modal cross-attention sub-network. The temporal convolutional sub-network consists of multiple stacked dilated temporal convolutional blocks, modeling long-range temporal dependencies through exponentially increasing dilation factors. The cross-modal cross-attention sub-network is positioned between adjacent repetition cycles of the temporal convolutional sub-network, performing multi-head attention operations with audio intermediate features as queries and visual features as keys and values, injecting lip movement information frame-by-frame into the audio separation path. The output of the cross-modal fusion separation module has a mask generation layer, which generates a target separation mask with a value range between 0 and 1 through Sigmoid activation. The target separation mask is then multiplied element-wise with the first feature representation to output the third feature representation.

[0104] The temporal decoding module is used to perform temporal decoding on the third feature representation to obtain the separated audio waveform of the target speech object. The temporal decoding module receives the third feature representation output by the cross-modal fusion separation module as input, and decodes the third feature representation frame by frame using a transposed convolutional filter bank symmetrically arranged with the temporal coding module. Adjacent decoded frames are then synthesized through overlapping addition to output the separated audio waveform of the target speech object. The transposed convolutional filter parameters of the temporal decoding module and the filter parameters of the temporal coding module are subject to the same loss function during training, but do not share weights.

[0105] During system operation, the temporal encoding module and the visual front-end module execute their respective input processing in parallel. After both are completed, the cross-modal fusion and separation module begins operation. After the cross-modal fusion and separation module completes its operation, the temporal decoding module performs the final waveform reconstruction. Feature tensors are transferred between modules through the storage space of the computing device. All operations are performed in the temporal encoding representation space, without introducing frequency domain transformation operations.

[0106] In one optional implementation, the cross-modal fusion separation module includes a temporal convolutional subnetwork and a cross-modal cross-attention subnetwork, wherein the cross-modal cross-attention subnetwork is positioned between adjacent repetition cycles of the temporal convolutional subnetwork. The temporal convolutional subnetwork comprises three repetition cycles, each containing eight dilated convolutional blocks, with the dilation factor increasing in each repetition. The process is exponentially increasing. The cross-modal cross-attention subnetwork consists of three layers, inserted after the output of the first, second, and third repetition cycles, respectively. Each layer employs four parallel attention heads for multi-head attention operations. The temporal convolutional subnetwork performs deep temporal modeling of the audio features, while the cross-modal cross-attention subnetwork is responsible for injecting visual guidance information into the audio separation path at each processing layer. The two work alternately to achieve layer-by-layer refinement of cross-modal fusion and separation.

[0107] Example 2

[0108] Optionally, the step of guiding the separation network to perform mask estimation of the first feature representation using the second feature representation through the cross-modal attention mechanism further includes a mask prior constraint mechanism guided by the sparsity of temporal coding. Specifically, the zero-value distribution pattern generated by ReLU activation in the first feature representation is used as a structural prior for mask estimation, and a mask zeroing constraint is directly applied to the coding positions known to be zero, so that the separation network only needs to perform effective mask estimation for non-zero coding positions.

[0109] The first feature represents After the ReLU activation in step S100, approximately 40% to 60% of the elements have a value of 0 (i.e., the corresponding filter response at that time frame is truncated to 0 as a non-positive value). These zero-value locations physically represent that the corresponding frequency component has no energy contribution at that moment, and the product of the zero value and any mask is always zero, regardless of the mask value. Therefore, the mask estimation of the separation network at these locations is redundant computation, meaning that its result does not affect the final third feature representation.

[0110] Define the encoding sparsity indicator matrix as follows:

[0111]

[0112] The mask prior constraint mechanism applies the sparse indicator matrix to the output of the mask generation layer:

[0113]

[0114] in The mask values ​​are used to separate the original estimated values ​​of the network. During the training phase, the constraint mechanism is implemented through gradient blocking, i.e., for... The position is not calculated to determine the gradient contribution of that position to the loss function, so that the parameter update of the separation network is driven only by the mask quality of the non-zero encoded position.

[0115] In conventional practices, the network performs a complete mask estimation operation on all encoding positions (including zero-value positions). The gradient signals of the network parameters at zero-value positions are mixed with the gradient signals at non-zero positions, resulting in two negative effects: First, although the mask estimation at zero-value positions does not change the final output during forward propagation (because the product of zero and any mask is always 0), it still generates non-zero gradient signals that are propagated back to the hidden layer parameters of the separation network during backpropagation. These gradient signals do not carry effective separation and discrimination information, constituting dilution noise for the effective gradient. Second, the cross-modal attention mechanism includes the zero-value time step in the denominator of the Softmax normalization when calculating attention weights, causing the attention share obtained by the effective time step to be diluted by the silent segment with no information. In this embodiment, the constraint mechanism explicitly applies mask zeroing to the zero-value positions during the forward propagation stage. More importantly, during the backpropagation stage, gradient blocking cuts off the gradient contribution path of the zero-value positions to the network parameters, so that the parameter update of the separation network is driven by the separation quality of the non-zero encoding positions (i.e., positions with effective audio energy). Simultaneously, the attention bias matrix excludes extremely sparse time steps from the effective set of attention queries, ensuring that the attention weight allocation of the cross-modal attention mechanism focuses on time periods containing effective speech activity, avoiding the interference of zero-feature patterns in silent segments as noise anchors on the spatial distribution of attention weights. This separates the effective decision space of the network from... The number of positions was reduced to approximately With a non-zero position, the network capacity is concentrated on the audio components that actually need to be judged, thereby improving the effective learning efficiency of mask estimation without increasing the number of model parameters.

[0116] Furthermore, the sparsity indicator matrix can also be used to construct the attention mask for the hidden layers of the separation network. Specifically, in the attention weight calculation of the cross-modal attention mechanism, for time steps where the encoded value is 0, the query vector at the corresponding position is set to zero or a very large negative bias is applied (making the Softmax output tend to 0) to prevent the zero-value position from interfering with the attention weight distribution as a "noise anchor point". The attention bias matrix is ​​as follows:

[0117]

[0118] in The sparsity threshold is set to [value]. (Rounded to 26), meaning that when more than 90% of the encoded channels at a given time step are zero, that time step is excluded from the attention query. Such extremely sparse time steps usually correspond to silent segments in audio, and their attention queries lack effective feature patterns; excluding them avoids introducing noisy attention.

[0119] For example, take time steps Assuming that the current moment is within the speech activity segment of the target speaker, 147 out of the 256 encoded channels are non-zero values ​​(sparseness rate 42.6%). This time step normally participates in attention calculation and mask estimation. (Take time step) Assuming that the moment is within the interval when both speakers are silent, only 18 out of 256 channels are non-zero values ​​(sparseness rate 93%). This time step is excluded from the attention query, and its mask value is directly set to zero by the constraint mechanism, eliminating the need for network inference. Globally, in a typical 4-second audio clip, approximately 5% to 15% of time steps are skipped due to extreme sparsity, correspondingly reducing the effective computational cost of the separation network.

[0120] Example 3

[0121] In this embodiment, the aforementioned process of guiding the separation network to perform mask estimation on the first feature representation using the second feature representation further includes a time-frequency boundary sharpening process based on the mask gradient field. Specifically, the target separation mask output by the separation network is post-processed as follows: discrete gradients are calculated along the time axis and channel axis respectively to construct a mask gradient field; the amplitude distribution of the gradient field is used to identify the boundary region between the target component and the interference component; a nonlinear sharpening transformation is applied to the mask values ​​within the boundary region to push the mask values ​​at the boundary towards the extreme directions of zero or one, reducing energy leakage caused by intermediate blur values. This sharpening transformation, as a parameter-free deterministic post-processing operation in the inference phase, is performed after the separation network completes mask estimation and before the element-wise multiplication of the mask and the encoded representation. It does not participate in the backpropagation process in the training phase and does not change the parameters of the separation network itself or the training objective.

[0122] Ideally, the target separation mask The mask should exhibit binarization characteristics, meaning the mask value at the location corresponding to the target component is close to 1, and the mask value at the location corresponding to the interference component is close to 0. However, in actual estimation, the mask value often remains in the middle range of 0.3 to 0.7 in the boundary region where the target component and the interference component intersect in the coding space. This results in the coding energy in this region being neither fully preserved nor fully suppressed, producing a trade-off effect between residual interference and target distortion.

[0123] The temporal gradient and channel gradient of the mask are:

[0124]

[0125]

[0126] The magnitude of the mask gradient field is:

[0127]

[0128] The gradient field magnitude is close to zero in the flat region of the mask (within the target region or the interference region), and takes a larger value in the boundary region of the target-interference transition. The boundary region indicator function is:

[0129]

[0130] Among them, the boundary detection threshold For locations marked as boundary regions, a nonlinear sharpening transformation is applied:

[0131]

[0132] Among them, sharpening index The sharpening transformation here is when... hour, Push the value toward 1 (e.g.) );when hour, Push the value towards 0 (e.g.) Transformation in Continuous at (both give a transformation of 0.5 themselves), in and The point is set as a fixed point to ensure that sharpening does not change the already determined extreme values.

[0133] It should be noted that the sharpening operation is applied only to the boundary regions marked by the gradient field. The mask value remains unchanged within flat regions to avoid information loss that may result from global sharpening. This is because the mask value is already close to 1 within the target component, making forced sharpening meaningless and potentially introducing quantization noise. Similarly, the mask value is already close to 0 within the interfering component, so further suppression is unnecessary. Therefore, applying sharpening only to the boundary regions can effectively address the energy leakage bottleneck.

[0134] It should be noted that although the mask gradient field sharpening is performed after the inference stage, compared with simple global thresholding or temperature scaling methods, this embodiment only applies a transformation to the boundary region marked by the gradient field without interfering with the converged flat region. The sharpening direction is automatically determined by the relative magnitude of the mask value and 0.5, without introducing additional hyperparameters or external signals. Furthermore, all operations of the sharpening transformation (gradient calculation, magnitude calculation, threshold determination, and piecewise function application) do not contain learnable parameters and do not depend on the training data distribution, and can be directly deployed on any trained mask estimation model.

[0135] Furthermore, unlike conventional post-processing methods such as simple global thresholding or temperature scaling, the time-frequency boundary sharpening of the mask gradient field in this embodiment can transfer the edge detection paradigm in image processing to the time-frequency coding space of the audio separation mask. By using the gradient field of the mask itself as a boundary discovery tool, a self-organizing path can be established within the mask estimation process, thereby achieving position selectivity (only the boundary region) and orientation adaptability (the sharpening direction is automatically determined by the 0.5 boundary) without introducing additional learnable parameters or external signals.

[0136] For example, take channel Time step to The mask value sequence is This presents a boundary pattern transitioning from the target region to the interference region. The temporal gradients are respectively... ,exist to The absolute value of the gradient is relatively large between these points. Assuming the channel gradient is close to zero at this location, the gradient field magnitude is approximately... .by The threshold for boundary detection, to The region is marked as a boundary area. A sharpening transformation is applied to the boundary area: Place After sharpening (Increased from 0.88 to 0.938, closer to 1); Place After sharpening (Increased from 0.71 to 0.843); Place After sharpening (From 0.45 down to 0.258, closer to 0); Place After sharpening (Downgraded from 0.12 to 0.062). The sharpened mask sequence becomes... The boundary transition is steeper, resulting in a clearer target-interference boundary and reducing energy leakage in the ambiguous middle region.

[0137] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0138] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware.

[0139] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A new media temporal audio separation method based on visual features, characterized in that, include: The mixed audio waveform is time-domain encoded to obtain the first feature representation; Visual features are extracted from the synchronized video containing the target speech object and time-aligned to obtain a second feature representation with the same frame rate as the first feature representation. The first feature representation and the second feature representation are input into a separation network, and the separation network is guided by the second feature representation to perform mask estimation on the first feature representation through a cross-modal attention mechanism to obtain a third feature representation; The third feature representation is then temporally decoded to obtain the separated audio waveform of the target speech object.

2. The method according to claim 1, characterized in that, The step of performing time-domain encoding on the mixed audio waveform to obtain a first feature representation includes: The mixed audio waveform is subjected to frame-by-frame convolution operation using a one-dimensional convolution filter bank; The output of the convolution operation is subjected to nonlinear activation processing to obtain the first feature representation.

3. The method according to claim 2, characterized in that, The convolution stride of the one-dimensional convolutional filter bank is half the filter length; the temporal decoding of the third feature representation to obtain the separated audio waveform of the target speech object includes: The third feature representation is decoded frame by frame by a transposed convolutional filter group that is symmetrically set with the one-dimensional convolutional filter group; The separated audio waveform is determined based on the decoded adjacent frames.

4. The method according to claim 1, characterized in that, The step of extracting visual features from synchronized video containing the target speech object and performing time alignment processing to obtain a second feature representation with the same frame rate as the first feature representation includes: Spatiotemporal features are extracted from image sequences of the lip region of the target speech object using a spatiotemporal convolutional network. The extracted spatiotemporal features are encoded into lip motion embedding sequences using a temporal modeling network; The lip motion embedding sequence is interpolated and upsampled along the time axis so that the number of frames in the upsampled sequence is consistent with the number of frames in the first feature representation, thereby obtaining the second feature representation.

5. The method according to claim 1, characterized in that, The separation network comprises multiple stacked dilated temporal convolutional blocks, the dilation factor of which increases exponentially in each repetition cycle.

6. The method according to claim 5, characterized in that, The cross-modal attention mechanism is set between adjacent repetition cycles of the dilated temporal convolution block; The cross-modal attention mechanism uses the intermediate features after the first feature representation has been processed by the separation network as the query, and the second feature representation as the key and value, to perform multi-head attention operations.

7. The method according to claim 1, characterized in that, The step of guiding the separation network to perform mask estimation on the first feature representation using the second feature representation through a cross-modal attention mechanism to obtain the third feature representation includes: Generate a mask matrix; The third feature representation is determined based on the mask matrix and the first feature representation.

8. A new media temporal audio separation system based on visual features, characterized in that, include: The time-domain coding module is used to perform time-domain coding on the mixed audio waveform to obtain the first feature representation; The visual front-end module is used to extract visual features from the synchronized video containing the target speech object and perform time alignment processing to obtain a second feature representation that is consistent with the frame rate of the first feature representation. The cross-modal fusion separation module is used to guide the separation network to perform mask estimation on the first feature representation using the second feature representation through a cross-modal attention mechanism, so as to obtain the third feature representation; The time-domain decoding module is used to perform time-domain decoding on the third feature representation to obtain the separated audio waveform of the target speech object.

9. The system according to claim 8, characterized in that, The cross-modal fusion and separation module includes a temporal convolutional sub-network and a cross-modal cross-attention sub-network, wherein the cross-modal cross-attention sub-network is positioned between adjacent repetition cycles of the temporal convolutional sub-network.

10. The system according to claim 8, characterized in that, The visual front-end module includes a spatiotemporal convolutional sub-network, a temporal modeling sub-network, and a temporal alignment sub-network. The spatiotemporal convolutional sub-network is used to extract spatiotemporal features from the image sequence of the lip region of the target speech object. The temporal modeling sub-network is used to encode the extracted spatiotemporal features into a lip motion embedding sequence. The temporal alignment sub-network is used to interpolate and upsample the lip motion embedding sequence along the time axis so that the number of frames in the upsampled sequence is consistent with the number of frames represented by the first feature.