Unsupervised mixed audio cross-modal separation method, system, device and storage medium
By employing an unsupervised hybrid audio cross-modal separation method, and utilizing loss function training based on in-video sound source object detection and visual guidance, the challenge of hybrid audio separation under unsupervised conditions is solved, achieving higher audio separation accuracy and independence.
Patent Information
- Application Number
- CN202411558075.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-11-04
AI Technical Summary
Existing hybrid audio separation techniques struggle to effectively separate multiple sound source signals under unsupervised conditions, especially in complex environments with noise and equipment quality issues. Furthermore, existing separation models rely on scarce real-source supervision information.
An unsupervised hybrid audio cross-modal separation method is adopted, which involves in-video sound source object detection, audio and video data preprocessing, cross-modal audio-visual semantic alignment learning, U-Net initial separation, independent adversarial learning, and visual guidance for unsupervised audio separation. Visual target objects are used as virtual supervision labels, and various loss functions are set to train the U-Net audio decoder to generate independent separated audio signals.
It improves the accuracy and independence of audio separation under unsupervised conditions. By balancing the correlation and independence of sound sources through independent adversarial learning, it enhances the independence and accuracy of the separated audio signals.
Smart Images

Figure CN119541523B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of audio processing, and particularly relates to an unsupervised mixed audio cross-modal separation method, system, device and storage medium. BACKGROUND
[0002] Mixed audio separation technology is an advanced technology for processing and separating multiple sound sources, aiming to extract independent sound signals from mixed audio, which has wide application in speech recognition, music processing, audio editing, audio enhancement and other fields. However, mixed audio separation technology also faces many challenges. First, audio signals are complex time series signals, and how to separate them without destroying the original audio quality is a difficult problem. Second, different sound source signals may have different spectral characteristics and temporal characteristics, which makes the separation process more complex. In addition, there may be various interference factors in practical applications, such as environmental noise, device quality, etc., which will affect the audio separation effect.
[0003] With the rapid development of artificial intelligence technology, audio separation technology has also made significant progress. Based on deep learning, the audio separation method no longer relies on traditional signal processing technology and pre-defined statistical model, but trains the separation model by collecting a large amount of data, learns the internal law of audio signals, and thus realizes more accurate sound source separation. However, existing separation models mostly rely on real single sound source as supervision information, while supervision information is often scarce in practical applications. SUMMARY
[0004] The main purpose of the present application is to overcome the shortcomings and deficiencies of the prior art, and to provide an unsupervised mixed audio cross-modal separation method, system, device and storage medium.
[0005] In order to achieve the above purpose, the present application adopts the following technical scheme:
[0006] An unsupervised mixed audio cross-modal separation method, comprising the following steps:
[0007] Video sound source object detection, for the video containing multiple sound sources of audio mixed signal, using target detection algorithm to extract the target object and visual features in the video;
[0008] Audio-video data preprocessing, processing the audio mixed signal to obtain mixed audio features;
[0009] Cross-modal audio-visual semantic alignment learning, including audio-visual semantic consistency learning and audio-visual cross-attention learning; through audio-visual semantic consistency learning, the semantic alignment is carried out and the semantic expression ability of audio-visual two modalities is enhanced; through audio-visual cross-attention learning, the feature semantics of audio-visual modalities are further refined;
[0010] Unsupervised mixed audio preliminary separation, U-Net is used as an audio decoder to obtain a plurality of separated audio signals;
[0011] Independent adversarial learning, a signal library is constructed by randomly extracting signals from the audio signals separated from different mixed audios, a loss function is set to narrow the joint distribution between the audio decoder generated signals and the samples in the randomly selected signal library, so as to train the U-Net audio decoder to generate independent separated audio signals;
[0012] Vision-guided unsupervised audio separation, the visual target object is used as a virtual supervision label, an audio-visual semantic consistency discrimination loss function is set, the U-Net audio decoder is trained to generate audio signals corresponding to the target object, and the mixed audio is accurately separated.
[0013] The application also includes an unsupervised mixed audio cross-modal separation system, which uses the unsupervised mixed audio cross-modal separation method provided by the application, and the system includes a sound source object detection module, an audio-video data preprocessing module, a cross-modal audio-visual semantic alignment learning module, an unsupervised mixed audio preliminary separation module, an independent adversarial learning module, and an unsupervised audio separation module.
[0014] The sound source object detection module uses a target detection algorithm to extract target objects and visual features in a video;
[0015] The audio-video data preprocessing module processes the audio mixed signal to obtain mixed audio features;
[0016] The cross-modal audio-visual semantic alignment learning module performs semantic alignment and enhances the semantic expression ability of the audio and visual modalities through audio-video semantic consistency learning; through audio-visual cross-attention learning, the feature semantics of the audio and visual modalities are further refined;
[0017] The unsupervised mixed audio preliminary separation module uses U-Net as an audio decoder to obtain a plurality of separated audio signals;
[0018] The independent adversarial learning module is used to construct a signal library by randomly extracting signals from the audio signals separated from different mixed audios, set a loss function to narrow the difference between the joint distribution of the audio decoder generated signals and the joint distribution of the samples in the randomly selected signal library, so as to train the U-Net audio decoder to generate independent separated audio signals;
[0019] The unsupervised audio separation module is used to use the visual target object as a virtual supervision label, set an audio-visual semantic consistency discrimination loss function, train the U-Net audio decoder to generate audio signals corresponding to the target object, and accurately separate the mixed audio.
[0020] The application further comprises a computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the unsupervised mixed audio cross-modal separation method provided by the application when executing the computer program.
[0021] The application further comprises a computer readable storage medium storing a computer program, which, when executed by a processor, implements the unsupervised mixed audio cross-modal separation method provided by the application.
[0022] Compared with the prior art, the application has the following advantages and beneficial effects:
[0023] 1. The application proposes an independent adversarial learning algorithm to balance the correlation and independence of the separated sound sources after unsupervised separation, thereby improving the accuracy of unsupervised separation; the application uses two processing methods to increase the independence of the separated audio signals; first, a mutual information loss function is set to minimize the mutual information of the separated audio to reduce the information overlap between the separated audio signals, thereby improving the independence of the separated audio; second, a reference library is formed by randomly selecting separated signals from different mixed audios; due to random selection, the audio signals in the library are relatively independent; the feature distribution of the separated audio is projected into the feature distribution of the audio signals in the reference library to improve the mutual independence of the separated audio signals. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 is a flowchart of the method of the application;
[0025] Figure 2 is a schematic diagram of cross-modal audio-visual speech semantic alignment in the application;
[0026] Figure 3 is a schematic diagram of unsupervised mixed audio preliminary separation in the application;
[0027] Figure 4 is a schematic diagram of independent adversarial learning in the application. DETAILED DESCRIPTION
[0028] The application will be described in further detail below with reference to the embodiments and the accompanying drawings, but the embodiments of the application are not limited thereto.
[0029] EMBODIMENT
[0030] As shown in Figure 1 , the application is an unsupervised mixed audio cross-modal separation method, comprising the following steps:
[0031] S1, video-in sound source object detection, for a video containing a plurality of sound sources, a target detection algorithm is used to extract the target objects and visual features in the video from the audio mixed signal;
[0032] S2. Audio and video data preprocessing: Processing the mixed audio signal to obtain mixed audio features; taking the separation of mixed audio from two videos as an example, this embodiment specifically involves:
[0033] Let the audio signals in the two videos be A. i1 A i2 The audio mix signal from the two videos is set as A. mix =A i1 +A i2 ;
[0034] Obtain the target object in the video and visual features Then, the mixed audio signal A is obtained through short-time Fourier transform (STFT) and Mel-frequency transform. mix Mel spectrum S mix and S mix Input audio encoder to obtain mixed audio features
[0035] S3, cross-modal audio-visual semantic alignment learning, including audio-video semantic consistency learning and audio-visual cross-attention learning;
[0036] Unlike traditional methods that directly align audio and visual features by providing real semantic labels, this embodiment uses an unsupervised learning approach. Through multimodal semantic consistency learning, it performs modal semantic alignment and enhances the semantic expressive power of both audio and visual modalities. Figure 2 As shown, the specific steps of audio-video semantic consistency learning are:
[0037] Fusion of target object visual features and audio mixing features Obtain fusion features
[0038] Contrastive learning is used to enhance the accuracy of visual features by fusing features corresponding to objects with the same sound source. Visual characteristics of the sound source object A positive sample pair is formed, and other combinations are treated as negative sample pairs. InfoNCE is used as the contrastive learning loss function to optimize the visual feature representation. The contrastive learning loss function is:
[0039]
[0040] Where τ is the temperature hyperparameter, Represents the negative sample space;
[0041] Audio-visual cross-attention learning specifically involves:
[0042] likeFigure 2 The enhanced fusion feature and visual features are used as the Query vector for interactive cross-modal attention learning, and the semantic consistent features of the audio-visual modal are further refined through spatial attention and channel attention learning to improve the semantic accuracy of feature expression, promote modal information fusion, and output the attention-enhanced fusion feature f c . The spatial attention aims to enhance the feature expression of the key spatial region. In the calculation process, the original feature is first pooled in the channel dimension, the multi-channel features are spliced, and the convolution and activation calculation are performed on the spliced features to generate a weight mask for each spatial position and weighted output, thereby enhancing the specific target region of interest and weakening the irrelevant background region. In the channel attention calculation, the spatial representation of the feature is compressed through average pooling to retain the channel expression information, and the weight coefficients of each feature channel are automatically obtained through the attention network learning method to strengthen the important channel features and suppress the non-important features.
[0043] S4, unsupervised mixed audio preliminary separation, using U-Net as an audio decoder to obtain a plurality of separated audio signals; as shown in Figure 3 , specifically:
[0044] U-Net is used as an audio decoder to generate a corresponding separation mask for the sound source target object By multiplying the mask and the sound mixed spectrum, the audio mel spectrum of the corresponding sound source target object is obtained The calculation formula is as follows:
[0045]
[0046] The spectrum is inversely transformed from the frequency domain to the time domain to obtain the time domain audio signal In order to avoid information loss in the audio separation process, a loss function L sep is set to evaluate the integrity and quality of the separated audio information generated by the U-Net audio decoder by setting the signal-to-noise ratio (SNR), and the loss function L sep is specifically as follows:
[0047]
[0048] Wherein, A i is the real data of the mixed audio in the video, and A i is obtained after unsupervised audio separation to obtain a plurality of separated audio signals The mixed audio obtained by integrating the separated audio signals is
[0049] the difference between the real mixed audio data A i can reflect the performance of the separation model. Therefore, by maximizing the loss L sep during the training of the U-Net audio decoder, the audio mixture A i and have the highest similarity, so that the loss of audio information during the separation process is zero or minimal.
[0050] S5, independent adversarial learning, constructing a signal library by randomly selecting signals from the audio signals separated from different mixed audios, setting a loss function to narrow the difference between the joint distribution of the audio decoder generated signals and the joint distribution of the samples in the randomly selected signal library, to train the U-Net audio decoder to generate independent separated audio signals;
[0051] The accuracy of the audio samples obtained after unsupervised preliminary separation needs to be evaluated due to the lack of real audio data as a reference, so the independent adversarial learning algorithm is also performed in this embodiment to balance the correlation and independence of the separated audio sources, thereby improving the accuracy of unsupervised separation.
[0052] In this embodiment, as shown in Figure 4 , the independent adversarial learning specifically includes:
[0053] For the separated audio signals , let denote two signals separated from one mixed audio, denote the joint probability distribution of the two signals.
[0054] A signal library is constructed by randomly selecting signals from the audio signals separated from different mixed audios, according to the law of large numbers, when the number of random selections is large enough, the signal distribution in the signal library is approximately independent distribution; the signal library is defined as Two signals are randomly selected from the signal library, and wherein, m1,m2∈{1,…,M},m1≠m2, the joint distribution of the two is defined as When the signal library samples are random enough, the two randomly selected signals are independent of each other, and the joint distribution is equivalent to the product of the marginal distribution of each signal, which is represented as follows:
[0055]
[0056] wherein, P m (·) represents the marginal distribution.
[0057] The independent adversarial learning aims to guide the audio decoder to generate relatively independent audio signals, realizing mixed audio separation. The joint distribution between the samples in the random sampling signal library is approximated to an independent distribution, which is taken as the target distribution, and a loss function L is set G to narrow the joint distribution of the audio decoder generated signals and the joint distribution between the samples in the random sampling signal library, so as to train the U-Net audio decoder to generate independent separated audio signals The loss function L G is expressed as:
[0058]
[0059] Wherein, the function T can adopt Kullback-Leibler (KL) divergence, Jensen-Shannon (JS) divergence, Wasserstein distance and Pearson χ 2 distance, etc.
[0060] In the audio separation process, in order to prevent the loss of audio information, the consistency of mixed audio before and after audio separation is maintained, specifically the separated audio signals are re-mixed, and the distribution of the re-mixed audio is aligned with the distribution of the original mixed audio in the same video, and the modified loss function L G is expressed as:
[0061]
[0062] The independent adversarial learning method mainly focuses on the independence of the separated audio signals, which may lead to over-separation, and may not cover the scenario of partial overlap of audio spectrum. Therefore, the audio separation operation needs to balance between the independence and integrity of the separated audio signals. In order to avoid over-separation, an independent decision loss function L D is set to expand the distribution difference between the joint distribution of the audio decoder generated signals and the joint distribution of the samples in the random sampling signal library, and to maintain a proper balance between the independence and correlation of the separated audio signals.
[0063] The decision loss function L D is specifically:
[0064]
[0065] Wherein, Φ represents the full convolution and activation operation on the distribution of the generated audio signals and the random sampling signals; the decision loss emphasizes the difference between the two distributions, so as to weaken the independence of the separated signals.
[0066] The loss function LG and L D The U-Net audio decoder is alternately trained for adversarial learning. D is used to expand the distribution of the difference between the element pairs, while the loss L G is used to minimize the probability distribution distance between them, for training the cross-modal audio separation decoder to generate mutually independent audio signals. Finally, the best balance between independence and correlation of mixed audio separation is achieved through adversarial training.
[0067] S6, visually guided unsupervised audio separation, taking the visual target object as a virtual supervision label, setting an audio-visual semantic consistency discriminant loss function, training the U-Net audio decoder to generate an audio signal corresponding to the target object, and realizing accurate separation of mixed audio; in this embodiment, specifically:
[0068] Taking the visual target object as a virtual supervision label, setting an audio-visual semantic consistency discriminant loss function L con , training the audio decoder to generate an audio signal corresponding to the target object, and accurately separating the mixed audio;
[0069] The audio-visual semantic consistency discriminant loss function L com is specifically:
[0070]
[0071] Wherein, the function represents a full-connection convolution and a category mapping, which is used for category judgment of visual information and separated audio information;
[0072] The definition of the complete loss function L used for training the U-Net audio decoder is:
[0073] L=L av +λ1L sep +λ2(L G +L D )+λ3L con ;
[0074] Wherein, λ1, λ2, λ3 represent loss function parameters.
[0075] In order to better illustrate the technical effect of the method of the present application, two large-scale audio-visual data sets MUSIC and VGGSound are used, four kinds of comparison methods are selected for comparison experiments, and source distortion ratio (SDR), source interference ratio (SIR) and source artifact ratio (SAR) are selected as evaluation standards; the four kinds of comparison methods are:
[0076] Sound-of-Pixels algorithm, see the document “H Zhao, C Gan, A Rouditchenko, C Vondrick, J H. McDermott and A Torralba. The sound of pixels. In Proceedings of the European Conference on Computer Vision (ECCV), pages 587-604, 2018.”
[0077] Co-Separation algorithm, see the document “R Gao and K Grauman. Co-separating sounds of visual objects. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV), pages 3878-3887, 2019.”
[0078] MP-Net algorithm, see the document “X Xu, B Dai and D Lin. Recursive visual sound separation using minus-plus net. In Proceedings of the IEEE / CVF International Conference on Computer Vision 691 (ICCV), pages 882-891, 2019.”
[0079] FCSN algorithm, see the document “S Ma, Y Ji, X Xu, and X Zhu. Vision-guided music source separation via a fine-grained cycle-separation network. In Proceedings of the ACM International Conference on Multimedia (ACM MM), pages 4202-4210, 2021.”
[0080] The results are shown in Tables 1, 2 and 3; Table 1 is a comparison table of unsupervised audio separation results of the present application and the comparative method on the mixed audio-video data set VGGSound. Table 2 is a comparison table of unsupervised audio separation accuracy of the present application and the comparative method on the mixed audio-video data set MUSIC. Among them, * algorithm + IAL means that the independent adversarial learning (IAL) of the present application is used as a plug-in combined with the existing algorithm, and the verification result of the unsupervised audio separation task. Table 3 is a comparison table of supervised audio separation accuracy of the present application and the comparative method on the mixed audio-video data set MUSIC.
[0081] Table 1 Comparison table of unsupervised audio separation results of the present application and the comparative method on the mixed audio-video data set VGGSound
[0082]
[0083] Table 2 Comparison table of unsupervised audio separation accuracy of the present application and the comparative method on the mixed audio-video data set MUSIC
[0084]
[0085] Table 3 Comparison table of supervised audio separation accuracy of the present application and the comparative method on the mixed audio-video data set MUSIC
[0086]
[0087] As shown in Tables 1, 2 and 3, the method of the present application has obvious advantages compared with these algorithms in the evaluation criteria SDR, SIR and SAR of sound source separation. The experimental results of the independent adversarial learning method used as an algorithm plug-in show that the independent adversarial learning has a significant improvement effect on audio separation. This proves the effectiveness of the audio separation algorithm based on independent adversarial learning proposed in the present application.
[0088] In another embodiment, an unsupervised mixed audio cross-modal separation system is provided, which uses the unsupervised mixed audio cross-modal separation method of the above-mentioned embodiment. The system includes a sound source object detection module, an audio-video data preprocessing module, a cross-modal audio-video semantic alignment learning module, an unsupervised mixed audio preliminary separation module, an independent adversarial learning module, and an unsupervised audio separation module.
[0089] The sound source object detection module uses a target detection algorithm to extract target objects and visual features in the video;
[0090] The audio-video data preprocessing module processes the audio mixed signal to obtain mixed audio features;
[0091] The cross-modal audio-visual semantic alignment learning module performs semantic alignment and enhances the semantic expression capability of the audio and visual modalities through audio-visual semantic consistency learning; and further refines the feature semantics of the audio and visual modalities through audio-visual cross-attention learning.
[0092] The unsupervised mixed audio preliminary separation module adopts a U-Net as an audio decoder to obtain a plurality of separated audio signals.
[0093] The independent adversarial learning module is configured to construct a signal library by randomly extracting signals from the audio signals separated from different mixed audios, set a loss function to narrow the difference between the joint distribution of the audio decoder generated signals and the joint distribution of the samples in the randomly extracted signal library, and train the U-Net audio decoder to generate independent separated audio signals.
[0094] The unsupervised audio separation module is configured to take the visual target object as a virtual supervision label, set an audio-visual semantic consistency discrimination loss function, train the U-Net audio decoder to generate audio signals corresponding to the target object, and accurately separate the mixed audio.
[0095] In another embodiment, a computer device is provided, which includes a memory and a processor, the memory stores a computer program, and the processor implements the unsupervised mixed audio cross-modal separation method of the above-mentioned embodiments when executing the computer program.
[0096] The hardware entities of the computer device include a processor, a memory and a communication interface; wherein the processor generally controls the overall operation of the computer device; the communication interface is configured to enable the computer device to communicate with other terminals or servers through a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed by the processor and data to be processed or having been processed by each module in the computer device (including but not limited to image data, audio data, voice communication data and video communication data), which can be realized by a FLASH or a RAM (Random Access Memory).
[0097] The processor, the communication interface and the memory can transmit data through a bus, which can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together.
[0098] In another embodiment, a computer readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the unsupervised mixed audio cross-modal separation method of the above-mentioned embodiments is implemented.
[0099] The storage medium can be transitory or non-transitory. Illustratively, the storage medium includes, but is not limited to, a variety of media that can store computer program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0100] Illustratively, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0101] It is also necessary to note that in the present specification, terms such as "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusions, so that processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0102] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An unsupervised mixed-audio cross-modal separation method, characterized in that, The method comprises the following steps: Video sound source object detection, a target detection algorithm is used to extract target objects and visual features in a video containing a plurality of sound sources; Audio-video data preprocessing, a mixed audio feature is obtained by processing an audio mixed signal; Cross-modal audio-visual semantic alignment learning, including audio-visual semantic consistency learning and audio-visual cross-attention learning; semantic alignment is performed through audio-visual semantic consistency learning, and the semantic expression capability of the audio and visual modalities is enhanced; through audio-visual cross-attention learning, the feature semantics of the audio and visual modalities are further refined; Unsupervised mixed audio preliminary separation, a U-Net is used as an audio decoder to obtain a plurality of separated audio signals; Independent adversarial learning, by randomly extracting signals from the audio signals separated from different mixed audio to construct a signal library, setting the loss function L G To narrow the joint distribution between the joint distribution of the audio decoder generated signal and the sample in the randomly sampled signal library, train the U-Net audio decoder to generate mutually independent separated audio signals; Vision-guided unsupervised audio separation, a visual target object is used as a virtual supervision label, an audio-visual semantic consistency discrimination loss function is set, a U-Net audio decoder is trained to generate an audio signal corresponding to the target object, and the mixed audio is accurately separated; Vision-guided unsupervised audio separation is specifically: A visual target object is used as a virtual supervision label, and an audio-visual semantic consistent discrimination loss function L is set con An audio decoder is trained to generate an audio signal corresponding to the target object, so as to accurately separate the mixed audio. Audio-visual semantic consistent discriminant loss function L con Specifically: wherein the function represents a full connection convolution and a class mapping, used for classifying the visual information and the separated audio information; represents a sound source target object, represents a separated audio signal; The definition of the complete loss function L used to train the U-Net audio decoder is: L = L av + λ1L sep + λ2(L G + L D )+ λ3L con ; wherein λ1, λ2, λ3 represent loss function parameters; L av is a contrastive learning loss function in the audio-visual semantic consistency learning process; L sep is a loss function, and the loss function L sep to evaluate the integrity and quality of the separated audio information generated by the U-Net audio decoder by setting the signal-to-noise ratio; L D is an independent decision loss function, which is used to expand the distribution difference between the joint distribution of the audio decoder generated signal and the sample in the random sampling signal library.
2. The unsupervised mixed-audio cross-modal separation method according to claim 1, wherein, Audio-video data preprocessing is specifically: The audio signals in the plurality of videos are denoted as A i1 , i2 , …, A in The audio mix signal in the plurality of videos is denoted as A mix = A i1 + A i2 + … + A in ; In obtaining a target object in a video and visual features After that, the mel spectrum S mix of the mixed audio signal A mix is obtained by short-time Fourier transform and mel spectrum transform, and S mix is input into an audio encoder to obtain mixed audio features 3. The unsupervised mixed-audio cross-modal separation method according to claim 2, wherein, Audio-visual semantic consistency learning is specifically: Fusing target object visual features And audio mixing features Get fused features Adopt contrastive learning to enhance the accuracy of visual features, and the fusion features corresponding to the same sound source object and the visual features of the sound source object Other combinations are used as negative sample pairs, and InfoNCE is used as the contrastive learning loss function to optimize the visual feature expression; the contrastive learning loss function is: where τ is a temperature hyperparameter, denotes the negative sample space; Audio-visual cross-attention learning is specifically: Enhanced fusion features and visual features The interactive is used as a query vector, interactive cross-modal attention learning is carried out, and the semantic consistent features of audio-visual modalities are further refined through spatial attention and channel attention learning, the semantic accuracy of feature expression is improved, the modal information fusion is promoted, and the enhanced fusion features f are output c .
4. The unsupervised mixed-audio cross-modal separation method according to claim 3, wherein, Unsupervised mixed audio preliminary separation is specifically: U-Net as an audio decoder to generate an audio mel-spectrogram of a corresponding sound source target object a corresponding separation mask by multiplying the mask with the sound mixed spectrum to obtain an audio mel-spectrogram of a corresponding sound source target object is represented as: spectrum performing inverse frequency-to-time domain transformation to obtain a time domain audio signal In order to avoid information loss in the audio separation process, a loss function L is set sep The integrity and quality of the separated audio information generated by the U-Net audio decoder are evaluated by setting the signal-to-noise ratio, and the loss function L sep Specifically: Wherein, A i is the mixed audio real data in the video, A i After unsupervised audio separation, a plurality of separated audio signals are obtained The separated audio signals are integrated, and the mixed audio obtained is 5. The unsupervised mixed-audio cross-modal separation method according to claim 4, wherein, Independent adversarial learning is specifically: for the resulting post-separation audio signal with denoting two signals separated from a mixed audio, denoting the joint probability distribution of the two signals; A signal library is constructed by randomly sampling signals from the separated audio signals of different mixed audio, according to Bernoulli's law of large numbers, when the number of random selection is large enough, the signal distribution in the signal library is approximately independent distribution; the signal library is defined as Two signals are randomly selected from the signal library, and wherein m1, m2∈{1,…,M}, m1≠m2, the joint distribution of the two is defined as When the signal library sample is random enough, the two randomly selected signals are independent of each other, and the joint distribution is equivalent to the product of the marginal distribution of the two signals, and is expressed as follows: where P m (·) denotes the edge distribution; A loss function L is set G to train the U-Net audio decoder to generate mutually independent separated audio signals A loss function L G is represented as: where the function T employs one of the Kullback-Leibler divergence, the Jensen-Shannon divergence, the Wasserstein distance, and the Pearson χ 2 distance In the audio separation process, in order to prevent the loss of audio information, the consistency of the mixed audio before and after the audio separation is maintained, specifically the separated audio signals are remixed, and the distribution of the remixed audio is aligned with the distribution of the original mixed audio in the same video, and the loss function L of the modified loss function G is represented as:
6. The unsupervised mixed-audio cross-modal separation method according to claim 5, wherein, To avoid over-separation, a separate decision loss function L is set D To enlarge the distribution difference between the joint distribution of the audio decoder generated signal and the joint distribution of the samples in the random drawn signal pool, a proper balance between the independence and correlation of the separated audio signals is maintained; Decision loss function L D Specifically: Where Φ represents the distribution of both the generated audio signal and the randomly selected signal. Perform full convolution and activation operations; the decision loss emphasizes the distribution of both. The differences weaken the independence of the separated signals.
7. An unsupervised mixed-audio cross-modal separation system, characterized in that, The system adopts the unsupervised mixed audio cross-modal separation method of any one of claims 1-6, and the system comprises a sound source object detection module, an audio-video data preprocessing module, a cross-modal audio-visual semantic alignment learning module, an unsupervised mixed audio preliminary separation module, an independent adversarial learning module, and an unsupervised audio separation module; The sound source object detection module extracts target objects and visual features in a video by using a target detection algorithm; The audio-video data preprocessing module processes a mixed audio signal to obtain a mixed audio feature; The cross-modal audio-visual semantic alignment learning module performs semantic alignment through audio-visual semantic consistency learning and enhances the semantic expression capability of the audio and visual modalities; through audio-visual cross-attention learning, the feature semantics of the audio and visual modalities are further refined; The unsupervised mixed audio preliminary separation module uses a U-Net as an audio decoder to obtain a plurality of separated audio signals; The independent adversarial learning module is used to construct a signal library by randomly extracting signals from the audio signals separated from different mixed audios, set a loss function to reduce the difference between the joint distribution of the signals generated by the audio decoder and the joint distribution of the samples in the randomly selected signal library, so as to train the U-Net audio decoder to generate independent separated audio signals; The unsupervised audio separation module uses a visual target object as a virtual supervision label, sets an audio-visual semantic consistency discrimination loss function, trains a U-Net audio decoder to generate an audio signal corresponding to the target object, and accurately separates the mixed audio.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method of any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Cross-modal voice separation method and system without visual information during testing
CN116978399A
KR1020391380000B1