Training method, electronic device and storage medium for audio-visual speech separation model
Quantitative tuning training of audio-visual speech separation model through the cross-direction multiplier method, solving the application problem of multimodal systems on low-resource devices, and achieving efficient speech separation performance of lightweight models on low-resource devices.
Patent Information
- Application Number
- CN202211573033.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-08
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-12-08
AI Technical Summary
Multimodal audio-visual speech separation systems require a large amount of computing resources and are difficult to apply to low-resource devices. In the prior art, STE models and conductivity quantizers have problems such as incomplete optimization goals as expected or training is difficult to converge.
The cross-direction multiplier method is used to quantize and tune the audio-visual speech separation model, and the cross-modal loss is mixed and accurate quantization is performed to reduce the calculation and storage costs. Combining the characteristics of different modal sensitivity to quantization, a lightweight audio-visual speech separation model is trained.
It realizes the application of lightweight audio-visual voice separation model on low-resource devices, maintains excellent voice separation performance, and balances the computing volume and performance.
Smart Images

Figure CN116312607B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent speech, and in particular to a training method, electronic device and storage medium for a video-audio speech separation model. Background Art
[0002] With the advancement of speech technology, multimodal audio-visual speech separation systems have demonstrated superior speech separation performance compared to pure speech processing systems. However, multimodal audio-visual speech separation systems require significant computational resources. On the one hand, audio-visual speech separation systems require a large number of parameters to model the modalities and their correlations, which increases memory usage. On the other hand, fusing information from both modalities requires computing larger feature maps, which in turn requires more floating-point operations.
[0003] In the process of implementing the present invention, the inventors discovered that there are at least the following problems in the related art:
[0004] Multimodal audio-visual speech separation systems require extensive computing resources, hindering their application on low-resource devices. Typically, to apply multimodal audio-visual speech separation systems to low-resource devices, the following methods are used: 1. STE (straight-through estimator). The STE model uses a quantizer during forward computation, but ignores the gradients provided by the quantizer during backpropagation, resulting in the optimization objective not being exactly the same as expected. 2. Differentiable quantizers use complex differentiable functions to approximate step functions for quantization, allowing gradients to be properly backpropagated. However, these methods can easily fail to converge to an optimal solution during training. Summary of the Invention
[0005] In order to at least solve the problem that the multimodal audio-visual speech separation system in the prior art is difficult to apply to low-resource devices. In a first aspect, an embodiment of the present invention provides a training method for an audio-visual speech separation model, comprising:
[0006] Inputting mixed training audio of multiple speakers into an audio-visual speech separation model to obtain predicted spectrograms of the multiple speakers;
[0007] Determining the predicted speaker audiovisual features of the predicted spectrogram and the reference speaker audiovisual features of the reference spectrogram of the mixed training audio;
[0008] Based on the cross-modal loss determined by the predicted speaker audio-visual features and the reference speaker audio-visual features, the audio-visual speech separation model is trained with mixed precision quantization conditions using the cross-directional multiplier method to obtain a lightweight audio-visual speech separation model.
[0009] In a second aspect, an embodiment of the present invention provides a training system for a video-audio speech separation model, comprising:
[0010] a prediction program module, configured to input mixed training audio of multiple speakers into an audio-visual speech separation model to obtain predicted spectrograms of the multiple speakers;
[0011] an audiovisual feature determination program module, configured to determine the audiovisual features of the predicted speaker of the predicted spectrogram and the audiovisual features of the reference speaker of the reference spectrogram of the mixed training audio;
[0012] A lightweight training program module is used to train the audio-visual speech separation model under mixed precision quantization conditions using the cross-modal loss determined based on the predicted speaker audio-visual features and the reference speaker audio-visual features through a cross-directional multiplier method to obtain a lightweight audio-visual speech separation model.
[0013] According to a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the method for training an audio-visual speech separation model of any embodiment of the present invention.
[0014] In a fourth aspect, an embodiment of the present invention provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the training method of the audio-visual speech separation model of any embodiment of the present invention are implemented.
[0015] The beneficial effects of the embodiments of the present invention are: based on the cross-directional multiplier method, the model is quantized and tuned to train a lightweight audio-visual speech separation model, and the model can be applied to low-resource devices with weak computing power. Moreover, through the multimodal model, the different quantization sensitivities of different modalities can be fully utilized to ensure the balance between the computational complexity and performance of the lightweight audio-visual speech separation model, and can be applied to low-resource devices to obtain better speech separation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is a flowchart of a method for training an audio-visual speech separation model provided by one embodiment of the present invention;
[0018] Figure 2 1 is a schematic diagram of the network structure of VisualVoice of a method for training a visual-audio speech separation model provided by one embodiment of the present invention;
[0019] Figure 3 1. It is a schematic diagram of a fixed quantization model and precision fine-tuning results of different parts of a training method for an audio-visual speech separation model provided by one embodiment of the present invention;
[0020] Figure 4 1 is a schematic diagram of mixed precision fine-tuning results of a training method for an audio-visual speech separation model provided by one embodiment of the present invention;
[0021] Figure 5 Schematic diagram of precise combination of layers based on KL divergence in a training method for an audiovisual speech separation model provided by one embodiment of the present invention;
[0022] Figure 6 1 is a schematic diagram of the structure of a training system for an audio-visual speech separation model provided by one embodiment of the present invention;
[0023] Figure 7 A schematic structural diagram of an embodiment of an electronic device for training an audio-visual speech separation model provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0025] like Figure 1 FIG2 is a flow chart of a method for training a video-audio speech separation model according to an embodiment of the present invention, comprising the following steps:
[0026] S11: Inputting mixed training audio of multiple speakers into an audio-visual speech separation model to obtain predicted spectrograms of the multiple speakers;
[0027] S12: Determine the audiovisual features of the predicted speaker of the predicted spectrogram and the audiovisual features of the reference speaker of the reference spectrogram of the mixed training audio;
[0028] S13: Based on the cross-modal loss determined by the predicted speaker audio-visual features and the reference speaker audio-visual features, the audio-visual speech separation model is trained with mixed precision quantization conditions using the cross-modal loss through a cross-directional multiplier method to obtain a lightweight audio-visual speech separation model.
[0029] In this embodiment, the audio-visual speech separation model of this method is the VisualVoice model, a multi-task learning framework. During the audio-visual speech separation process, it receives the mixed speech of K speakers and the lip videos of the K speakers while speaking. It also provides K facial images as additional input for calculating the speaker identity embedding and cross-modal loss. For the lips, the video containing the lip area is processed by a lip reading network, which consists of a 3D convolutional network and ShuffleNet V2. Finally, a temporal convolutional network is used to extract the lip feature sequence. For the face, ResNet-18 is used to extract the speaker identity embedding from the facial input. For the voice, a U-Net convolutional network is used to process the complex spectrogram of the input audio mixture. The hidden audio features in the U-Net are connected with the lip feature sequence and facial features in the channel dimension to obtain a fused audio-visual feature of the lips, face, and voice. These features are applied to the input spectrogram to obtain the predicted spectrograms of the K speakers.
[0030] In step S11, during training, mixed training audio of multiple speakers and a reference spectrogram of the mixed training audio are prepared. To train the audio-visual speech separation model, the mixed training audio of multiple speakers is input into the audio-visual speech separation model, and the audio-visual speech separation model is used to predict the spectrograms of the multiple speakers.
[0031] For step S12, a speech attribute analysis network module is introduced to determine the reference speaker audio-visual features and the predicted speaker audio-visual features from the reference spectrogram and the predicted spectrogram, which include speaker identity embeddings to learn the correlation between audio and video embeddings.
[0032] Regarding step S13, although VisualVoice has achieved leading performance on many audio-visual datasets including VoxCeleb2 and LRS2, its computational cost and model size are usually unacceptable for small smart devices with low computational requirements. Therefore, this method uses quantization technology to reduce the computational and storage costs of VisualVoice, thereby obtaining a lightweight version of it. Lightweight refers to the degree of dependence of the software architecture on the environment. For example, small smart devices such as voice recorders usually have low computing power due to hardware limitations. If they are equipped with complex neural networks for speech processing without an Internet connection, such low-computing smart devices will find it difficult to execute. In order to enable such low-computing smart devices to have speech processing capabilities, they are usually equipped with lightweight language models.
[0033] The neural network quantization definition of VisualVoice's audio-visual speech separation model is as follows: given a neural network with parameters W = {W1, W2, ..., W L} neural network N, where Represents the parameters of the i-th layer. The goal of training is to find the scaling factor And each W i Quantized integer matrix Defined as in, represents the bit accuracy of layer i, Represents the valid set of precisions.
[0034] As an implementation method, the function for training the audio-visual speech separation model with mixed precision quantization conditions using the cross-modal loss through the cross-directional multiplier method is:
[0035] L ρ (W,G,λ)=f(W)+ρ / 2||WG-λ|| 2 -ρ / 2||λ||
[0036] Where W is the first to Lth layer parameters of the audio-visual speech separation model, G = {α i Q i} L i=1 , the α scaling factor, the Q is a quantization integer matrix, i is the i-th layer of the audio-visual speech separation model, λ is the Lagrange multiplier, and ρ is a hyperparameter of the audio-visual speech separation model. In this embodiment, the method uses an optimization algorithm based on ADMM (Alternating Direction Methods of Multipliers) for the first time to quantize the audio-visual multimodal speech separation system. The neural network quantization is regarded as a non-convex optimization problem with discrete constraints. This allows the training process to naturally extend from fixed-precision fine-tuning to mixed-precision quantization.
[0037] The specific algorithm is:
[0038]
[0039] By tuning and quantizing the above algorithm, the memory usage and computational cost of the neural network of the audio-visual speech separation model are greatly reduced.
[0040] While the aforementioned fine-tuned quantization can reduce the memory usage and computational cost of the neural network used in the audio-visual speech separation model, extreme quantization can also degrade the model's performance. This study discovered that different layers of the model have varying sensitivities to quantization error, allowing a mixed-precision quantization strategy to effectively improve the performance of the audio-visual speech separation model based on a lightweight model.
[0041] As an implementation method, the audio-visual speech separation model is trained using the cross-modal loss through a cross-directional multiplier method based on a mixed precision quantization condition determined by the first search space sensitivity, the second search space sensitivity, and the training performance sensitivity.
[0042] The mixed precision quantization condition includes: searching for the first search space sensitivity of each layer parameter in the audio-visual speech separation model based on the Hessian trace.
[0043] In this embodiment, the best accuracy combination can be searched based on the Hessian trace. The sensitivity of the network layer to quantization can be determined by multiplying the trace of the parameter Hessian matrix by the squared quantization error. It is assumed that the impact of quantization of each layer on the audio-visual speech separation model is independent. Therefore, the best combination can be found by searching for the combination with the smallest sum of the sensitivity scores of all layers. The second-order gradient of the parameters is calculated, and the final sensitivity score formula is:
[0044]
[0045] Among them, H i It's W i The Hessian matrix of By W i The average number of parameters in H i The Hessian trace can be calculated using an open source toolkit. The higher layers maintain a higher accuracy of the audio-visual speech separation model and further limit the search space. According to experiments, better training results can be obtained based on the lightweight speech separation model.
[0046] As another implementation, the mixed precision quantization condition includes: searching for a second search space sensitivity of each layer parameter in the audio-visual speech separation model based on relative entropy.
[0047] In this embodiment, relative entropy, also known as the KL-Kullback-Leibler divergence, determines the KL deviation based on the outputs of the quantized model and the full-precision model. When the neural network of the audio-visual speech separation model has a large number of layers, the search space becomes too large, resulting in an unacceptable time consumption. To further reduce the model's search space, the following greedy search algorithm is used:
[0048]
[0049]
[0050] Where X is the calibration data determined using the audiovisual features of the reference speaker of the reference spectrogram, and gw(*) represents the mask calculation process of the network with parameters W. Represents the parameters of the network, where the i-th layer is quantized to b i Bit, Represents the quantization parameter of this layer. In order to limit the search process, the model size is penalized using the adjustable parameter φ.
[0051] As another implementation, the mixed precision quantization condition includes: selecting the training performance sensitivity of each layer parameter in the audio-visual speech separation model based on prior knowledge.
[0052] In this implementation, a partial quantization model is relied upon to ensure the performance of the audio-visual speech separation model. More precisely, the entire network is divided into several components based on prior knowledge. The audio-visual speech separation model neural network quantizes each component to a lower precision for experimentation. Based on the experimental results on the validation set, the sensitivity of each component to quantization is determined, and manual selection is performed based on these experimental results. This ensures the performance of the lightweight audio-visual speech separation model after training.
[0053] It can be seen from this implementation that a lightweight audio-visual speech separation model is trained by quantizing and tuning the model based on the cross-directional multiplier method. The model can be applied to low-resource devices with weak computing power. In addition, the multimodal model can fully utilize the different quantization sensitivities of different modalities to ensure a balance between the computational complexity and performance of the lightweight audio-visual speech separation model, and can be applied to low-resource devices to obtain better speech separation performance.
[0054] We experimented with this method, using ESPNet SE (an end-to-end speech enhancement and separation toolkit designed to replicate similar performance to VisualVoice. The multi-speaker speech mixtures in the training set were pre-generated, and a corresponding validation set was also prepared.
[0055] For fixed-precision quantization experiments, 1000 samples were randomly selected from the original training set to form a small set for fine-tuning. To perform mixed-precision combination search, only 4 samples from the validation set were used. The model was always fine-tuned for 30 epochs, and the final result was the model with the best validation performance among the resulting models. In the ADMM-based QAT (quantization-aware training), the learning rates η1 and η2 were set to 5.0×10 -6 and 5.0×10 -7 , and ρ is set to 100. Since the sound attribute analysis network does not participate in the inference phase, it is not quantized during training. 0 , 6×10 1 , 5×10 2 Models of equivalent size to fixed 6, 4, and 3 bit models were obtained in a KL divergence-based search. For greedy search, the lip network, then the face network, and finally the U-Net were searched in that order. In all experiments, activations of all layers were quantized to 8 bits using a min-max strategy.
[0056] In addition to applying fixed-precision QAT to the entire VisualVoice network, it is also divided into different parts and quantized separately, such as Figure 2 As shown in Figure 2. While LipNet and FaceNet represent lip movement analysis networks and facial attribute analysis networks, respectively, U-Net is divided into two different schemes. One is the left-right scheme, where the left side contains the first half of the U-Net and the right side contains the rest. The other is the inside-outside scheme, where the inside contains 8 symmetrical smaller layers in the middle of the U-Net and the outside contains 8 larger layers at the input and output of the U-Network. These two schemes were designed based on prior knowledge of the U-Net structure. The results are shown in Figure 2. Figure 3 Different characteristics of the neural network part with fixed quantization are shown, where “Q-Part” represents the quantized part.
[0057] From the results, it's clear that FaceNet is the least sensitive component of the entire network, as quantizing it to 3 bits barely impacts performance. LipNet shows similar tolerance at 4 bits. However, it exhibits significant degradation when quantized to 3 bits. As for U-Net, the left and right components show no significant difference before quantization to 2 bits. However, when split according to the inner-outer scheme, the inner components are more robust to quantization. This is considered a common phenomenon in the U-Net architecture. This is because U-Net relies on skip connections between symmetric outer layers to propagate low-level information. For the inner layers, high-precision computation is unnecessary because the information flow is highly compressed. Another interesting observation is that for some components, 2-bit quantization results outperform some higher-precision quantization results. This may be because, at extreme quantization levels, high precision does not necessarily mean small quantization error. Since uniform quantization is performed on the system, this phenomenon suggests that retaining only the sign of the weights (and zeros) may be more effective than quantizing them to a discrete set.
[0058] In order to obtain better performance, three mixed precision fine-tuning strategies are further applied to the quantization system trained above. The results are as follows Figure 4As shown in the figure, the final performance is evaluated using three metrics: SDR, compression ratio, and bit operations (BOP), where the "Eq.bits" column represents the equivalent fixed bit width setting of the current bit, where they have approximately the same size. For the manual strategy, we choose the highest possible accuracy for each part while maintaining the same model size as the fixed precision model. The effective SDR based on the systems listed in Table 1 is selected, and its trend is the same as the results in the table. Specifically, for the 6-bit equivalent manual selection, it is {6, 4, 6, 8} for LipNet, FaceNet, UNetInner, and UNet Outer respectively. For the 4-bit selection, it is {4, 3, 3, 8}, and for the 3-bit selection, it is {3, 2, 3, 6}. It can be clearly seen in the figure that the three strategies give comparable results on the 6-bit equivalent setting. At the same time, they outperform the 8-bit fixed precision setting in all three metrics, demonstrating the effectiveness of mixed precision quantization. On the 4-bit equivalent setting, the manual strategy slightly outperforms the Hessian trace-based selection and KL divergence-based greedy search strategies by about 1 dB in SDR, and is also slightly better in terms of compression ratio. This shows that the prior knowledge of the network partitioned by this method corresponds well to the characteristics of each layer. However, when it comes to the 3-bit equivalent setting, the first two strategies cannot maintain an acceptable SDR. At the same time, the greedy search based on KL divergence proposed in this method is able to obtain 7.2dB better results in SDR while maintaining the competitive model size and BOP. This may be because the KL divergence is calculated directly from the output mask, so it is expected to focus on optimizing the final accuracy rather than minimizing the quantization error.
[0059] In order to analyze the quantized network in detail, Figure 5 The KL divergence based on the 3-bit equivalent system shows the combination of the accuracy of each layer. The trend shown in the figure is similar to that in Figure 3The observations in are very consistent. UNetInner follows FaceNet and is also specified as low precision. It can be observed that when quantized to 3 bits, LipNet exhibits a relatively severe degradation of more than 4dB, while the above two are only affected by less than 2dB. This explains why some layers in LipNet require high precision while other layers can remain at low precision. Finally, UNet Outer is the most sensitive part and is specified as the highest precision. Based on the above observations, designing precision combinations through automatic search and prior knowledge may be a promising strategy in more specific practical applications. By searching using some direct and automatic methods (such as those based on KL divergence), one can first find an approximately optimal combination. Then, a more suitable combination can be manually defined based on the search results, into which prior knowledge can be integrated. By applying the above process, one can expect to squeeze out every last bit from the quantized network while maintaining the highest possible performance.
[0060] Overall, this method aims to reduce the size and computational complexity of VisualVoice, an audio-visual speech separation system. We attempt to train a quantized version of it using a quantization-aware training method based on ADMM. To further optimize the trade-off between space, speed, and size, we employ three strategies to generate a mixed-precision quantized network. Experimental results show that under relatively loose constraints, the manual selection strategy provides the best prediction accuracy, comparable to or even better than higher-precision results. Furthermore, our method's greedy search strategy based on KL divergence demonstrates excellent performance in the extreme case of 3-bit quantization, outperforming the other two strategies by approximately 8dB and the fixed-precision quantization results by approximately 13dB.
[0061] like Figure 6 The figure shows a structural diagram of a training system for an audio-visual speech separation model provided by one embodiment of the present invention. The system can execute the training method for the audio-visual speech separation model described in any of the above embodiments and be configured in a terminal.
[0062] The present embodiment provides a training system 10 for an audiovisual speech separation model, including a prediction program module 11 , an audiovisual feature determination program module 12 , and a lightweight training program module 13 .
[0063] Among them, the prediction program module 11 is used to input the mixed training audio of multiple speakers into the audio-visual speech separation model to obtain the predicted spectrograms of the multiple speakers; the audio-visual feature determination program module 12 is used to determine the predicted speaker audio-visual features of the predicted spectrogram and the reference speaker audio-visual features of the reference spectrogram of the mixed training audio; the lightweight training program module 13 is used to train the audio-visual speech separation model with mixed precision quantization conditions based on the cross-modal loss determined by the predicted speaker audio-visual features and the reference speaker audio-visual features through the cross-directional multiplier method to obtain a lightweight audio-visual speech separation model.
[0064] An embodiment of the present invention further provides a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions can execute the training method of the audio-visual speech separation model in any of the above method embodiments;
[0065] As an embodiment, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are configured as follows:
[0066] Inputting mixed training audio of multiple speakers into an audio-visual speech separation model to obtain predicted spectrograms of the multiple speakers;
[0067] Determining the predicted speaker audiovisual features of the predicted spectrogram and the reference speaker audiovisual features of the reference spectrogram of the mixed training audio;
[0068] Based on the cross-modal loss determined by the predicted speaker audio-visual features and the reference speaker audio-visual features, the audio-visual speech separation model is trained with mixed precision quantization conditions using the cross-directional multiplier method to obtain a lightweight audio-visual speech separation model.
[0069] A non-volatile computer-readable storage medium can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present invention. One or more program instructions stored in the non-volatile computer-readable storage medium, when executed by a processor, perform the method for training the audio-visual speech separation model in any of the above-described method embodiments.
[0070] Figure 7 This is a hardware structure diagram of an electronic device for a training method of an audio-visual speech separation model provided by another embodiment of the present application, such as Figure 7 As shown, the device includes:
[0071] One or more processors 710 and memory 720, Figure 7A processor 710 is used as an example. The apparatus for the method of training a video-audio speech separation model may further include: an input device 730 and an output device 740.
[0072] The processor 710, the memory 720, the input device 730 and the output device 740 may be connected via a bus or other means. Figure 7 The bus connection is taken as an example.
[0073] Memory 720, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the audio-visual speech separation model training method in the embodiments of the present application. Processor 710 executes the various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in memory 720, thereby implementing the audio-visual speech separation model training method in the above-mentioned method embodiment.
[0074] The memory 720 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data, etc. In addition, the memory 720 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 720 may optionally include a memory remotely located relative to the processor 710, and these remote memories may be connected to the mobile device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0075] The input device 730 can receive input digital or character information. The output device 740 can include a display device such as a display screen.
[0076] The one or more modules are stored in the memory 720 and, when executed by the one or more processors 710 , perform the training method of the audio-visual speech separation model in any of the above method embodiments.
[0077] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.
[0078] The non-volatile computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0079] An embodiment of the present invention also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the training method of the audio-visual speech separation model of any embodiment of the present invention.
[0080] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:
[0081] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and their primary purpose is to provide voice and data communications. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0082] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers and have computing and processing capabilities, and generally also have mobile Internet access. These terminals include PDAs, MIDs, and UMPC devices, such as tablet computers.
[0083] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0084] (4) Other electronic devices with data processing functions.
[0085] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise" include not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or device. In the absence of further limitations, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device that includes the elements.
[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A training method for an audiovisual speech separation model, comprising: Inputting mixed training audio of multiple speakers into an audio-visual speech separation model to obtain predicted spectrograms of the multiple speakers; Determining the predicted speaker audiovisual features of the predicted spectrogram and the reference speaker audiovisual features of the reference spectrogram of the mixed training audio; Based on the cross-modal loss determined by the predicted speaker audio-visual features and the reference speaker audio-visual features, the audio-visual speech separation model is trained with mixed precision quantization conditions using the cross-modal loss through a cross-directional multiplier method to obtain a lightweight audio-visual speech separation model, wherein the function for training the audio-visual speech separation model with mixed precision quantization conditions using the cross-modal loss through a cross-directional multiplier method is: L ρ (W,G,λ)=f(W)+ρ / 2||WG-λ|| 2 -ρ / 2||λ||, W is the first to Lth layer parameters of the audio-visual speech separation model, G={α i Q i } L i=1 , α is a scaling factor, Q is a quantized integer matrix, i is the i-th layer of the audio-visual speech separation model, λ is a Lagrange multiplier, and ρ is a hyperparameter of the audio-visual speech separation model.
2. The method according to claim 1, wherein The mixed precision quantization condition includes: searching for the first search space sensitivity of each layer parameter in the audio-visual speech separation model based on the Hessian trace.
3. The method according to claim 1, wherein The mixed precision quantization condition includes: searching for the second search space sensitivity of each layer parameter in the audio-visual speech separation model based on relative entropy.
4. The method according to claim 1, wherein The mixed precision quantization condition includes: selecting the training performance sensitivity of each layer parameter in the audio-visual speech separation model based on prior knowledge.
5. The method according to any one of claims 2 to 4, wherein The training of the audio-visual speech separation model using the cross-modal loss to perform mixed precision quantization conditions includes: The audio-visual speech separation model is trained using the cross-modal loss through a cross-directional multiplier method based on a mixed precision quantization condition determined by the first search space sensitivity, the second search space sensitivity and the training performance sensitivity.
6. A training system for an audiovisual speech separation model, comprising: a prediction program module, configured to input mixed training audio of multiple speakers into an audio-visual speech separation model to obtain predicted spectrograms of the multiple speakers; an audiovisual feature determination program module, configured to determine the audiovisual features of the predicted speaker of the predicted spectrogram and the audiovisual features of the reference speaker of the reference spectrogram of the mixed training audio; A lightweight training program module is used to train the audio-visual speech separation model under mixed precision quantization conditions using the cross-modal loss determined based on the predicted speaker audio-visual features and the reference speaker audio-visual features through a cross-directional multiplier method to obtain a lightweight audio-visual speech separation model, wherein the function for training the audio-visual speech separation model under mixed precision quantization conditions using the cross-modal loss through the cross-directional multiplier method is: L ρ (W,G,λ)=f(W)+ρ / 2||WG-λ|| 2 -ρ / 2||λ||, W is the first to Lth layer parameters of the audio-visual speech separation model, G={α i Q i } L i=1 , α is a scaling factor, Q is a quantized integer matrix, i is the i-th layer of the audio-visual speech separation model, λ is a Lagrange multiplier, and ρ is a hyperparameter of the audio-visual speech separation model.
7. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1 to 5.
8. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Single-channel voice separation method based on deep learning
CN111292762A
Audio-visual network-based multi-mode voice separation method and device
CN112863538A