Pre-training method of audio encoder, audio detection method and device
Patent Information
- Application Number
- CN202211595442.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-12-13
AI Technical Summary
随着人工智能的发展,各种神经网络逐渐应用于各个领域中,当神经网络应用于信息提取时,往往先通过音频、图像、文本这三种模态中两两进行对比学习来进行模型的预训练,然而这种方法一方面需要大量的样本数据进行训练,训练速度较慢,另一反面,由于音频中往往包含大量的噪声,在进行对比学习时会无可避免的影响预训练模型的精度
[0068] The pre-training method for the audio encoder provided in this embodiment can first train a target image encoder and a target text encoder through contrastive learning. Then, the sample multimodal features, which fuse the first image features extracted by the target image encoder and the first text features extracted by the target text encoder, are used as supervision data to train the initial audio encoder to be trained. This breaks down the multiple contrastive learning processes, reducing the learning difficulty and improving training efficiency. On the other hand, the sample multimodal features are used as supervision data and are not affected by audio noise, making them more accurate. Therefore, the accuracy of the trained target audio encoder is also correspondingly higher.
Smart Images

Figure CN116030798B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically, to a pre-training method for an audio encoder, an audio detection method, and an apparatus. Background Technology
[0002] With the development of information technology, audio, images, and text have become the three main modes of information dissemination, making the extraction of information from these modalities particularly important. With the development of artificial intelligence, various neural networks are gradually being applied to various fields. When neural networks are applied to information extraction, they are often pre-trained by comparing each of the three modalities (audio, image, and text) pairwise. However, this method requires a large amount of sample data for training, resulting in a slow training speed. Furthermore, since audio often contains a lot of noise, the accuracy of the pre-trained model inevitably suffers during comparative learning. Summary of the Invention
[0003] This disclosure provides at least one method for pre-training an audio encoder, an audio detection method, and an apparatus.
[0004] In a first aspect, embodiments of this disclosure provide a pre-training method for an audio encoder, comprising:
[0005] Obtain a first sample data group, which includes a first sample image, a first sample text, and sample audio from the same multimedia resource;
[0006] The first sample image is input into a pre-trained target image encoder to determine the first image feature; the first sample text is input into a pre-trained target text encoder to determine the first text feature; and the sample audio is input into an initial audio encoder to be trained to determine the sample audio feature; wherein the target image encoder and the target text encoder are obtained based on contrastive learning training;
[0007] The first image features and the first text features are fused to obtain the sample multimodal features;
[0008] Based on the sample multimodal features and the sample audio features, the initial audio encoder to be trained is trained, so as to perform audio detection based on the trained target audio encoder.
[0009] In one optional implementation, the audio detection based on the trained target audio encoder includes:
[0010] Based on the target audio encoder, an audio detection model incorporating the target audio encoder is constructed;
[0011] The audio detection model is fine-tuned based on sample audio.
[0012] After acquiring the audio to be detected, the audio detection result corresponding to the audio to be detected is determined based on the finely tuned audio detection model.
[0013] In an optional implementation, the method further includes training the target image encoder and the target text encoder according to the following method:
[0014] Acquire a second sample data set, which includes a second sample image and a second sample text from the same multimedia resource;
[0015] The second sample image is input into the initial image encoder to be trained to determine the second image feature corresponding to the second sample image; and the second sample text is input into the initial text encoder to be trained to determine the second text feature corresponding to the second sample text.
[0016] Based on the second image features and the second text features, feature similarity is determined, and the initial image encoder and the initial text encoder are trained based on the feature similarity to obtain the target image encoder and the target text encoder.
[0017] In an optional implementation, the method further includes determining the multimedia resource according to the following method:
[0018] Retrieve multiple multimedia resources to be filtered;
[0019] Based on the popularity information of the multimedia resources to be screened, the multimedia resources are determined from the plurality of multimedia resources to be screened.
[0020] In one optional implementation, the multimedia resource includes sample videos;
[0021] The first sample image is a preset number of frame-segmented images obtained by performing frame-segmentation processing on the sample video.
[0022] The first sample text is the title of the sample video;
[0023] The sample audio is an audio file of a preset length from the sample video.
[0024] In one optional implementation, the step of inputting the first sample image into a pre-trained target image encoder to determine the first image features includes:
[0025] The preset number of frame-scraped images are input into the target image encoder to obtain the initial image features corresponding to each frame-scraped image.
[0026] The initial image features corresponding to each of the extracted frames are fused to obtain the first image feature.
[0027] In one optional implementation, the second sample data group includes positive sample pairs and negative sample pairs;
[0028] The method further includes determining the positive sample pairs and the negative sample pairs according to the following method:
[0029] Acquire multiple second sample images and second sample texts originating from the same multimedia resource;
[0030] Second sample images and second sample texts from the same multimedia resource are used as positive sample pairs; second sample images and second sample texts from different multimedia resources are combined as negative sample pairs.
[0031] Secondly, this disclosure also provides an audio detection method, including:
[0032] Obtain the audio to be detected;
[0033] The audio to be detected is input into a target audio encoder trained by a pre-training method of the audio encoder based on the first aspect or any possible implementation of the first aspect, and the audio features corresponding to the audio to be detected are determined.
[0034] The audio detection result corresponding to the audio to be detected is determined based on the audio features.
[0035] Thirdly, embodiments of this disclosure provide a pre-training apparatus for an audio encoder, comprising:
[0036] The first acquisition module is used to acquire a first sample data group, which includes a first sample image, a first sample text, and sample audio from the same multimedia resource.
[0037] The feature extraction module is used to input the first sample image into a pre-trained target image encoder to determine the first image features; input the first sample text into a pre-trained target text encoder to determine the first text features; and input the sample audio into an initial audio encoder to be trained to determine the sample audio features; wherein the target image encoder and the target text encoder are obtained based on contrastive learning training;
[0038] The fusion module is used to fuse the first image features and the first text features to obtain sample multimodal features;
[0039] The training module is used to train the initial audio encoder to be trained based on the sample multimodal features and the sample audio features, so as to perform audio detection based on the trained target audio encoder.
[0040] In one optional implementation, the device further includes a detection module for:
[0041] Based on the target audio encoder, an audio detection model incorporating the target audio encoder is constructed;
[0042] The audio detection model is fine-tuned based on sample audio.
[0043] After acquiring the audio to be detected, the audio detection result corresponding to the audio to be detected is determined based on the finely tuned audio detection model.
[0044] In an optional implementation, the training module is further configured to train the target image encoder and the target text encoder according to the following method:
[0045] Acquire a second sample data set, which includes a second sample image and a second sample text from the same multimedia resource;
[0046] The second sample image is input into the initial image encoder to be trained to determine the second image feature corresponding to the second sample image; and the second sample text is input into the initial text encoder to be trained to determine the second text feature corresponding to the second sample text.
[0047] Based on the second image features and the second text features, feature similarity is determined, and the initial image encoder and the initial text encoder are trained based on the feature similarity to obtain the target image encoder and the target text encoder.
[0048] In an optional implementation, the first acquisition module is further configured to determine the multimedia resource according to the following method:
[0049] Retrieve multiple multimedia resources to be filtered;
[0050] Based on the popularity information of the multimedia resources to be screened, the multimedia resources are determined from the plurality of multimedia resources to be screened.
[0051] In one optional implementation, the multimedia resources include sample videos;
[0052] The first sample image is a preset number of frame-segmented images obtained by performing frame-segmentation processing on the sample video.
[0053] The first sample text is the title of the sample video;
[0054] The sample audio is an audio file of a preset length from the sample video.
[0055] In an optional implementation, the feature extraction module, when inputting the first sample image into a pre-trained target image encoder to determine the first image features, is used to:
[0056] The preset number of frame-scraped images are input into the target image encoder to obtain the initial image features corresponding to each frame-scraped image.
[0057] The initial image features corresponding to each of the extracted frames are fused to obtain the first image feature.
[0058] In one optional implementation, the second sample data group includes positive sample pairs and negative sample pairs;
[0059] The training module is further configured to determine the positive sample pairs and the negative sample pairs according to the following method:
[0060] Acquire multiple second sample images and second sample texts originating from the same multimedia resource;
[0061] Second sample images and second sample texts from the same multimedia resource are used as positive sample pairs; second sample images and second sample texts from different multimedia resources are combined as negative sample pairs.
[0062] Fourthly, embodiments of this disclosure also provide an audio detection device, comprising:
[0063] The second acquisition module is used to acquire the audio to be detected;
[0064] The encoding module is used to input the audio to be detected into a target audio encoder trained based on the pre-training method of the audio encoder described in the first aspect or any possible implementation of the first aspect, and to determine the audio features corresponding to the audio to be detected.
[0065] The determination module is used to determine the audio detection result corresponding to the audio to be detected based on the audio features.
[0066] Fifthly, embodiments of this disclosure also provide a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the first aspect described above, or any possible implementation of the first aspect, or the steps of the second aspect described above are executed.
[0067] In a sixth aspect, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the first aspect described above, or any possible implementation of the first aspect, or performs the steps of the second aspect described above.
[0068] The pre-training method for the audio encoder provided in this embodiment can first train a target image encoder and a target text encoder through contrastive learning. Then, the sample multimodal features, which fuse the first image features extracted by the target image encoder and the first text features extracted by the target text encoder, are used as supervision data to train the initial audio encoder to be trained. This breaks down the multiple contrastive learning processes, reducing the learning difficulty and improving training efficiency. On the other hand, the sample multimodal features are used as supervision data and are not affected by audio noise, making them more accurate. Therefore, the accuracy of the trained target audio encoder is also correspondingly higher.
[0069] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0070] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0071] Figure 1 A flowchart illustrating a pre-training method for an audio encoder provided in an embodiment of this disclosure is shown;
[0072] Figure 2 A flowchart illustrating the overall process of the pre-training method for the audio encoder provided in this embodiment is shown.
[0073] Figure 3 A flowchart of an audio detection method provided by an embodiment of this disclosure is shown;
[0074] Figure 4 A schematic diagram of a pre-training apparatus for an audio encoder provided in an embodiment of the present disclosure is shown;
[0075] Figure 5 A schematic diagram of an audio detection device provided in an embodiment of this disclosure is shown;
[0076] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0077] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0078] In related technologies, when neural networks are applied to information extraction, the model is often pre-trained by performing pairwise comparative learning between the three modalities: audio, image, and text. For example, comparative learning can be performed between the encoders of audio and image, between the encoders of image and text, and between the encoders of audio and text. However, this method requires multiple comparative learning sessions, necessitates a large amount of sample data, and results in slow training speed. On the other hand, since audio contains a lot of noise, the encoder extracting audio features can affect the accuracy of the features learned by other encoders when participating in comparative learning, thus affecting the accuracy of the encoder itself.
[0079] Based on this, this disclosure provides a pre-training method, apparatus, computer device, and storage medium for an audio encoder. It first trains a target image encoder and a target text encoder through contrastive learning. Then, it uses the sample multimodal features, which fuse the first image features extracted by the target image encoder and the first text features extracted by the target text encoder, as supervisory data to train the initial audio encoder to be trained. This approach breaks down the multiple contrastive learning processes, reducing the learning difficulty and improving training efficiency. Furthermore, using the sample multimodal features as supervisory data ensures that the multimodal features are unaffected by audio noise and are relatively accurate, resulting in a correspondingly higher accuracy of the trained target audio encoder.
[0080] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0081] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0082] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0083] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0084] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0085] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0086] To facilitate understanding of this embodiment, a detailed description of the pre-training method for an audio encoder disclosed in this disclosure is provided first. The execution entity of the pre-training method for the audio encoder provided in this disclosure is generally a computer device with certain computing capabilities. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, the pre-training method for the audio encoder can be implemented by a processor calling computer-readable instructions stored in memory.
[0087] See Figure 1 The diagram shows a flowchart of a pre-training method for an audio encoder provided in this embodiment of the present disclosure. The method includes steps 101 to 104, wherein:
[0088] Step 101: Obtain the first sample data group, which includes a first sample image, a first sample text, and sample audio from the same multimedia resource.
[0089] Step 102: Input the first sample image into a pre-trained target image encoder to determine the first image feature; input the first sample text into a pre-trained target text encoder to determine the first text feature; and input the sample audio into an initial audio encoder to be trained to determine the sample audio feature; wherein the target image encoder and the target text encoder are obtained based on contrastive learning training.
[0090] Step 103: Fuse the first image features and the first text features to obtain sample multimodal features.
[0091] Step 104: Based on the sample multimodal features and the sample audio features, train the initial audio encoder to be trained, so as to perform audio detection based on the trained target audio encoder.
[0092] The following is a detailed description of the steps described above.
[0093] Regarding step 101,
[0094] The multimedia resources refer to resources that include images, text, and audio. For example, the multimedia resources can refer to videos. This disclosure does not limit other multimedia resources that can simultaneously contain images, text, and audio.
[0095] For example, when the multimedia resource includes a sample video, the first sample image may be all the video frames of the sample video, or it may be a preset number of frame images obtained after frame extraction processing; the first sample text is the title of the sample video, and the sample audio may be all the audio of the sample video, or it may be audio of a preset length in the sample video.
[0096] The reason for obtaining the first sample image, first sample text, and sample audio from the same multimedia resource is that the first sample image, first sample text, and sample audio from the same multimedia resource have a high correlation. The target text encoder and target image encoder described below are trained through contrastive learning. Therefore, for the first sample image and first sample text from the same multimedia resource, the features extracted by the target text encoder and the features extracted by the target image encoder are similar. The learning objective of the initial audio encoder should also be to extract similar features. The premise of similar features is that the first sample image, first sample text, and sample audio are themselves related. Therefore, it is necessary to obtain the first sample image, first sample text, and sample audio from the same multimedia resource to ensure that the first sample image, first sample text, and the sample video are related to each other.
[0097] In one possible implementation, multimedia resources with higher popularity generally have a higher correlation between the images, text, and audio they contain. Therefore, when determining the multimedia resource, for example, multiple multimedia resources to be screened can be obtained first, and then the multimedia resource can be determined from the multiple multimedia resources to be screened based on the popularity information of the multiple multimedia resources to be screened.
[0098] In another possible implementation, after acquiring multiple multimedia resources to be screened, when determining the multimedia resources from the multiple multimedia resources to be screened, in addition to based on the popularity information of the multimedia resources to be screened, it is also possible to base it on an existing trained multimodal pre-trained model.
[0099] Specifically, as described in the background of this disclosure, in related technologies, a pre-trained multimodal model can be trained through pairwise modal comparison learning. After training is completed, for any multimedia resource to be screened, the first feature of the sample image, the second feature of the sample text, and the third feature of the sample audio contained in the multimedia resource to be screened can be extracted by the multimodal model. Then, the feature similarity between the first feature, the second feature, and the third feature is calculated, and the multimedia resource to be screened with a feature similarity exceeding a preset similarity threshold is taken as the multimedia resource.
[0100] In practical applications, after acquiring multimedia resources, at least one first sample data group can be determined based on each multimedia resource. For example, a multimedia resource may include multiple images, multiple audio segments, and multiple text segments. Based on the multiple images, multiple audio segments, and multiple text segments, multiple first sample data groups can be determined.
[0101] Regarding step 102,
[0102] The target image encoder is an encoder used for image feature extraction, and the target text encoder is an encoder used for text feature extraction.
[0103] In one possible implementation, if the multimedia resource corresponding to the first sample data group is a sample video, and the first sample image is a preset number of frame-segmented images obtained after frame-segmentation processing of the sample video, then when the first sample image is input into a pre-trained target image encoder to determine the first image feature, the preset number of frame-segmented images can be input into the target image encoder respectively to obtain the initial image features corresponding to each frame-segmented image, and then the initial image features corresponding to each frame-segmented image are fused to obtain the first image feature.
[0104] For example, when fusing the initial images corresponding to each of the extracted frames, the initial images corresponding to each of the extracted frames can be pooled to obtain the first image features.
[0105] The reason why the first sample number is a predetermined number of frame images is that the video lengths of different sample videos may be different, and the audio encoder needs to ensure the consistency of feature size when encoding, so the number of images is fixed.
[0106] Similarly, when the sample audio is audio from a sample video, the audio length may be different in different sample videos. In order to ensure the consistency of feature size, the length of the input sample audio needs to be fixed when the audio encoder is encoding.
[0107] For example, the input sample audio can be audio of a preset length following the start time of the audio of the sample video.
[0108] In determining the features of the sample audio, one possible implementation is to directly input the sample audio into the initial audio encoder to be trained to determine the features of the sample audio; in another possible implementation, the spectrogram corresponding to the sample audio can be determined first, and then the audio spectrum can be extracted from the spectrogram and input into the initial audio encoder (for example, the audio spectrum can be extracted by Short Time Fourier Transform (SFT)) to determine the features of the sample audio.
[0109] It should be noted that if the sample audio is input to the initial audio encoder during training, then the input to the target audio encoder after the initial audio encoder is trained must also be audio; if the audio spectrum extracted from the spectrogram of the sample audio is input to the initial audio encoder during training, then the input to the target audio encoder after the initial audio encoder is trained must also be the audio spectrum extracted from the spectrogram of the audio.
[0110] The initial image encoder and its specific training process will be described below, and will not be elaborated here.
[0111] Regarding steps 103 and 104,
[0112] The sample multimodal features are fused with the first sample image and the first sample text features. Therefore, the sample multimodal features can be used to characterize the features of multiple modalities of the multimedia resources. In one possible implementation, when fusing the first image features and the first text features, the first image features and the first text features can be pooled to obtain the sample multimodal features.
[0113] In one possible implementation, when training the initial audio encoder to be trained based on the sample multimodal features and the sample audio features, the feature similarity between the sample multimodal features and the sample audio features can be calculated first, and then the initial audio encoder to be trained can be adjusted based on the feature similarity.
[0114] Here, since the training objective of the initial audio encoder should be to make the sample audio features and sample multimodal features similar (i.e., to increase the feature similarity between sample audio features and sample multimodal features), the network parameter values of the initial audio encoder can be adjusted according to this training objective.
[0115] In another possible implementation, when training the initial audio encoder to be trained based on the sample multimodal features and the sample audio features, the loss value of this training can be calculated based on the sample multimodal features and the sample audio features, and the network parameter values of the initial audio encoder can be adjusted based on the loss value.
[0116] For example, the loss between the sample audio features and the sample multimodal features can be a cosine similarity loss, which can be calculated using, for example, the following formula:
[0117] loss(x3,x4) = 1 - cos(x3,x4)
[0118] Where x3 represents the sample audio features, x4 represents the sample multimodal features, and loss(x3,x4) represents the cosine similarity loss.
[0119] Alternatively, the loss between the sample audio features and the sample multimodal features can be the mean square error (MSE) loss, which can be calculated, for example, using the following formula:
[0120] loss(x3,x4)=L={l1,…,l d} T ,l i =x 3i -x 4i
[0121] Where d represents the dimension of the sample audio features and the sample multimodal features, x 3i Let represent the value of the i-th dimension of the sample audio feature, and let represent the value of the i-th dimension of the sample multimodal feature.
[0122] It should be noted that, since the target image encoder and the target text encoder are pre-trained, and the sample multimodal features are used as supervisory data to train the initial audio encoder, the accuracy of the sample multimodal features is high and is not affected by the accuracy of the audio encoder.
[0123] Optionally, after the target audio encoder model is trained, a multimodal pre-trained model can be obtained that includes the target image encoder, the target text encoder, and the target audio encoder. When the multimodal pre-trained model is applied to the downstream data processing process, each target encoder can be used individually or in combination.
[0124] For example, an audio detection model containing the target audio encoder can be constructed based on the target audio encoder, and then the audio detection model can be fine-tuned based on sample audio. In this way, after obtaining the audio to be detected, the audio detection result corresponding to the audio to be detected can be determined based on the fine-tuned audio detection model.
[0125] Here, the audio detection model may include, in addition to the target audio encoder, a decoder, a classifier, etc., and the specific network structure can be set according to the audio detection target.
[0126] Since the target audio encoder uses sample multimodal features as its supervision data during training, and these features are not affected by noise in the sample audio, the trained target audio encoder extracts relatively high-quality features, resulting in higher accuracy of the determined audio detection results.
[0127] The training process of the target image encoder and the target text encoder described above will be introduced below. In one possible implementation, the target image encoder and the target text encoder can be trained through the following steps:
[0128] Step a1: Obtain the second sample data group, which includes a second sample image and a second sample text from the same multimedia resource;
[0129] The second sample data group and the first sample data group can have the same data type, for example, both can be videos. In this case, the second sample data group can be a portion of the data in the first sample data group. Alternatively, the second sample data group and the first sample data group can have different data types. For example, the first sample data group can be a video, and the second sample data group can be text and images.
[0130] Specifically, the method for determining the multimedia resources corresponding to the second sample data group is similar to / the same as the method for determining the multimedia resources corresponding to the first sample data group. The specific steps are described above and will not be repeated here.
[0131] Step a2: Input the second sample image into the initial image encoder to be trained to determine the second image feature corresponding to the second sample image; and input the second sample text into the initial text encoder to be trained to determine the second text feature corresponding to the second sample text.
[0132] Step a3: Determine the feature similarity based on the second image features and the second text features, and train the initial image encoder and the initial text encoder based on the feature similarity to obtain the target image encoder and the target text encoder.
[0133] The second sample data group may also include positive sample pairs and negative sample pairs. When determining the positive sample pairs and negative sample pairs in the second sample data group, similar to determining the positive sample pairs and negative sample pairs in the first sample data group, we can first obtain the second sample image and the second sample text from the same multimedia resource, and then use the second sample image and the second sample text from the same multimedia resource as the positive sample pair; and combine the second sample images and the second sample text from different multimedia resources as the negative sample pair.
[0134] Here, if M multimedia resources are acquired, then the corresponding second sample data group will have M positive sample pairs and M negative sample pairs. 2 -M items.
[0135] Furthermore, since the second sample data group includes positive sample pairs and negative sample pairs, theoretically, the features between positive sample pairs should be similar, while the features between negative sample pairs should be dissimilar. Based on this, when training the initial image encoder and the initial text encoder based on the feature similarity, the training objective should be: for positive sample pairs, increase the feature similarity between the second image features and the second text features; for negative sample pairs, decrease the feature similarity between the second image features and the second text features.
[0136] Therefore, when training the initial image encoder and the initial text encoder, the network parameter values of the initial image encoder and the initial text encoder can be adjusted based on the above training objectives, and the training steps can be executed cyclically until the cutoff condition is met.
[0137] The cutoff condition may be, for example, that the number of training sessions reaches a preset number, and / or that the feature similarity of positive sample pairs is greater than a first similarity, and the feature similarity of negative sample pairs is less than a second preset similarity.
[0138] The following section provides an overall overview of the pre-training method for the aforementioned audio encoder, based on the overall process. (See also...) Figure 2 The diagram shows the overall flowchart of the pre-training method for the audio encoder provided in this disclosure, which includes the following steps:
[0139] After obtaining the feature training samples, image features and text features are extracted separately. Then, comparative learning is performed based on the image features and text features. After the comparative learning is completed, the multimodal features of the samples are determined. At the same time, the feature training samples can be used as audio input to extract the audio features of the samples. Then, the audio features of the samples are used to fit the multimodal features of the samples.
[0140] Here, the feature training samples input at different stages are different sample data. During contrastive learning, the input feature training samples are the sample data from the second sample data group mentioned above. When determining the multimodal features of the samples and extracting the audio features of the samples, the input feature training samples are the data from the first sample data group. The fitting of the multimodal features of the samples can be understood as, in step 104 above, making the audio features of the samples similar to the multimodal features of the samples.
[0141] Based on the same concept, this disclosure also provides an audio detection method, see [link to relevant documentation]. Figure 3 The diagram shown is a flowchart of an audio detection method provided in this disclosure, which includes the following steps:
[0142] Step 301: Obtain the audio to be detected.
[0143] Step 302: Input the audio to be detected into the target audio encoder trained by the pre-training method of the audio encoder described in the above embodiment, and determine the audio features corresponding to the audio to be detected.
[0144] Step 303: Determine the audio detection result corresponding to the audio to be detected based on the audio features.
[0145] Here, the audio detection results may include, for example, audio-text detection results, audio classification results, etc., and the specific settings need to be tailored to the application scenario. A detailed description of the above steps is provided in the embodiments above and will not be repeated here.
[0146] The pre-training method for the audio encoder provided in this embodiment can first train a target image encoder and a target text encoder through contrastive learning. Then, the sample multimodal features, which fuse the first image features extracted by the target image encoder and the first text features extracted by the target text encoder, are used as supervision data to train the initial audio encoder to be trained. This breaks down the multiple contrastive learning processes, reducing the learning difficulty and improving training efficiency. On the other hand, the sample multimodal features are used as supervision data and are not affected by audio noise, making them more accurate. Therefore, the accuracy of the trained target audio encoder is also correspondingly higher.
[0147] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0148] Based on the same inventive concept, this disclosure also provides a pre-training device for an audio encoder corresponding to the pre-training method of the audio encoder. Since the principle of the device in this disclosure for solving the problem is similar to the pre-training method of the audio encoder described above in this disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0149] Reference Figure 4 The diagram shown is an architectural schematic of a pre-training device for an audio encoder provided in this embodiment of the present disclosure. The device includes: a first acquisition module 401, a feature extraction module 402, a fusion module 403, a training module 404, and a detection module 405; wherein,
[0150] The first acquisition module 401 is used to acquire a first sample data group, which includes a first sample image, a first sample text and sample audio from the same multimedia resource.
[0151] The feature extraction module 402 is used to input the first sample image into a pre-trained target image encoder to determine the first image features; input the first sample text into a pre-trained target text encoder to determine the first text features; and input the sample audio into an initial audio encoder to be trained to determine the sample audio features; wherein the target image encoder and the target text encoder are obtained based on contrastive learning training;
[0152] The fusion module 403 is used to fuse the first image features and the first text features to obtain sample multimodal features;
[0153] The training module 404 is used to train the initial audio encoder to be trained based on the sample multimodal features and the sample audio features, so as to perform audio detection based on the trained target audio encoder.
[0154] In an optional embodiment, the device further includes a detection module 405, used for:
[0155] Based on the target audio encoder, an audio detection model incorporating the target audio encoder is constructed;
[0156] The audio detection model is fine-tuned based on sample audio.
[0157] After acquiring the audio to be detected, the audio detection result corresponding to the audio to be detected is determined based on the finely tuned audio detection model.
[0158] In an optional implementation, the training module 404 is further configured to train the target image encoder and the target text encoder according to the following method:
[0159] Acquire a second sample data set, which includes a second sample image and a second sample text from the same multimedia resource;
[0160] The second sample image is input into the initial image encoder to be trained to determine the second image feature corresponding to the second sample image; and the second sample text is input into the initial text encoder to be trained to determine the second text feature corresponding to the second sample text.
[0161] Based on the second image features and the second text features, feature similarity is determined, and the initial image encoder and the initial text encoder are trained based on the feature similarity to obtain the target image encoder and the target text encoder.
[0162] In an optional implementation, the first acquisition module 401 is further configured to determine the multimedia resource according to the following method:
[0163] Retrieve multiple multimedia resources to be filtered;
[0164] Based on the popularity information of the multimedia resources to be screened, the multimedia resources are determined from the plurality of multimedia resources to be screened.
[0165] In one optional implementation, the multimedia resources include sample videos;
[0166] The first sample image is a preset number of frame-segmented images obtained by performing frame-segmentation processing on the sample video.
[0167] The first sample text is the title of the sample video;
[0168] The sample audio is an audio file of a preset length from the sample video.
[0169] In an optional implementation, the feature extraction module 402, when inputting the first sample image into a pre-trained target image encoder to determine the first image features, is used to:
[0170] The preset number of frame-scraped images are input into the target image encoder to obtain the initial image features corresponding to each frame-scraped image.
[0171] The initial image features corresponding to each of the extracted frames are fused to obtain the first image feature.
[0172] In one optional implementation, the second sample data group includes positive sample pairs and negative sample pairs;
[0173] The training module 404 is further configured to determine the positive sample pairs and the negative sample pairs according to the following method:
[0174] Acquire multiple second sample images and second sample texts originating from the same multimedia resource;
[0175] Second sample images and second sample texts from the same multimedia resource are used as positive sample pairs; second sample images and second sample texts from different multimedia resources are combined as negative sample pairs.
[0176] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0177] Based on the same inventive concept, this disclosure also provides an audio detection device corresponding to the audio detection method. Since the principle of the device in this disclosure for solving the problem is similar to that of the audio detection method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0178] Reference Figure 5 The diagram shown is an architectural schematic of an audio detection device provided in an embodiment of this disclosure. The device includes: a second acquisition module 501, an encoding module 502, and a determination module 503; wherein,
[0179] The second acquisition module 501 is used to acquire the audio to be detected;
[0180] The encoding module 502 is used to input the audio to be detected into the target audio encoder trained by the pre-training method of the audio encoder described in the above embodiments, and to determine the audio features corresponding to the audio to be detected.
[0181] The determining module 503 is used to determine the audio detection result corresponding to the audio to be detected based on the audio features.
[0182] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0183] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 6The diagram shows the structure of a computer device 600 provided in this embodiment of the present disclosure, including a processor 601, a memory 602, and a bus 603. The memory 602 stores execution instructions and includes main memory 6021 and external memory 6022. The main memory 6021, also called internal memory, is used to temporarily store computational data in the processor 601 and data exchanged with external memory 6022 such as a hard disk. The processor 601 exchanges data with the external memory 6022 through the main memory 6021. When the computer device 600 is running, the processor 601 and the memory 602 communicate through the bus 603, causing the processor 601 to execute the following instructions:
[0184] Obtain a first sample data group, which includes a first sample image, a first sample text, and sample audio from the same multimedia resource;
[0185] The first sample image is input into a pre-trained target image encoder to determine the first image feature; the first sample text is input into a pre-trained target text encoder to determine the first text feature; and the sample audio is input into an initial audio encoder to be trained to determine the sample audio feature; wherein the target image encoder and the target text encoder are obtained based on contrastive learning training;
[0186] The first image features and the first text features are fused to obtain the sample multimodal features;
[0187] Based on the sample multimodal features and the sample audio features, the initial audio encoder to be trained is trained, so as to perform audio detection based on the trained target audio encoder.
[0188] In one optional implementation, the instructions executed by the processor 601, wherein the audio detection based on the trained target audio encoder, includes:
[0189] Based on the target audio encoder, an audio detection model incorporating the target audio encoder is constructed;
[0190] The audio detection model is fine-tuned based on sample audio.
[0191] After acquiring the audio to be detected, the audio detection result corresponding to the audio to be detected is determined based on the finely tuned audio detection model.
[0192] In an optional implementation, the instructions executed by the processor 601 further include training the target image encoder and the target text encoder according to the following method:
[0193] Acquire a second sample data set, which includes a second sample image and a second sample text from the same multimedia resource;
[0194] The second sample image is input into the initial image encoder to be trained to determine the second image feature corresponding to the second sample image; and the second sample text is input into the initial text encoder to be trained to determine the second text feature corresponding to the second sample text.
[0195] Based on the second image features and the second text features, feature similarity is determined, and the initial image encoder and the initial text encoder are trained based on the feature similarity to obtain the target image encoder and the target text encoder.
[0196] In one optional implementation, the instructions executed by the processor 601 further include determining the multimedia resource according to the following method:
[0197] Retrieve multiple multimedia resources to be filtered;
[0198] Based on the popularity information of the multimedia resources to be screened, the multimedia resources are determined from the plurality of multimedia resources to be screened.
[0199] In one optional implementation, the instructions executed by the processor 601 include a sample video as one of the multimedia resources.
[0200] The first sample image is a preset number of frame-segmented images obtained by performing frame-segmentation processing on the sample video.
[0201] The first sample text is the title of the sample video;
[0202] The sample audio is an audio file of a preset length from the sample video.
[0203] In an optional implementation, the instructions executed by the processor 601, wherein inputting the first sample image into a pre-trained target image encoder to determine the first image features, include:
[0204] The preset number of frame-scraped images are input into the target image encoder to obtain the initial image features corresponding to each frame-scraped image.
[0205] The initial image features corresponding to each of the extracted frames are fused to obtain the first image feature.
[0206] In one optional implementation, the instructions executed by the processor 601 include a second sample data group comprising positive sample pairs and negative sample pairs.
[0207] The method further includes determining the positive sample pairs and the negative sample pairs according to the following method:
[0208] Acquire multiple second sample images and second sample texts originating from the same multimedia resource;
[0209] Second sample images and second sample texts from the same multimedia resource are used as positive sample pairs; second sample images and second sample texts from different multimedia resources are combined as negative sample pairs.
[0210] Alternatively, processor 601 can execute the following instructions:
[0211] Obtain the audio to be detected;
[0212] The audio to be detected is input into a target audio encoder trained by the pre-training method of the audio encoder described in the above embodiments to determine the audio features corresponding to the audio to be detected.
[0213] The audio detection result corresponding to the audio to be detected is determined based on the audio features.
[0214] This disclosure also provides a computer-readable storage medium storing a computer program. When a processor runs the computer program, it executes the steps of the audio encoder pre-training method and audio detection method described in the above method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0215] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the audio encoder pre-training method and audio detection method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0216] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0217] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0218] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0219] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0220] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0221] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A pre-training method for an audio encoder, comprising: Obtain a first sample data group, which includes a first sample image, a first sample text, and sample audio from the same multimedia resource; The first sample image is input into a pre-trained target image encoder to determine the first image feature; the first sample text is input into a pre-trained target text encoder to determine the first text feature; and the sample audio is input into an initial audio encoder to be trained to determine the sample audio feature; wherein the target image encoder and the target text encoder are obtained based on contrastive learning training; The first image features and the first text features are fused to obtain sample multimodal features; Based on the sample multimodal features and the sample audio features, the initial audio encoder to be trained is trained, and audio detection is performed based on the trained target audio encoder to determine the audio detection result.
2. The method according to claim 1, wherein, The audio detection based on the trained target audio encoder, and the determination of the audio detection result, include: Based on the target audio encoder, an audio detection model incorporating the target audio encoder is constructed; The audio detection model is fine-tuned based on sample audio. After acquiring the audio to be detected, the audio detection result corresponding to the audio to be detected is determined based on the finely tuned audio detection model.
3. The method according to claim 1, further comprising training the target image encoder and the target text encoder according to the following method: Acquire a second sample data set, which includes a second sample image and a second sample text from the same multimedia resource; The second sample image is input into the initial image encoder to be trained to determine the second image feature corresponding to the second sample image; and the second sample text is input into the initial text encoder to be trained to determine the second text feature corresponding to the second sample text. Based on the second image features and the second text features, feature similarity is determined, and the initial image encoder and the initial text encoder are trained based on the feature similarity to obtain the target image encoder and the target text encoder.
4. The method according to claim 1 or 3, further comprising determining the multimedia resource according to the following method: Retrieve multiple multimedia resources to be filtered; Based on the popularity information of the multiple multimedia resources to be screened, the multimedia resources are determined from the multiple multimedia resources to be screened.
5. The method according to claim 1, wherein, The multimedia resources include sample videos; The first sample image is a preset number of frame-segmented images obtained by performing frame-segmentation processing on the sample video. The first sample text is the title of the sample video; The sample audio is an audio file of a preset length from the sample video.
6. The method according to claim 5, wherein, The step of inputting the first sample image into a pre-trained target image encoder to determine the first image features includes: The preset number of frame-scraped images are input into the target image encoder to obtain the initial image features corresponding to each frame-scraped image. The initial image features corresponding to each of the extracted frames are fused to obtain the first image feature.
7. The method according to claim 3, wherein, The second sample data group includes positive sample pairs and negative sample pairs; The method further includes determining the positive sample pairs and the negative sample pairs according to the following method: Acquire second sample images and second sample texts from multiple multimedia resources; The second sample image and the second sample text from the same multimedia resource are used as the positive sample pair; the second sample image and the second sample text from different multimedia resources are combined as the negative sample pair.
8. An audio detection method, comprising: Obtain the audio to be detected; The audio to be detected is input into a target audio encoder trained by the pre-training method of the audio encoder according to any one of claims 1 to 7, and the audio features corresponding to the audio to be detected are determined. The audio detection result corresponding to the audio to be detected is determined based on the audio features.
9. A pre-training device for an audio encoder, comprising: The first acquisition module is used to acquire a first sample data group, which includes a first sample image, a first sample text, and sample audio from the same multimedia resource. The feature extraction module is used to input the first sample image into a pre-trained target image encoder to determine the first image features; input the first sample text into a pre-trained target text encoder to determine the first text features; and input the sample audio into an initial audio encoder to be trained to determine the sample audio features; wherein the target image encoder and the target text encoder are obtained based on contrastive learning training; The fusion module is used to fuse the first image features and the first text features to obtain sample multimodal features; The training module is used to train the initial audio encoder to be trained based on the multimodal features of the samples and the audio features of the samples, so as to perform audio detection based on the trained target audio encoder and determine the audio detection result.
10. An audio detection device, comprising: The second acquisition module is used to acquire the audio to be detected; The encoding module is used to input the audio to be detected into a target audio encoder trained by the pre-training method of the audio encoder according to any one of claims 1 to 7, and to determine the audio features corresponding to the audio to be detected; The determination module is used to determine the audio detection result corresponding to the audio to be detected based on the audio features.
11. A computer device, comprising: The computer device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the pre-training method of the audio encoder as described in any one of claims 1 to 7, or the steps of the audio detection method as described in claim 8.
12. A computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of a pre-training method for an audio encoder as claimed in any one of claims 1 to 7, or the steps of an audio detection method as claimed in claim 8.