A model training method and related device
Through clustering processing and pseudo-label training, the feature encoding capabilities of video encoding networks and audio encoding networks are improved, the problem of unsatisfactory encoding capabilities in the prior art is solved, and the effective application of encoding networks in downstream tasks is realized.
Patent Information
- Application Number
- CN202210452459.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-04-27
AI Technical Summary
The feature encoding capabilities of the video encoding network and audio encoding network trained in the prior art are not ideal and are difficult to effectively apply to downstream tasks.
By acquiring multiple training samples, the first predictive feature of the training sample is determined using the first coding network, and the categories are determined based on clustering processing, pseudo-labels are configured for the second segment, and the second predictive feature and pseudo-labels of the second coding network are used to train the sample, collaborative training of the video encoding network and the audio encoding network is realized.
The resource consumption of labeling training samples is avoided, the feature encoding capability of the encoding network is improved, and the effective application of the encoding network in downstream tasks is ensured.
Smart Images

Figure CN115130650B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method and related device. Background Art
[0002] In practical applications, the interaction between vision and hearing can make the human perception function more complete and accurate; for example, when people watch a video, they usually need to rely on sound to understand the content in the video picture. Based on this, when performing related tasks (such as classification tasks, etc.) on a video, it is often necessary to comprehensively consider the image features and audio features of the video; currently, mainly through a video encoding network to determine the image features of the video according to the video picture, and through an audio encoding network to determine the audio features of the video according to the audio of the video.
[0003] In the related art, usually, a contrastive learning method is adopted to train the above-mentioned video encoding network and audio encoding network. Specifically, the synchronized video segment and audio segment in a video can be used as positive samples, and the video segments and audio segments in different videos, or the asynchronous video segments and audio segments in the same video can be used as negative samples; then, a binary classification model for identifying positive samples and negative samples is trained, and the video encoding network and audio encoding network included in the binary classification model will also be correspondingly trained in this process.
[0004] However, the feature encoding capabilities of the video encoding network and audio encoding network trained by the above method are not ideal, and the image features and audio features encoded by the two are often difficult to be well applied to downstream tasks. The reason is that the difference between the positive samples and negative samples used in the above training method is usually very obvious. During the model training process, the trained binary classification model can easily and accurately distinguish positive samples and negative samples, while the video encoding network and audio encoding network therein are not fully trained. Summary of the Invention
[0005] Embodiments of this application provide a model training method and related device, which can ensure that the trained video encoding network and audio encoding network have better feature encoding capabilities, so as to be better applied to downstream tasks.
[0006] In view of this, the first aspect of this application provides a model training method, and the method includes:
[0007] Obtain a plurality of training samples; the training samples include video segments and their corresponding audio segments;
[0008] For each of the training samples, through a first encoding network, according to the first segment in the training sample, determine a first predicted feature corresponding to the training sample; the first encoding network is either a video encoding network or an audio encoding network.
[0009] Perform clustering processing based on the first prediction features corresponding to each of the multiple training samples to determine the category to which the first segment in each training sample belongs; and for each training sample, configure a corresponding pseudo-label for the second segment in the training sample according to the category to which the first segment in the training sample belongs; the second segment is different from the first segment.
[0010] For each training sample, through a second encoding network, determine the second prediction feature corresponding to the training sample according to the second segment in the training sample; determine the category prediction result corresponding to the second segment in the training sample according to the second prediction feature corresponding to the training sample; the second encoding network is any one of the video encoding network and the audio encoding network and is different from the first encoding network.
[0011] Train the second encoding network based on the category prediction results and pseudo-labels corresponding to the second segments in the multiple training samples.
[0012] A second aspect of the present application provides a model training device, the device includes:
[0013] A training sample acquisition module, configured to acquire a plurality of training samples; the training samples include video segments and their corresponding audio segments.
[0014] A first feature prediction module, configured to, for each training sample, through a first encoding network, determine the first prediction feature corresponding to the training sample according to the first segment in the training sample; the first encoding network is any one of the video encoding network and the audio encoding network.
[0015] A first feature clustering module, configured to perform clustering processing based on the first prediction features corresponding to each of the multiple training samples to determine the category to which the first segment in each training sample belongs; and for each training sample, configure a corresponding pseudo-label for the second segment in the training sample according to the category to which the first segment in the training sample belongs; the second segment is different from the first segment.
[0016] A second network prediction module, configured to, for each training sample, through a second encoding network, determine the second prediction feature corresponding to the training sample according to the second segment in the training sample; determine the category prediction result corresponding to the second segment in the training sample according to the second prediction feature corresponding to the training sample; the second encoding network is any one of the video encoding network and the audio encoding network and is different from the first encoding network.
[0017] A second network training module, configured to train the second encoding network based on the class prediction results and pseudo-labels respectively corresponding to the second segments in the multiple training samples.
[0018] A third aspect of the present application provides a computer device, which includes a processor and a memory:
[0019] The memory is used to store a computer program;
[0020] The processor is configured to execute the steps of the model training method as described in the first aspect above according to the computer program.
[0021] A fourth aspect of the present application provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the steps of the model training method as described in the first aspect above.
[0022] A fifth aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the model training method as described in the first aspect above.
[0023] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0024] The embodiment of the present application provides a model training method. In this method, first, a plurality of training samples including video segments and their corresponding audio segments are obtained; then, for each training sample, through a first encoding network (which can be either a video encoding network or an audio encoding network), according to the first segment in the training sample (which is the segment in the video segment and the audio segment that is suitable for being processed by the first encoding network), the first prediction feature corresponding to the training sample is determined; next, clustering processing is performed based on the first prediction features corresponding to the plurality of training samples, the category to which the first segment in each training sample belongs is determined, and accordingly, a corresponding pseudo-label is configured for the second segment (the other segment in the video segment and the audio segment except the first segment) in the training sample; furthermore, for each training sample, through a second encoding network (which is the other encoding network in the video encoding network and the audio encoding network except the first encoding network), according to the second segment in the training sample, the second prediction feature corresponding to the training sample is determined, and according to the second prediction feature corresponding to the training sample, the category prediction result corresponding to the second segment in the training sample is determined; finally, the second encoding network can be trained according to the category prediction results and pseudo-labels corresponding to the second segments in the plurality of training samples. When training the video encoding network and the audio encoding network through the above method, the clustering result of the encoding features generated by one of the encoding networks is used to determine the supervised signal available when training the other encoding network; on the one hand, the annotation of training samples is avoided, that is, the processing resources consumed by annotating training samples are saved, and at the same time, the problem that the performance of the trained encoding network is not good due to the defects in the constructed training samples can also be avoided; on the other hand, since there is a corresponding relationship between the video segment and the audio segment in the training sample, therefore, based on the feature clustering result corresponding to one type of segment in the training sample, a corresponding pseudo-label is configured for the other type of segment in the training sample, which can ensure the reliability of the configured pseudo-label to a certain extent. Correspondingly, using the pseudo-label as a supervised signal to train the other encoding network can ensure reliable training of the encoding network, that is, ensure that the trained video encoding network or audio encoding network can have better feature encoding ability and can be better applied to downstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of the application scenario of the model training method provided by the embodiment of the present application;
[0026] Figure 2 It is a schematic flowchart of the model training method provided by the embodiment of the present application;
[0027] Figure 3 It is a schematic diagram of the implementation principle of co-training the video encoding network and the audio encoding network provided by the embodiment of the present application;
[0028] Figure 4 Schematic diagram of the implementation principle of applying a video encoding network and an audio encoding network to a target classification task provided by an embodiment of the present application;
[0029] Figure 5 Schematic diagram of the implementation principle of applying a video encoding network and an audio encoding network to a background audio generation task provided by an embodiment of the present application;
[0030] Figure 6 Schematic diagram of the implementation principle of a model training method provided by an embodiment of the present application;
[0031] Figure 7 Schematic diagram of the structure of a model training device provided by an embodiment of the present application;
[0032] Figure 8 Schematic diagram of the structure of a terminal device provided by an embodiment of the present application;
[0033] Figure 9 Schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners
[0034] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0035] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0036] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0037] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0038] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0039] The solution provided in the embodiments of this application relates to the machine learning technology of artificial intelligence. In addition, the embodiments of this application can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.
[0040] In related technologies, the video coding network and audio coding network trained by using the contrast learning method generally have poor performance, and their feature coding capabilities are not ideal. The image features and audio features encoded by them are often difficult to be well applied to downstream tasks.
[0041] To solve the above problems, the embodiments of this application provide a model training method. When training the video coding network and audio coding network through this method, the clustering result of the coding features generated by one of the coding networks is used to assist in training the other coding network, so as to achieve the effect of co-training the video coding network and audio coding network and improve the performance of the trained video coding network and audio coding network.
[0042] Specifically, in the model training method provided in the embodiments of the present application, multiple training samples including video segments and their corresponding audio segments are first obtained. Then, for each training sample, through a first encoding network (which can be either a video encoding network or an audio encoding network), according to the first segment in the training sample (which is the segment in the video segment and the audio segment that is suitable for being processed by the first encoding network), the first prediction feature corresponding to the training sample is determined. Next, clustering processing is performed based on the first prediction features corresponding to the multiple training samples respectively, to determine the category to which the first segment in each training sample belongs, and accordingly, a corresponding pseudo-label is configured for the second segment (the other segment in the video segment and the audio segment except the first segment) in the training sample. Furthermore, for each training sample, through a second encoding network (which is the other encoding network in the video encoding network and the audio encoding network except the first encoding network), according to the second segment in the training sample, the second prediction feature corresponding to the training sample is determined, and according to the second prediction feature corresponding to the training sample, the category prediction result corresponding to the second segment in the training sample is determined. Finally, the second encoding network can be trained according to the category prediction results and pseudo-labels corresponding to the second segments in the multiple training samples respectively.
[0043] In the above model training method, positive and negative samples are not constructed during the process of training the video encoding network and the audio encoding network. Therefore, problems caused by the use of positive and negative samples in the model training method based on contrast learning in the related art can be avoided. When training the video encoding network and the audio encoding network through the above method, the clustering result of the encoding features generated by one of the encoding networks is used to determine the supervised signal available when training the other encoding network. On the one hand, the annotation of training samples is avoided, that is, the processing resources consumed by annotating training samples are saved, and at the same time, the problem that the performance of the trained encoding network is poor due to the defects in the constructed training samples can also be avoided. On the other hand, since there is a corresponding relationship between the video segment and the audio segment in the training sample, therefore, based on the clustering result of the features corresponding to one type of segment in the training sample, a corresponding pseudo-label is configured for the other type of segment in the training sample, which can ensure the reliability of the configured pseudo-label to a certain extent. Correspondingly, using the pseudo-label as a supervised signal to train the other encoding network can ensure reliable training of the encoding network, that is, ensure that the trained video encoding network or audio encoding network can have better feature encoding ability and can be better applied to downstream tasks.
[0044] It should be understood that the model training method provided in the embodiments of this application can be executed by a computer device with image processing capabilities and audio processing capabilities, and this computer device can be a terminal device or a server. Among them, the terminal device includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc.; the server can specifically be an application server or a Web server. In actual deployment, it can be an independent server or a cluster server or a cloud server composed of multiple physical servers. The data involved in the embodiments of this application can be stored on the blockchain.
[0045] To facilitate the understanding of the model training method provided in the embodiments of this application, the following takes the execution entity of this model training method as a server as an example to make an exemplary introduction to the application scenario of this model training method.
[0046] See Figure 1 , Figure 1 which is a schematic diagram of the application scenario of the model training method provided in the embodiments of this application. As Figure 1 shown, this application scenario includes a server 110 and a database 120; the server 110 can access the database 120 through a network, or the database 120 can also be integrated in the server 110. Among them, the server 110 is used to execute the method provided in the embodiments of this application to train a video coding network or an audio coding network; the database 120 stores several video segments and audio segments with corresponding relationships.
[0047] In actual applications, the server 110 can obtain multiple groups of video segments and audio segments with corresponding relationships from the database 120, and then use the obtained video segments and audio segments with corresponding relationships as training samples. Exemplarily, the video segments and audio segments with corresponding relationships can be obtained from the same audio-visual video. For example, a video segment corresponding to a certain playback time period can be intercepted from a certain audio-visual video as the video segment in the training sample, and the corresponding video audio in this audio-visual video during this playback time period can be intercepted as the audio segment in this training sample. In this way, the video segment and the audio segment corresponding to the same playback time period in this audio-visual video are the video segment and the audio segment with corresponding relationships.
[0048] After the server 110 obtains multiple training samples including video segments and audio segments with corresponding relationships, for each training sample, the first coding network 111 can determine the first predicted feature corresponding to this training sample according to the first segment in this training sample. It should be noted that the first coding network 111 can be either a video coding network or an audio coding network; correspondingly, the first segment can be the segment in the video segment and the audio segment included in the training sample that is suitable for being processed by the first coding network 111.
[0049] After the server 110 obtains the first prediction features corresponding to multiple training samples through the first encoding network 111, it can perform clustering processing on the first prediction features corresponding to multiple training samples respectively to determine the category to which the first segment in each training sample belongs. In addition, since there is a corresponding relationship between the video segment and the audio segment included in each training sample, after determining the category to which the first segment in each training sample belongs, a corresponding pseudo-label can be further configured for the second segment in the training sample according to the category to which the first segment in the training sample belongs, and this pseudo-label can be used as a supervised signal during subsequent training of the second encoding network 112. It should be noted that the second encoding network 112 is either a video encoding network or an audio encoding network, and the second encoding network 112 is different from the first encoding network 111; correspondingly, the second segment is the segment in the video segment and the audio segment included in the training sample that is suitable for being processed by the second encoding network 112, and the second segment is different from the first segment.
[0050] After the server 110 configures the pseudo-labels corresponding to the second segments in each training sample in the above manner, it can further train the second encoding network 112 based on the second segments in each training sample and their corresponding pseudo-labels. Specifically, for each training sample, the server 110 can determine the second prediction feature corresponding to the training sample through the second encoding network 112 according to the second segment in the training sample; then, according to the second prediction feature corresponding to the training sample, the category prediction result corresponding to the second segment in the training sample can be determined.
[0051] After the server 110 obtains the category prediction results corresponding to the second segments in each training sample in the above manner, it can train the second encoding network 112 based on the category prediction results and pseudo-labels corresponding to the second segments in each training sample; it should be understood that when the second encoding network 112 is a video encoding network, the training of the video encoding network can be achieved through the above manner, and when the second encoding network 112 is an audio encoding network, the training of the audio encoding network can be achieved through the above manner.
[0052] It should be understood that in practical applications, in addition to training the second encoding network 112 based on the above method, the server 110 can also train the first encoding network 111 based on the above method. Specifically, the server 110 can also configure corresponding pseudo-labels for the first segments in the training samples through clustering processing; that is, the server 110 can perform clustering processing based on the second prediction features corresponding to multiple training samples, determine the category to which the second segment in each training sample belongs, and for each training sample, configure a corresponding pseudo-label for the first segment in the training sample according to the category to which the second segment in the training sample belongs. Furthermore, the server 110 can, for each training sample, determine the category prediction result corresponding to the first segment in the training sample according to the first prediction feature corresponding to the training sample (i.e., the encoding feature determined by the first encoding network 111 based on the first segment in the training sample); furthermore, train the first encoding network 111 based on the category prediction results and pseudo-labels corresponding to the first segments in multiple training samples.
[0053] It should be understood that Figure 1 The application scenarios shown are only examples. In practical applications, the model training method provided by the embodiments of the present application can also be applied to other scenarios; for example, the server 110 can also obtain multiple training samples through other channels (such as determining training samples based on the uploaded audio-visual videos of specific objects). No limitations are imposed on the application scenarios of the model training method provided by the embodiments of the present application.
[0054] The model training method provided by the present application will be introduced in detail below through method embodiments.
[0055] See Figure 2 , Figure 2 which is a schematic flowchart of the model training method provided by the embodiments of the present application. For ease of description, the following embodiments will still be described by taking the server as the execution subject of the model training method. As Figure 2 shown, the model training method includes the following steps:
[0056] Step 201: Obtain multiple training samples; the training samples include video segments and their corresponding audio segments.
[0057] In the embodiments of the present application, before training the video encoding network or the audio encoding network, the server needs to first obtain multiple unsupervised training samples for training the video encoding network or the audio encoding network, and the obtained training samples include video segments and audio segments with corresponding relationships.
[0058] It should be noted that the above video encoding network is a neural network for encoding video features based on multiple video frames with temporal relationships in a video. The above audio encoding network is a neural network for encoding audio features based on the audio of the video. The video clip and the audio clip with a corresponding relationship in the training samples can be obtained based on the same audio-video. For example, a video clip within a certain playing time period can be intercepted from a certain audio-video, and the audio clip within the same playing time period in the audio-video can be intercepted. In this way, the video clip and the audio clip corresponding to the same playing time period intercepted based on the same audio-video are the video clip and the audio clip with a corresponding relationship in the training samples. Of course, in practical applications, the video clip and the audio clip with a corresponding relationship in the training samples can also be obtained through other means. For example, a corresponding background audio clip can be marked for a certain video clip, and the video clip and the marked background audio clip can also be used as the video clip and the audio clip with a corresponding relationship in the training samples. This application does not make any limitations on the obtaining methods of the video clip and the audio clip with a corresponding relationship in the training samples.
[0059] Exemplarily, the server can obtain the above multiple training samples based on a relevant database. Here, "multiple" specifically refers to greater than or equal to two. For example, the server can obtain a large number of audio-videos from a relevant database, and then intercept the video clip and the audio clip corresponding to the same playing time period from the audio-videos to form training samples. The server can also obtain the above training samples based on the audio-video sent by the terminal device. For example, the server can receive the audio-video uploaded by the terminal device, and then intercept the video clip and the audio clip corresponding to the same playing time period from the audio-video to form training samples. Another example is that the server can receive the video clip uploaded by the terminal device and the background audio clip configured for the video clip, and then use the video clip and the background audio clip to form training samples. The server can also directly obtain the above multiple training samples based on an open-source training video dataset (such as the AudioSet dataset). Such open-source training video datasets usually include a large amount of video data. The server can obtain a specific proportion (such as 90%) of the video data from the training video dataset to construct training samples, and obtain the remaining (such as 10%) of the video data in the training video dataset to construct test samples. By intercepting the video clip and the audio clip corresponding to the same playing time period in the video data, training samples and test samples are obtained. This application does not make any limitations on the methods and channels for the server to obtain training samples.
[0060] Step 202: For each of the training samples, through a first encoding network, determine a first predicted feature corresponding to the training sample according to a first segment in the training sample; the first encoding network is either the video encoding network or the audio encoding network.
[0061] After the server obtains multiple training samples, for each training sample, the first encoding network can perform feature encoding processing on the first segment in the training sample, so as to obtain the first prediction feature corresponding to the training sample. It should be understood that the feature encoding processing here is to convert the first segment in the training sample into machine-recognizable numerical information, and the converted numerical information can to a certain extent reflect the characteristics of the first segment itself.
[0062] It should be noted that the above first encoding network is any one of the video encoding network and the audio encoding network to be trained, the first segment is the segment in the video segment and the audio segment included in the training sample that is suitable for being processed by the first encoding network, and the first prediction feature is the predicted encoding feature corresponding to the first segment in the training sample. For example, when the first encoding network is a video encoding network, the first segment is the video segment in the training sample, and the first prediction feature is the predicted video encoding feature corresponding to the video segment in the training sample; when the first encoding network is an audio encoding network, the first segment is the audio segment in the training sample, and the first prediction feature is the predicted audio encoding feature corresponding to the audio segment in the training sample.
[0063] In a possible implementation manner, when the first encoding network is a video encoding network, before the server performs feature encoding processing on the video segment in the training sample through the video encoding network, it can first preprocess the video segment in the training sample. Exemplarily, the server can randomly sample a specific number of video frames (such as 16 frames) from the video segment in the training sample; and perform scaling processing on the sampled video frames. For example, without changing the aspect ratio of the video frame, make the shorter side of the video frame 256 pixels; then, the server can intercept a specific area (such as 224 pixels × 224 pixels) from the video frame obtained through the scaling processing; furthermore, based on the areas intercepted from each frame of the sampled video frames, an image tensor with a size of 16×3×224×224 can be constructed, where 16 represents the number of video frames sampled from the video segment included in the training sample, 3 represents the Red Green Blue (RGB) channel values, and 224×224 represents the area size intercepted from each frame of the video frame; the image tensor obtained in this way can be used as the input data of the video encoding network.
[0064] Furthermore, the server can input the image tensor obtained through the above preprocessing into the video encoding network. The video encoding network can analyze and process the input image tensor, and can correspondingly output the corresponding predicted video encoding feature, that is, the first prediction feature corresponding to the training sample.
[0065] It should be noted that the above video encoding network can use an R(2+1)D network with a preset number of layers (such as 18 layers). This network combines two-dimensional convolution and three-dimensional convolution, and can use two-dimensional convolution to extract spatial information and three-dimensional convolution to synthesize spatio-temporal information, thus being more conducive to learning the features of video segments. In addition, the video encoding network can also be other network structures, such as SlowFast network, Dilated 3D ConvNet (I3D), Action Recognition Network (C3D), Video Swin Transformer, etc. The present application does not make any limitation on the structure of the video encoding network herein.
[0066] In another possible implementation, when the first encoding network is an audio encoding network, before the server performs feature encoding processing on the audio segments in the training samples through the audio encoding network, it can first preprocess the audio segments in the training samples. Exemplarily, the server can perform a short-time Fourier transform on the audio segments and perform a logarithmic operation on the result obtained from the short-time Fourier transform to obtain a logarithmic spectrogram with time on the horizontal axis, frequency on the vertical axis, and intensity logarithm as the value (for example, it can be a tensor with a size of 40×100). This logarithmic spectrogram can be used as the input data for the audio encoding network.
[0067] Furthermore, the server can input the logarithmic spectrogram obtained through the above preprocessing into the audio encoding network. By analyzing and processing the input logarithmic spectrogram, the audio encoding network can correspondingly output the corresponding predicted audio encoding features, that is, the first predicted features corresponding to the training samples.
[0068] It should be noted that the above audio encoding network can use an 18-layer ResNet based on two-dimensional convolution, or can also use structures such as a temporal convolutional network, a recurrent neural network (RNN), etc. The present application also does not make any limitation on the structure of the audio encoding network herein.
[0069] Step 203: Perform clustering processing based on the first predicted features corresponding to the multiple training samples respectively, and determine the category to which the first segment in each training sample belongs; and for each training sample, configure a corresponding pseudo-label for the second segment in the training sample according to the category to which the first segment in the training sample belongs; the second segment is different from the first segment.
[0070] After the server completes the encoding process of the first segment in each training sample through the first encoding network and obtains the first prediction features corresponding to each training sample, it can further perform clustering processing on the first prediction features corresponding to each training sample to determine the category to which the first segment in each training sample belongs. Furthermore, for each training sample, the server can configure a corresponding pseudo-label for the second segment in the training sample according to the category to which the first segment in the training sample belongs.
[0071] It should be noted that the above-mentioned second segment is another segment in the training sample except the first segment. For example, when the first segment is a video segment in the training sample, the second segment is an audio segment in the training sample; when the first segment is an audio segment in the training sample, the second segment is a video segment in the training sample. The second segment is the processing object of the second encoding network, and the second encoding network is another encoding network other than the first encoding network in the video encoding network and the audio encoding network to be trained. For example, when the first encoding network is a video encoding network, the second encoding network is an audio encoding network; when the first encoding network is an audio encoding network, the second encoding network is a video encoding network.
[0072] It should be understood that the pseudo-label configured for the second segment here is equivalent to the category labeled for the second segment. Since the performance of the first encoding network may not be reliable and stable enough during the model training stage, the category labeled for the second segment through clustering processing of the first prediction features generated by the first encoding network may also not be reliable and stable enough, so it is called a pseudo-label. Since there is a corresponding relationship between the video segment and the audio segment in the training sample, the category to which the first segment in the training sample belongs can also represent the category to which the second segment in the training sample belongs. This pseudo-label can be used as a supervised signal when training the second encoding network based on the second segment in the training sample.
[0073] Exemplarily, the server can use the K-Means clustering algorithm to perform clustering processing on the first prediction features corresponding to each training sample to obtain a preset number of clustering clusters. Among them, each clustering cluster corresponds to a category, and the category corresponding to the clustering cluster to which the first prediction feature corresponding to the training sample belongs is the category to which the first segment in the training sample belongs. Furthermore, for each training sample, the server can use the category to which the first segment in the training sample belongs as the pseudo-label corresponding to the second segment in the training sample. For example, the identifier of the category to which the first segment in the training sample belongs can be used as the pseudo-label corresponding to the second segment in the training sample.
[0074] For ease of understanding, the following uses the first encoding network as a video encoding network, the first segment as a video segment in the training samples, and the first predicted feature as the predicted video encoding feature of the video segment in the training samples as an example to introduce the above-mentioned pseudo-label configuration process. Assume that for the video segment vi in the i-th (i is an integer greater than or equal to 1) training sample, the video encoding network encodes to obtain the corresponding predicted video encoding feature F(vi); the server uses the K-Means algorithm to cluster the predicted video encoding features corresponding to each training sample, and divides the predicted video encoding features corresponding to each training sample into 256 categories; among them, the category to which the predicted video encoding feature corresponding to the video segment vi in the i-th training sample belongs is y vi , that is, the category to which the video segment vi in the i-th training sample belongs is y vi ; based on this, for the audio segment ai in the i-th training sample, the server can configure the corresponding pseudo-label for the audio segment ai as y vi , and this pseudo-label y vi can be used as the supervised signal when training the audio encoding network.
[0075] It should be understood that in practical applications, when the server clusters the first predicted features corresponding to each training sample, in addition to the K-Means clustering algorithm, other clustering algorithms can also be used, and this application does not make any limitations on the clustering algorithm used.
[0076] Step 204: For each of the training samples, through a second encoding network, determine the second predicted feature corresponding to the training sample according to the second segment in the training sample; according to the second predicted feature corresponding to the training sample, determine the class prediction result corresponding to the second segment in the training sample; the second encoding network is any one of the video encoding network and the audio encoding network, and is different from the first encoding network.
[0077] After the server obtains multiple training samples, it can also, for each training sample, determine the second predicted feature corresponding to the training sample through the second encoding network according to the second segment in the training sample, and then classify the second segment in the training sample according to the second predicted feature corresponding to the training sample to obtain the class prediction result corresponding to the second segment in the training sample.
[0078] It should be noted that the above second prediction feature is the prediction coding feature corresponding to the second segment in the training sample. For example, when the second segment is a video segment in the training sample and the second coding network is a video coding network, the second prediction feature is the predicted video coding feature corresponding to the video segment in the training sample. When the second segment is an audio segment in the training sample and the second coding network is an audio coding network, the second prediction feature is the predicted audio coding feature corresponding to the audio segment in the training sample.
[0079] As introduced in step 202 above, before performing feature encoding processing on the first segment in the training sample through the first coding network, it is necessary to preprocess the first segment in the training sample to obtain input data suitable for processing by the first coding network. The preprocessing methods for video segments and audio segments in the above step 202 have been introduced in detail. Similarly, before the server performs feature encoding processing on the second segment in the training sample through the second coding network, it is also necessary to preprocess the second segment in the training sample. When the second segment is a video segment in the training sample, the preprocessing method for video segments introduced in step 202 can be used for preprocessing. When the second segment is an audio segment in the training sample, the preprocessing method for audio segments introduced in step 202 can be used for preprocessing. For details, please refer to the relevant introduction content above and will not be elaborated here.
[0080] Exemplarily, the server can preprocess the second segment in each training sample through the corresponding preprocessing method to obtain input data suitable for processing by the second coding network. Furthermore, for each training sample, the server can input the input data obtained by preprocessing the second segment in the training sample into the second coding network. The second coding network analyzes and processes the input data and will correspondingly output the second prediction feature corresponding to the training sample. Then, the server can use the classifier to further predict the class prediction result corresponding to the second segment in the training sample according to the second prediction feature corresponding to the training sample. The class prediction result is the class to which the second segment in the training sample belongs predicted based on the second prediction feature corresponding to the training sample.
[0081] For the sake of easy understanding, the following takes the first coding network as a video coding network, the second coding network as an audio coding network, the first segment as a video segment in the training sample, and the second segment as an audio segment in the training sample as an example to introduce the above process of determining the class prediction result. Assume that through the processing of step 203, it is determined that the class to which the video segment vi in the i-th training sample belongs is y vi , and accordingly, the corresponding pseudo-label y is configured for the audio segment ai in the i-th training sample vi; For the i-th training sample, the server can perform feature encoding processing on the audio segment ai in the i-th training sample through an audio encoding network to obtain the corresponding predicted audio encoding feature G(ai), and determine the class prediction result corresponding to the audio segment ai through a classifier based on the predicted audio encoding feature G(ai).
[0082] It should be understood that in practical applications, the server can first execute step 202 and step 203, and then execute step 204, or first execute step 204, and then execute step 202 and step 203, or execute step 204, step 202 and step 203 simultaneously. This application does not make any limitation on the execution order between step 202 and step 203, and step 204. It should be noted that there is a timing correlation between step 202 and step 203. Step 202 needs to be executed first, and then step 203 is executed based on the execution result of step 202. Therefore, step 202 and step 203 can be regarded as a whole, and this whole is juxtaposed with step 204.
[0083] Step 205: Train the second encoding network based on the class prediction results and pseudo-labels corresponding to the second segments in the multiple training samples.
[0084] After the server obtains the pseudo-labels corresponding to the second segments in each training sample through step 203, and the class prediction results corresponding to the second segments in each training sample through step 204, it can construct a loss function based on the class prediction results and pseudo-labels corresponding to the second segments in each training sample, and adjust the model parameters in the second encoding network based on this loss function to achieve the purpose of training the second encoding network.
[0085] Exemplarily, the server can construct a cross-entropy loss function according to the class prediction results and pseudo-labels corresponding to the second segments in each training sample, and adjust the model parameters in the second encoding network with the goal of reducing this cross-entropy loss function. It should be understood that in practical applications, the server can also construct other types of loss functions. This application does not make any limitation on the type of loss function constructed when training the second encoding network.
[0086] It should be understood that when the second encoding network is a video encoding network, the server can achieve the training of the video encoding network through the above method; when the second encoding network is an audio encoding network, the server can achieve the training of the audio encoding network through the above method.
[0087] In order to achieve the collaborative training of the video encoding network and the audio encoding network, in the embodiments of the present application, the first encoding network can also be trained in a manner similar to that for training the second encoding network, that is, to achieve the collaborative training of the first encoding network and the second encoding network, so as to simultaneously train and obtain a video encoding network and an audio encoding network that can be put into practical application.
[0088] That is, the server can perform clustering processing based on the second prediction features corresponding to each training sample to determine the category to which the second segment in each training sample belongs; and for each training sample, according to the category to which the second segment in the training sample belongs, configure a corresponding pseudo-label for the first segment in the training sample. Then, for each training sample, the server can determine the category prediction result corresponding to the first segment in the training sample according to the first prediction feature corresponding to the training sample. Furthermore, based on the category prediction results and pseudo-labels corresponding to the first segments in multiple training samples, the first encoding network is trained.
[0089] For ease of understanding, the following takes the first encoding network as the video encoding network, the second encoding network as the audio encoding network, the first segment as the video segment in the training sample, and the second segment as the audio segment in the training sample as an example to exemplarily introduce the implementation manner of the above-mentioned collaborative training of the video encoding network and the audio encoding network.
[0090] Figure 3 It is a schematic diagram of the implementation principle of the collaborative training of the video encoding network and the audio encoding network provided by the embodiments of the present application. As Figure 3 shown in the training process on the left side, the server can first fix the video encoding network and train the audio encoding network. When training the audio encoding network, the server can use the fixed video encoding network to perform feature encoding processing on the video segments vi (i = 1, 2, 3,...) in each training sample to obtain the predicted video encoding features F(vi) (i.e., the first predicted encoding features) corresponding to each training sample; then, use the K-Means clustering algorithm to perform clustering processing on the predicted video encoding features F(vi) corresponding to each training sample, divide the predicted video encoding features corresponding to each training sample into 256 categories, and determine the category y to which the video segment in each training sample belongs vi , and configure a corresponding pseudo-label y for the audio segment ai in each training sample vi . Then, the server can use the currently trained audio encoding network to perform feature encoding processing on the audio segments ai in each training sample to obtain the predicted audio encoding features G(ai) (i.e., the second predicted encoding features) corresponding to each training sample, and determine the category prediction result corresponding to the audio segment in each training sample through the classifier according to the predicted audio encoding feature G(ai) Furthermore, the server can, according to the class prediction results corresponding to the audio segments in each training sample and the pseudo-label y vi construct a loss function and train the audio encoding network based on this loss function.
[0091] As Figure 3 shown in the training process on the right in ai , the server can fix the audio encoding network and train the video encoding network. When training the video encoding network, the server can use the fixed audio encoding network to perform feature encoding processing on the audio segments ai (i = 1, 2, 3,...) in each training sample to obtain the predicted audio encoding features G(ai) (i.e., the second predicted encoding features) corresponding to each training sample; then, use the K-Means clustering algorithm to perform clustering processing on the predicted audio encoding features G(ai) corresponding to each training sample, divide the predicted audio encoding features corresponding to each training sample into 256 categories, and determine the category y to which the audio segment in each training sample belongs ai . Then, for each training sample, the server can configure the corresponding pseudo-label y for the video segment vi. Furthermore, the server can, according to the class prediction results corresponding to the video segments in each training sample and the pseudo-label y ai construct a loss function and train the video encoding network based on this loss function.
[0092] In this way, the server can alternately train the video encoding network and the audio encoding network in the above manner, using the clustering result of the predicted features encoded by one of the encoding networks as the supervised signal used when training the other encoding network, to achieve the collaborative and efficient training of the video encoding network and the audio encoding network.
[0093] Optionally, considering that each time clustering processing is performed based on the predicted features (including the first predicted features and the second predicted features) corresponding to all training samples, it requires a large amount of computing resources and a long computing time; therefore, in order to save computing resources and computing time, in the embodiments of the present application, before performing clustering processing based on the predicted features corresponding to each training sample, the encoding network of the predicted features used for clustering processing can be tested, and when it is determined through testing that the performance of this encoding network meets the preset requirements, then clustering processing is performed based on the predicted features generated by this encoding network.
[0094] In a possible implementation, before clustering based on the first prediction features corresponding to each training sample, the server may obtain a plurality of test samples, where the test samples include video segments and their corresponding audio segments. Then, for each test sample, through the first encoding network, based on the first segment in the test sample, determine the first prediction feature corresponding to the test sample; and based on the first prediction feature corresponding to the test sample, determine the class prediction result corresponding to the first segment in the test sample. Next, based on the class prediction results and pseudo-labels corresponding to the first segments in each test sample, construct a first reference loss function; here, the pseudo-labels corresponding to the first segments in the test samples are determined by clustering the second prediction features corresponding to each test sample, and the second prediction features are determined by the second encoding network according to the second segments in the test samples. Furthermore, determine whether the first reference loss function satisfies the first preset loss condition; if so, perform clustering based on the first prediction features corresponding to each training sample to determine the class to which the first segment in each training sample belongs; if not, continue to train the first encoding network based on the plurality of training samples.
[0095] Specifically, before clustering based on the first prediction features corresponding to each training sample, the server may first use the test samples to test the first encoding network that generates the first prediction features. It should be noted that the test samples here are samples used to test the video encoding network and the audio encoding network during the training process, and the test samples also include video segments and audio segments with corresponding relationships; the acquisition method of the test samples is similar to the acquisition method of the training samples introduced above. Exemplarily, the server may use 90% of the video data in the open-source training video dataset (such as the AudioSet dataset) as the training samples provided, and use 10% of the video data in the training video dataset as the test samples provided.
[0096] The implementation manner in which the server determines the first prediction feature corresponding to the test sample through the first encoding network and determines the class prediction result corresponding to the first segment in the test sample based on the first prediction feature corresponding to the test sample is the same as the implementation manner introduced above in which the first prediction feature corresponding to the training sample is determined through the first encoding network and the class prediction result corresponding to the first segment in the training sample is determined based on the first prediction feature corresponding to the training sample, and will not be elaborated here.
[0097] After the server obtains the category prediction results corresponding to the first segments in each test sample, it can construct a first reference loss function based on the category prediction results corresponding to the first segments in each test sample and the pseudo-labels. For example, a cross-entropy loss function can be constructed. It should be noted that the method for determining the pseudo-labels corresponding to the first segments in the test samples here is the same as the method for determining the pseudo-labels corresponding to the first segments in the training samples in the previous text. Specifically, the second coding network trained in the previous training round can be used to determine the second prediction features corresponding to each test sample according to the second segments in each test sample, and then clustering processing can be performed based on the second prediction features corresponding to each test sample to determine the category to which the second segment in each test sample belongs. For each test sample, the category to which the second segment in the test sample belongs is used as the pseudo-label corresponding to the first segment in the test sample.
[0098] Furthermore, the server can determine whether the first reference loss function satisfies a first preset loss condition. Here, the first preset loss condition is a condition used to measure whether the current performance of the first coding network meets the requirements of the current training round for the first coding network. For example, the first preset loss condition can be that the decrease amplitude corresponding to the first reference loss function (which is the decrease amplitude of the first reference loss function determined this time relative to the first reference loss function determined last time) is less than the preset decrease amplitude, and the first preset loss condition can also be whether the loss value corresponding to the first reference loss function is less than the preset loss value, etc. The present application does not make any limitation on the first preset loss condition here. If the first reference loss function satisfies the first preset loss condition, it can be considered that the current first coding network has met the requirements of the current training round for the first coding network, and the first prediction features currently determined by the first coding network are relatively reliable. Therefore, clustering processing can be performed based on the first prediction features corresponding to each training sample determined by the first coding network. Correspondingly, the clustering result obtained through the clustering processing is also relatively reliable, which can ensure the reliability of the pseudo-labels configured for the second segments in the training samples. On the contrary, if the first reference loss function does not satisfy the first preset loss condition, it can be considered that the current first coding network does not meet the requirements of the current training round for the first coding network, and the first prediction features currently determined by the first coding network are not reliable. Correspondingly, the first coding network needs to be continuously trained based on each training sample.
[0099] In another possible implementation, before performing clustering processing based on the second prediction features corresponding to each training sample, the server may obtain a plurality of test samples, where the test samples include video segments and their corresponding audio segments. Then, for each test sample, through the second encoding network, based on the second segment in the test sample, determine the second prediction feature corresponding to the test sample; and based on the second prediction feature corresponding to the test sample, determine the class prediction result corresponding to the second segment in the test sample. Next, based on the class prediction results and pseudo-labels corresponding to the second segments in each test sample, construct a second reference loss function; here, the pseudo-labels corresponding to the second segments in the test samples are determined by performing clustering processing on the first prediction features corresponding to each test sample, and the first prediction feature is determined by the first encoding network based on the first segment in the test sample. Furthermore, determine whether the second reference loss function satisfies the second preset loss condition; if so, perform clustering processing based on the second prediction features corresponding to each training sample to determine the class to which the second segment in each training sample belongs; if not, continue to train the second encoding network based on the plurality of training samples.
[0100] Similarly, before performing clustering processing based on the second prediction features corresponding to each training sample, the server may also first test the second encoding network that generates the second prediction feature by using the test samples. The specific implementation of testing the second encoding network is similar to the implementation of testing the first encoding network in the above text. For details, please refer to the relevant introduction content in the above text and will not be elaborated here.
[0101] It should be understood that the second prediction loss condition used when testing the second encoding network is a condition for measuring whether the current performance of the second encoding network meets the requirements of the current training round for the second encoding network; for example, the second preset loss condition may be that the decrease amplitude corresponding to the second reference loss function (which is the decrease amplitude of the second reference loss function determined this time relative to the second reference loss function determined last time) is less than the preset decrease amplitude, and the second preset loss condition may also be whether the loss value corresponding to the second reference loss function is less than the preset loss value, etc. The present application does not make any limitation on the second preset loss condition here.
[0102] In this way, before performing clustering processing based on the prediction features (including the first prediction feature and the second prediction feature) corresponding to each training sample, the server tests the encoding network that generates the prediction feature by using the test samples, which can ensure the reliability of the prediction features used in the clustering processing, thereby facilitating reducing the number of clustering processing operations required in the process of co-training the video encoding network and the audio encoding network, and further improving the model training efficiency and reducing the waste of computing resources.
[0103] In the embodiments of the present application, in order to ensure that both the trained video coding network and the audio coding network have excellent performance, the server can iteratively train the video coding network and the audio coding network for multiple training rounds. Specifically, when the trained first coding network meets the first training end condition in the current training round, and the trained second coding network meets the second training end condition in the current training round, it can be determined that the model training for the current training round is completed; then, it is detected whether the number of completed training rounds has reached the preset number of training times; if so, it is determined that the training for the first coding network and the second coding network is completed; if not, the model training for the next training round is continued.
[0104] Specifically, when the server alternately trains the first coding network and the second coding network, it can detect whether the currently trained coding network meets the training end condition corresponding to this coding network in the current training round. Taking the example of first fixing the first coding network and training the second coding network, and then fixing the second coding network and training the first coding network during the alternating training process; when the server trains the second coding network, it can detect whether the trained second coding network meets the second training end condition in the current training round. If it meets, it can be determined that the training for the second coding network in the current training round is completed, and the training for the first coding network is started. If it does not meet, the second coding network needs to be continuously trained based on the training samples; here, the second training end condition is a condition used to measure whether to stop the training for the second coding network in the current training round. The second training end condition can be, for example, that the second reference loss function mentioned above meets the second preset loss condition, or the second training end condition can be, for example, that the number of training times for the second coding network in the current training round reaches the preset number of times, etc. The present application does not make any limitation on the second training end condition here.
[0105] After the server completes the training for the second coding network in the current training round, it can train the first coding network. During the process of training the first coding network, the server can detect whether the trained first coding network meets the first training end condition in the current training round. If it meets, it can be determined that the training for the first coding network in the current training round is completed, that is, it can be determined that the model training for the current round (including the training for the first coding network and the second coding network) is completed. If it does not meet, the first coding network needs to be continuously trained based on the training samples; here, the first training end condition is a condition used to measure whether to stop the training for the first coding network in the current training round. The first training end condition can be, for example, that the first reference loss function mentioned above meets the first preset loss condition, or the first training end condition can be, for example, that the number of training times for the first coding network in the current training round reaches the preset number of times, etc. The present application does not make any limitation on the first training end condition here.
[0106] After the server determines that the model training for the current training round is completed, it can detect whether the number of completed training rounds has reached a preset number of training times (such as 10 times); if it has reached, it can be determined that the training for the first encoding network and the second encoding network is completed, that is, it is determined that the training for the video encoding network and the audio encoding network is completed; if it has not reached, it can start the model training for the next training round.
[0107] In this way, through the above method, on the basis of alternately training the first encoding network and the second encoding network, both the first encoding network and the second encoding network are trained for multiple rounds, and it is ensured that the first encoding network and the second encoding network meet the corresponding training end conditions in each round of training, so that the trained first encoding network and the second encoding network can have better model performance.
[0108] Optionally, in the embodiments of the present application, in order to further improve the model training efficiency and ensure that the trained video encoding network and audio encoding network have better model performance, the embodiments of the present application can further introduce the idea of knowledge distillation in the process of constructing the loss function.
[0109] Knowledge distillation refers to using the features learned by an encoding network to influence the training of another encoding network, so as to achieve information transfer between the trained video encoding network and audio encoding network. In the embodiments of the present application, using the category to which a segment in the training sample belongs to configure a corresponding pseudo-label for another segment in the training sample can also essentially play a role in information transfer, but the efficiency of this information transfer method is relatively low. The reason is that the configured pseudo-label is essentially a "hard label", which can only represent that the data belongs to a certain category, and cannot further reflect the relationship between this data and other data, nor can it reflect the relationship between each category. For example, assume that the category to which a certain video segment belongs is "playing the piano". Then, since the relationship between "playing the piano" and actions such as "playing the guitar" and "playing the accordion" is relatively close, the predicted video encoding features corresponding to this video segment should be relatively close to the predicted video encoding features corresponding to the video segments belonging to categories such as "playing the guitar" and "playing the accordion". On the contrary, since the relationship between "playing the piano" and the action of "playing basketball" is relatively far, the predicted video encoding features corresponding to this video segment should be relatively distant from the predicted video encoding features corresponding to the video segments belonging to "playing basketball"; while the pseudo-label cannot reflect the proximity of the relationship between the above data and data, that is, during the model training process, the relationship between the video segment configured with the pseudo-label "playing the piano" and the video segment configured with "playing the guitar" is the same as the relationship between the video segment configured with the pseudo-label "playing the piano" and the video segment configured with "playing basketball"; and this is not conducive to improving the expression ability of the audio-visual encoding network to the audio-visual features it has learned.
[0110] Based on this, in the embodiments of the present application, when the server trains the second encoding network, it can train the second encoding network based on the category prediction results corresponding to the first segments in multiple training samples, the category prediction results corresponding to the second segments in multiple training samples, and the pseudo-labels. Similarly, when the server trains the first encoding network, it can also train the first encoding network based on the category prediction results corresponding to the second segments in multiple training samples, the category prediction results corresponding to the first segments in multiple training samples, and the pseudo-labels.
[0111] Specifically, when the server trains the second encoding network, it can determine the differences between the distribution of the class prediction results corresponding to the first segments in each training sample and the distribution of the class prediction results corresponding to the second segments in each training sample, and determine the differences between the class prediction results corresponding to the second segments in each training sample and the pseudo-labels, construct a loss function, and then train the second encoding network based on this loss function. Similarly, when the server trains the first encoding network, it can determine the differences between the distribution of the class prediction results corresponding to the first segments in each training sample and the distribution of the class prediction results corresponding to the second segments in each training sample, and determine the differences between the class prediction results corresponding to the first segments in each training sample and the pseudo-labels, construct a loss function, and then train the first encoding network based on this loss function.
[0112] In a possible implementation manner, the server can train the second encoding network in the following way: for each training sample, construct the basic loss function corresponding to the training sample according to the class prediction result and the pseudo-label corresponding to the second segment in the training sample; and, construct the distillation loss function corresponding to the training sample according to the class prediction results corresponding to the first segment and the second segment in the training sample respectively; furthermore, the server can train the second encoding network based on the basic loss function and the distillation loss function corresponding to each training sample respectively.
[0113] Specifically, the server can construct a cross-entropy loss function according to the difference between the class prediction result corresponding to the second segment in each training sample and the pseudo-label as the basic loss function corresponding to the training sample; in addition, for each training sample, the server can also construct the distillation loss function corresponding to the training sample according to the difference between the class prediction result corresponding to the first segment in the training sample and the class prediction result corresponding to the second segment in the training sample. Furthermore, the server can construct a comprehensive loss function according to the basic loss function and the distillation loss function corresponding to each training sample respectively; for example, the server can construct an overall basic loss function according to the basic loss function corresponding to each training sample, construct an overall distillation loss function according to the distillation loss function corresponding to each training sample, and then, the server can perform a weighted process on the above overall basic loss function and the overall distillation loss function to obtain a comprehensive loss function. Finally, the server can adjust the model parameters of the second encoding network based on this comprehensive loss function.
[0114] It should be understood that the training method of the first encoding network is similar to the training method of the second encoding network described above, and will not be elaborated here.
[0115] As an example, when constructing the distillation loss function corresponding to the training samples, the server may construct at least one of a response-based distillation loss function and a relationship-based distillation loss function. Specifically, the server may construct a first distillation loss function (i.e., the response-based distillation loss function) according to the difference between the class prediction results corresponding to the first segment and the second segment in the training sample; the server may construct a second distillation loss function (i.e., the relationship-based distillation loss function) according to the difference between the class prediction result corresponding to the first segment in the training sample and the class prediction result corresponding to the first segment in other training samples, and the difference between the class prediction result corresponding to the second segment in the training sample and the class prediction result corresponding to the second segment in other training samples.
[0116] Exemplarily, the class prediction result may include the probability that the segment belongs to each class. In addition to the positive label (i.e., the class to which the segment belongs, and the probability that the segment belongs to this class is the highest), the negative label (i.e., the class to which the segment does not belong, and the probability that the segment belongs to this class is not the highest) also contains a large amount of knowledge and relationships between classes inferred by the model. For example, for a well-trained classifier, since there is a high similarity between the three classes of "playing the piano", "playing the guitar", and "playing the accordion", the prediction result of the classifier for a certain segment in the class of "playing the piano" (i.e., the probability of predicting that the segment belongs to the class of "playing the piano") should have a high response with the prediction results for the segment in the classes of "playing the guitar" and "playing the accordion".
[0117] Based on the above knowledge, the embodiment of the present application constructs a response-based distillation loss function, that is, the first distillation loss function, for transmitting knowledge from the encoding network of one modality to the encoding network of another modality. Specifically, for each training sample, the server may construct the first distillation loss function corresponding to the training sample according to the difference between the class prediction result corresponding to the first segment in the training sample and the class prediction result corresponding to the second segment in the training sample. When training the second encoding network, the server needs to construct the overall first distillation loss function according to the first distillation loss functions corresponding to each training sample. Specifically, the server may construct the overall first distillation loss function L through the following formula (1): Response :
[0118]
[0119] where N is the total number of training samples, is the class prediction result corresponding to the video segment vi in the i-th training sample, is the class prediction result corresponding to the audio segment ai in the i-th training sample.
[0120] In addition to the relationships between categories, the relationships between training samples are also important knowledge to be learned. Transmitting the relationships between training samples learned by one encoding network to be trained to another encoding network, narrowing the distance between training samples of the same class and widening the distance between training samples of different classes can help improve the training speed of the two encoding networks.
[0121] Based on the above knowledge, the embodiment of the present application constructs a relationship-based distillation loss function, that is, the second distillation loss function, for transmitting the learned relationships between training samples from an encoding network of one modality to an encoding network of another modality. Specifically, for each training sample, the server can construct the second distillation loss function corresponding to the training sample according to the difference between the class prediction result corresponding to the first segment in the training sample and the class prediction results corresponding to the first segments in other training samples, and the difference between the class prediction result corresponding to the second segment in the training sample and the class prediction results corresponding to the second segments in other training samples. When training the second encoding network, the server needs to construct an overall second distillation loss function according to the second distillation loss functions corresponding to each training sample. Specifically, the server can construct the overall second distillation loss function L through the following formula (2) Relation :
[0122]
[0123] where N is the total number of training samples, is the class prediction result corresponding to the video segment vi in the i-th training sample, is the class prediction result corresponding to the video segment vj in the j-th training sample, is the class prediction result corresponding to the audio segment ai in the i-th training sample, is the class prediction result corresponding to the audio segment aj in the j-th training sample.
[0124] In this way, by the above method, introducing at least one of the first distillation loss function and the second distillation loss function can improve the training efficiency of the video encoding network and the audio encoding network based on the introduced distillation loss function, and improve the feature encoding ability of the trained video encoding network and audio encoding network.
[0125] The video encoding network and audio encoding network trained by the model training method provided by the embodiments of the present application can be further applied to downstream tasks to utilize the image features encoded by the video encoding network and the audio features encoded by the audio encoding network to assist in the implementation of downstream tasks. In the embodiments of the present application, the downstream tasks may include any task implemented based on the image features encoded by the video encoding network and the audio features encoded by the audio encoding network, such as video classification tasks, action recognition tasks, video background audio generation tasks, etc. The present application does not make any limitation on the downstream tasks herein.
[0126] In a possible implementation manner, the video encoding network and audio encoding network trained by the model training method provided by the embodiments of the present application can be applied to a target classification task, where the target classification task is a task of classifying a video based on the video frames in the video and the audio of the video. The target classification task may, for example, be an action recognition task (i.e., a task of identifying the category to which the actions present in the video belong), or may, for example, be a general video classification task (i.e., the category to which the video belongs, such as game videos, food videos, pet videos, beauty videos, etc.).
[0127] When applying the trained video encoding network and audio encoding network to the target classification task, a small number of pre-annotated first annotation samples can be used to train a classification model including the video encoding network and the audio encoding network. That is, the server can obtain a plurality of first annotation samples corresponding to the target classification task. The first annotation samples include video segments and audio segments with corresponding relationships, as well as classification labels. Then, the server can, through the video encoding network in the classification model to be trained, determine the image features corresponding to the first annotation sample according to the video segments in the first annotation sample, and through the audio encoding network in the classification model, determine the audio features corresponding to the first annotation sample according to the audio segments in the first annotation sample. Furthermore, through the classifier in the classification model, determine the category prediction result corresponding to the first annotation sample according to the image features and audio features corresponding to the first annotation sample. Finally, train the classification model based on the category prediction result corresponding to the first annotation sample and the classification label in the first annotation sample.
[0128] It should be noted that the first labeled sample is a supervised training sample for training a classification model for performing a target classification task; the first labeled sample includes video segments and audio segments with a corresponding relationship. Here, the video segments and audio segments with a corresponding relationship are obtained in a similar manner to the video segments and audio segments with a corresponding relationship included in the training sample introduced in step 201 above; the first labeled sample also includes a classification label, which is used to characterize the category to which the video segments and audio segments in the first labeled sample belong. The category characterized by the classification label is a certain classification category in the target classification task. For example, when the target classification task is an action recognition task, the category characterized by the classification label is the category to which the actions existing in the video belong.
[0129] In addition, since the classification model to be trained includes a video coding network and an audio coding network that have been trained by the model training method provided in the embodiments of the present application, and the video coding network and the audio coding network have excellent feature coding capabilities after training, when training the classification model, only a small number of first labeled samples under the target classification task are used to fine-tune the classification model, that is, only a small number of first labeled samples need to be obtained. In this way, the resources consumed for labeling the first labeled samples can be reduced.
[0130] Figure 4 It is a schematic diagram of the implementation principle of applying a video coding network and an audio coding network to a target classification task provided by the embodiments of the present application. As Figure 4As shown in the figure, after the server obtains the first labeled sample, it can first preprocess the video segment and audio segment in the first labeled sample to obtain input data that can be processed by the video encoding network and the audio encoding network. The preprocessing method has been introduced in detail in step 202 and will not be elaborated here. After the server obtains the input data corresponding to the video segment in the first labeled sample through preprocessing, it inputs the input data into the video encoding network of the classification model to be trained. By analyzing and processing the input data, the video encoding network can correspondingly output the image features corresponding to the first labeled sample. Similarly, after the server obtains the input data corresponding to the audio segment in the first labeled sample through preprocessing, it inputs the input data into the audio encoding network of the classification model to be trained. By analyzing and processing the input data, the audio encoding network can correspondingly output the audio features corresponding to the first labeled sample. Furthermore, the image features and audio features corresponding to the first labeled sample can be concatenated to obtain fused features, and the classifier in the classification model can determine the class prediction result corresponding to the first labeled sample based on the fused features. Finally, based on the class prediction result corresponding to the first labeled sample and the classification label in the first labeled sample, a loss function is constructed, and the classification model is trained based on the loss function to adjust the parameters of the video encoding network, audio encoding network, and classifier in the classification model.
[0131] In this way, applying the video encoding network and audio encoding network trained by the method provided in the embodiments of the present application to the classification model for implementing the target classification task can reduce the labeled samples required for training the classification model. It is only necessary to perform supervised fine-tuning on the classification model based on a small number of labeled samples under the target classification task, and a good model training effect can be achieved through short-term training, that is, it can be ensured that the trained classification model has a high accuracy rate.
[0132] As an example, the classification model trained in the above manner can be applied to target classification tasks such as action recognition and video classification. At this time, when performing the target classification task based on the classification model, the server can determine the image features corresponding to the first video segment to be processed through the video encoding network in the classification model; and through the audio encoding network in the classification model, based on the first audio segment corresponding to the first video segment to be processed, determine the audio features corresponding to the first audio segment to be processed; furthermore, through the classifier in the classification model, based on the image features corresponding to the first video segment to be processed and the audio features corresponding to the first audio segment to be processed, determine the classification result of the first video segment to be processed in the target classification task.
[0133] It should be noted that the first video segment to be processed and the first audio segment to be processed are the objects to be processed when performing the target classification task. That is, the first video segment to be processed is the video segment to be classified, and the first audio segment to be processed is the audio segment corresponding to the first video segment to be processed, such as the audio of the first video segment to be processed or the background audio corresponding to the first video segment to be processed.
[0134] When the server performs the target classification task, it can first preprocess the first video segment to be processed and the first audio respectively to obtain the input data that can be processed by the video encoding network and the audio encoding network. Then, through the video encoding network in the trained classification model, according to the input data corresponding to the first video segment to be processed, determine the image features corresponding to the first video segment to be processed; and, through the audio encoding network in the trained classification model, according to the input data corresponding to the first audio segment to be processed, determine the audio features corresponding to the first audio segment to be processed. Furthermore, through the classifier in the trained classification model, according to the concatenated features of the image features corresponding to the first video segment to be processed and the audio features corresponding to the first audio segment to be processed, obtain the classification result of the first video segment to be processed in the target classification task; for example, when the target classification task is an action recognition task, the classification result should be the category to which the action existing in the first video segment to be processed belongs, and for another example, when the target classification task is a general video classification task, the classification result should be the category to which the first video segment to be processed belongs, such as game videos, food videos, pet videos, beauty videos, and so on.
[0135] As an example, the classification model trained in the above manner can be applied to the action temporal localization task. The so-called action temporal localization task refers to detecting the category to which the actions existing in a video segment with a long duration belong and determining the occurrence time sequence of each detected action. When performing the action temporal localization task based on this classification model, the server can determine the sub-candidate video segments with preset actions existing in the second video segment to be processed and determine the arrangement order of each sub-candidate video segment. Then, for each sub-candidate video segment, the server can, through the video encoding network in the classification model, according to the sub-candidate video segment, determine the image features corresponding to the sub-candidate video segment; and through the audio encoding network in the classification model, according to the sub-candidate audio segment corresponding to the sub-candidate video segment, determine the audio features corresponding to the sub-candidate audio segment. Furthermore, through the classifier in the classification model, according to the image features corresponding to the sub-candidate video segment and the audio features corresponding to the sub-candidate audio segment, determine the action recognition result corresponding to the sub-candidate video segment. Finally, according to the action recognition results corresponding to each sub-candidate video segment and the arrangement order of each sub-candidate video segment, determine the action temporal localization result.
[0136] It should be noted that the second video segment to be processed is the object to be processed when performing the action timing localization task. The duration of this second video segment to be processed is usually long, and it may involve multiple actions, and it may include sub-segments where no action occurs.
[0137] When the server performs the action timing localization task for the second video segment to be processed, it can first use a regressor to detect sub-candidate video segments in the second video segment to be processed where actions may exist, and determine the start time and end time corresponding to each sub-candidate video segment, so as to determine the time arrangement order of each sub-candidate video segment accordingly. For each sub-candidate video segment, the server can first preprocess the sub-candidate video segment and its corresponding sub-candidate audio segment to obtain input data that can be processed by the video encoding network and the audio encoding network; then, the server can process the input data corresponding to the sub-candidate video segment through the video encoding network in the trained classification model to obtain the image features corresponding to the sub-candidate video segment; and, process the input data corresponding to the sub-candidate audio segment through the audio encoding network in the classification model to obtain the audio features corresponding to the sub-candidate audio segment; furthermore, through the classifier in the classification model, according to the concatenated features of the image features corresponding to the sub-candidate video segment and the audio features corresponding to the sub-candidate audio segment, determine the action recognition result corresponding to the sub-candidate video segment, that is, determine the action category to which the actions existing in the sub-candidate video segment belong. Finally, the server can arrange the action recognition results corresponding to each sub-candidate video segment in the time arrangement order of each sub-candidate video segment to obtain the action timing localization result corresponding to the second video segment to be processed.
[0138] In another possible implementation, the video encoding network and the audio encoding network trained by the model training method provided in the embodiments of the present application can be applied to the background audio generation task, which is a task for generating the corresponding background audio for a video segment.
[0139] When applying the trained video encoding network and audio encoding network to the background audio generation task, a small number of pre-annotated second annotation samples can be used to train the feature conversion network and the background audio generation model respectively; the feature conversion network here is a neural network for converting the image features corresponding to a video segment into corresponding audio features, and the background audio generation model here is a neural network model for generating the corresponding background audio according to audio features.
[0140] That is, the server can obtain multiple second annotation samples corresponding to the background audio generation task. The second annotation samples include video segments and annotated background audio segments with corresponding relationships. Then, the server can, through a pre-trained video encoding network, determine the image features corresponding to the video segment according to the video segment in the second annotation sample; and, through a pre-trained audio encoding network, determine the audio features corresponding to the annotated background audio segment according to the annotated background audio segment in the second annotation sample. Furthermore, the server can, through a feature transformation network to be trained, perform feature transformation processing on the image features corresponding to the video segment in the second annotation sample to obtain reference audio transformation features; and based on the audio features corresponding to the annotated background audio segment in the second annotation sample and the reference audio transformation features, train the feature transformation network. In addition, the server can, through a background audio generation model to be trained, generate a predicted background audio segment according to the audio features corresponding to the annotated background audio segment; and based on the annotated background audio segment and the predicted background audio segment, train the background audio generation model.
[0141] It should be noted that the second annotation sample is a supervised training sample for training the feature transformation network and the background audio generation model for performing the background audio generation task; the second annotation sample includes video segments and annotated background audio segments with corresponding relationships, and the annotated background audio segment here is the background audio annotated in combination with the content of the video segment.
[0142] Figure 5 This is a schematic diagram of the implementation principle of applying a video encoding network and an audio encoding network to the background audio generation task provided by an embodiment of the present application. As Figure 5 shown, for a certain second annotation sample, the server can first perform preprocessing on the video segment and the annotated background audio segment in the second annotation sample respectively to obtain the input data corresponding to the video segment and the annotated background audio segment respectively. Then, the server can process the input data corresponding to the video segment in the second annotation sample through a pre-trained video encoding network to obtain the image features corresponding to the video segment; and process the input data corresponding to the annotated background audio segment in the second annotation sample through a pre-trained audio encoding network to obtain the audio features corresponding to the annotated background audio segment.
[0143] Furthermore, the server can train the feature transformation network and the background audio generation model respectively based on the image features corresponding to the video segment in the second annotation sample and the audio features corresponding to the annotated background audio segment in the second annotation sample.
[0144] When specifically training the feature conversion network, the server can first use the feature conversion network to be trained to perform feature conversion processing on the image features corresponding to the video segments in the second labeled sample, so as to obtain the corresponding reference audio conversion features. Furthermore, the server can construct a loss function based on the difference between the reference audio conversion features and the audio features corresponding to the labeled background audio segments in the second labeled sample, and train the feature conversion network based on this loss function.
[0145] When specifically training the background audio generation model, the server can first use the background audio generation model to be trained to generate a predicted background audio segment according to the audio features corresponding to the labeled background audio segments in the second labeled sample. Furthermore, the server can construct a loss function based on the difference between the labeled background audio segment and the predicted background audio segment, and train the background audio generation model based on this loss function.
[0146] In this way, when using the video encoding network and audio encoding network trained by the method provided in the embodiments of the present application to train the feature conversion network and background audio generation model for realizing the background audio generation task, it can be ensured that the background audio generated based on the trained feature conversion network and background audio generation model is more matched with the video.
[0147] As an example, the feature conversion network and background audio generation model trained in the above manner can implement the background audio generation task in the following way: through the video encoding network, determine the image features corresponding to the target video segment according to the target video segment for which the background audio is to be generated; then, through the feature conversion network, perform feature conversion processing on the image features corresponding to the target video segment to obtain the audio conversion features corresponding to the target video segment; furthermore, through the background audio generation model, generate the background audio segment corresponding to the target video segment according to the audio conversion features corresponding to the target video segment.
[0148] Specifically, for the target video segment for which the background audio is to be generated, the server can first preprocess the target video segment to obtain the input data corresponding to the target video segment. Then, the server can input the input data corresponding to the target video segment into the pre-trained video encoding network. The video encoding network can analyze and process the input data and correspondingly output the image features corresponding to the target video segment. Furthermore, the server can input the image features corresponding to the target video segment into the pre-trained feature conversion network. The feature conversion network can perform feature conversion processing on the image features and correspondingly obtain the audio conversion features corresponding to the target video segment. Finally, the server can process the audio conversion features corresponding to the target video segment through the pre-trained background audio generation model to obtain the background audio segment corresponding to the target video segment.
[0149] It should be understood that the video encoding network and the audio encoding network trained by the method provided in the embodiments of the present application can be applied to other downstream tasks in addition to the above-mentioned downstream tasks. The present application does not make any limitation on the downstream tasks to which the video encoding network and the audio encoding network can be applied.
[0150] When training the video encoding network and the audio encoding network by the above model training method, the clustering result of the encoding features generated by one of the encoding networks is used to determine the supervised signal available when training the other encoding network. On the one hand, the annotation of training samples is avoided, that is, the processing resources consumed by annotating training samples are saved, and at the same time, the problem that the performance of the trained encoding network is poor due to the defects in the constructed training samples can also be avoided. On the other hand, since there is a corresponding relationship between the video segments and the audio segments in the training samples, based on the clustering result of the features corresponding to one type of segment in the training samples, a corresponding pseudo-label is configured for the other type of segment in the training samples, which can ensure the reliability of the configured pseudo-label to a certain extent. Correspondingly, using the pseudo-label as a supervised signal to train the other encoding network can ensure reliable training of the encoding network, that is, ensure that the trained video encoding network or audio encoding network can have better feature encoding ability and can be better applied to downstream tasks.
[0151] To facilitate further understanding of the model training method provided in the embodiments of the present application, the following Figure 6 is an overall exemplary introduction to the model training method in combination with the schematic diagram of the implementation principle shown.
[0152] In the embodiments of the present application, a video encoding network and audio encoding can be trained based on the AudioSet dataset. The AudioSet dataset includes more than two million videos. In the embodiments of the present application, 90% of the videos can be selected as training samples, and the remaining 10% of the videos can be used as test samples. The duration of each video in this dataset is about 10 seconds, and the frame rate is 30 frames per second. In the embodiments of the present application, video segments and audio segments with corresponding relationships and each having a duration of 2 seconds can be intercepted from each video to construct training samples and test samples. For the video segments in the training samples and test samples, the server can randomly sample 16 video frames from the 2-second video segments. For each video frame, it can be scaled without changing its aspect ratio so that the shorter side is 256 pixels, and then a region with a size of 224×224 pixels can be randomly intercepted from the scaled video frame. In this way, after the above preprocessing, a video tensor vi with a size of 16×3×224×224 corresponding to each video segment can be obtained as the input of the video encoding network. For the audio segments in the training samples and test samples, the server can perform short-time Fourier transform on the 2-second audio segments and take the logarithm of the short-time Fourier transform results to obtain a logarithmic spectrogram ai with time on the horizontal axis, frequency on the vertical axis, and intensity value (which is an audio tensor with a size of 40×100) as the input of the audio encoding network.
[0153] When training the video encoding network and the audio encoding network, the server can train them alternately. Specifically, the video encoding network can be fixed first to train the audio encoding network. That is, the video tensors vi (i = 1, 2, 3,...) corresponding to the video segments in each training sample are processed in sequence through the video encoding network to obtain the predicted video encoding features corresponding to each training sample. Then, the K-Means clustering algorithm is used to perform clustering processing on the predicted video encoding features corresponding to each training sample, and the predicted video encoding features are divided into 256 categories to determine the category y to which the video segment in each training sample belongs. vi And the category to which the video segment in each training sample belongs is used as the pseudo label corresponding to the audio segment in this training sample. Furthermore, the audio tensors ai corresponding to the audio segments in each training sample can be processed in sequence through the audio encoding network to obtain the predicted audio encoding features corresponding to each training sample, and the category prediction results corresponding to the audio segments in each training sample are determined according to the predicted audio encoding features corresponding to each training sample. Finally, the pseudo labels corresponding to the audio segments in each training sample can be used as supervised signals, and a cross-entropy loss function can be constructed based on the category prediction results and pseudo labels corresponding to the audio segments in each training sample to train the audio encoding network.
[0154] After completing the training of the audio encoding network in the current training round, the server can fix the audio encoding network and train the video encoding network. That is, the server can use the K-Means clustering algorithm to cluster the predicted audio encoding features corresponding to each training sample (determined by the audio encoding network trained in the current training round), divide the predicted audio encoding features into 256 categories, and determine the category y to which the audio segment in each training sample belongs. ai And use the category to which the audio segment in each training sample belongs as the pseudo-label corresponding to the video segment in this training sample. Then, the video tensors vi corresponding to the video segments in each training sample can be processed sequentially through the video encoding network to obtain the predicted video encoding features corresponding to each training sample, and based on the predicted video encoding features corresponding to each training sample, determine the category prediction results corresponding to the video segments in each training sample. Furthermore, the pseudo-labels corresponding to the video segments in each training sample can be used as supervised signals. Based on the category prediction results and pseudo-labels corresponding to the video segments in each training sample, a cross-entropy loss function can be constructed to train the video encoding network.
[0155] The server can repeat the steps of alternately training the video encoding network and the audio encoding network 10 times, that is, perform 10 rounds of training on the video encoding network and the audio encoding network respectively.
[0156] Considering that clustering based on the predicted encoding features corresponding to all training samples (including predicted video encoding features and predicted audio encoding features) is very time-consuming, therefore, before each clustering process, the encoding network used to generate the predicted encoding features for clustering can be tested using each test sample first. Before the loss function based on the test sample stops decreasing, continue to train the encoding network. After the encoding network is fully trained, then perform clustering based on the predicted encoding features generated by the encoding network to train another encoding network based on the results of the clustering process. In this way, the number of clustering processes can be reduced and the model training efficiency can be improved.
[0157] In addition, in the embodiments of the present application, during the process of training the video encoding network and the audio encoding network, the idea of knowledge distillation can also be introduced to transfer information between the trained video encoding network and audio encoding network, so that the encoding network of one modality can learn the knowledge of the other modality, thereby accelerating network training and improving performance.
[0158] Specifically, the category prediction results usually include the probabilities of a segment belonging to each category. Besides the positive label (i.e., the category to which the segment belongs and the probability of the segment belonging to this category is the highest), the negative labels (i.e., the categories to which the segment does not belong and the probability of the segment belonging to these categories is not the highest) also contain a lot of knowledge inferred by the model and the relationships between categories. For example, for a well-trained classifier, since there is a high similarity among the three categories of "playing the piano", "playing the guitar", and "playing the accordion", the prediction result of the classifier for a certain segment in the category of "playing the piano" (i.e., the probability of predicting that the segment belongs to the category of "playing the piano") should have a high response with the prediction results for the segment in the categories of "playing the guitar" and "playing the accordion". Based on this, in the embodiments of the present application, when training the video encoding network and the audio encoding network, a response-based distillation loss function as shown below can be constructed:
[0159]
[0160] where N is the total number of training samples, is the category prediction result corresponding to the video segment vi in the i-th training sample, is the category prediction result corresponding to the audio segment ai in the i-th training sample.
[0161] Besides the relationships between categories, the relationships between training samples are also important knowledge to be learned. Transmitting the relationships between training samples learned by one encoding network to be trained to another encoding network, narrowing the distance between similar training samples and widening the distance between different training samples can help improve the training speed of the two encoding networks. Based on this, in the embodiments of the present application, when training the video encoding network and the audio encoding network, a relationship-based distillation loss function as shown below can be constructed:
[0162]
[0163] where N is the total number of training samples, is the category prediction result corresponding to the video segment vi in the i-th training sample, is the category prediction result corresponding to the video segment vj in the j-th training sample, is the category prediction result corresponding to the audio segment ai in the i-th training sample, is the category prediction result corresponding to the audio segment aj in the j-th training sample.
[0164] When the server trains the video encoding network and the audio encoding network, if it constructs the cross-entropy loss function (constructed based on the difference between the class prediction result and the pseudo-label), as well as the above-mentioned response-based distillation loss function and relationship-based distillation loss function at the same time, then these three can be weighted to obtain a comprehensive loss function, and the video encoding network and the audio encoding network are trained based on this comprehensive loss function.
[0165] The video encoding network and the audio encoding network trained in the above way can be applied to classification tasks (such as action recognition tasks, video classification tasks, etc.). In this case, the server can obtain the labeled training samples under the corresponding classification tasks. Such labeled training samples may include video segments and audio segments with corresponding relationships, as well as the corresponding classification labels. Then, the server can determine the image features corresponding to the labeled training samples according to the video segments in the labeled training samples through the video encoding network in the classification model to be trained; and, through the audio encoding network in the classification model, determine the audio features corresponding to the labeled training samples according to the audio segments in the labeled training samples. Furthermore, the server can determine the class prediction result corresponding to the labeled training samples through the classifier in the classification model according to the video features and audio features corresponding to the labeled training samples; finally, the server can train the classification model according to the class prediction result corresponding to the labeled training samples and the classification labels in the labeled training samples.
[0166] More specifically, the video encoding network and the audio encoding network trained in the above way can be applied to classification tasks in game scenarios (such as game operation recognition tasks, game video classification tasks, etc.). For example, when applied to the game operation recognition task, the server can obtain the video segments and audio segments with corresponding relationships in the game video, as well as the corresponding classification labels (used to represent the game operations included in the game video) to construct labeled training samples, and then, based on the obtained labeled classification samples, train the game operation recognition model including the video encoding network and the audio encoding network. Another example, when applied to the game video classification task, the server can obtain the video segments and audio segments with corresponding relationships in the game video, as well as the corresponding classification labels (used to represent the category to which the game video belongs) to construct labeled training samples, and then, based on the obtained labeled classification samples, train the game video classification model including the video encoding network and the audio encoding network.
[0167] The video encoding network and the audio encoding network trained through the above method can be applied to background audio generation tasks (such as background music generation tasks, etc.). In this case, the server can obtain the labeled training samples under the corresponding background audio generation task. Such labeled training samples may include video segments and their corresponding labeled background audio segments. Then, the server can determine the image features corresponding to the labeled training samples according to the video segments in the labeled training samples through the pre-trained video encoding network; and, through the pre-trained audio encoding network, determine the audio features corresponding to the labeled training samples according to the labeled background audio segments in the labeled training samples. Furthermore, the server can perform conversion processing on the image features corresponding to the labeled training samples through the feature conversion network to be trained, obtain the audio conversion features corresponding to the labeled training samples, and train the feature conversion network according to the distance (such as cosine distance, etc.) between the audio conversion features and the audio features corresponding to the labeled training samples. In addition, the server can also generate predicted background audio segments according to the audio features corresponding to the labeled training samples through the background audio generation model to be trained, and train the background audio generation model according to the predicted background audio segments and the labeled background audio segments in the labeled training samples.
[0168] More specifically, the video encoding network and the audio encoding network trained through the above method can be applied to the background audio generation task in the game scenario, that is, generate background audio matching the operation rhythm in the game video. In this case, the server can obtain the labeled training samples including game video segments and their corresponding labeled background audio segments; then, respectively through the pre-trained video encoding network and audio encoding network, generate the image features and audio features corresponding to the labeled training samples according to the game video segments and the labeled background audio segments in the labeled training samples; furthermore, train the feature conversion network for converting image features into audio features based on the image features and audio features corresponding to the labeled training samples; and, train the background audio generation model for predicting the background audio of the game video based on the audio features corresponding to the labeled training samples and the labeled background audio segments.
[0169] For the model training method described above, the present application also provides a corresponding model training device to enable the above model training method to be applied and implemented in practice.
[0170] See Figure 7 , Figure 7 is a schematic structural diagram of a model training device 700 corresponding to the model training method shown above. As Figure 2 shown, the model training device 700 includes: Figure 7 shown, the model training device 700 includes:
[0171] A training sample acquisition module 701, configured to acquire a plurality of training samples; the training samples include video segments and their corresponding audio segments;
[0172] A first feature prediction module 702, configured to, for each of the training samples, determine a first predicted feature corresponding to the training sample according to a first segment in the training sample through a first encoding network; the first encoding network is either a video encoding network or an audio encoding network;
[0173] A first feature clustering module 703, configured to perform clustering processing based on the first predicted features respectively corresponding to the plurality of training samples, determine the category to which the first segment in each training sample belongs; and for each training sample, configure a corresponding pseudo label for a second segment in the training sample according to the category to which the first segment in the training sample belongs; the second segment is different from the first segment;
[0174] A second network prediction module 704, configured to, for each of the training samples, determine a second predicted feature corresponding to the training sample according to a second segment in the training sample through a second encoding network; determine a category prediction result corresponding to the second segment in the training sample according to the second predicted feature corresponding to the training sample; the second encoding network is either the video encoding network or the audio encoding network, and is different from the first encoding network;
[0175] A second network training module 705, configured to train the second encoding network based on the category prediction results and pseudo labels respectively corresponding to the second segments in the plurality of training samples.
[0176] Optionally, the apparatus further includes:
[0177] A second feature clustering module, configured to perform clustering processing based on the second predicted features respectively corresponding to the plurality of training samples, determine the category to which the second segment in each training sample belongs; and for each training sample, configure a corresponding pseudo label for a first segment in the training sample according to the category to which the second segment in the training sample belongs;
[0178] A first network prediction module, configured to, for each of the training samples, determine a category prediction result corresponding to the first segment in the training sample according to the first predicted feature corresponding to the training sample;
[0179] A first network training module, configured to train the first encoding network based on the category prediction results and pseudo labels respectively corresponding to the first segments in the plurality of training samples.
[0180] Optionally, the apparatus further includes a first network testing module; the first network testing module is configured to:
[0181] Obtain a plurality of test samples; the test samples include video segments and their corresponding audio segments;
[0182] For each of the test samples, through the first encoding network, determine a first predicted feature corresponding to the test sample according to a first segment in the test sample; according to the first predicted feature corresponding to the test sample, determine a class prediction result corresponding to the first segment in the test sample;
[0183] Construct a first reference loss function based on the class prediction results and pseudo-labels respectively corresponding to the first segments in the plurality of test samples; the pseudo-label corresponding to the first segment in the test sample is determined by performing clustering processing on the second predicted features respectively corresponding to the plurality of test samples, and the second predicted features are determined by the second encoding network according to a second segment in the test sample;
[0184] Judge whether the first reference loss function satisfies a first preset loss condition; if so, perform clustering processing based on the first predicted features respectively corresponding to the plurality of training samples to determine the class to which the first segment in each training sample belongs; if not, continue to train the first encoding network based on the plurality of training samples.
[0185] Optionally, the apparatus further includes a second network testing module; the second network testing module is configured to:
[0186] Obtain a plurality of test samples; the test samples include video segments and their corresponding audio segments;
[0187] For each of the test samples, through the second encoding network, determine a second predicted feature corresponding to the test sample according to a second segment in the test sample; according to the second predicted feature corresponding to the test sample, determine a class prediction result corresponding to the second segment in the test sample;
[0188] Construct a second reference loss function based on the class prediction results and pseudo-labels respectively corresponding to the second segments in the plurality of test samples; the pseudo-label corresponding to the second segment in the test sample is determined by performing clustering processing on the first predicted features respectively corresponding to the plurality of test samples, and the first predicted features are determined by the first encoding network according to a first segment in the test sample;
[0189] Determine whether the second reference loss function satisfies the second preset loss condition; if so, perform clustering processing based on the second prediction features corresponding to the multiple training samples respectively, and determine the category to which the second segment in each training sample belongs; if not, continue to train the second encoding network based on the multiple training samples.
[0190] Optionally, the apparatus further includes a training round detection module; the training round detection module is configured to:
[0191] When the trained first encoding network satisfies the first training end condition in the current training round, and the trained second encoding network satisfies the second training end condition in the current training round, determine that the model training for the current training round is completed;
[0192] Detect whether the number of completed training rounds has reached the preset number of training times;
[0193] If so, determine that the training for the first encoding network and the second encoding network is completed; if not, continue to perform the model training for the next training round.
[0194] Optionally, the second network training module 705 is specifically configured to:
[0195] Train the second encoding network based on the class prediction results corresponding to the first segments in the multiple training samples, the class prediction results corresponding to the second segments in the multiple training samples, and the pseudo labels.
[0196] Optionally, the second network training module 705 is specifically configured to:
[0197] For each training sample, construct a basic loss function corresponding to the training sample according to the class prediction result corresponding to the second segment in the training sample and the pseudo label; and construct a distillation loss function corresponding to the training sample according to the class prediction results corresponding to the first segment and the second segment in the training sample;
[0198] Train the second encoding network based on the basic loss functions and distillation loss functions corresponding to the multiple training samples respectively.
[0199] Optionally, the second network training module 705 is specifically configured to construct a distillation loss function corresponding to a training sample by at least one of the following methods:
[0200] Construct a first distillation loss function according to the difference between the class prediction results corresponding to the first segment and the second segment in the training sample;
[0201] Construct a second distillation loss function based on the differences between the class prediction results corresponding to the first segments in the training samples and the class prediction results corresponding to the first segments in other training samples, and the differences between the class prediction results corresponding to the second segments in the training samples and the class prediction results corresponding to the second segments in other training samples.
[0202] Optionally, the apparatus further includes a classification model training module, which is configured to:
[0203] Obtain a plurality of first labeled samples corresponding to the target classification task; the first labeled samples include video segments and audio segments with corresponding relationships, as well as classification labels;
[0204] Through the video encoding network in the classification model to be trained, determine the image features corresponding to the first labeled sample according to the video segment in the first labeled sample; through the audio encoding network in the classification model, determine the audio features corresponding to the first labeled sample according to the audio segment in the first labeled sample;
[0205] Through the classifier in the classification model, determine the class prediction result corresponding to the first labeled sample according to the image features and audio features corresponding to the first labeled sample;
[0206] Train the classification model based on the class prediction result corresponding to the first labeled sample and the classification label in the first labeled sample.
[0207] Optionally, the apparatus further includes a first classification model application module, which is configured to:
[0208] After completing the training of the classification model, through the video encoding network in the classification model, determine the image features corresponding to the first video segment to be processed according to the first video segment to be processed; through the audio encoding network in the classification model, determine the audio features corresponding to the first audio segment to be processed according to the first audio segment corresponding to the first video segment to be processed;
[0209] Through the classifier in the classification model, determine the classification result of the first video segment to be processed in the target classification task according to the image features corresponding to the first video segment to be processed and the audio features corresponding to the first audio segment to be processed.
[0210] Optionally, the apparatus further includes a second classification model application module, which is configured to:
[0211] When the target classification task is an action temporal location task, after completing the training of the classification model, for the second video segment to be processed, determine the sub-candidate video segments with preset actions in the second video segment to be processed, and determine the arrangement order of each sub-candidate video segment;
[0212] For each sub-candidate video segment, through the video encoding network in the classification model, determine the image features corresponding to the sub-candidate video segment according to the sub-candidate video segment; through the audio encoding network in the classification model, determine the audio features corresponding to the sub-candidate audio segment according to the sub-candidate audio segment corresponding to the sub-candidate video segment; through the classifier in the classification model, determine the action recognition result corresponding to the sub-candidate video segment according to the image features corresponding to the sub-candidate video segment and the audio features corresponding to the sub-candidate audio segment;
[0213] Determine the action temporal location result according to the action recognition results corresponding to each sub-candidate video segment and the arrangement order of each sub-candidate video segment.
[0214] Optionally, the device further includes a background audio generation model training module; the background audio generation model training module is used for:
[0215] Obtain a plurality of second annotation samples corresponding to the background audio generation task; the second annotation samples include video segments and annotated background audio segments with corresponding relationships;
[0216] Through the video encoding network, determine the image features corresponding to the video segment according to the video segment in the second annotation sample; through the audio encoding network, determine the audio features corresponding to the annotated background audio segment according to the annotated background audio segment in the second annotation sample;
[0217] Through the feature conversion network to be trained, perform feature conversion processing on the image features corresponding to the video segment in the second annotation sample to obtain reference audio conversion features; and based on the audio features corresponding to the annotated background audio segment in the second annotation sample and the reference audio conversion features, train the feature conversion network;
[0218] Through the background audio generation model to be trained, generate a predicted background audio segment according to the audio features corresponding to the annotated background audio segment; based on the annotated background audio segment and the predicted background audio segment, train the background audio generation model.
[0219] Optionally, the device further includes a background audio generation model application module; the background audio generation model application module is used for:
[0220] After completing the training of the feature conversion network and the background audio generation model, through the video encoding network, according to the target video segment for which the background audio is to be generated, determine the image features corresponding to the target video segment;
[0221] Through the feature conversion network, perform feature conversion processing on the image features corresponding to the target video segment to obtain the audio conversion features corresponding to the target video segment;
[0222] Through the background audio generation model, generate the background audio segment corresponding to the target video segment according to the audio conversion features corresponding to the target video segment.
[0223] When the above model training device trains the video encoding network and the audio encoding network, it uses the clustering result of the encoding features generated by one of the encoding networks to determine the supervised signal available when training the other encoding network. On the one hand, it avoids annotating training samples, that is, it saves the processing resources consumed by annotating training samples, and at the same time, it can also avoid the problem that the performance of the trained encoding network is not good due to defects in the constructed training samples. On the other hand, since there is a corresponding relationship between the video segments and the audio segments in the training samples, therefore, based on the feature clustering result corresponding to one type of segment in the training sample, configuring the corresponding pseudo-label for the other type of segment in the training sample can ensure the reliability of the configured pseudo-label to a certain extent. Correspondingly, using this pseudo-label as the supervised signal to train the other encoding network can ensure reliable training of this encoding network, that is, ensure that the trained video encoding network or audio encoding network can have better feature encoding capabilities and can be better applied to downstream tasks.
[0224] The embodiment of the present application also provides a computer device for training a model. This computer device can specifically be a terminal device or a server. Below, the terminal device and the server provided by the embodiment of the present application will be introduced from the perspective of hardware implementation.
[0225] See Figure 8 , Figure 8 is a schematic structural diagram of the terminal device provided by the embodiment of the present application. As Figure 8 shown, for the sake of convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. This terminal can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, PDA), a point of sales (POS), an in-vehicle computer, etc. Taking the terminal as a computer as an example:
[0226] Figure 8The block diagram of a part of the structure of a computer related to the terminal provided by an embodiment of the present application is shown. Refer to Figure 8 , the computer includes: a Radio Frequency (RF) circuit 810, a memory 820, an input unit 830 (including a touch panel 831 and other input devices 832), a display unit 840 (including a display panel 841), a sensor 850, an audio circuit 860 (which can be connected to a speaker 861 and a microphone 862), a wireless fidelity (WiFi) module 870, a processor 880, and a power supply 890, etc. Those skilled in the art can understand that Figure 8 the computer structure shown in
[0227] does not constitute a limitation on the computer, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0228] The processor 880 is the control center of the computer, connecting various parts of the entire computer through various interfaces and lines, and executing various functions of the computer and processing data by running or executing the software programs and / or modules stored in the memory 820, and calling the data stored in the memory 820. Optionally, the processor 880 may include one or more processing units; preferably, the processor 880 may integrate an application processor and a modulation and demodulation processor, where the application processor mainly processes the operating system, user interface, and application programs, etc., and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above modulation and demodulation processor may not be integrated into the processor 880 either.
[0229] In the embodiment of the present application, the processor 880 included in the terminal is further used to execute the steps of any implementation manner of the model training method provided by the embodiment of the present application.
[0230] See Figure 9 , Figure 9Schematic diagram of the structure of a server 900 provided by an embodiment of the present application. The server 900 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 922 (for example, one or more processors) and a memory 932, and one or more storage media 930 (for example, one or more mass storage devices) for storing application programs 942 or data 944. Among them, the memory 932 and the storage media 930 may be transient storage or persistent storage. The program stored in the storage media 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 922 may be configured to communicate with the storage media 930 and execute a series of instruction operations in the storage media 930 on the server 900.
[0231] The server 900 may further include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input / output interfaces 958, and / or one or more operating systems, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0232] The steps performed by the server in the above embodiments may be based on the Figure 9 server structure shown.
[0233] Among them, the CPU 922 is used to execute the steps of any implementation manner of the model training method provided by the embodiment of the present application.
[0234] The embodiment of the present application further provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute any implementation manner of the model training method described in the foregoing embodiments.
[0235] The embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes any implementation manner of the model training method described in the foregoing embodiments.
[0236] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0237] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0238] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0239] In addition, in each embodiment of the present application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0240] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store computer programs.
[0241] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0242] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than limiting them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A model training method, characterized in that, The method includes: Obtaining a plurality of training samples; the training samples include video clips and their corresponding audio clips; For each of the training samples, through a first encoding network, according to the first clip in the training sample, determining a first predicted feature corresponding to the training sample; the first encoding network is either a video encoding network or an audio encoding network; Performing clustering processing based on the first predicted features corresponding to the plurality of training samples respectively, determining the category to which the first clip in each training sample belongs; and for each training sample, according to the category to which the first clip in the training sample belongs, configuring a corresponding pseudo-label for the second clip in the training sample; the second clip is different from the first clip; For each of the training samples, through a second encoding network, according to the second clip in the training sample, determining a second predicted feature corresponding to the training sample; according to the second predicted feature corresponding to the training sample, determining a category prediction result corresponding to the second clip in the training sample; the second encoding network is either the video encoding network or the audio encoding network, and is different from the first encoding network; For each of the training samples, according to the first predicted feature corresponding to the training sample, determining a category prediction result corresponding to the first clip in the training sample; Training the second encoding network based on the category prediction results and pseudo-labels corresponding to the second clips in the plurality of training samples respectively, including: for each training sample, constructing a basic loss function corresponding to the training sample according to the category prediction result and pseudo-label corresponding to the second clip in the training sample; and constructing a distillation loss function corresponding to the training sample according to the category prediction results corresponding to the first clip and the second clip in the training sample respectively; training the second encoding network based on the basic loss functions and distillation loss functions corresponding to the plurality of training samples respectively.
2. The method according to claim 1, wherein The method further includes: Performing clustering processing based on the second predicted features corresponding to the plurality of training samples respectively, determining the category to which the second clip in each training sample belongs; and for each training sample, according to the category to which the second clip in the training sample belongs, configuring a corresponding pseudo-label for the first clip in the training sample; For each of the training samples, according to the first predicted feature corresponding to the training sample, determining a category prediction result corresponding to the first clip in the training sample; Training the first encoding network based on the category prediction results and pseudo-labels corresponding to the first clips in the plurality of training samples respectively.
3. The method according to claim 1, wherein The method further includes: Obtaining a plurality of test samples; the test samples include video clips and their corresponding audio clips; For each of the test samples, through the first encoding network, according to the first clip in the test sample, determining a first predicted feature corresponding to the test sample; according to the first predicted feature corresponding to the test sample, determining a category prediction result corresponding to the first clip in the test sample; Construct a first reference loss function based on the class prediction results and pseudo-labels corresponding to the first segments in the multiple test samples; the pseudo-labels corresponding to the first segments in the test samples are determined by clustering the second prediction features corresponding to the multiple test samples, and the second prediction features are determined by the second encoding network according to the second segments in the test samples; Determine whether the first reference loss function satisfies a first preset loss condition; if so, perform clustering on the first prediction features corresponding to the multiple training samples to determine the class to which the first segment in each training sample belongs; if not, continue to train the first encoding network based on the multiple training samples.
4. The method according to claim 2, wherein The method further includes: Obtain multiple test samples; the test samples include video segments and their corresponding audio segments; For each of the test samples, use the second encoding network to determine the second prediction feature corresponding to the test sample according to the second segment in the test sample; determine the class prediction result corresponding to the second segment in the test sample according to the second prediction feature corresponding to the test sample; Construct a second reference loss function based on the class prediction results and pseudo-labels corresponding to the second segments in the multiple test samples; the pseudo-labels corresponding to the second segments in the test samples are determined by clustering the first prediction features corresponding to the multiple test samples, and the first prediction features are determined by the first encoding network according to the first segments in the test samples; Determine whether the second reference loss function satisfies a second preset loss condition; if so, perform clustering on the second prediction features corresponding to the multiple training samples to determine the class to which the second segment in each training sample belongs; if not, continue to train the second encoding network based on the multiple training samples.
5. The method according to any one of claims 2 to 4, characterized in that The method further includes: When the trained first encoding network satisfies the first training end condition in the current training round and the trained second encoding network satisfies the second training end condition in the current training round, determine that the model training for the current training round is completed; Detect whether the number of completed training rounds has reached a preset number of training times; If so, determine that the training of the first encoding network and the second encoding network is completed; if not, continue to perform model training for the next training round.
6. The method according to claim 1, wherein Constructing the distillation loss function corresponding to the training sample according to the class prediction results corresponding to the first segment and the second segment in the training sample includes at least one of the following: Construct a first distillation loss function according to the difference between the class prediction results corresponding to the first segment and the second segment in the training sample; Construct a second distillation loss function based on the difference between the class prediction results corresponding to the first segments in the training samples and the class prediction results corresponding to the first segments in other training samples, and the difference between the class prediction results corresponding to the second segments in the training samples and the class prediction results corresponding to the second segments in other training samples.
7. The method according to claim 1 or 2, characterized in that, The method further includes: Obtain a plurality of first labeled samples corresponding to the target classification task; the first labeled samples include video segments and audio segments with corresponding relationships, as well as classification labels; Through the video encoding network in the classification model to be trained, determine the image features corresponding to the first labeled samples according to the video segments in the first labeled samples; through the audio encoding network in the classification model, determine the audio features corresponding to the first labeled samples according to the audio segments in the first labeled samples; Through the classifier in the classification model, determine the class prediction results corresponding to the first labeled samples according to the image features and audio features corresponding to the first labeled samples; Train the classification model based on the class prediction results corresponding to the first labeled samples and the classification labels in the first labeled samples.
8. The method according to claim 7, wherein After completing the training of the classification model, the method further includes: Through the video encoding network in the classification model, determine the image features corresponding to the first video segment to be processed according to the first video segment to be processed; through the audio encoding network in the classification model, determine the audio features corresponding to the first audio segment to be processed according to the first audio segment to be processed corresponding to the first video segment to be processed; Through the classifier in the classification model, determine the classification result of the first video segment to be processed in the target classification task according to the image features corresponding to the first video segment to be processed and the audio features corresponding to the first audio segment to be processed.
9. The method according to claim 7, wherein When the target classification task is an action temporal localization task, after completing the training of the classification model, the method further includes: For the second video segment to be processed, determine the sub-candidate video segments with preset actions in the second video segment to be processed, and determine the arrangement order of each sub-candidate video segment; For each sub-candidate video segment, through the video encoding network in the classification model, determine the image features corresponding to the sub-candidate video segment according to the sub-candidate video segment; through the audio encoding network in the classification model, determine the audio features corresponding to the sub-candidate audio segment according to the sub-candidate audio segment corresponding to the sub-candidate video segment; through the classifier in the classification model, determine the action recognition result corresponding to the sub-candidate video segment according to the image features corresponding to the sub-candidate video segment and the audio features corresponding to the sub-candidate audio segment; Determine the action temporal localization result according to the action recognition results corresponding to each sub-candidate video segment and the arrangement order of each sub-candidate video segment.
10. The method according to claim 1 or 2, characterized in that, The method further includes: Obtain multiple second annotation samples corresponding to the background audio generation task; the second annotation samples include video segments and annotated background audio segments with corresponding relationships. Through the video encoding network, determine the image features corresponding to the video segments according to the video segments in the second annotation samples; through the audio encoding network, determine the audio features corresponding to the annotated background audio segments according to the annotated background audio segments in the second annotation samples. Through the feature conversion network to be trained, perform feature conversion processing on the image features corresponding to the video segments in the second annotation samples to obtain reference audio conversion features; and based on the audio features corresponding to the annotated background audio segments in the second annotation samples and the reference audio conversion features, train the feature conversion network. Through the background audio generation model to be trained, generate a predicted background audio segment according to the audio features corresponding to the annotated background audio segments; based on the annotated background audio segments and the predicted background audio segments, train the background audio generation model.
11. The method according to claim 10, wherein After completing the training of the feature conversion network and the background audio generation model, the method further includes: Through the video encoding network, determine the image features corresponding to the target video segment according to the target video segment for which the background audio is to be generated. Through the feature conversion network, perform feature conversion processing on the image features corresponding to the target video segment to obtain the audio conversion features corresponding to the target video segment. Through the background audio generation model, generate the background audio segment corresponding to the target video segment according to the audio conversion features corresponding to the target video segment.
12. A model training device, characterized in that, The device includes: A training sample acquisition module, configured to obtain multiple training samples; the training samples include video segments and their corresponding audio segments. A first feature prediction module, configured to, for each of the training samples, through a first encoding network, determine the first predicted feature corresponding to the training sample according to the first segment in the training sample; the first encoding network is either the video encoding network or the audio encoding network. A first feature clustering module, configured to perform clustering processing based on the first predicted features corresponding to the multiple training samples respectively, determine the category to which the first segment in each training sample belongs; and for each training sample, configure a corresponding pseudo-label for the second segment in the training sample according to the category to which the first segment in the training sample belongs; the second segment is different from the first segment. A second network prediction module, configured to, for each of the training samples, through a second encoding network, determine the second predicted feature corresponding to the training sample according to the second segment in the training sample; and determine the category prediction result corresponding to the second segment in the training sample according to the second predicted feature corresponding to the training sample; the second encoding network is either the video encoding network or the audio encoding network, and is different from the first encoding network. The first network prediction module is used to determine the class prediction result corresponding to the first segment in each training sample according to the first prediction feature corresponding to the training sample; The second network training module is used to train the second encoding network based on the class prediction results and pseudo-labels corresponding to the second segments in the multiple training samples, including: for each training sample, constructing a basic loss function corresponding to the training sample according to the class prediction result and pseudo-label corresponding to the second segment in the training sample; and constructing a distillation loss function corresponding to the training sample according to the class prediction results corresponding to the first segment and the second segment in the training sample; training the second encoding network based on the basic loss functions and distillation loss functions corresponding to the multiple training samples respectively.
13. The device according to claim 12, characterized in that, The apparatus further includes a second feature clustering module and a first network training module; The second feature clustering module is used to perform clustering processing based on the second prediction features corresponding to the multiple training samples respectively, to determine the class to which the second segment in each training sample belongs; and for each training sample, configuring a corresponding pseudo-label for the first segment in the training sample according to the class to which the second segment in the training sample belongs; The first network prediction module is further used to determine the class prediction result corresponding to the first segment in each training sample according to the first prediction feature corresponding to the training sample; The first network training module is used to train the first encoding network based on the class prediction results and pseudo-labels corresponding to the first segments in the multiple training samples respectively.
14. The device according to claim 12, characterized in that The apparatus further includes a first network testing module, and the first network testing module is used for: Obtaining a plurality of test samples; the test samples include video segments and their corresponding audio segments; For each test sample, through the first encoding network, determining the first prediction feature corresponding to the test sample according to the first segment in the test sample; Determining the class prediction result corresponding to the first segment in the test sample according to the first prediction feature corresponding to the test sample; Constructing a first reference loss function based on the class prediction results and pseudo-labels corresponding to the first segments in the multiple test samples respectively; the pseudo-label corresponding to the first segment in the test sample is determined by performing clustering processing on the second prediction features corresponding to the multiple test samples respectively, and the second prediction feature is determined by the second encoding network according to the second segment in the test sample; Judging whether the first reference loss function satisfies a first preset loss condition; if so, performing clustering processing based on the first prediction features corresponding to the multiple training samples respectively to determine the class to which the first segment in each training sample belongs; if not, continuing to train the first encoding network based on the multiple training samples.
15. The device according to claim 13, characterized in that The apparatus further includes a second network testing module, and the second network testing module is used for: Obtaining a plurality of test samples; the test samples include video segments and their corresponding audio segments; For each of the test samples, through the second encoding network, determine the second prediction feature corresponding to the test sample according to the second segment in the test sample; According to the second prediction feature corresponding to the test sample, determine the class prediction result corresponding to the second segment in the test sample; Based on the class prediction results and pseudo-labels respectively corresponding to the second segments in the multiple test samples, construct a second reference loss function; the pseudo-label corresponding to the second segment in the test sample is determined by clustering the first prediction features respectively corresponding to the multiple test samples, and the first prediction feature is determined by the first encoding network according to the first segment in the test sample; Determine whether the second reference loss function satisfies the second preset loss condition; if so, perform clustering on the second prediction features respectively corresponding to the multiple training samples to determine the class to which the second segment in each training sample belongs; if not, continue to train the second encoding network based on the multiple training samples.
16. The device according to any one of claims 13 to 15, characterized in that, The apparatus further includes a training round detection module, and the training round detection module is configured to: When the trained first encoding network satisfies the first training end condition in the current training round, and the trained second encoding network satisfies the second training end condition in the current training round, determine that the model training for the current training round is completed; Detect whether the number of completed training rounds has reached a preset number of training times; If so, determine that the training of the first encoding network and the second encoding network is completed; if not, continue to perform model training for the next training round.
17. The device according to claim 12, characterized in that, The second network training module is specifically configured to construct a distillation loss function corresponding to the training sample by at least one of the following methods: Construct a first distillation loss function according to the difference between the class prediction results respectively corresponding to the first segment and the second segment in the training sample; Construct a second distillation loss function according to the difference between the class prediction result corresponding to the first segment in the training sample and the class prediction results corresponding to the first segments in other training samples, and the difference between the class prediction result corresponding to the second segment in the training sample and the class prediction results corresponding to the second segments in other training samples.
18. The device according to claim 12 or 13, characterized in that The apparatus further includes a classification model training module, and the classification model training module is configured to: Obtain a plurality of first labeled samples corresponding to the target classification task; the first labeled samples include video segments and audio segments with corresponding relationships, and classification labels; Through the video encoding network in the classification model to be trained, determine the image feature corresponding to the first labeled sample according to the video segment in the first labeled sample; through the audio encoding network in the classification model, determine the audio feature corresponding to the first labeled sample according to the audio segment in the first labeled sample; Through the classifier in the classification model, determine the class prediction result corresponding to the first labeled sample according to the image feature and the audio feature corresponding to the first labeled sample; Train the classification model based on the class prediction result corresponding to the first labeled sample and the classification label in the first labeled sample.
19. The device according to claim 18, characterized in that, The device further includes a first classification model application module, and the first classification model application module is configured to: After completing the training of the classification model, through the video encoding network in the classification model, determine the image features corresponding to the first video segment to be processed according to the first video segment to be processed; through the audio encoding network in the classification model, according to the first audio segment to be processed corresponding to the first video segment to be processed, determine the audio features corresponding to the first audio segment to be processed; Through the classifier in the classification model, determine the classification result of the first video segment to be processed in the target classification task according to the image features corresponding to the first video segment to be processed and the audio features corresponding to the first audio segment to be processed.
20. The device according to claim 18, characterized in that The device further includes a second classification model application module, and the second classification model application module is configured to: When the target classification task is an action timing localization task, after completing the training of the classification model, for a second video segment to be processed, determine the sub-candidate video segments with a preset action in the second video segment to be processed, and determine the arrangement order of each sub-candidate video segment; For each sub-candidate video segment, through the video encoding network in the classification model, determine the image features corresponding to the sub-candidate video segment according to the sub-candidate video segment; through the audio encoding network in the classification model, according to the sub-candidate audio segment corresponding to the sub-candidate video segment, determine the audio features corresponding to the sub-candidate audio segment; through the classifier in the classification model, determine the action recognition result corresponding to the sub-candidate video segment according to the image features corresponding to the sub-candidate video segment and the audio features corresponding to the sub-candidate audio segment; Determine the action timing localization result according to the action recognition results corresponding to each sub-candidate video segment and the arrangement order of each sub-candidate video segment.
21. The device according to claim 12 or 13, characterized in that, The device further includes a background audio generation model training module, and the background audio generation model training module is configured to: Obtain a plurality of second labeled samples corresponding to the background audio generation task; the second labeled samples include video segments and labeled background audio segments with corresponding relationships; Through the video encoding network, determine the image features corresponding to the video segment according to the video segment in the second labeled sample; through the audio encoding network, determine the audio features corresponding to the labeled background audio segment according to the labeled background audio segment in the second labeled sample; Through the feature conversion network to be trained, perform feature conversion processing on the image features corresponding to the video segment in the second labeled sample to obtain reference audio conversion features; and based on the audio features corresponding to the labeled background audio segment in the second labeled sample and the reference audio conversion features, train the feature conversion network; Using the background audio generation model to be trained, according to the audio features corresponding to the labeled background audio segments, generate predicted background audio segments; Based on the labeled background audio segments and the predicted background audio segments, train the background audio generation model.
22. The device according to claim 21, characterized in that, The device further includes a background audio generation model application module, and the background audio generation model application module is configured to: After completing the training of the feature conversion network and the background audio generation model, through the video encoding network, according to the target video segment for which the background audio is to be generated, determine the image features corresponding to the target video segment; Through the feature conversion network, perform feature conversion processing on the image features corresponding to the target video segment to obtain the audio conversion features corresponding to the target video segment; Through the background audio generation model, according to the audio conversion features corresponding to the target video segment, generate the background audio segment corresponding to the target video segment.
23. A computer device, characterized in that, The device includes a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the model training method according to any one of claims 1 to 11 according to the computer program.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the model training method according to any one of claims 1 to 11.
25. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or the instruction is executed by the processor, the model training method according to any one of claims 1 to 11 is implemented.