Training method of media classification model, media data classification method and device
By introducing feature conversion and classification of modal information into the video classification model, the training loss function is constructed, and the problem of poor training effect of video classification model in the existing technology is solved, and the accuracy of video classification is improved.
Patent Information
- Application Number
- CN202210504251.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-10
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-05-10
AI Technical Summary
The training effect of existing video classification models is poor, resulting in low video classification accuracy, mainly due to limited supervision information.
By obtaining training data, including sample media data, sample modal information and classification tags, the first and second network structures in the media classification model are used to transform and classify the sample media data and modal information, and a training loss function is constructed to train the media classification model.
The training effect of the media classification model is improved, the accuracy of video classification is enhanced, and by mining supervision information from the modal information corresponding to the media data, the problem of insufficient supervision information when training using only classification tags is overcome.
Smart Images

Figure CN117093733B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a media classification model training method, a media data classification method and a device. Background Art
[0002] With the development of science and technology, various video platforms are becoming more and more popular, and the number of videos contained in each video platform is increasing. In order to better promote the development of video recommendation, search and distribution services, video classification technology has been proposed. Video classification technology is to determine video tags. Among them, video tags can not only accurately describe the characteristics of the video, but also help describe the interests and habits of the target object, and can provide a comprehensive and accurate basis for video recommendation, search and distribution services.
[0003] In traditional technology, when performing video classification, video classification is mainly performed through video classification models. Among them, the video classification model is obtained by training the classification label of the video as supervision information. Therefore, the supervision information used in the training process is relatively limited, resulting in poor training effect of the video classification model, and further resulting in low accuracy of video classification through the video classification model. Summary of the invention
[0004] Based on this, it is necessary to provide a media classification model training method, a media data classification method and a device that can improve the training effect of the media classification model in response to the above-mentioned technical problems.
[0005] On the one hand, the present application provides a method for training a media classification model, the method comprising:
[0006] Acquire training data, where the training data includes sample media data, sample modality information corresponding to the sample media data, and a classification label to which the sample media data belongs;
[0007] Performing feature conversion on sample media latent features of sample media data through a first network structure in a media classification model to be trained to obtain sample auxiliary latent features, and performing classification according to the sample media latent features and the sample auxiliary latent features to obtain a first prediction result;
[0008] Performing feature conversion on the sample modal features of the sample modal information through the second network structure in the media classification model to be trained to obtain modal reference features, and performing classification based on the modal reference features to obtain a second prediction result;
[0009] According to the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature, a training loss function is constructed;
[0010] The media classification model to be trained is trained by training the loss function and stops when the training stop condition is reached to obtain a trained media classification model.
[0011] On the other hand, the present application also provides a training device for a media classification model, the device comprising:
[0012] An acquisition module, used to acquire training data, the training data including sample media data, sample modality information corresponding to the sample media data, and a classification label to which the sample media data belongs;
[0013] A first feature conversion module, used to perform feature conversion on sample media latent features of sample media data through a first network structure in a media classification model to be trained, to obtain sample auxiliary latent features;
[0014] A first classification module, used for classifying according to the sample media latent features and the sample auxiliary latent features to obtain a first prediction result;
[0015] A second feature conversion module, used for performing feature conversion on the sample modal features of the sample modal information through a second network structure in the media classification model to be trained to obtain a modal reference feature;
[0016] A second classification module, used for performing classification based on the modal reference feature to obtain a second prediction result;
[0017] A construction module, used to construct a training loss function according to the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature;
[0018] The training module is used to train the media classification model to be trained by using a training loss function, and stop when the training stop condition is reached to obtain a trained media classification model.
[0019] In one of the embodiments, the sample media data includes a sample video; the acquisition module is further used to obtain a video embedding feature sequence of the sample video; determine the timing of each video embedding feature in the video embedding feature sequence; and superimpose the video embedding features in sequence according to the timing of each video embedding feature in the video embedding feature sequence to obtain the sample media latent feature.
[0020] In one of the embodiments, the acquisition module is also used to extract multiple image frames from the sample video; extract the features of each image frame at multiple dimensional levels; combine and process the multiple features corresponding to each image frame to obtain the video embedding features corresponding to each image frame; and use the feature sequence formed by the video embedding features corresponding to each image frame as the video embedding feature sequence of the sample video.
[0021] In one embodiment, the first network structure includes a feature conversion substructure, which is composed of at least one fully connected layer; the first feature conversion module is used to perform full-connection processing on the sample media latent features through the feature conversion substructure to obtain the sample auxiliary latent features.
[0022] In one of the embodiments, the first classification module is used to combine the sample media latent features and the sample auxiliary latent features to obtain sample combination features; and to perform classification based on the sample combination features to obtain a first prediction result.
[0023] In one of the embodiments, the second feature conversion module is used to perform at least one full-connection processing on the sample modal features through the second network structure in the media classification model to be trained to obtain the modal reference features.
[0024] In one embodiment, the device further comprises:
[0025] The third classification module is used to classify based on the sample media latent features through the first network structure to obtain a third prediction result; accordingly, the construction module is used to construct a training loss function according to the differences between the first prediction result, the second prediction result, the third prediction result and the classification label, and the difference between the sample auxiliary latent features and the modal reference features.
[0026] In one embodiment, a construction module is used to determine a first loss based on a difference between a first prediction result and a classification label; determine a second loss based on a difference between a second prediction result and a classification label; determine a third loss based on a difference between a sample auxiliary latent feature and a modality reference feature; and construct a training loss function based on the first loss, the second loss, and the third loss.
[0027] In one of the embodiments, the construction module is also used to calculate the similarity between the sample auxiliary latent features corresponding to the corresponding training data and the modal reference features for each training data; sum the similarities corresponding to each training data, and use the sum result as the third loss.
[0028] In one embodiment, the training loss function includes a first training loss function and a second training loss function; accordingly, the construction module is used to construct the first training loss function according to the difference between the first prediction result and the classification label; and construct the second training loss function according to the difference between the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature;
[0029] A training module is used to execute a second training process based on a second training loss function, and to execute a first training process based on a first training loss function, and the second training process is executed alternately with the first training process; wherein the first training process is a process of adjusting parameters of a first network structure based on a first loss function and training samples of a current batch, and the second training process is a process of adjusting parameters of a second network structure based on a second loss function and training samples of a current batch.
[0030] In one embodiment, the device further comprises:
[0031] The third classification module is used to classify based on the latent features of the sample media through the first network structure to obtain a third prediction result; accordingly, the construction module is used to construct a first training loss function according to the difference between the first prediction result and the classification label, and the difference between the third prediction result and the classification label.
[0032] In one embodiment, the device may also perform classification of media data by using the first network structure in the trained media classification model; accordingly, the device further includes:
[0033] The classification application module is used to obtain the target media data to be classified, extract the media latent features in the target media data; perform feature conversion on the media latent features to obtain auxiliary latent features for characterizing the modal information of the target media data; and perform classification based on the media latent features and the auxiliary latent features to obtain the category to which the target media data to be classified belongs.
[0034] On the other hand, the present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the training method of the media classification model when executing the computer program.
[0035] On the other hand, the present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the training method of the media classification model when the computer program is executed by a processor.
[0036] On the other hand, the present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned media classification model training method when executed by a processor.
[0037] The training method, apparatus, computer equipment, storage medium and computer program product of the above-mentioned media classification model construct a training loss function for training the media classification model based on the losses of each substructure in the first network structure and the second network structure. These losses include losses calculated based on the difference between the sample auxiliary latent features and the modal reference features. Among them, the sample auxiliary latent features are obtained by learning the feature representation corresponding to its modal information from the sample media latent features, and the modal reference features can be used as training labels. Therefore, it is equivalent to mining the supervisory information used for training from the modal information corresponding to the media data. At the same time, the modal information is indeed associated with the content of the media data and is helpful for the classification of the media data, thereby overcoming the problem that the supervisory information is too weak when only the classification label is used to train the media classification model, resulting in poor training effect. Therefore, it is beneficial to improve the training effect of the media classification model, and the classification accuracy can also be improved when the media classification model is subsequently used to classify the media data.
[0038] On the other hand, the present application provides a method for classifying media data, the method comprising:
[0039] Obtain target media data to be classified, and extract media hidden features in the target media data;
[0040] Perform feature conversion on the media latent features to obtain auxiliary latent features for characterizing the modal information of the target media data;
[0041] Combining media latent features and auxiliary latent features to obtain target combined features;
[0042] Classification is performed based on the target combination features, and the category to which the target media data to be classified belongs is output.
[0043] On the other hand, the present application provides a media data classification device, the device comprising:
[0044] An acquisition module is used to acquire target media data to be classified and extract media hidden features in the target media data;
[0045] A feature conversion module, used to perform feature conversion on media latent features to obtain auxiliary latent features for characterizing modal information of target media data;
[0046] A combination module is used to combine the media latent features and the auxiliary latent features to obtain the target combined features;
[0047] The classification module is used to perform classification based on target combination features and output the category to which the target media data to be classified belongs.
[0048] In one of the embodiments, a feature conversion module is used to perform feature conversion on media latent features based on a pre-trained feature conversion substructure to obtain auxiliary latent features for characterizing modal information of target media data; the feature conversion substructure is obtained by training with modal reference features as label information, and the modal reference features are features obtained in the process of classifying sample modal information of sample media data during the training phase.
[0049] On the other hand, the present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the media data classification method provided above when executing the computer program.
[0050] On the other hand, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the media data classification method provided above are implemented.
[0051] On the other hand, the present application also provides a computer program product, including a computer program, which implements the steps of the media data classification method provided above when executed by a processor.
[0052] The above-mentioned media data classification method, apparatus, computer equipment, storage medium and computer program product can obtain auxiliary latent features for characterizing the modal information of the target media data by performing feature conversion on the media latent features in the target media data to be classified. By combining the media latent features and the auxiliary latent features, a more comprehensive and accurate target combined feature that characterizes the target media data can be obtained, and the target media data can be accurately classified based on the combined feature. In this way, the modal information of other dimensions of the target media data can be introduced without increasing the complexity of reasoning, which can greatly improve the classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 A diagram of an application environment of a method for training a media classification model in one embodiment;
[0054] Figure 2 A schematic diagram of a flow chart of a method for training a media classification model in one embodiment;
[0055] Figure 3 is a schematic diagram of the structure of a first network structure in an embodiment;
[0056] Figure 4 is a schematic diagram of the structure of a second network structure in an embodiment;
[0057] Figure 5 A schematic diagram of the structure of LSTM in one embodiment;
[0058] Figure 6 A schematic diagram of the structure of a forget gate in an LSTM in one embodiment;
[0059] Figure 7 A schematic diagram of the structure of an input gate in an LSTM in one embodiment;
[0060] Figure 8 A schematic diagram of cell state update in LSTM in one embodiment;
[0061] Fig. 9 Schematic diagram of the structure of the output gate in LSTM in one embodiment;
[0062] Fig.10 is a flowchart of a method for training a media classification model in another embodiment;
[0063] Fig.11 is a flow chart of a method for classifying media data in one embodiment;
[0064] Fig.12 A schematic diagram of the structure of a media classification model during training in one embodiment;
[0065] Fig.13 A schematic diagram of the structure of a media classification model in an application process in an embodiment;
[0066] Fig.14 is a structural block diagram of a training device for a media classification model in one embodiment;
[0067] Fig.15 is a structural block diagram of a media data classification device in one embodiment;
[0068] Fig.16 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0070] First, the terms and technologies involved in the embodiments of this application are briefly explained:
[0071] Media data: refers to data presented through media, which can include text, images, video, audio, etc.
[0072] Modal information: refers to information related to the content of media data. For example, if the media data is a video, it is understood that a video usually does not only have video data, but also usually has information related to the content of the video, such as a video title or video introduction. This information can be called modal information. It is understandable that different media data will have different corresponding modal information. For another example, if the media data is an image, the modal information can be an image introduction.
[0073] Hidden features: In addition to the input layer and the output layer, the neural network also has some processing layers in the middle. The features of the outputs of these processing layers are called hidden features; the intermediate processing layers can include convolutional layers, pooling layers, and fully connected layers.
[0074] Media classification: refers to determining the type label corresponding to a given piece of media data. For example, taking the media data as a video, you can get the corresponding classification of the video, such as daily life videos, pet videos, funny videos, and film review videos.
[0075] In addition, in the embodiments of the present application, the training process of the media classification model and the subsequent classification application process mainly involve artificial intelligence (AI) and machine learning technology, and are designed based on speech technology, natural language processing technology and machine learning (ML) in artificial intelligence.
[0076] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence.
[0077] Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology mainly includes computer vision technology, natural language processing technology, and machine learning / deep learning. With the research and progress of artificial intelligence technology, artificial intelligence has been studied and applied in many fields, such as common smart homes, smart customer service, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, robots and smart medical care. It is believed that with the development of technology, artificial intelligence will be applied in more fields and play an increasingly important role.
[0078] Machine learning is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Compared with data mining, which seeks mutual characteristics between big data, machine learning focuses more on algorithm design, allowing computers to automatically "learn" rules from data and use the rules to predict unknown data.
[0079] Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. Reinforcement learning (RL), also known as reinforcement learning, evaluation learning, or enhanced learning, is one of the paradigms and methodologies of machine learning. It is used to describe and solve the problem of how an agent can maximize its rewards or achieve specific goals through learning strategies during its interaction with the environment.
[0080] In some embodiments, in combination with the above explanations, the media classification model training method or media data classification method provided in the embodiments of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 can communicate with the server 104 directly or indirectly through a wired or wireless network, and the embodiment of the present application does not specifically limit this. In addition, the terminal 102 or the server 104 can be used separately to execute the training method of the media classification model in the embodiment of the present application, and can also be used separately to execute the media data classification method in the embodiment of the present application; it can also be that the two are used in conjunction to execute the training method of the media classification model in the embodiment of the present application, or the two are used in conjunction to execute the media data classification method in the embodiment of the present application.
[0081] For independent execution, one implementation process when the server 104 independently executes the training method of the media classification model is now taken as an example. Specifically, the server 104 can obtain and store the training data in advance, perform feature conversion processing on the training data through the network structure in the internally stored media classification model, and construct a training loss function based on the features obtained after the feature conversion, so as to implement the training of the media classification model based on the training loss function.
[0082] For collaborative execution, taking one of the implementation processes of the method for training a media classification model in which the two collaborate to execute as an example, the terminal 102 can upload training data to the server 104, and the server 104 can perform feature conversion on the uploaded training data based on the training data uploaded by the terminal 102 through the network structure in the media classification model stored in the server 104, and construct a training loss function based on the features obtained after the feature conversion, thereby implementing the training of the media classification model based on the training loss function. The data storage system can store the training data obtained by the server 104, and can also store the media classification model, so as to subsequently train the media classification model based on the training data. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other servers.
[0083] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. Applications can be run on the terminal, such as video applications, or audio applications, etc., for presenting media data. The server 104 can be a background server corresponding to software or web pages, applets, etc., or a server specifically used for media classification, which is not specifically limited in the embodiments of the present application. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0084] In some embodiments, combined with the above-mentioned explanations of terms, technical explanations and implementation environment descriptions, such as Figure 2 As shown, a method for training a media classification model is provided, and the method is applied to a computer device (the computer device may be specifically Figure 1 The terminal or server in the example is used to illustrate, including the following steps:
[0085] Step 202: Acquire training data, where the training data includes sample media data, sample modality information corresponding to the sample media data, and a classification label to which the sample media data belongs.
[0086] Among them, the sample modal information mainly refers to information given from another perspective and associated with the content of the sample media data. The type of the sample modal information can be associated with the type of the sample media data. For example, if the sample media data is video data, the modal information can be a video title or video introduction associated with the video content, that is, text data. Alternatively, for the audio in the video, it will naturally be associated with the video content, and thus can also be used as modal information. For another example, if the sample media data is image data, the modal information can be an image introduction associated with the image content, that is, text data.
[0087] The classification label to which the sample media data belongs is mainly used to indicate which category the sample media data belongs to. Taking the sample media data as a video as an example, the classification label to which the video belongs can be a film and television drama, a TV series, a sports game, an educational lecture, and a live broadcast with goods, etc. The data form of the classification label can be a digital number or a field identifier, which is not specifically limited in the embodiments of the present application.
[0088] It should be noted that, as can be seen from the above examples, for a certain type of sample media data, the corresponding sample modal information may be more than one type. For example, for a video, the video title, video introduction and audio in the video can all serve as its corresponding modal information. Therefore, in this step, the type of sample modal information corresponding to each sample modal information in the training data acquired by the computer device may not be unique, and the embodiments of the present application do not make specific limitations on this. In addition, the training data used in the embodiments of the present application may come from Imagenet, which refers to a large-scale general object recognition open source data set. Of course, there may be other sources of training data in the actual implementation process, and the embodiments of the present application do not make specific limitations on this.
[0089] Taking the computer device as a server as an example, it can be understood that the training media classification model in the embodiment of the present application can be used to classify the media data subsequently stored by the server, and the media data stored by the server usually comes from the terminal upload. That is, in some application scenarios, the terminal will actually upload media data to the server, and the server needs to store the media data uploaded by the terminal. Therefore, when the server obtains the training data in this step, it can be real-time acquisition of the media data uploaded by the terminal as training data. That is, whenever the terminal uploads media data to the server, in addition to the media classification model classifying the media data, the server can also store the media data uploaded by the terminal for the training process. Through the above process, the media classification model can continuously learn new samples, thereby improving the generalization ability of the media classification model, and then improving the accuracy of the subsequent classification of the media classification model.
[0090] Step 204: perform feature conversion on the sample media latent features of the sample media data through the first network structure in the media classification model to be trained to obtain sample auxiliary latent features, perform classification according to the sample media latent features and the sample auxiliary latent features, and obtain a first prediction result.
[0091] Before executing this step, the computer device may first convert the sample media data into sample media latent features that can be processed by the media classification model, that is, the process of obtaining the sample media latent features of the sample media data. For example, taking the sample media data as a video, since the video data cannot directly participate in the calculation process in the model, it can be converted into a video feature vector, that is, the sample media latent features, by the computing device.
[0092] In one of the embodiments, the first prediction result is the classification result of the first network structure in the media classification model for classifying the sample media data. It should be noted that the classification result actually corresponding to a sample media data may be one or more, and the embodiment of the present application does not specifically limit this. For example, if the content in a video is a performer playing a musical instrument, the category corresponding to the sample video may be a musical instrument performance; or, if the content in a video is someone singing while walking, the category corresponding to the sample video may be daily life or singing. In addition, the first prediction result in this step, and the second prediction result in the subsequent step, may belong to the same type range as the classification label, such as film and television dramas, TV dramas, sports competitions, educational lectures, and live broadcasts, etc., and the embodiment of the present application does not specifically limit this.
[0093] The media classification model mentioned in this step may be a neural network model, which usually includes many processing base layers, and there are connection relationships between these processing base layers. It can be understood that different processing base layers can form a network structure with local complete processing functions by connecting to each other. For example, in some neural networks, there may be a network structure that converts raw data into feature vectors, such as a network structure that converts text data into feature vectors, and this network structure is also composed of multiple processing base layers. In an embodiment of the present application, a first network structure with local complete processing functions can be formed in the media classification model. Corresponding to the various processing processes mentioned in this step, the first network structure can realize the function of feature conversion of sample media latent features, and classify sample media data according to sample media latent features and sample auxiliary latent features.
[0094] As can be seen from the above content, the sample media data needs to be converted into sample media hidden features. Therefore, in addition to realizing the various functions mentioned above, the first network structure can also be used to realize the function of converting sample media data into sample media hidden features. Of course, in the actual implementation process, this function may not be realized by the first network structure, but by other network structures in the media classification model, and the embodiment of the present application does not specifically limit this.
[0095] Taking the first network structure as an example of realizing the function of converting sample media data into sample media latent features, the purpose of converting sample media data into sample media latent features is mainly to convert sample media data into data that can be processed by the media classification model. Therefore, the first network structure may include a processing base layer for converting sample media data into sample media latent features. The feature conversion of sample media latent features is mainly to enable the sample media latent features to learn the feature representation corresponding to the sample modal information, thereby forming latent features carrying the feature representation corresponding to its modal information, that is, sample auxiliary latent features. Therefore, the first network structure may include a processing base layer for performing feature conversion on sample media latent features. The above two functions are mainly used to obtain latent features for classification, and obtaining latent features naturally requires classification based on latent features. Therefore, the first network structure may also include a processing base layer for implementing classification to obtain a first prediction result.
[0096] In summary, the first network structure may include a processing base layer for realizing the above functions. The specific processing methods corresponding to the processing base layers for realizing the above functions, that is, the methods for obtaining latent features of sample media, the methods for feature conversion, and the methods for classification processing, may be associated with the type of the processing base layer and its internal specific structure. In combination with the above-mentioned contents, the connection relationship between the processing base layers for realizing the above functions in the first network structure can be referred to Figure 3 .
[0097] It should be noted that Figure 3 It is only exemplary. Whether the first network structure also includes Figure 3 Processing base outside of the examples, Figure 3 Whether the processing base layer in the Figure 3 Whether there are other processing bases between the processing bases in the examples can be set based on actual needs. The embodiments of the present application do not specifically limit whether the first network structure also includes other processing bases, the types of processing bases that implement the above functions in the first network structure, the specific internal structures of the processing bases, and the connection relationships between them. For example, the type of the processing base can be a fully connected layer, and the internal structure and connection relationship of the processing base can be multiple fully connected layers connected to each other.
[0098] Step 206: Perform feature conversion on the sample modal features of the sample modal information through the second network structure in the media classification model to be trained to obtain modal reference features, perform classification based on the modal reference features, and obtain a second prediction result.
[0099] Before executing this step, the sample modal information can be converted into sample modal features that can be processed by the media classification model, that is, the process of obtaining sample modal features of the sample modal information. For example, taking the sample modal information as a media title as an example, since the media title is actually text data, and text data cannot directly participate in the calculation process in the model, it can be converted into a text feature vector, that is, a sample modal feature.
[0100] In one embodiment, the second network structure can realize the function of feature conversion of sample modal features. Combined with the content of the subsequent steps, it can be seen that the reason why it is called "modal reference feature" is that on the one hand, this feature is actually mainly converted based on modal information, so the word "modal" is introduced; on the other hand, this feature will actually be used as a training label corresponding to the sample auxiliary hidden feature in the future, so the word "reference" is introduced.
[0101] Similar to the explanation in step 204, it can be understood that the second network structure may include a processing base layer for realizing feature conversion of sample modal features, and may also include a processing base layer for realizing classification to obtain a second prediction result. In this step, the second prediction result is the classification result of the second network structure in the media classification model that classifies the sample media data based on the sample modal information. The specific processing method corresponding to the processing base layer that realizes the above-mentioned functions in the second network structure, that is, the method of feature conversion and classification processing of sample modal features, can also be associated with the type of processing base layer and its internal specific structure. In combination with the above-mentioned content, the connection relationship between the processing base layers used to realize the above-mentioned functions in the second network structure can be referred to. Figure 4 .
[0102] It should be noted that Figure 4 It is only exemplary. Whether the second network structure also includes Figure 4 Processing base outside of the examples, Figure 4 Whether the processing base layer in the Figure 4Whether there are other processing bases between the processing bases in the example can be set based on actual needs. The embodiment of the present application does not specifically limit whether the second network structure also includes other processing bases, the type of processing bases that implement the above functions in the second network structure, the specific internal structure of the processing bases, and the connection relationship between each other. For example, the type of the processing base can be a fully connected layer, and the internal structure and connection relationship of the processing base can be multiple fully connected layers connected to each other.
[0103] It should also be noted that the above content mentions that the sample modal information needs to be converted into sample modal features. Figure 4 In the second network structure, the processing base layer for converting the sample modality information into the sample modality feature can also be included. Figure 4 In the embodiment of the present application, the sample modal information is connected to the "processing base layer for realizing the feature conversion function" to input the sample modal features to the processing base layer, and the embodiment of the present application does not specifically limit this. Of course, in the actual implementation process, the function of converting the sample modal information into the sample modal features can also be implemented by other network structures in the media classification model instead of the second network structure, and the embodiment of the present application does not specifically limit this.
[0104] In addition, in the subsequent steps, the modal reference features will be used as training labels corresponding to the sample auxiliary latent features. Figure 3 The first network structure shown in Figure 4 There may also be an association relationship between the second network structures shown in , which is specifically reflected in that the modal reference features generated in the second network structure can be transferred to the first network structure as training labels.
[0105] Step 208: construct a training loss function based on the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modal reference feature.
[0106] Among them, the difference between the first prediction result and the classification label can be used to represent the loss of the substructure in the first network structure that is classified according to the sample media latent features and the sample auxiliary latent features. The difference between the second prediction result and the classification label can be used to represent the loss of the substructure in the second network structure that is classified based on the modal reference features. The difference between the sample auxiliary latent features and the modal reference features can be used to represent the loss of the substructure in the first network structure that performs feature conversion on the sample media latent features. Based on the three losses mentioned above, a training loss function can be constructed. The embodiment of the present application does not specifically limit the way to construct a training loss function based on the three losses mentioned above, including but not limited to: performing weighted summation of the three losses to obtain a training loss function.
[0107] The specific calculation method of the above three losses can be set based on actual needs, such as using mean square error loss, L1 loss, L2 loss, exponential loss, negative log-likelihood loss or square loss, etc. The embodiment of the present application does not make specific limitations on this.
[0108] It should be noted that, since the computer equipment used to train the media classification model usually has limited processing resources, and the training data is usually massive, the training data is usually trained in batches during the actual implementation, that is, there are usually multiple batches of training data and each batch has multiple training data. Therefore, the corresponding loss can be calculated for each training data, and the training loss function is obtained by accumulating the loss corresponding to each training data, so the value of the training loss function can be related to the batch of training data and the number of training data in each batch. That is, the variables in the training loss function can include the total batches of training data and the total number of training data in each batch.
[0109] Step 210: Train the media classification model to be trained by using a training loss function, and stop when a training stop condition is reached to obtain a trained media classification model.
[0110] Specifically, when training a media classification model, the training object may be the parameters of the network structure in the media classification model. For example, in combination with the explanations in the above steps, it can be seen that the first network structure may include a substructure for classification based on sample media latent features and sample auxiliary latent features, the second network structure may include a substructure for classification based on modal reference features, and the first network structure may also include a substructure for feature conversion of sample media latent features. In this step, the parameters in the three substructures mentioned above may be trained.
[0111] It should be noted that, in combination with the content of step 204, it can be known that the sample media data can be converted into sample media latent features through the first network structure. Therefore, the first network structure can also include a substructure for converting the sample media data into sample media latent features. And in combination with the content of step 206, it can be known that the computer device can convert the sample modal information into sample modal features through the second network structure. Therefore, the second network structure can also include a substructure for converting the sample modal information into sample modal features. For the two substructures mentioned above, when training the media classification model, the parameters in at least one of the two substructures can be trained simultaneously based on demand, and the embodiments of the present application do not make specific limitations on this.
[0112] It should also be noted that when the computer device trains the first network structure and the second network structure in the media classification model, the first network structure and the second network structure can use the same training data in each training process. At the same time, during the training process, the computer device can use the same training data to simultaneously train the first network structure and the second network structure in the media classification model. Specifically, the computer device can use a batch of training data to train the first network structure and the second network structure at the same time, and obtain the loss of each substructure mentioned in the above content under the batch of training data. Then, the computer device can change a batch of training data, and then train the first network structure and the second network structure at the same time. Since the loss of each substructure under each batch of training data can be known, the value corresponding to the training loss function can naturally be obtained. The computer device repeats the above training process until the training stop condition is reached.
[0113] Of course, in addition to the above-mentioned simultaneous training method, an alternating training method can also be used. Specifically, the computer device can use a batch of training data to first train the second network structure, and then use the same batch of training data to train the first network structure. At the same time, the loss of each substructure mentioned in the above content under the batch of training data can be obtained. Then, the computer device can change a batch of training data and perform alternating training again. Since the loss of each substructure under each batch of training data can be known, the numerical value corresponding to the training loss function can naturally be obtained. By repeating the above training process, stop until the training stop condition is reached. It can be understood that the training stop condition mentioned in this step is the end condition of the training process. Regarding the setting method of the training stop condition, the embodiment of the present application does not specifically limit this, including but not limited to: the numerical value of the training loss function converges or no longer decreases.
[0114] The above-mentioned media classification model training method constructs a training loss function for training the media classification model based on the losses of each substructure in the first network structure and the second network structure. These losses include losses calculated based on the difference between the sample auxiliary latent features and the modal reference features. Among them, the sample auxiliary latent features are obtained by learning the feature representation corresponding to its modal information from the sample media latent features, and the modal reference features can be used as training labels. Therefore, it is equivalent to mining the supervisory information used for training from the modal information corresponding to the media data. At the same time, the modal information is indeed associated with the content of the media data and is helpful for the classification of the media data, thereby overcoming the problem that the supervisory information is too weak when only the classification labels are used to train the media classification model, resulting in poor training results. Therefore, it is beneficial to improve the training effect of the media classification model, and the classification accuracy can also be improved when the media classification model is subsequently used to classify the media data.
[0115] The above embodiments mentioned that the sample media data needs to be converted into sample media latent features that can be processed by the media classification model, and this process can be implemented by the first network structure. Based on this description, the process of converting the sample video into the sample media latent features is now described by taking the sample media data as a sample video as an example. In some embodiments, the process of obtaining the sample media latent features includes: obtaining a video embedding feature sequence of the sample video; determining the time sequence of each video embedding feature in the video embedding feature sequence; and superimposing the video embedding features in sequence according to the time sequence of each video embedding feature in the video embedding feature sequence to obtain the sample media latent features.
[0116] It is understandable that a video is actually composed of multiple image frames with a time sequence. For some of the multiple image frames, each of these image frames can be converted into a video embedding feature. Specifically, the computer device can realize the conversion between image frames and features through a convolutional neural network or a deep neural network, and the embodiment of the present application does not specifically limit the conversion method. Among them, the corresponding video embedding features of image frames that are related in image content will be closer or have the same rules in some dimensions.
[0117] It is understandable that, corresponding to the time sequence corresponding to multiple image frames, each video embedding feature in the video embedding feature sequence also has a time sequence, and the time sequence corresponding to each video embedding feature can be represented by a moment. In order to adapt to the processing of time series data, in an embodiment of the present application, a recurrent neural network can be used to perform superposition processing on the video embedding feature sequence to obtain sample media hidden features. Furthermore, considering the problem of gradient disappearance in the recurrent neural network, LSTM (Long Short-Term Memory) can also be used in an embodiment of the present application to perform superposition processing on the video embedding feature sequence.
[0118] For ease of understanding, we take the example of superimposing the video embedding feature sequence through the long short-term memory network. Figure 5 The structure of the medium and long short-term memory network explains the process of superposition processing. Figure 5 As shown, Figure 5 The situation where three packaging structures are connected in sequence is given in FIG. Figure 5 Each encapsulation structure in is a layer structure in the long short-term memory network, and each video embedding feature in the video embedding feature sequence corresponds to a layer structure in the long short-term memory network. Taking each video embedding feature as an example with 2048 dimensions, the video embedding feature at time t can be expressed as x t , and x tIt can be further expressed as [Batch, t, 2048], where "t" indicates the time at which the video embedding feature is, "2048" indicates the dimension of the video embedding feature, and "Batch" indicates which batch of videos the sample video corresponding to the video embedding feature belongs to. t represents the hidden features output at time t. Correspondingly, x t-1 represents the video embedding feature at time t-1, h t-1 represents the hidden features output at the t-1th moment; x t+1 represents the video embedding feature at time t+1, h t+1 Represents the hidden features output at time t+1.
[0119] In the embodiment of the present application, each layer structure in LSTM can be further divided into LSTM1 and LSTM2. For a certain layer structure, the LSTM1 layer in the layer structure is used to output the hidden features corresponding to the layer structure, and the LSTM2 layer is used to output the cell state corresponding to the layer structure.
[0120] Each layer of LSTM has three gate structures: forget gate, input gate and output gate. For the layer structure corresponding to the tth time, the structure of the forget gate can be referred to Figure 6 .like Figure 6 As shown in the figure, the bold part is the forget gate structure. t That is, the video embedding feature at the tth moment, C t-1 Refers to the cell state output at time t-1, h t-1 Refers to the hidden features output at the t-1th moment. t Refers to the forgetting vector, each position of which has a value between 0 and 1. By comparing the forgetting vector with C t-1 Perform bitwise multiplication, C t-1 Some values in the vector will become smaller, which is equivalent to the "information" being forgotten, and the forgetting vector naturally indicates the extent to which C is forgotten. t-1 Information in. t It is based on h t-1 and C t-1 The specific calculation process can refer to the following calculation formula (1):
[0121] f t =σ(W f [h t-1 ,x t ]+b f ); (1)
[0122] In the above formula (1), σ represents the activation function, W f and b fRespectively represent the weight and bias of the forget gate. It should be noted that the output latent feature can be 1024-dimensional data, and the output cell state can be 512-dimensional data. The embodiment of the present application does not specifically limit the dimension of the output feature.
[0123] The structure of the input gate can be referred to Figure 7 ,like Figure 7 As shown in the figure, the bold part is the input gate structure. t It is equivalent to the information enhancement vector, and the value of each position is related to f t The same is 0 to 1, and Refers to the updated value of the unit state. t Can be used to control Which features in are used to update C t , thereby achieving the selective recording of new information into the cell state. t and The specific calculation process of each can refer to the following calculation formulas (2) and (3):
[0124] i t =σ(W i ·[h t-1 ,x t ]+b i ); (2)
[0125]
[0126] In the above formulas (2) and (3), tanh represents the activation function, W i and b i Represent the weight and bias of the input gate, W c and b c denote the weight and bias of the unit state respectively.
[0127] The process of cell state update can be referred to Figure 8 ,like Figure 8 As shown in the figure, the bold part is the process of updating the cell state. Among them, the processing process of the forget gate and the input gate is mainly to prepare for the cell state update. The specific calculation process corresponding to the cell state update can be referred to the following formula (4):
[0128]
[0129] In the above formula (4), the meaning of each parameter can refer to the above description. It can be understood that f t With C t-1 Multiplication is mainly to indicate which information in the cell state at the previous moment needs to be forgotten. t and The multiplication is mainly to indicate which new information in the cell state needs to be recorded in the cell state, that is, the new candidate value. By accumulating "forget" and "update", the cell state at the tth moment can be obtained.
[0130] Finally, the structure of the output gate can be referred to Fig. 9 ,like Fig. 9 As shown in the figure, the bold part is the output gate structure. The specific calculation process corresponding to the output process of the output gate can be referred to the following formulas (5) and (6):
[0131] o t =σ(W o [h t-1 ,x t ]+b o ); (5)
[0132] h t =o t tanh(C t ); (6)
[0133] In the above formulas (5) and (6), o t Mainly used to determine the cell state C t Which part of W will be output. o and b o Respectively represent the weight and bias of the output gate, h t Refers to the hidden features output at the tth moment. Since the cell state needs to accumulate and sum the information of all moments in real time, it can be understood as long-range information. The hidden features output at each moment are determined by multiple factors at the current moment, such as the current moment input xt, the hidden features output at the previous moment, and the cell state, so they can be understood as short-range information. LSTM uses long-range information and short-range information to superimpose features in sequence to achieve long-term and short-term memory.
[0134] It should be noted that, in the actual implementation process, the sample media latent feature obtained by the embodiment of the present application can be the latent feature output by the LSTM at the last moment. At this time, the number of latent features obtained is one. Of course, in the actual implementation process, the sample media latent feature obtained by the embodiment of the present application can be more than one, such as the latent feature output by the LSTM at other moments can also be used as the sample media latent feature, and the embodiment of the present application does not specifically limit this.
[0135] It should also be noted that the above process mainly takes the sample media data as a sample video as an example to illustrate the process of converting the sample media data into the sample media latent features. It can be understood that LSTM can be applied to sample media data with time series characteristics, but not all sample media data have time series characteristics, such as images. Therefore, for sample media data with time series characteristics, the relevant processing method of the above LSTM can be used to convert the sample media data into sample media latent features. For sample media data that do not have time series characteristics, such as images, the sample media data can be directly converted into media embedding features, and then further processed (such as directly performing a layer of convolution processing) to obtain sample media latent features.
[0136] In the above embodiment, the video embedding features are sequentially superimposed according to the time sequence of each video embedding feature in the video embedding feature sequence, and the obtained sample media latent features are used as the basis for subsequent classification processing. Since the time sequence of the image frames in the video is directly related to the video content, the classification result is naturally associated with the time sequence of the input image frames. The sample media latent features obtained after superposition not only retain the sequence characteristics of the image frames, but also have a long-term memory function, which is conducive to improving the training effect of the media classification model. When the media latent features obtained by the superposition process are subsequently used to classify the media data, the classification accuracy can also be improved.
[0137] In some embodiments, obtaining a video embedding feature sequence of a sample video includes: extracting multiple image frames from the sample video; extracting features of each image frame at multiple dimensional levels; combining and processing multiple features corresponding to each image frame to obtain a video embedding feature corresponding to each image frame; and using a feature sequence formed by the video embedding features corresponding to each image frame as a video embedding feature sequence of the sample video.
[0138] Specifically, the embodiment of the present application does not specifically limit the method of extracting multiple image frames from the sample video, including but not limited to the method of uniformly extracting frames with a fixed length or randomly extracting frames. For example, the computer device can extract a frame from the video every 10 frames. Of course, other methods can also be used, such as the method of extracting frames with an indefinite length or continuously.
[0139] It is understandable that, compared to continuous frame extraction or taking all video frames in a video as input, random frame extraction or fixed-length uniform frame extraction can enhance the randomness of training data in the time dimension, thereby improving the generalization ability of subsequent media classification models. In addition, the computer device can also enhance the randomness of training data in the spatial dimension, such as cropping the image frames or adding noise to the image frames, etc., which is not specifically limited in the embodiments of the present application. Among them, when the computer device adds noise to the image frame, Gaussian noise can be added. Of course, other types of noise, such as white noise, can also be added, which is not specifically limited in the embodiments of the present application.
[0140] After obtaining the image frame from the video in the above manner, it can be understood that for each image frame, it is necessary to convert it into a feature that can be processed. Therefore, for any image frame, the computer device can process the image frame through the processing base layers connected in sequence, so that each processing base layer can extract the features of the image frame at each dimensional level. Finally, the features extracted from each dimensional level are combined by a processing base layer to obtain the video embedding features corresponding to the image frame. The feature sequence formed by the video embedding features corresponding to each image frame is the video embedding feature sequence of the sample video. Among them, the processing base layer used to extract features at each dimensional level can be a convolution layer, and the processing base layer used for combined processing can be a pooling layer. Of course, in the actual implementation process, other types of processing base layers can also be used, and the embodiments of the present application do not specifically limit this.
[0141] In the above embodiment, compared with processing all the image frames in the sample video, extracting multiple image frames from the sample video can effectively reduce the processing amount, thereby improving processing efficiency and saving processing resources. In addition, since the features of each image frame at multiple dimensional levels can be extracted respectively, while ensuring that the features can carry image information, the image as high-dimensional raw data can be mapped to features of multiple dimensions with lower dimensions, thereby effectively reducing the processing amount, thereby improving processing efficiency and saving processing resources.
[0142] It can be seen from the above embodiments that the first network structure needs to realize the function of converting the sample media latent features into the sample auxiliary latent features, and this function can be realized by the substructure in the first network structure. Based on this description, in some embodiments, the first network structure includes a feature conversion substructure, and the feature conversion substructure is composed of at least one fully connected layer; accordingly, the sample media latent features are subjected to feature conversion to obtain the sample auxiliary latent features, including: through the feature conversion substructure, the sample media latent features are subjected to full connection processing to obtain the sample auxiliary latent features.
[0143] Specifically, the dimension of the sample auxiliary latent feature can also be 512 dimensions, and the embodiment of the present application does not specifically limit the feature dimension output by the feature conversion substructure. In the actual implementation process, the feature conversion substructure may only include one fully connected layer. Of course, other structures can also be used, such as fully connected layers at both ends, and at least one substructure composed of an activation layer and a fully connected layer in the middle. Among them, the number of intermediate connected substructures can also be set according to needs, and the embodiment of the present application does not make specific limitations on this. At this time, through the feature conversion substructure, it is not limited to fully connecting the sample media latent features, but also can be processed by activation function. Among them, the activation layer can be specifically a Relu activation function layer, and the embodiment of the present application does not make specific limitations on this.
[0144] It should be noted that the above embodiments are mainly described from the perspective of introducing only one type of modal information when classifying media data. For example, when a computer device classifies a video, it only introduces sample video titles as sample modal information when training a media classification model. However, it is understandable that in the actual implementation process, it is not limited to introducing only one type of sample modal information for training. Therefore, in step 202, when the computer device obtains training data, it can obtain a sample media data and obtain multiple types of sample modal information corresponding to the sample media data.
[0145] In view of the fact that different types of modal information usually do not use the same model, in an embodiment of the present application, for the situation where one sample media data corresponds to multiple types of sample modal information, multiple feature conversion substructures can also be set in the first network structure. Among them, each type of sample modal information corresponds to a feature conversion substructure, and different types of sample modal information correspond to different feature conversion substructures. By fully connecting the sample media latent features through different feature conversion substructures, the sample media latent features can learn the feature representation of the corresponding type of sample modal information. Through multiple feature conversion substructures, multiple sample auxiliary latent features can be obtained.
[0146] It should be noted that, since the second network structure needs to provide modal reference features as training labels, and considering that different types of modal information usually do not use the same model, the second network structure may include multiple substructures for performing feature conversion on sample modal features of sample modal information. Each type of sample modal information may correspond to a substructure in the second network structure for converting itself into a modal reference feature.
[0147] In the above embodiment, by only including at least one fully connected layer of feature conversion substructure, the sample media latent features can learn the feature representation corresponding to the sample modal information. Compared with the feature representation of modal information obtained by processing the modal information through a large and complex complete deep neural network model, the data processing time in the training process can be effectively reduced, thereby effectively improving the training efficiency of the media classification model. In addition, when the media classification model is used to classify the media data later, it is not necessary to extract the modal features of the modal information through a large and complex complete deep neural network model, but it is possible to directly obtain the feature representation of the corresponding modal information based on the media latent features, thereby also improving the classification efficiency. Finally, since it is not limited to introducing only one type of sample modal information for training, it can be understood that the more types of sample modal information are introduced, the more types of supervisory information used to train the media classification model, which is conducive to improving the training effect of the media classification model, and the classification accuracy can also be improved when the media classification model is used to classify the media data later.
[0148] It can be seen from the contents of the above embodiments that the first network structure can realize the function of classification processing. Therefore, in some embodiments, classifying according to the sample media latent features and the sample auxiliary latent features to obtain the first prediction result includes: combining the sample media latent features and the sample auxiliary latent features to obtain the sample combination features; and classifying based on the sample combination features to obtain the first prediction result.
[0149] Specifically, the combination processing method can be splicing or weighted processing, which is not specifically limited in the embodiments of the present application. Among them, the splicing method can be direct head-to-tail splicing, and the splicing order of the sample media latent features and the sample auxiliary latent features during splicing can be arbitrarily selected, which is not specifically limited in the embodiments of the present application. In the case where the combination processing method is direct splicing, the dimensions of the sample media latent features and the sample auxiliary latent features can be the same or different, which is not specifically limited in the embodiments of the present application.
[0150] In addition, regarding the method of classification based on sample combination features, the embodiments of the present application do not make any specific limitations on this, including but not limited to: fully connecting the sample combination features through a fully connected layer to obtain a fully connected processing result; processing the fully connected processing result through an activation function to obtain a first prediction result.
[0151] It should be noted that, as can be seen from the above embodiments, the sample modal information can be of more than one type, and the sample auxiliary latent features obtained therefrom can also be of more than one type. In the case where the sample auxiliary latent features are of multiple types, the embodiment of the present application actually combines the sample media latent features and multiple types of sample auxiliary latent features for processing. At this time, the combined processing method can also be splicing or weighted processing, which is not specifically limited in the embodiment of the present application.
[0152] In the above embodiment, compared with taking sample modal information as input and redesigning the media classification model, on the basis of maintaining the original media classification model structure, the sample media latent features can learn the feature representation corresponding to the sample modal information, thereby avoiding changing the structure of the media classification model and saving workload. In addition, taking sample modal information as input and redesigning the media classification model will inevitably increase the complexity of the media classification model. On the one hand, it will increase the storage resources occupied by the model storage, and on the other hand, the training and subsequent use processes will also occupy more processing resources due to the more complex model, resulting in reduced processing efficiency. By not changing the structure of the media classification model, the storage resources occupied by the model storage can be reduced as much as possible, and the processing efficiency during training and subsequent use can also be improved.
[0153] Finally, during the classification, the sample media latent features and the sample auxiliary latent features that are learned to represent the features corresponding to the sample modal information are integrated. These two latent features correspond to the sample media data and the sample modal information respectively. Therefore, the sample media data and the sample modal information can be simultaneously integrated into the training process, which can help improve the training effect of the media classification model. The classification accuracy can also be improved when the media classification model is used to classify the media data subsequently.
[0154] It can be seen from the contents of the above embodiments that the second network structure can realize the function of feature conversion. Therefore, in some embodiments, the sample modal features of the sample modal information are converted by the second network structure in the media classification model to be trained to obtain modal reference features, including: performing at least one full connection process on the sample modal features by the second network structure in the media classification model to be trained to obtain modal reference features.
[0155] Before executing the embodiment of the present application, the computer device may obtain the sample modal features obtained by converting the sample modal information. It can be seen from the content of the above embodiment that the function of converting the sample modal information into the sample modal features may also be implemented by the second network structure, or may not be implemented by the second network structure, and the embodiment of the present application does not specifically limit this. In the embodiment of the present application, for the substructure in the second network structure used to convert the sample modal features into modal reference features, the substructure can be composed of at least one fully connected layer, and each fully connected layer is used to perform a fully connected process. Taking the example of the substructure composed of two fully connected layers, the input of the first fully connected layer is the sample modal features, and the output can be a 1024-dimensional feature; the input of the second fully connected layer is a 1024-dimensional feature, and the output can be a 512-dimensional feature. The embodiment of the present application does not specifically limit the feature dimension of each fully connected processing output.
[0156] In the above embodiment, since the sample modal features can be fully connected at least once, the influence of the feature position on the subsequent classification can be greatly reduced. Therefore, it is beneficial to improve the training effect of the media classification model, and the classification accuracy can also be improved when the media classification model is used to classify media data later.
[0157] It is understandable that, in addition to the method mentioned in the above embodiment, that is, in addition to the method of using both the sample media latent features and the sample auxiliary latent features to classify the sample media data, in actual implementation, the sample media data can also be directly classified using only the sample media latent features. Therefore, the first network structure can also include a substructure for directly classifying according to the sample media latent features. It is also understandable that, under the premise that the classification label is known, the substructure can naturally have corresponding losses.
[0158] Based on the above description, the embodiment of the present application can also introduce the loss of the substructure when constructing the training loss function. Therefore, in some embodiments, the method also includes: through the first network structure, classification based on the sample media latent features to obtain a third prediction result; accordingly, according to the differences between the first prediction result and the second prediction result and the classification label, and the difference between the sample auxiliary latent features and the modal reference features, a training loss function is constructed, including: according to the differences between the first prediction result, the second prediction result, and the third prediction result and the classification label, and the difference between the sample auxiliary latent features and the modal reference features, a training loss function is constructed.
[0159] From the contents of the above embodiments, it can be seen that the difference between the first prediction result and the classification label can be used to represent the loss of the substructure in the first network structure that performs classification based on the sample media latent features and the sample auxiliary latent features. The difference between the second prediction result and the classification label can be used to represent the loss of the substructure in the second network structure that performs classification based on the modal reference features. The difference between the sample auxiliary latent features and the modal reference features can be used to represent the loss of the substructure in the first network structure that performs feature conversion on the sample media latent features.
[0160] What is newly added in the embodiment of the present application, that is, the difference between the third prediction result and the classification label, can be used to represent the loss of the substructure in the first network structure that directly classifies according to the implicit features of the sample media. Among them, the loss can also be set based on actual needs, such as using exponential loss, negative log-likelihood loss or square loss, etc., which is not specifically limited in the embodiment of the present application. In addition, the above-mentioned substructures can all be constructed by at least one fully connected layer, and the dimension of their output features can be related to the total types of classification labels. For example, if the total types of classification labels are 10, the feature dimension of the substructure output can also be 10, which is not specifically limited in the embodiment of the present application.
[0161] The computer device can construct a training loss function based on the four losses mentioned above. The embodiment of the present application does not specifically limit the way to construct a training loss function based on the four losses mentioned above, including but not limited to: weighted summing the four losses to obtain a training loss function. Among them, the variables in the training loss function can also include the total batches of training data and the total number of training data in each batch. In addition, in the embodiment of the present application, the subsequent computer device can also train the media classification model based on the training loss function, and its training process can refer to the content of the above embodiment, which will not be repeated here.
[0162] In the above embodiment, in addition to mining and expanding the supervisory information used for training from the modal information corresponding to the media data, the loss calculated based on the difference between the third prediction result obtained by classification based on the sample media latent features and the classification label is also added as supervisory information, thereby enriching the type of supervisory information to overcome the problem of poor training effect caused by too weak supervisory information when only one loss is used to train the media classification model. In addition, since the difference between the third prediction result obtained by classification based on the sample media latent features and the classification label can reflect the classification effect when classification is performed only based on the sample media data, using it as a loss to construct a training loss function can make the trained media classification model more generalizable.
[0163] Based on the various losses mentioned in the above embodiments, in some embodiments, a training loss function is constructed according to the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modal reference feature, including: determining the first loss according to the difference between the first prediction result and the classification label; determining the second loss according to the difference between the second prediction result and the classification label; determining the third loss according to the difference between the sample auxiliary latent feature and the modal reference feature; and constructing a training loss function based on the first loss, the second loss and the third loss.
[0164] Specifically, for the i-th sample media data in a batch of training data, the computer device obtains the sample media latent features of the i-th sample media data; performs feature conversion on the sample media latent features through the first network structure to obtain sample auxiliary latent features; performs classification according to the sample media latent features and the sample auxiliary latent features to obtain a first prediction result; the difference between the first prediction result and the classification label, the first loss determined can refer to the following formula (7):
[0165]
[0166] In the above formula (7), c represents the cth classification label, and there are M classification labels in total. ic Indicates whether the classification label of the i-th sample media data is the c-th classification label. If it is the c-th classification label, y ic The value of is 1, otherwise it is 0. ic Indicates the prediction probability that the first prediction result of the i-th sample media data is the c-th classification label.
[0167] Since the second loss is also the difference between the prediction result and the classification label, the second loss can also be calculated based on the same method in formula (7), or other methods can be used for calculation, and the embodiments of the present application do not specifically limit this. Since the sample auxiliary latent features and the modal reference features are essentially feature vectors, and the similarity between the feature vectors can reflect the degree of difference between the feature vectors, in the actual implementation process, the third loss can be calculated based on the similarity between the sample auxiliary latent features and the modal reference features. Among them, there can be a variety of algorithms for similarity, such as Euclidean distance or cosine similarity, etc., and the embodiments of the present application do not specifically limit this.
[0168] In the above embodiment, a training loss function is constructed to train the media classification model by calculating the loss based on the difference between the auxiliary latent features of the sample and the modal reference features. Since the supervisory information used for training can be mined and expanded from the modal information corresponding to the media data, the problem of poor training effect caused by too weak supervisory information when training the media classification model using only classification labels can be overcome. Therefore, it is beneficial to improve the training effect of the media classification model, and the classification accuracy can also be improved when the media classification model is used to classify the media data later.
[0169] It can be seen from the above embodiments that the third loss can be calculated based on the similarity. Therefore, in some embodiments, the third loss is determined based on the difference between the sample auxiliary latent feature and the modal reference feature, including: for each training data, calculating the similarity between the sample auxiliary latent feature and the modal reference feature corresponding to the corresponding training data; summing the similarities corresponding to each training data, and taking the summed result as the third loss.
[0170] Specifically, for the i-th sample media data in a batch of training data, the sample auxiliary latent feature corresponding to the i-th sample media data can be recorded as y i p , the modal reference feature corresponding to the i-th sample media data can be recorded as y i Therefore, taking the Euclidean distance as an example to represent the similarity, the calculation process of the third loss corresponding to the i-th sample media data can refer to the following formula (8):
[0171] (y i -y i p ) 2 ; (8)
[0172] In the above embodiment, a training loss function is constructed to train the media classification model by calculating the loss based on the difference between the auxiliary latent features of the sample and the modal reference features. Since the supervisory information used for training can be mined and expanded from the modal information corresponding to the media data, the problem of poor training effect caused by too weak supervisory information when training the media classification model using only classification labels can be overcome. Therefore, it is beneficial to improve the training effect of the media classification model, and the classification accuracy can also be improved when the media classification model is used to classify the media data later.
[0173] In some embodiments, the training loss function includes a first training loss function and a second training loss function, and constructs a training loss function according to the difference between the first prediction result and the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modal reference feature, including: constructing a first training loss function according to the difference between the first prediction result and the classification label; constructing a second training loss function according to the difference between the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modal reference feature;
[0174] Correspondingly, the media classification model to be trained is trained by means of a training loss function, including: executing a second training process based on a second training loss function, executing a first training process based on a first training loss function, and the second training process is executed alternately with the first training process; wherein the first training process is a process of adjusting the parameters of a first network structure based on the first loss function and the training samples of the current batch, and the second training process is a process of adjusting the parameters of a second network structure based on the second loss function and the training samples of the current batch.
[0175] Specifically, in combination with the description of the first loss in the above example, the computer device may first construct a first loss function. The first training loss function may refer to the following formula (9):
[0176]
[0177] In the above formula (9), N represents the total number of training data in the batch of training data, and loss1 represents the first training loss function. Other parameters can refer to the explanations in the above formula and will not be repeated here.
[0178] Combined with the description of the second loss in the above example, referring to formula (9), the computer device can construct a corresponding loss function based on the difference between the second prediction result and the classification label. Combined with the description of the third loss in the above example, the computer device can also construct a corresponding loss function based on the difference between the sample auxiliary latent features and the modal reference features. The two constructed loss functions can be integrated to obtain a second training loss function. Among them, the second training loss function can be recorded as loss2. The integration method can be averaging or weighted summing, which is not specifically limited in the embodiments of the present application. According to the difference between the sample auxiliary latent features and the modal reference features, the corresponding loss function constructed can refer to the following formula (10):
[0179]
[0180] In the above formula (10), the definitions of various parameters can be referred to in the above formula and will not be repeated here.
[0181] After obtaining the first training loss function loss1 and the first training loss function loss2, the media classification model can be trained according to loss1 and loss2. As can be seen from the above embodiments, the training process can adopt a simultaneous training or alternating training method. In the embodiment of the present application, an alternating training method can be adopted.
[0182] Specifically, alternating training can be achieved by alternating the second training process with the first training process. It can be understood that the first training loss function is mainly associated with certain substructures in the first network structure, such as a substructure for feature conversion of sample media latent features, and a substructure for classification based on sample media latent features and sample auxiliary latent features. Therefore, the first training process adjusts the parameters in the first network structure, which is actually adjusting the parameters in the associated substructure in the first network structure. The second training loss function is mainly associated with certain substructures in the second network structure, such as a substructure for feature conversion of sample modal features of sample modal information, and a substructure for classification based on modal reference features. Therefore, the second training process adjusts the parameters in the second network structure, which is actually adjusting the parameters in the associated substructure in the second network structure.
[0183] To facilitate understanding of the first training process and the second training process, the alternating execution process is now described in conjunction with each batch of training data: for a certain batch of training data, the computer device first uses the batch of training data to train the second network structure, and then uses the batch of training data to train the first network structure to obtain the respective values of the first training loss function and the second training loss function.
[0184] According to the respective values of the first training loss function and the second training loss function, the computer device determines whether the training stop condition is reached. If not, a batch of training data is updated, the second network structure is trained using the updated batch training data, and the first network structure is trained using the updated batch training data to obtain the respective values of the first training loss function and the second training loss function. The computer device repeats the above process of updating the batch training data and calculating the respective values of the first training loss function and the second training loss function until the respective values of the first training loss function and the second training loss function calculated in a certain training process reach the training stop condition, and the training process ends.
[0185] Among them, when the second network structure is trained using a certain batch of training data, the computer device can adjust the parameters of the second training loss function in the substructure associated with the second network structure, while the parameters of the first training loss function in the substructure associated with the first network structure can remain unchanged. When the first network structure is trained using the same batch of training data, the computer device can adjust the parameters of the first training loss function in the substructure associated with the first network structure, while the parameters of the second training loss function in the substructure associated with the second network structure can remain unchanged. In addition, regarding the method of adjusting parameters during the training process, the embodiments of the present application do not specifically limit this, including but not limited to: adjusting the parameters using a stochastic gradient descent algorithm. Of course, other parameter adjustment methods can also be used, such as a batch gradient descent algorithm or a small batch gradient descent algorithm.
[0186] When the stochastic gradient descent algorithm is used to adjust the parameters, the learning rate can be set according to the demand, and the embodiment of the present application does not specifically limit this. For example, the learning rate initialization can be 0.005, and after completing the training process corresponding to 10 batches of training data, the learning rate can be increased by 10%. In addition, the initialization of the parameters can be realized based on Gaussian distribution. Of course, other initialization methods such as Xavier initialization or MSRA initialization can also be used, and the embodiment of the present application does not specifically limit this. When Gaussian distribution is used to realize parameter initialization, a Gaussian distribution with a variance of 0.01 and a mean of 0 can be used, and the embodiment of the present application does not specifically limit this.
[0187] In combination with the content of the above embodiment, when the process of alternating execution of the first training process and the second training process is controlled based on the first training loss function and the second training loss function, the training stop condition may be that the weighted sum result between loss1 and loss2 no longer decreases or converges, and the embodiment of the present application does not make specific limitations on this. Specifically, the training stop condition may be that (loss1+0.5loss2) no longer decreases. It should be noted that the weight corresponding to loss1 in the above training stop condition is 1, and the weight corresponding to loss2 is 0.5. In the actual implementation process, the weighted weights can be set according to the needs, and the sum of the weights may not be 1, and the embodiment of the present application does not make specific limitations on this.
[0188] In addition, the weighted weights of loss1 and loss2 may not remain constant, and may be dynamically adjusted as the training process progresses. For example, if you need to pay more attention to the training effect corresponding to the modal information at the beginning of the training, you can increase the weight corresponding to loss2. If you need to pay more attention to the training effect corresponding to the media data in the later stage of the training, you can increase the weight corresponding to loss1. Among them, both "the beginning of training" and "the later stage of training" can be determined by the rounds of training. Of course, it can also be determined by the length of training, and the embodiments of the present application do not make specific limitations on this. For example, if a total of 100 rounds of training are conducted, the first 10 rounds can be considered to be the beginning of training, and the last 10 rounds can be considered to be the later stage of training.
[0189] In the actual implementation process, in addition to the alternating training mentioned above, simultaneous training can also be performed, that is, each batch of training data is used to simultaneously train the parameters in the first network structure and the second network structure. At this time, the training stop condition can also be that the weighted sum result between loss1 and loss2 no longer decreases or converges, such as the mean between loss1 and loss2 no longer decreases or converges.
[0190] In the above embodiment, considering that the first network structure mainly processes media data and the second network structure mainly processes modal information, and media data usually carries more information than modal information, this leads to different convergence speeds of the first network structure and the second network structure. If the parameters in the first network structure and the second network structure are trained simultaneously without alternating training, the training tasks corresponding to the first network structure and the second network structure will compete with each other, and the competition will lead to slow parameter gradient descent, which will lead to slow convergence of the training loss function, and then lead to the overall training process consuming a lot of time and low training efficiency. Alternating training can avoid competition between training tasks and improve the training convergence speed. In addition, the first network structure and the second network structure can each fully learn the feature representations of different branches, so that the first network structure and the second network structure with different convergence speeds can each achieve better and more stable training effects, and can achieve joint optimization, and can also improve the classification accuracy when the media classification model is used to classify media data later.
[0191] In the above embodiment, it is mentioned that a third prediction result can be obtained by classification based on the sample media latent features. The difference between the third prediction result and the classification label can be used to represent the loss of the substructure in the first network structure that directly classifies based on the sample media latent features. At the same time, the loss can also be used as supervisory information to train the media classification model.
[0192] Based on this description, in some embodiments, the method also includes: performing classification based on latent features of the sample media through a first network structure to obtain a third prediction result; accordingly, constructing a first training loss function based on the difference between the first prediction result and the classification label, including: constructing a first training loss function based on the difference between the first prediction result and the classification label, and the difference between the third prediction result and the classification label.
[0193] Specifically, referring to formula (9) mentioned in the above embodiment, the computer device can construct a corresponding loss function according to the difference between the first prediction result and the classification label. Referring to formula (9), the corresponding loss function can also be constructed according to the difference between the third prediction result and the classification label. Of course, in the actual implementation process, the loss function corresponding to the third prediction result can also be constructed in a manner different from that in formula (9), such as using exponential loss, negative log-likelihood loss or square loss, etc., which is not specifically limited in the embodiment of the present application.
[0194] Based on the two loss functions constructed by the above process, the computer device can construct a first training loss function loss1. The embodiment of the present application does not specifically limit the method of constructing loss1 based on the two loss functions, including but not limited to: taking the average of the two loss functions and taking the average as loss1. Of course, loss1 can also be constructed by weighted summation.
[0195] In the above embodiment, in addition to mining and expanding the supervisory information used for training from the modal information corresponding to the media data, the loss calculated based on the difference between the third prediction result obtained by classification based on the sample media latent features and the classification label is also added as supervisory information, thereby enriching the type of supervisory information to overcome the problem of poor training effect caused by too weak supervisory information when only one loss is used to train the media classification model. In addition, since the difference between the third prediction result obtained by classification based on the sample media latent features and the classification label can reflect the classification effect when classification is performed only based on the sample media data, using it as a loss to construct a training loss function can make the trained media classification model more generalizable.
[0196] The above embodiment is mainly a process of training a media classification model. In the actual implementation process, the media classification model can also be applied to classify media data. Therefore, in some embodiments, the method can also include a step of performing media data classification through the first network structure in the trained media classification model. This step includes: obtaining the target media data to be classified, extracting the media latent features in the target media data; performing feature conversion on the media latent features to obtain auxiliary latent features for characterizing the modal information of the target media data; and performing classification based on the media latent features and the auxiliary latent features to obtain the category to which the target media data to be classified belongs.
[0197] Specifically, the target media data to be classified may be uploaded to the computer device by the terminal. After receiving the target media data, the computer device may classify the target media data. The process of extracting media latent features, converting the media latent features, and subsequent classification through the media classification model can refer to the relevant description of the training process of the media classification model in the above embodiment, which will not be repeated here.
[0198] In the above embodiment, by performing feature conversion on the media latent features, auxiliary latent features for characterizing the modal information of the target media data can be obtained. Compared with obtaining the feature representation of the modal information by processing the modal information through a large and complex complete deep neural network model, the data processing time in the training process can be effectively reduced, thereby effectively improving the training efficiency of the media classification model. In addition, when classifying media data using a media classification model, it is not necessary to extract the feature representation of the modal information through a large and complex complete deep neural network model, but it is possible to directly obtain the feature representation of the corresponding modal information based on the media latent features, thereby also improving the classification efficiency.
[0199] Finally, compared to taking sample modal information as input and redesigning the media classification model, on the basis of maintaining the original media classification model structure, the sample media hidden features can learn the feature representation corresponding to the sample modal information, thereby avoiding large-scale changes to the structure of the media classification model to save workload. In addition, taking sample modal information as input and redesigning the media classification model will inevitably increase the complexity of the media classification model. On the one hand, this will increase the storage resources occupied by the model storage, and on the other hand, the training and subsequent use processes will also occupy more processing resources due to the more complex model, resulting in reduced processing efficiency. In the embodiment of the present application, only the substructure for realizing feature conversion needs to be added, and there is no need to change the structure of the media classification model on a large scale, thereby reducing the storage resources occupied by the model storage as much as possible, and can also improve the processing efficiency during training and subsequent use.
[0200] For ease of understanding, the training process mentioned in the embodiment of the present application is now described by taking the media data as video, the modal information as one type, the first network structure including a feature conversion substructure, the feature conversion substructure being composed of at least one fully connected layer, the training loss function being constructed by combining four losses, the training loss function including a first training loss function and a second training loss function, and the training process using an alternating training method as an example. Fig.10 In a specific embodiment, the training method and subsequent application method of the media classification model specifically include the following steps:
[0201] Step 1002: Acquire training data, where the training data includes sample videos, sample modality information corresponding to the sample videos, and classification labels to which the sample videos belong.
[0202] Step 1004: extract multiple image frames from the sample video, extract features of each image frame at multiple dimensional levels, combine the multiple features corresponding to each image frame to obtain the video embedding features corresponding to each image frame, and use the feature sequence formed by the video embedding features corresponding to each image frame as the video embedding feature sequence of the sample video.
[0203] Step 1006: determine the time sequence of each video embedding feature in the video embedding feature sequence, and perform superposition processing on the video embedding features in sequence according to the time sequence of each video embedding feature in the video embedding feature sequence to obtain the sample video latent feature.
[0204] Step 1008: Fully connect the sample video latent features through the feature conversion substructure in the first network structure of the media classification model to be trained to obtain sample auxiliary latent features, and combine the sample video latent features and the sample auxiliary latent features through the first network structure to obtain sample combined features, and classify based on the sample combined features to obtain a first prediction result.
[0205] Step 1010: extract features of the sample modal information at multiple dimensional levels, combine the multiple features to obtain sample modal features corresponding to the sample modal information, perform full connection processing on the sample modal features at least once through the second network structure in the video classification model to be trained, obtain modal reference features, perform classification based on the modal reference features, and obtain a second prediction result.
[0206] Step 1012: Classify based on the latent features of the sample video through the first network structure to obtain a third prediction result, construct a first training loss function according to the difference between the first prediction result and the classification label, and the difference between the third prediction result and the classification label, and construct a second training loss function according to the difference between the second prediction result and the classification label, and the difference between the sample auxiliary latent features and the modal reference features.
[0207] Among them, the media classification model to be trained may include the first network structure and the second network structure mentioned above, and the first network structure may include the feature conversion substructure mentioned above. The difference between the sample auxiliary latent feature and the modal reference feature can be used to represent the loss of the feature conversion substructure, the difference between the second prediction result and the classification label can be used to represent the loss of the second network structure, the difference between the first prediction result and the classification label can be used to represent the loss of the substructure in the first network structure used to implement classification based on sample combination features, and the difference between the third prediction result and the classification label can be used to represent the loss of the substructure in the first network structure used to implement classification based on sample video latent features.
[0208] Step 1014: Execute a second training process for training the second network structure in the media classification model to be trained based on the second training loss function, and execute a first training process for training the first network structure in the media classification model to be trained based on the first training loss function, and the second training process is executed alternately with the first training process; stop when the training stop condition is reached, and obtain a trained media classification model.
[0209] Among them, the first training process is a process of adjusting the parameters of the first network structure based on the first loss function and the training samples of the current batch, and the second training process is a process of adjusting the parameters of the second network structure based on the second loss function and the training samples of the current batch.
[0210] Step 1016: obtain the target video data to be classified, extract the video latent features in the target video data through the trained media classification model, perform feature conversion on the video latent features, obtain auxiliary latent features for characterizing the modal information of the target video data, and perform classification based on the video latent features and the auxiliary latent features to obtain the category to which the target video data to be classified belongs.
[0211] The training method of the above-mentioned media classification model constructs a training loss function for training the media classification model based on the losses of each substructure in the first network structure and the second network structure. These losses include losses calculated based on the difference between the sample auxiliary latent features and the modal reference features. Among them, the sample auxiliary latent features are obtained by learning the feature representation corresponding to the modal information of the sample video latent features, and the modal reference features can be used as training labels. Therefore, it is equivalent to mining the supervisory information used for training from the modal information corresponding to the video data. At the same time, the modal information is indeed associated with the content of the video data and is helpful for the classification of the video data, thereby overcoming the problem that the supervisory information is too weak when only the classification label is used to train the media classification model, resulting in poor training effect. Therefore, it is beneficial to improve the training effect of the media classification model, and the classification accuracy can also be improved when the media classification model is subsequently used to classify the video data.
[0212] Secondly, by performing feature conversion on the latent features of the video, auxiliary latent features for characterizing the modal information of the target video data can be obtained. Compared with obtaining the feature representation of the modal information by processing the modal information through a large and complex complete deep neural network model, the data processing time in the training process can be effectively reduced, thereby effectively improving the training efficiency of the media classification model. In addition, when classifying video data using a media classification model, it is not necessary to extract the feature representation of the modal information through a large and complex complete deep neural network model. Instead, the feature representation of the corresponding modal information can be directly obtained based on the latent features of the video, thereby also improving the classification efficiency.
[0213] In addition, compared to taking sample modal information as input and redesigning the media classification model, on the basis of maintaining the original media classification model structure, the sample video hidden features can learn the feature representation corresponding to the sample modal information, thereby avoiding large-scale changes to the structure of the media classification model to save workload. In addition, taking sample modal information as input and redesigning the media classification model will inevitably increase the complexity of the media classification model. On the one hand, this will increase the storage resources occupied by the model storage, and on the other hand, the training and subsequent use processes will also occupy more processing resources due to the more complex model, resulting in reduced processing efficiency. In the embodiment of the present application, only the substructure for realizing feature conversion needs to be added, and there is no need to change the structure of the media classification model on a large scale, thereby reducing the storage resources occupied by the model storage as much as possible, and can also improve the processing efficiency during training and subsequent use.
[0214] Finally, considering that the first network structure mainly processes media data and the second network structure mainly processes modal information, and media data usually carries more information than modal information, this leads to different convergence speeds of the first network structure and the second network structure. If the parameters in the first network structure and the second network structure are trained simultaneously without alternating training, the training tasks corresponding to the first network structure and the second network structure will compete with each other, and the competition will lead to slow parameter gradient descent, which will lead to slow convergence of the training loss function, and then the overall training process will take a lot of time and low training efficiency. Alternating training can avoid competition between training tasks and improve the training convergence speed. In addition, the first network structure and the second network structure can each fully learn the feature representations of different branches, so that the first network structure and the second network structure with different convergence speeds can each achieve better and more stable training effects, and can achieve joint optimization, and can also improve the classification accuracy when the media classification model is used to classify media data later.
[0215] The above embodiment mainly involves the process of training the media classification model. In the actual implementation process, the media classification model can also be used to classify media data. Fig.11 As shown, a media data classification method is provided, and the method is applied to a computer device (the computer device can be specifically Figure 1 The terminal or server in the example is used to illustrate, including the following steps:
[0216] Step 1102: Obtain target media data to be classified, and extract media latent features in the target media data.
[0217] Step 1104: perform feature conversion on the media latent features to obtain auxiliary latent features for characterizing the modal information of the target media data.
[0218] Step 1106: Combine the media latent features and the auxiliary latent features to obtain the target combined features.
[0219] Step 1108: Classify based on the target combination features and output the category to which the target media data to be classified belongs.
[0220] The specific implementation process may refer to the relevant description in the embodiment of the training method of the media classification model, which will not be repeated here.
[0221] The above-mentioned media data classification method can obtain auxiliary latent features for characterizing the modal information of the target media data by performing feature conversion on the media latent features. Since when the media classification model is used to classify the media data, it is not necessary to extract the feature representation of the modal information through a large and complex complete deep neural network model, but the feature representation of the corresponding modal information can be directly obtained based on the media latent features, thereby improving the classification efficiency.
[0222] In addition, compared to taking sample modal information as input and redesigning the media classification model, on the basis of maintaining the original media classification model structure, the sample media hidden features can learn the feature representation corresponding to the sample modal information, thereby avoiding large-scale changes to the structure of the media classification model to save workload. In addition, taking sample modal information as input and redesigning the media classification model will inevitably increase the complexity of the media classification model. On the one hand, this will increase the storage resources occupied by the model storage, and on the other hand, the training and subsequent use processes will also occupy more processing resources due to the more complex model, resulting in reduced processing efficiency. In the embodiment of the present application, only the substructure for realizing feature conversion needs to be added, and there is no need to change the structure of the media classification model on a large scale, thereby reducing the storage resources occupied by the model storage as much as possible, and improving the processing efficiency during the use of the model.
[0223] In some embodiments, feature conversion is performed on media latent features to obtain auxiliary latent features for characterizing modal information of target media data, including: based on a pre-trained feature conversion substructure, feature conversion is performed on media latent features to obtain auxiliary latent features for characterizing modal information of target media data; the feature conversion substructure is obtained by training with modal reference features as label information, and the modal reference features are features obtained in the process of classifying sample modal information of sample media data during the training phase.
[0224] The specific implementation process may refer to the relevant description in the embodiment of the training method of the media classification model, which will not be repeated here.
[0225] In the above embodiment, by only including at least one fully connected layer of feature conversion substructure, the sample media latent features can learn the feature representation corresponding to the sample modal information. Since the media classification model is used to classify the media data later, it is not necessary to extract the modal features of the modal information through a large and complex complete deep neural network model, but the feature representation of the corresponding modal information can be directly obtained based on the media latent features, thereby improving the classification efficiency.
[0226] For ease of understanding, take the media data as video, the modal information as title, and the training method as alternating training as an example. Fig.12The specific structure of the media classification model shown in the figure is used to illustrate the implementation process of the media classification model training method and the media data classification method provided in the embodiment of the present application:
[0227] exist Fig.12 In the training process, “video” refers to the sample video. First, extract the image frames from the sample video, such as Fig.12 A total of 6 frames are extracted. The 6 frames of images are input into an image embedding model composed of a convolutional neural network CNN (Convolutional Neural Networks) and a first fully connected layer FC (Full Connection), and an image embedding sequence can be output. The image embedding sequence corresponds to the video embedding feature sequence mentioned in the above embodiment. This process is mainly an image feature extraction process. In the actual implementation process, image feature extraction can be achieved by using the resnet101 model pre-trained by imagenet, and resnet101 can include Fig.12 CNN and FC shown in. Fig.12 The internal structures of CNN and FC shown in the figure can be referred to in Table 1.
[0228] Table 1
[0229]
[0230] In Table 1 above, pool represents the pooling layer, Max pool represents sampling without using learning parameters, and blocks represents residual blocks. stride represents the participation of learning parameters in the operation, and stride = 2 means that half of the features will be discarded during the convolution process, thus reducing the convolution processing by half and improving the processing speed.
[0231] Input the image embedding sequence into the long short-term memory network LSTMs to obtain the sample video hidden features mentioned in the above embodiment. It should be noted that the output here can be the sample video hidden features output by LSTMs at the last moment. For example, there are 6 image frames corresponding to 6 moments. Since the sample video hidden features output at the last moment have the most complete features learned, the sample video hidden features output at the 6th moment can be used. Input the sample video hidden features into Fig.12 The third classification layer Fc_class3 in can be classified based only on the latent features of the sample video to obtain the third prediction result. The structure of LSTMs and Fc_class3 can be referred to in Table 2 below.
[0232] Table 2
[0233] Processing base name Input / Output Dimensions Layer Type LSTM1 6x2048 / 6x1024 LSTM LSTM2 6x1024 / 6x512 LSTM Fc_class3 1x10 Fc connection
[0234] pass Fig.12 The feature conversion substructure FC_title in the example can perform feature conversion on the sample video latent features to obtain the sample auxiliary latent features. The structure of FC_title can be referred to in Table 3 below.
[0235] Table 3
[0236] Processing base name Output Dimensions Layer Type FC_title 1x512 Fc connection
[0237] Of course, it can be known from the above embodiments that FC_title can also be other structures, such as fully connected layers at both ends, and at least one substructure consisting of an activation layer and a fully connected layer connected in the middle.
[0238] pass Fig.12 The combined processing layer concat in the example can combine the sample video latent features and the sample auxiliary latent features to obtain the sample combined features. Fig.12 The first classification layer FC_class1 in classifies the sample combination features to obtain the first prediction result. The structure of concat and FC_class1 can be referred to in Table 4 below.
[0239] Table 4
[0240] Processing base name Output Dimensions Layer Type concat 1x1024 Fc connection Fc_class1 1x10 Fc connection
[0241] And enter the title of the sample video into Fig.12 The title embedding model in can obtain the sample modal features mentioned in the above embodiment. Fig.12 The second fully connected layer tFC1 and the third fully connected layer tFC2 in can convert the sample modal features to obtain modal reference features. Fig.12 As shown, the output of tFC2 will be connected to FC_title, that is, the modal reference feature output by tFC2 will be used as the training label of the sample auxiliary hidden feature output by FC_title. Fig.12 The second classification layer Fc_class2 in the classification layer classifies the modal reference features to obtain the second prediction result. The structures of tFC1, tFC2 and FC_class2 can be referred to in Table 5 below.
[0242] Table 5
[0243] Processing base name Output Dimensions Layer Type FC1 1x1024 Fc connection tFC2 1x512 Fc connection Fc_class2 1x10 Fc connection
[0244] The above content and various tables illustrate the various substructures in the media classification model. It can be understood that, in combination with the content in the above embodiments, tFC1, tFC2 and Fc_class2 can all belong to the second network structure in the media classification model, and concat and Fc_class1 can all belong to the first network structure in the media classification model. Of course, in the actual implementation process, LSTMs can also belong to the first network structure, the resnet101 model can also belong to the first network structure, and the title Embedding model can belong to the second network structure. As for FC_title, it can belong to the first network structure alone, the second network structure alone, or both the first network structure and the second network structure.
[0245] Combined with the above description of the structural division in the media classification model, the specific process of alternating training is now explained. Fig.12 The first loss mentioned in the above embodiment may be included and recorded as Loss class1; the second loss mentioned in the above embodiment may be included and recorded as Loss class2; the third loss mentioned in the above embodiment may be included and recorded as Loss title-embedding. In addition, Fig.12 The loss constructed by the third prediction result mentioned in the above embodiment may also be included. For the convenience of explanation, it can be called the fourth loss and recorded as Loss class3.
[0246] In combination with the content in the above embodiment, according to the first loss and the fourth loss, a first training loss function can be constructed and recorded as loss1. According to the second loss and the third loss, a second training loss function can be constructed and recorded as loss2. When the second training process is performed based on loss2 (corresponding to Fig.12 The second training process in the second training process may be specifically a process of adjusting the parameters in tFC1, tFC2, Fc_class2 and FC_title. When the first training process is performed based on loss1 (corresponding to Fig.12 The first training process may specifically be a process of adjusting parameters in LSTMs, concat, Fc_class1, and Fc_class3.
[0247] It should be noted that, when executing the first training process, except for the substructures covered by the first training process, the parameters in other substructures in the media classification model can be fixed. Similarly, when executing the second training process, except for the substructures covered by the second training process, the parameters in other substructures in the media classification model can also be fixed. In addition, FC_title can only participate in the second training process mentioned above, or can participate in the first training process and the second training process at the same time, and the embodiment of the present application does not make specific limitations on this. Among them, the training stop condition can be that loss2+0.5loss1 no longer decreases.
[0248] After training the media classification model, the media classification model can be used to implement video classification. Fig.13 The application process is described as follows: the target video to be classified is processed through the image embedding model to obtain the video embedding feature sequence of the target video, that is, the image embedding sequence output. The video embedding feature sequence is processed through the long short-term memory network LSTMs to obtain the video latent features. The video latent features are converted through the feature conversion substructure FC_title to obtain auxiliary latent features. The video latent features and auxiliary latent features are concatenated through the combination processing layer concat to obtain the combined features. The combined features are classified through the first classification layer Fc_class1 to obtain the category to which the target video belongs. The application process of the above video classification can be specifically referred to Fig.13 .
[0249] In addition, due to Fig.12 and Fig.13 The image embedding sequence output by the image embedding model can actually be used to characterize the features of the video, so that in practical applications, the similarity between different videos can be calculated based on the image embedding sequences of different videos. Based on the similarity between different videos, functions such as video clustering and video recommendation can be implemented later.
[0250] From the above application process, we can see that not all substructures involved in the training process are used in the application process. Therefore, in the actual implementation process, some substructures can be divided into modules to achieve functional decoupling. Fig.12In the example, the second fully connected layer tFC1, the third fully connected layer tFC2 and the second classification layer Fc_class2 can be used together as a title classification auxiliary module. It can be understood that the substructure composed of the three layers is mainly to enable the hidden features of the video to learn the feature representation corresponding to the title, so that the title can assist in video classification. In addition, the long short-term memory network LSTMs, the feature conversion substructure FC_title, the combination processing layer concat, the first classification layer Fc_class1 and the third classification layer Fc_class3 can be used together as a sequence representation learning module. It can be understood that the substructure composed of the above five layers is mainly to obtain various feature representations of the video for subsequent classification.
[0251] The embodiment of the present application also provides an application scenario, in which the training method of the above-mentioned media classification model is applied, and the computer device is used as an example for explanation. Specifically, the application of the training method of the media classification model in the application scenario is as follows:
[0252] The server obtains the sample video, the sample title corresponding to the sample video, and the classification label to which the sample video belongs, such as film and television drama, variety show, sports competition, educational lecture, live streaming and other categories. The server performs feature conversion on the sample video latent features of the sample video through the first network structure in the media classification model to obtain the sample auxiliary latent features, and classifies according to the sample video latent features and the sample auxiliary latent features to obtain the first prediction result.
[0253] The server performs feature conversion on the sample title features of the sample title through the second network structure in the media classification model to obtain title reference features, performs classification based on the title reference features, and obtains a second prediction result.
[0254] The server constructs a training loss function based on the differences between the first prediction result, the second prediction result and the classification label, and the differences between the sample auxiliary latent features and the title reference features. The server trains the media classification model through the training loss function and stops when the training stop condition is reached, thereby obtaining a trained media classification model.
[0255] The user then makes a target video and uploads it to the server through the terminal. After receiving the target video made by the user, the server can extract the video latent features of the target video through the media classification model, perform feature conversion on the video latent features, and obtain auxiliary latent features for characterizing the title information of the target video. The server classifies the target video based on the video latent features and the auxiliary latent features through the media classification model to obtain the category to which the target video belongs. Based on the category to which the target video belongs, the server can store the target video in the storage space allocated for the same category.
[0256] The embodiment of the present application also provides an application scenario, and the application scenario applies the above-mentioned media classification model training method. Specifically, the application of the media classification model training method in the application scenario is as follows:
[0257] The terminal is pre-arranged with a media classification model trained based on the above training process. Taking the training process of the media classification model completed by the terminal as an example, the process of the terminal training the media classification model can be as follows:
[0258] The terminal obtains a sample image, a sample profile corresponding to the sample image, and a classification label to which the sample image belongs, such as sports, animals, food, and people. The terminal performs feature conversion on the sample image latent features of the sample image through the first network structure in the media classification model to obtain sample auxiliary latent features, and performs classification according to the sample image latent features and the sample auxiliary latent features to obtain a first prediction result.
[0259] The terminal performs feature conversion on the sample profile features of the sample profile through the second network structure in the media classification model to obtain profile reference features, performs classification based on the profile reference features, and obtains a second prediction result.
[0260] The terminal constructs a training loss function based on the differences between the first prediction result, the second prediction result and the classification label, and the differences between the sample auxiliary latent features and the profile reference features. The terminal trains the media classification model through the training loss function and stops when the training stop condition is reached to obtain a trained media classification model.
[0261] When the user subsequently takes a photo and saves it locally on the terminal, the terminal can extract the image latent features of the target image through the media classification model, perform feature conversion on the image latent features, and obtain auxiliary latent features used to characterize the profile information of the target image. The terminal classifies the target image based on the image latent features and auxiliary latent features through the media classification model to obtain the category to which the target image belongs. Based on the category to which the target image belongs, the terminal can classify and store the target image locally.
[0262] It should be noted that the above application scenarios are illustrative application scenarios, which are used to help understand the solutions of the present application and are not used to limit the actual application scenarios of the present application.
[0263] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0264] Based on the same inventive concept, the embodiment of the present application also provides a media classification model training device for implementing the above-mentioned media classification model training method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in the embodiments of the training device for one or more media classification models provided below can refer to the limitations of the media classification model training method above, and will not be repeated here.
[0265] In some embodiments, Fig.14 As shown, a training device 1400 for a media classification model is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an acquisition module 1402, a first feature conversion module 1404, a first classification module 1406, a second feature conversion module 1408, a second classification module 1410, a construction module 1412 and a training module 1414, wherein:
[0266] An acquisition module 1402 is used to acquire training data, where the training data includes sample media data, sample modality information corresponding to the sample media data, and a classification label to which the sample media data belongs;
[0267] A first feature conversion module 1404 is used to perform feature conversion on the sample media latent features of the sample media data through the first network structure in the media classification model to be trained to obtain sample auxiliary latent features;
[0268] A first classification module 1406, configured to perform classification according to the sample media latent features and the sample auxiliary latent features to obtain a first prediction result;
[0269] A second feature conversion module 1408 is used to perform feature conversion on the sample modal features of the sample modal information through a second network structure in the media classification model to be trained to obtain a modal reference feature;
[0270] A second classification module 1410, configured to perform classification based on the modal reference feature to obtain a second prediction result;
[0271] A construction module 1412 is used to construct a training loss function according to the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature;
[0272] The training module 1414 is used to train the media classification model to be trained by using a training loss function, and stop when a training stop condition is reached to obtain a trained media classification model.
[0273] In some embodiments, the sample media data includes a sample video; the acquisition module 1402 is also used to obtain a video embedding feature sequence of the sample video; determine the timing of each video embedding feature in the video embedding feature sequence; and superimpose the video embedding features in sequence according to the timing of each video embedding feature in the video embedding feature sequence to obtain the sample media latent features.
[0274] In some embodiments, the acquisition module 1402 is also used to extract multiple image frames from the sample video; extract the features of each image frame at multiple dimensional levels; combine and process the multiple features corresponding to each image frame to obtain the video embedding features corresponding to each image frame; and use the feature sequence formed by the video embedding features corresponding to each image frame as the video embedding feature sequence of the sample video.
[0275] In some embodiments, the first network structure includes a feature conversion substructure, which is composed of at least one fully connected layer; the first feature conversion module 1404 is used to perform fully connected processing on the sample media latent features through the feature conversion substructure to obtain the sample auxiliary latent features.
[0276] In some embodiments, the first classification module 1410 is used to combine the sample media latent features and the sample auxiliary latent features to obtain sample combination features; and perform classification based on the sample combination features to obtain a first prediction result.
[0277] In some embodiments, the second feature conversion module 1408 is used to perform at least one full-connection processing on the sample modal features through the second network structure in the media classification model to be trained to obtain modal reference features.
[0278] In some embodiments, the apparatus further comprises:
[0279] The third classification module is used to classify based on the sample media latent features through the first network structure to obtain a third prediction result; accordingly, the construction module 1412 is used to construct a training loss function according to the differences between the first prediction result, the second prediction result, the third prediction result and the classification label, and the difference between the sample auxiliary latent features and the modal reference features.
[0280] In some embodiments, module 1412 is constructed to determine a first loss based on a difference between a first prediction result and a classification label; determine a second loss based on a difference between a second prediction result and a classification label; determine a third loss based on a difference between a sample auxiliary latent feature and a modality reference feature; and construct a training loss function based on the first loss, the second loss, and the third loss.
[0281] In some embodiments, the construction module 1412 is also used to calculate the similarity between the sample auxiliary latent features corresponding to the corresponding training data and the modal reference features for each training data; sum the similarities corresponding to each training data, and use the sum result as the third loss.
[0282] In some embodiments, the training loss function includes a first training loss function and a second training loss function; accordingly, the construction module 1412 is used to construct the first training loss function according to the difference between the first prediction result and the classification label; and to construct the second training loss function according to the difference between the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature;
[0283] The training module 1414 is used to execute the second training process based on the second training loss function, and to execute the first training process based on the first training loss function, and the second training process is executed alternately with the first training process; wherein the first training process is a process of adjusting the parameters of the first network structure based on the first loss function and the training samples of the current batch, and the second training process is a process of adjusting the parameters of the second network structure based on the second loss function and the training samples of the current batch.
[0284] In some embodiments, the apparatus further comprises:
[0285] The third classification module is used to classify based on the latent features of the sample media through the first network structure to obtain a third prediction result; accordingly, the construction module 1412 is used to construct a first training loss function according to the difference between the first prediction result and the classification label, and the difference between the third prediction result and the classification label.
[0286] In some embodiments, the apparatus may also perform classification of media data by using the first network structure in the trained media classification model; accordingly, the apparatus further includes:
[0287] The classification application module is used to obtain the target media data to be classified, extract the media latent features in the target media data; perform feature conversion on the media latent features to obtain auxiliary latent features for characterizing the modal information of the target media data; and perform classification based on the media latent features and the auxiliary latent features to obtain the category to which the target media data to be classified belongs.
[0288] The training device of the above-mentioned media classification model constructs a training loss function for training the media classification model based on the losses of each substructure in the first network structure and the second network structure. These losses include losses calculated based on the difference between the sample auxiliary latent features and the modal reference features. Among them, the sample auxiliary latent features are obtained by learning the feature representation corresponding to its modal information from the sample media latent features, and the modal reference features can be used as training labels. Therefore, it is equivalent to mining the supervisory information used for training from the modal information corresponding to the media data. At the same time, the modal information is indeed associated with the content of the media data and is helpful for the classification of the media data, thereby overcoming the problem that the supervisory information is too weak and the training effect is poor when only the classification label is used to train the media classification model. Therefore, it is beneficial to improve the training effect of the media classification model, and the classification accuracy can also be improved when the media classification model is used to classify the media data in the future.
[0289] For the specific definition of the training device for the media classification model, please refer to the definition of the training method for the media classification model in the above text, which will not be repeated here. Each module in the above-mentioned training device for the media classification model can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0290] Based on the same inventive concept, the embodiment of the present application also provides a media data classification device for implementing the media data classification method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more media data classification device embodiments provided below can refer to the limitations of the media data classification method above, and will not be repeated here.
[0291] In some embodiments, Fig.15 As shown, a media data classification device 1500 is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an acquisition module 1502, a feature conversion module 1504, a combination module 1506 and a classification module 1508, wherein:
[0292] An acquisition module 1502 is used to acquire target media data to be classified and extract media latent features in the target media data;
[0293] A feature conversion module 1504 is used to perform feature conversion on the media latent features to obtain auxiliary latent features for characterizing modal information of target media data;
[0294] A combination module 1506, used for combining the media latent features and the auxiliary latent features to obtain a target combined feature;
[0295] The classification module 1508 is used to perform classification based on the target combination features and output the category to which the target media data to be classified belongs.
[0296] In some embodiments, the feature conversion module 1504 is used to perform feature conversion on media latent features based on a pre-trained feature conversion substructure to obtain auxiliary latent features for characterizing the modal information of the target media data; the feature conversion substructure is obtained by training with the modal reference feature as label information, and the modal reference feature is a feature obtained in the process of classifying the sample modal information of the sample media data during the training phase.
[0297] The above-mentioned media data classification device can obtain auxiliary latent features for characterizing the modal information of the target media data by performing feature conversion on the media latent features. Since when the media classification model is used to classify the media data, it is not necessary to extract the feature representation of the modal information through a large and complex complete deep neural network model, but the feature representation of the corresponding modal information can be directly obtained based on the media latent features, thereby improving the classification efficiency.
[0298] Finally, compared to taking sample modal information as input and redesigning the media classification model, on the basis of maintaining the original media classification model structure, the sample media hidden features can learn the feature representation corresponding to the sample modal information, thereby avoiding large-scale changes to the structure of the media classification model to save workload. In addition, taking sample modal information as input and redesigning the media classification model will inevitably increase the complexity of the media classification model. On the one hand, it will increase the storage resources occupied by the model storage, and on the other hand, the training and subsequent use processes will also occupy more processing resources due to the more complex model, resulting in reduced processing efficiency. In the embodiment of the present application, only the substructure for realizing feature conversion needs to be added, and there is no need to change the structure of the media classification model on a large scale, thereby reducing the storage resources occupied by the model storage as much as possible, and improving the processing efficiency during the use of the model.
[0299] For the specific definition of the media data classification device, please refer to the definition of the media data classification method above, which will not be repeated here. Each module in the above-mentioned media data classification device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0300] In one embodiment, a computer device is provided. The computer device may be a terminal or a server. The internal structure diagram thereof may be as follows: Fig.16 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store training data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a training method for a media classification model or a media data classification method is implemented.
[0301] Those skilled in the art will understand that Fig.16 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0302] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0303] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0304] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0305] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0306] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited thereto. The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0307] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for training a media classification model, characterized in that: The method comprises: Acquire training data, the training data including sample media data, sample modal information corresponding to the sample media data, and a classification label to which the sample media data belongs; the sample modal information is information given from another perspective and associated with the content of the sample media data; Performing feature conversion on the sample media latent features of the sample media data through the first network structure in the media classification model to be trained to obtain sample auxiliary latent features, and performing classification according to the sample media latent features and the sample auxiliary latent features to obtain a first prediction result; Performing feature conversion on the sample modality features of the sample modality information through the second network structure in the media classification model to be trained to obtain modality reference features, and performing classification based on the modality reference features to obtain a second prediction result; Constructing a training loss function according to the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature; The media classification model to be trained is trained using the training loss function and stopped when a training stop condition is reached to obtain a trained media classification model.
2. The method according to claim 1, characterized in that The sample media data includes a sample video; and the process of acquiring the hidden features of the sample media includes: Obtaining a video embedding feature sequence of the sample video; Determining the timing of each video embedding feature in the video embedding feature sequence; According to the time sequence of each video embedding feature in the video embedding feature sequence, the video embedding features are sequentially superimposed to obtain sample media latent features.
3. The method according to claim 2, characterized in that The step of obtaining the video embedding feature sequence of the sample video includes: Extracting a plurality of image frames from the sample video; Extract the features of each image frame at multiple dimensional levels respectively; Combining and processing multiple features corresponding to each image frame to obtain a video embedding feature corresponding to each image frame; A feature sequence formed by the video embedding features corresponding to each image frame is used as the video embedding feature sequence of the sample video.
4. The method according to claim 1, characterized in that The first network structure includes a feature conversion substructure, and the feature conversion substructure is composed of at least one fully connected layer; The performing feature conversion on the sample media latent features of the sample media data to obtain sample auxiliary latent features includes: The sample media latent features of the sample media data are fully connected through the feature conversion substructure to obtain sample auxiliary latent features.
5. The method according to claim 1, characterized in that The classifying according to the sample media latent features and the sample auxiliary latent features to obtain a first prediction result includes: Combining the sample media latent features and the sample auxiliary latent features to obtain sample combined features; Classification is performed based on the sample combination features to obtain a first prediction result.
6. The method according to claim 1, characterized in that The step of performing feature conversion on the sample modality features of the sample modality information through the second network structure in the media classification model to be trained to obtain modality reference features includes: The sample modal features are fully connected at least once through the second network structure in the media classification model to be trained to obtain modal reference features.
7. The method according to claim 1, characterized in that The method further comprises: Classifying based on the latent features of the sample media using the first network structure to obtain a third prediction result; The constructing of a training loss function according to the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature, comprises: A training loss function is constructed according to the differences between the first prediction result, the second prediction result, the third prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature.
8. The method according to claim 1, characterized in that: The constructing of a training loss function according to the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature, comprises: Determining a first loss according to a difference between the first prediction result and the classification label; Determining a second loss according to a difference between the second prediction result and the classification label; Determining a third loss according to a difference between the sample auxiliary latent feature and the modality reference feature; Based on the first loss, the second loss and the third loss, a training loss function is constructed.
9. The method according to claim 8, characterized in that The determining of the third loss according to the difference between the sample auxiliary latent feature and the modality reference feature comprises: For each training data, the similarity between the sample auxiliary latent features corresponding to the corresponding training data and the modal reference features is calculated; The similarities corresponding to each training data are summed up, and the sum result is used as the third loss.
10. The method according to claim 1, characterized in that The training loss function includes a first training loss function and a second training loss function; the training loss function is constructed according to the difference between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature, including: Constructing a first training loss function according to the difference between the first prediction result and the classification label; constructing a second training loss function according to the difference between the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature; The training of the media classification model to be trained by using the training loss function includes: Executing a second training process based on the second training loss function, executing a first training process based on the first training loss function, and the second training process and the first training process are executed alternately; The first training process is a process of adjusting the parameters of the first network structure based on the first training loss function and the training samples of the current batch, and the second training process is a process of adjusting the parameters of the second network structure based on the second training loss function and the training samples of the current batch.
11. The method according to claim 10, characterized in that The method further comprises: Classifying based on the latent features of the sample media using the first network structure to obtain a third prediction result; The constructing a first training loss function according to the difference between the first prediction result and the classification label includes: A first training loss function is constructed based on the difference between the first prediction result and the classification label, and the difference between the third prediction result and the classification label.
12. The method according to any one of claims 1 to 11, characterized in that The method further includes a step of performing media data classification by using a first network structure in the trained media classification model, the step comprising: Acquire target media data to be classified, and extract media hidden features in the target media data; Performing feature conversion on the media latent features to obtain auxiliary latent features for characterizing modal information of the target media data; Classification is performed based on the media latent features and the auxiliary latent features to obtain the category to which the target media data to be classified belongs.
13. A method for classifying media data, characterized in that: The method comprises: Acquire target media data to be classified, and extract media hidden features in the target media data; Based on the feature conversion substructure in the first network structure of the pre-trained media classification model, the media latent features are converted to obtain auxiliary latent features for characterizing the modal information of the target media data; the media classification model is obtained by training the training method of the media classification model according to any one of claims 1 to 12; Combining the media latent feature and the auxiliary latent feature to obtain a target combined feature; Classification is performed based on the target combination feature, and the category to which the target media data to be classified belongs is output.
14. A training device for a media classification model, characterized in that: The device comprises: an acquisition module, configured to acquire training data, wherein the training data includes sample media data, sample modal information corresponding to the sample media data, and a classification label to which the sample media data belongs; the sample modal information is information given from another perspective and associated with the content of the sample media data; A first feature conversion module, configured to perform feature conversion on the sample media latent features of the sample media data through a first network structure in the media classification model to be trained, so as to obtain sample auxiliary latent features; A first classification module, used for classifying according to the sample media latent features and the sample auxiliary latent features to obtain a first prediction result; a second feature conversion module, configured to perform feature conversion on the sample modal features of the sample modal information through a second network structure in the media classification model to be trained, so as to obtain modal reference features; A second classification module, used for performing classification based on the modal reference feature to obtain a second prediction result; A construction module, used to construct a training loss function according to the differences between the first prediction result, the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature; The training module is used to train the media classification model to be trained by using the training loss function, and stop when the training stop condition is reached to obtain a trained media classification model.
15. The device according to claim 14, characterized in that The sample media data includes a sample video; the acquisition module is also used to obtain a video embedding feature sequence of the sample video; determine the timing of each video embedding feature in the video embedding feature sequence; and superimpose the video embedding features in sequence according to the timing of each video embedding feature in the video embedding feature sequence to obtain a sample media latent feature.
16. The device according to claim 15, characterized in that The acquisition module is also used to extract multiple image frames from the sample video; extract the features of each image frame at multiple dimensional levels; Combining and processing multiple features corresponding to each image frame to obtain a video embedding feature corresponding to each image frame; A feature sequence formed by the video embedding features corresponding to each image frame is used as the video embedding feature sequence of the sample video.
17. The device according to claim 14, characterized in that The first network structure includes a feature conversion substructure, and the feature conversion substructure is composed of at least one fully connected layer; The first feature conversion module is further used to perform full connection processing on the sample media latent features of the sample media data through the feature conversion substructure to obtain sample auxiliary latent features.
18. The device according to claim 14, characterized in that The first classification module is further used to combine the sample media latent features and the sample auxiliary latent features to obtain sample combination features; and to perform classification based on the sample combination features to obtain a first prediction result.
19. The device according to claim 14, characterized in that The second feature conversion module is also used to perform at least one full-connection process on the sample modal features through the second network structure in the media classification model to be trained to obtain modal reference features.
20. The device according to claim 14, characterized in that The device also includes: A third classification module, configured to perform classification based on the latent features of the sample media through the first network structure to obtain a third prediction result; The construction module is used to construct a training loss function according to the differences between the first prediction result, the second prediction result, the third prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modal reference feature.
21. The device according to claim 14, characterized in that The construction module is also used to determine a first loss based on the difference between the first prediction result and the classification label; determine a second loss based on the difference between the second prediction result and the classification label; determine a third loss based on the difference between the sample auxiliary latent feature and the modal reference feature; and construct a training loss function based on the first loss, the second loss and the third loss.
22. The device according to claim 21, characterized in that The construction module is also used to calculate the similarity between the sample auxiliary latent features corresponding to the corresponding training data and the modal reference features for each training data; sum the similarities corresponding to each training data, and use the sum result as the third loss.
23. The device according to claim 14, characterized in that The training loss function includes a first training loss function and a second training loss function; the construction module is further used to construct the first training loss function according to the difference between the first prediction result and the classification label; and to construct the second training loss function according to the difference between the second prediction result and the classification label, and the difference between the sample auxiliary latent feature and the modality reference feature; The training module is also used to execute a second training process based on the second training loss function, and to execute a first training process based on the first training loss function, and the second training process is executed alternately with the first training process; wherein the first training process is a process of adjusting parameters of a first network structure based on the first training loss function and training samples of a current batch, and the second training process is a process of adjusting parameters of a second network structure based on the second training loss function and training samples of a current batch.
24. The device according to claim 23, characterized in that The device also includes: A third classification module, configured to perform classification based on the latent features of the sample media through the first network structure to obtain a third prediction result; The construction module is also used to construct a first training loss function according to the difference between the first prediction result and the classification label, and the difference between the third prediction result and the classification label.
25. The device according to any one of claims 14 to 24, characterized in that The device also includes a classification application module, which is used to perform media data classification through the first network structure in the trained media classification model, and the classification application module is also used to obtain target media data to be classified and extract media hidden features in the target media data; Performing feature conversion on the media latent features to obtain auxiliary latent features for characterizing modal information of the target media data; Classification is performed based on the media latent features and the auxiliary latent features to obtain the category to which the target media data to be classified belongs.
26. A media data classification device, characterized in that: The device comprises: An acquisition module, used to acquire target media data to be classified and extract media hidden features in the target media data; A feature conversion module, configured to perform feature conversion on the media latent features based on a feature conversion substructure in a first network structure in a pre-trained media classification model, so as to obtain auxiliary latent features for characterizing modal information of the target media data; the media classification model is obtained by training the media classification model training device according to any one of claims 14 to 25; A combination module, used for combining the media latent feature and the auxiliary latent feature to obtain a target combined feature; The classification module is used to perform classification based on the target combination feature and output the category to which the target media data to be classified belongs.
27. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 13 are implemented.
28. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
29. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Video classification method and device, computer readable storage medium and electronic equipment
CN112364810A
Video classification model training method and device, video classification method and device, equipment and medium
CN113449700A
Federal learning modeling optimization method and device, readable storage medium and program product
CN113516255A
Multimedia resource classification model training method and multimedia resource recommendation method
CN113590849A
Video type determination method and related device
CN114064972A