Method for training multimedia resource classification model and multimedia resource classification method
Patent Information
- Application Number
- CN202510178573.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2026-08-18
AI Technical Summary
这种实现方式需要训练大量分类模型,耗费训练资源和计算资源,且分类效率不高
[0027]This application obtains sample description information corresponding to sample multimedia resources; then, it inputs the sample description information into a pre-constructed multimedia resource classification model, and uses multiple expert classification networks in the model to perform multi-level label prediction processing on the sample description information, obtaining the sample-level expert prediction results output by each expert classification network. Each expert classification network can capture the features of the sample description information from different dimensions, thereby improving the accuracy of the classification results. Next, based on the differences between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels, the sample network weights corresponding to each expert classification network are determined. Based on the sample network weights corresponding to each expert classification network and the sample-level expert prediction results output by each expert classification network, the sample-level comprehensive classification result is determined. Since the sample network weights change in real time with the expert prediction results during model training, these sample network weights can represent the classification accuracy of the expert classification networks in real time. Therefore, the fused sample-level comprehensive classification result can accurately integrate the outputs of multiple expert classification networks. Finally, based on the differences between the hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchy relationships among the comprehensive classification results of multiple samples, the multimedia resource classification model is trained to obtain the target multimedia resource classification model. This model considers both the differences between the comprehensive classification results and the labels, and the consistency among the comprehensive classification results. On the one hand, it ensures the accuracy of the classification results; on the other hand, it eliminates the need to train a corresponding next-level classification model for each item in each category hierarchy, saving training and computational resources. Therefore, the method in this application can improve the accuracy and efficiency of multi-level classification of multimedia resources.
Smart Images

Figure CN122594949A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method for a multimedia resource classification model and a multimedia resource classification method. Background Technology
[0002] Hierarchical multi-label classification is an important task in the fields of natural language processing and computer vision. As the types of data increase, the category system will gradually expand and become more subdivided. The greater the degree of subdivision, the more difficult the recognition becomes and the lower the recognition accuracy.
[0003] In existing technologies, hierarchical multi-label classification tasks are typically divided into multiple basic multi-classification tasks. Taking a two-level category hierarchy as an example, the category system is flattened, and the model directly predicts the second-level categories. Then, the results of the predicted second-level categories are used to trace back to the first-level categories to obtain the hierarchical classification results. However, this approach has low classification accuracy because the model struggles to learn the accurate feature representations of each sub-category. Alternatively, a classification model can first predict the first-level categories, and then the second-level classification model corresponding to the first-level category classification results can be used to further subdivide the categories to obtain the second-level classification results. This approach requires training a large number of classification models, consuming significant training and computational resources, and has low classification efficiency. Summary of the Invention
[0004] This application provides a training method for a multimedia resource classification model and a multimedia resource classification method, which can improve the accuracy and efficiency of multi-level classification of multimedia resources.
[0005] On the one hand, this application provides a training method for a multimedia resource classification model, the multimedia resource classification model including multiple expert classification networks, the method comprising:
[0006] Obtain sample description information corresponding to the sample multimedia resources; the sample description information is labeled with multi-level classification tags of the preset category level corresponding to the sample multimedia resources;
[0007] The sample description information is input into the multimedia resource classification model, and the sample description information is processed by multiple expert classification networks to perform multi-level label prediction, so as to obtain the sample-level expert prediction results output by each expert classification network.
[0008] Based on the difference between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels, the sample network weights corresponding to each expert classification network are determined.
[0009] Based on the sample network weights corresponding to each expert classification network and the sample-level expert prediction results output by each expert classification network, the comprehensive classification result at the sample level is determined; the comprehensive classification result at the sample level includes the comprehensive classification results corresponding one-to-one with the preset category level;
[0010] The multimedia resource classification model is trained based on the differences between the sample hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchical associations among multiple sample comprehensive classification results, to obtain the target multimedia resource classification model.
[0011] On the other hand, this application provides a multimedia resource classification method, which includes:
[0012] Obtain the target description information corresponding to the target multimedia resource;
[0013] The target description information is input into the target multimedia resource classification model for multi-level label prediction processing to obtain the target hierarchical comprehensive classification result; the target multimedia resource classification model is trained according to the training method of the multimedia resource classification model described above.
[0014] On the other hand, this application provides a training device for a multimedia resource classification model, the multimedia resource classification model including multiple expert classification networks, the device comprising:
[0015] The sample acquisition module is used to acquire sample description information corresponding to sample multimedia resources; the sample description information is marked with multi-level classification tags of the preset category level corresponding to the sample multimedia resources;
[0016] The expert prediction module is used to input the sample description information into the multimedia resource classification model, and use the multiple expert classification networks to perform multi-level label prediction processing on the sample description information to obtain the sample-level expert prediction results output by each expert classification network.
[0017] The network weight determination module is used to determine the sample network weights corresponding to each expert classification network based on the differences between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels.
[0018] The comprehensive classification result determination module is used to determine the comprehensive classification result at the sample level based on the sample network weights corresponding to each expert classification network and the sample-level expert prediction results output by each expert classification network; the comprehensive classification result at the sample level includes the comprehensive classification results of samples that correspond one-to-one with the preset category level;
[0019] The model training module is used to train the multimedia resource classification model based on the differences between the sample hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchical associations among multiple sample comprehensive classification results, to obtain the target multimedia resource classification model.
[0020] On the other hand, this application provides a multimedia resource classification device, which includes:
[0021] The acquisition module is used to acquire the target description information corresponding to the target multimedia resource;
[0022] The classification module is used to input the target description information into the target multimedia resource classification model for multi-level label prediction processing to obtain the target hierarchical comprehensive classification result; the target multimedia resource classification model is trained according to the training method of the multimedia resource classification model described above.
[0023] On the other hand, an electronic device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the training method of the multimedia resource classification model as described above, or the multimedia resource classification method as described above.
[0024] On the other hand, a computer storage medium is provided, which stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the training method of the multimedia resource classification model as described above, or the multimedia resource classification method as described above.
[0025] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, executes the computer instructions, and causes the computer device to perform a training method for implementing the multimedia resource classification model as described above, or a multimedia resource classification method as described above.
[0026] The training method and multimedia resource classification method of the multimedia resource classification model provided in this application have the following technical effects:
[0027] This application obtains sample description information corresponding to sample multimedia resources; then, it inputs the sample description information into a pre-constructed multimedia resource classification model, and uses multiple expert classification networks in the model to perform multi-level label prediction processing on the sample description information, obtaining the sample-level expert prediction results output by each expert classification network. Each expert classification network can capture the features of the sample description information from different dimensions, thereby improving the accuracy of the classification results. Next, based on the differences between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels, the sample network weights corresponding to each expert classification network are determined. Based on the sample network weights corresponding to each expert classification network and the sample-level expert prediction results output by each expert classification network, the sample-level comprehensive classification result is determined. Since the sample network weights change in real time with the expert prediction results during model training, these sample network weights can represent the classification accuracy of the expert classification networks in real time. Therefore, the fused sample-level comprehensive classification result can accurately integrate the outputs of multiple expert classification networks. Finally, based on the differences between the hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchy relationships among the comprehensive classification results of multiple samples, the multimedia resource classification model is trained to obtain the target multimedia resource classification model. This model considers both the differences between the comprehensive classification results and the labels, and the consistency among the comprehensive classification results. On the one hand, it ensures the accuracy of the classification results; on the other hand, it eliminates the need to train a corresponding next-level classification model for each item in each category hierarchy, saving training and computational resources. Therefore, the method in this application can improve the accuracy and efficiency of multi-level classification of multimedia resources. Attached Figure Description
[0028] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a schematic diagram of the application environment of the multimedia resource classification model provided in the embodiments of this application;
[0030] Figure 2 This is a schematic diagram of the structure for text classification using a Bert encoder provided in an embodiment of this application;
[0031] Figure 3 This is a flowchart illustrating a training method for a multimedia resource classification model provided in an embodiment of this application;
[0032] Figure 4This is a schematic diagram of the structure of the multimedia resource classification model provided in the embodiments of this application;
[0033] Figure 5 This is a flowchart illustrating the process of determining the sample network weights corresponding to the expert classification network, provided in an embodiment of this application.
[0034] Figure 6 This is a schematic diagram of the process for determining the target loss and training the model provided in an embodiment of this application;
[0035] Figure 7 This is a schematic diagram of the structure of the training device for the multimedia resource classification model provided in this application embodiment;
[0036] Figure 8 This is a schematic diagram of the structure of the multimedia resource classification device provided in the embodiments of this application;
[0037] Figure 9 This is a hardware structure block diagram of a server for a training method of a multimedia resource classification model provided in an embodiment of this application. Detailed Implementation
[0038] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0039] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0040] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0041] Please see Figure 1 , Figure 1 This is a schematic diagram of the application environment of the multimedia resource classification model provided in this application embodiment. The application environment may include at least a server 100 and a terminal 200.
[0042] In an optional embodiment, server 100 can be used to train a multimedia resource classification model and deploy the trained target multimedia resource classification model. Server 100 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0043] In an optional embodiment, terminal 200 can be used to display the hierarchical classification results of multimedia resources. For video multimedia resources, the hierarchical tags of the video can be displayed, such as game-mobile game, dance-square dance, etc. For document multimedia resources, the hierarchical tags of the document can be displayed, such as news-finance, books-children's books, etc. It should be noted that multimedia resources can also be images, applications, or other forms, and this application does not limit the specific type of multimedia resources. Specifically, terminal 200 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, in-vehicle terminals, smart TVs, etc.; it can also be software running on the above-mentioned electronic devices, such as applications, applets, etc. The operating system running on the electronic device in this embodiment of the application can include, but is not limited to, Android, iOS, Linux, Windows, etc.
[0044] Hierarchical multi-label classification is an important task in the fields of natural language processing and computer vision. As the types of data increase, the classification system gradually expands and becomes more subdivided. The greater the subdivision, the more difficult the recognition becomes and the lower the recognition accuracy.
[0045] In existing technologies, multi-label hierarchical classification tasks are typically divided into several basic multi-classification tasks. Taking a two-level category hierarchy as an example, existing technologies flatten the category system, allowing the model to directly predict the second-level categories, and then use the predicted second-level category results to trace back to the first-level categories to obtain the hierarchical classification results. In practical applications, the BERT encoder is usually used to directly output the classification results of the second-level categories. The BERT encoder (Bidirectional Encoder Representations from Transformers) is a pre-trained natural language processing model capable of performing text classification tasks. Figure 2 This is a schematic diagram illustrating the structure of text classification using a BERT encoder, as provided in an embodiment of this application. Figure 2 As shown, the text to be processed is first segmented into multiple words. Then, a [CLS] tag is added to the beginning of the text and a [SEP] tag is added to the end of each sentence. Finally, it is converted into a fixed-length embedded representation, such as... Figure 2 E in CLS E1…En, E SEP Next, these embedded representations are input to the encoding layer and encoded to obtain the output vector, such as... Figure 2 O in CLS O1…On, O SEP Since this is a text classification task, O(n) was finally used. CLS The text's secondary category classification is predicted using a classification layer (usually a fully connected layer). Then, the primary category classification is retrieved to obtain the hierarchical classification result.
[0046] However, the hierarchical classification results obtained in this way are particularly dependent on the accuracy of the secondary classification results. Using only a single classifier like the BERT model for hierarchical classification makes it difficult to capture features across multiple dimensions, and it does not consider the relationships between different levels when predicting classification results, leading to low classification accuracy.
[0047] In another example, existing technologies first use a primary classification model to predict primary categories, and then use a secondary classification model corresponding to the primary category results to predict secondary categories, thus obtaining hierarchical classification results. However, this approach requires training a classification model for each item in each category level to perform the next level of classification. For example, the primary category "games" needs a corresponding classification model to be trained to further distinguish between "mobile games, PC games, and mini-games"; similarly, the secondary category "dance" also needs a corresponding classification model to be trained... When there are many category levels and many items in each category level, it often requires a lot of training resources.
[0048] The following describes a training method and a classification method for a multimedia resource classification model based on this application. The multimedia resource classification model in this application is a hierarchical classification model that can output multi-level classification labels and does not require the use of multiple classifiers with hierarchical relationships, thus effectively improving the accuracy and efficiency of multi-level classification of multimedia resources. Figure 3 This is a flowchart illustrating a training method for a multimedia resource classification model provided in this application. This specification provides the operational steps of the method described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or server product execution, the method can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment) as shown in the embodiments or drawings. Specifically, the multimedia resource classification model of this application includes multiple expert classification networks, such as... Figure 3 As shown, the method may include:
[0049] S301: Obtain sample description information corresponding to the sample multimedia resources;
[0050] The sample description information is labeled with multi-level classification tags corresponding to the preset category level of the sample multimedia resources;
[0051] The sample description information is used to describe the semantic content of the sample multimedia resource, and this application does not limit the specific form of the sample description information. For example, when the sample multimedia resource is a video, the sample description information can be the video title. The video title is one of the main components of the video content, and based on the video title, the semantic content of the video can be understood, so as to classify the video multimedia resource into an accurate subcategory. As another example, when the sample multimedia resource is text, the sample description information can be a summary of the text; when the sample multimedia resource is an image, the sample description information can be the descriptive text corresponding to the image.
[0052] It should be noted that this application does not limit the number of preset category levels. The preset category levels can be set according to actual needs, such as 2, 3, 4, etc. In this application, the preset category level is 2. Assuming the sample description information is "xxx (mini-game name), a guide to getting 600 points", the multi-level category tag corresponding to this sample multimedia resource can be "games-mini-games"; assuming the sample description information is "square dancing is healthier", the multi-level category tag corresponding to this sample multimedia resource can be "dance-square dance"; for example, if the sample description information is "xxx (game character) is economically suppressed and can't get ahead, here's your phone to play!", the multi-level category tag corresponding to this sample multimedia resource can be "games-mobile games"; if the sample description information is "the rapid iteration speed of mobile phones is the main reason they are called consumables", the multi-level category tag corresponding to this sample multimedia resource can be "digital-mobile phones".
[0053] S303: Input the sample description information into the multimedia resource classification model, and use the multiple expert classification networks to perform multi-level label prediction processing on the sample description information to obtain the sample-level expert prediction results output by each expert classification network;
[0054] The multimedia resource classification model of this application includes multiple expert classification networks, each of which can output hierarchical expert prediction results for a preset category level. That is, the sample-level expert prediction results include sample expert prediction results that correspond one-to-one with the preset category level. Taking the second-level preset category level as an example, for sample description information, each expert classification network can simultaneously output prediction results for the first-level category and prediction results for the second-level category.
[0055] It should be noted that the network structures of the expert classification networks can vary. For example, expert classification network 1 can be a Convolutional Neural Network (CNN), expert classification network 2 can be a Long Short-Term Memory (LSTM) network, expert classification network 3 can be a Feedforward Neural Network (FNN), and so on. Expert classification networks can also use other types of network structures, which will not be listed here. By setting up multiple types of expert classification networks, information from different dimensions of sample descriptions can be captured and comprehensively judged, thereby improving classification accuracy.
[0056] S305: Based on the difference between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels, determine the sample network weights corresponding to each expert classification network.
[0057] It should be noted that each level in the multi-level classification labels can be one-hot encoded, resulting in multiple label encoding vectors. The expert classification network can simultaneously output the confidence score of the sample description information belonging to each category item in the preset category level. For example, for sample description information A, expert classification network 1 outputs a confidence score of 0.8 for sample description information A belonging to the first-level category "Games", 0.1 for the first-level category "Dance", 0.1 for the first-level category "Digital", 0.6 for the second-level category "Mini Games", and so on. Based on the above confidence scores, prediction vectors can be formed according to the category level. By comparing the difference between the label encoding vector and the prediction vector corresponding to the same category level, the prediction difference value of the expert classification network at that category level can be determined, thereby determining the sample network weight of the expert classification network at that category level.
[0058] S307: Based on the sample network weights corresponding to each expert classification network and the sample-level expert prediction results output by each expert classification network, determine the comprehensive classification result at the sample level;
[0059] The comprehensive classification result of the sample hierarchy includes the comprehensive classification result of the sample corresponding one-to-one with the preset category hierarchy;
[0060] It should be noted that the sample network weights change dynamically in real time during model training based on the expert prediction results at the sample level, in order to represent the real-time contribution of each expert classification network to the final comprehensive classification result.
[0061] The comprehensive classification result of the samples is the classification result corresponding to each preset category level. The comprehensive classification result of samples at multiple category levels is the comprehensive classification result of the sample level.
[0062] By fusing the sample-level expert prediction results output by each expert classification network, a comprehensive sample-level classification result corresponding to the sample multimedia resources is obtained. In this application, a weighted summation method can be used to fuse the sample-level expert prediction results output by each expert classification network.
[0063] S309: Based on the differences between the comprehensive classification results of the sample levels and the multi-level classification labels, as well as the category hierarchy relationships among the comprehensive classification results of multiple samples, train the multimedia resource classification model to obtain the target multimedia resource classification model.
[0064] Because the expert prediction results at the sample level include the comprehensive classification results of samples at the same preset category level, there may be situations where the prediction result for a second-level category is not below the prediction result for a first-level category. For example, the predicted comprehensive classification result at the sample level might be "Dance-Mobile Games". Therefore, in order to further learn the constraint relationships between category levels, this application, based on determining the differences between labels, further introduces category level association relationships to train the multimedia resource classification model, thereby obtaining an accurate target multimedia resource classification model.
[0065] This application embodiment obtains sample description information corresponding to sample multimedia resources; then, it inputs the sample description information into a pre-constructed multimedia resource classification model, and uses multiple expert classification networks in the multimedia resource classification model to perform multi-level label prediction processing on the sample description information, obtaining the sample-level expert prediction results output by each expert classification network; each expert classification network can capture the features of the sample description information from different dimensions, thereby improving the accuracy of the classification results. Next, based on the differences between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels, the sample network weights corresponding to each expert classification network are determined; based on the sample network weights corresponding to each expert classification network and the sample-level expert prediction results output by each expert classification network, the sample-level comprehensive classification result is determined; since the sample network weights change in real time with the expert prediction results during model training, the sample network weights can represent the classification accuracy of the expert classification networks in real time, and the fused sample-level comprehensive classification result can accurately integrate the outputs of multiple expert classification networks. Finally, based on the differences between the hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchy relationships among the comprehensive classification results of multiple samples, the multimedia resource classification model is trained to obtain the target multimedia resource classification model. This model considers both the differences between the comprehensive classification results and the labels, and the consistency among the comprehensive classification results. On the one hand, it ensures the accuracy of the classification results; on the other hand, it eliminates the need to train a classification model for each item in each category hierarchy, saving training and computational resources. Therefore, the method in this application can improve the accuracy and efficiency of multi-level classification of multimedia resources.
[0066] It should be noted that, before inputting the sample description information into the multimedia resource classification model, the method further includes:
[0067] The sample description information is subjected to semantic feature extraction processing to obtain sample description semantic features;
[0068] It should be noted that the semantic feature extraction process described above utilizes a pre-trained semantic extraction network. This pre-trained semantic extraction network can be trained independently and does not participate in the training process of the multimedia resource classification model. Alternatively, it can be understood that the multimedia resource classification model can include the pre-trained semantic extraction network, but its network parameters are frozen during training.
[0069] Pre-trained semantic extraction networks can be Figure 2 The BERT encoder in the code can also be other network structures.
[0070] Specifically, the step of extracting semantic features from the sample description information to obtain sample description semantic features may include:
[0071] The sample description information is segmented into words to obtain a character sequence;
[0072] The character sequence is subjected to text embedding encoding to obtain the embedding feature corresponding to each character;
[0073] Semantic features are extracted from the embedding features corresponding to each character to obtain the character features corresponding to each character;
[0074] The character features corresponding to each character are concatenated to obtain the semantic features of the sample description.
[0075] In this application, the textual sample description information is first segmented into a series of characters (tokens). Then, a semantic extraction network is used to perform text embedding encoding on each token in the sample description information, obtaining the embedding feature corresponding to each character. It should be noted that text embedding encoding can include character encoding (TokenEmbedding), segment encoding (Segment Embedding), and position encoding (Position Embedding). The embedding feature corresponding to each character is obtained by adding the character encoding result, segment encoding result, and position encoding result. Next, the above embedding features are input into the encoding layer for semantic feature extraction, obtaining the semantic features of the sample description.
[0076] In this embodiment, semantic features are extracted from the sample description information in advance to learn accurate sample description semantic features, which facilitates subsequent classification prediction by various expert classification networks.
[0077] The step of using the multiple expert classification networks to perform multi-level label prediction processing on the sample description information to obtain the sample-level expert prediction results output by each expert classification network includes: using the multiple expert classification networks to perform multi-level label prediction processing on the semantic features of the sample description respectively to obtain the sample-level expert prediction results output by each expert classification network.
[0078] Figure 4 This is a schematic diagram of the structure of the multimedia resource classification model provided in the embodiments of this application.
[0079] like Figure 4 As shown, the multiple expert classification networks include at least one convolutional neural network, which captures input information linking information through its receptive field. The process of using the multiple expert classification networks to perform multi-level label prediction processing on the semantic features of the sample descriptions, obtaining the sample-level expert prediction results output by each expert classification network, includes:
[0080] Based on the dimension of the character feature corresponding to each character in the sample description semantic features, the convolution kernel size is determined and a convolution kernel is constructed according to the determined kernel size. The sample description semantic features can be viewed as a k*m matrix vector representation. The character feature T of each character can be represented as [e1, e2, ... em], and the dimension of the character feature corresponding to each character is m. k is the number of character tokens, i.e., the sentence length of the sample description information. The width of the convolution kernel is equal to the dimension of the character feature, so the convolution kernel size can be h*m, where h is the character span, which can be understood as the receptive field size. The character span h can take values of 2, 3, 5, etc.
[0081] The convolution kernel is used to perform a sliding convolution operation on the semantic features describing the sample to obtain a convolutional feature set.
[0082] Perform pooling operations on each convolutional feature in the convolutional feature set to obtain a pooled feature set;
[0083] Dimensionality reduction is performed on each pooling feature in the pooling feature set to obtain the prediction feature set;
[0084] It should be noted that a fully connected layer can be used to reduce the dimensionality of each pooling feature in the pooling feature set. Specifically, a discard layer can be used to randomly set each element in the pooling feature set to 0, thereby reducing the occurrence of overfitting.
[0085] Multi-level label prediction processing is performed on the predicted feature set to obtain sample-level expert prediction results corresponding to the convolutional neural network.
[0086] The convolutional neural network performs classification prediction to obtain the category probability distribution, which can be represented as:
[0087] P0 = CNN([T0, Ti…Tk])
[0088] P0 represents the sample-level expert prediction result output by the convolutional neural network, Ti represents the character feature of the i-th character, and k is the number of characters.
[0089] In this embodiment, by introducing a convolutional neural network as an expert classification network, it is possible to capture the multi-scale semantic features of the input information through different receptive fields, as well as the combined features between characters, thereby obtaining accurate classification results and providing a foundation for the subsequent integration of classification results from various expert classification networks.
[0090] like Figure 4 As shown, the multiple expert classification networks include at least one Long Short-Term Memory (LSTM) network, which captures information based on the coherence of continuous textual expression. The process of using the multiple expert classification networks to perform multi-level label prediction processing on the semantic features of the sample descriptions, obtaining the sample-level expert prediction results output by each expert classification network, includes:
[0091] The number of stacked layers of the long short-term memory (LSTM) in the LSM network is determined based on the number of characters represented in the semantic features of the sample description.
[0092] The semantic features of the sample description are input into a long short-term memory layer with the number of stacked layers for feature extraction processing in sequence, and the output of the last long short-term memory layer is determined as the predicted feature set.
[0093] Multi-level label prediction processing is performed on the predicted feature set to obtain the sample-level expert prediction results corresponding to the Long Short-Term Memory Network.
[0094] The sample-level expert prediction results corresponding to this Long Short-Term Memory network can be expressed as follows:
[0095] P1 = LSTM([T0, Ti…Tk])
[0096] P1 represents the sample-level expert prediction result output by the Long Short-Term Memory network, Ti represents the character feature of the i-th character, and k is the number of characters.
[0097] In this embodiment, by introducing a long short-term memory network, effective semantic features are captured from the perspective of character coherence. This method can obtain the continuous expression features of the input information, thereby obtaining accurate classification results and providing a foundation for the subsequent comprehensive classification results of various expert classification networks.
[0098] like Figure 4As shown, the plurality of expert classification networks includes at least one decision tree network, which can serve as the output classification result for statistical features. In one implementation, the decision tree network can be an extreme gradient boosting (XGB) algorithm, or other statistical feature decision networks; this application does not limit the specific type of decision tree network. The step of using the plurality of expert classification networks to perform multi-level label prediction processing on the sample description semantic features to obtain the sample-level expert prediction results output by each expert classification network includes:
[0099] Obtain the number of characters represented in the semantic features of the sample description, i.e., the number of characters k mentioned above;
[0100] Obtain the term frequency (TF) and inverse document frequency (IDF) features of the characters represented in the semantic features of the sample description; the calculation process of the term frequency and inverse document frequency features is existing technology and will not be described here.
[0101] Determine the proportion of different types of characters in the semantic features of the sample description, that is, the proportion of numbers, Chinese characters, and English words in the sample description information;
[0102] The quantity, the word frequency feature, the inverse document frequency feature, and the proportion of different types of characters are used as decision features and input into the decision tree network for multi-level label prediction processing to obtain the sample-level expert prediction result P2 corresponding to the decision tree network.
[0103] Based on the term frequency features and the inverse document frequency features, the term frequency-inverse document frequency (TF-IDF) feature of each character can be determined. In one implementation, the average and total TF-IDF values, as well as the maximum and minimum values, of the characters represented in the sample description semantic features can be obtained as character frequency features. The number of characters, the character frequency features, and the proportion of the different types of characters are then input into the decision tree network as decision features for multi-level label prediction processing.
[0104] The expert prediction results for the sample level corresponding to this decision tree network can be represented as follows:
[0105] In this embodiment, statistical features of sample description information are extracted and a decision tree network is introduced for classification prediction. This provides multi-dimensional features for the subsequent comprehensive classification results of various expert classification networks, which is beneficial to obtaining more accurate comprehensive classification results.
[0106] like Figure 4As shown, the multiple expert classification networks also include at least one pre-processor neural network. Compared to other models, the pre-processor network can better capture coarse-grained sentence expression information. The sample-level expert prediction result corresponding to this pre-processor neural network can be expressed as:
[0107] P3 = FNN([T0, Ti…Tk])
[0108] P3 represents the sample-level expert prediction result output by the front-end neural network, Ti represents the character feature of the i-th character, and k is the number of characters.
[0109] In this application, CNN, LSTM, FNN and XGB networks are used to capture features of different dimensions of sample description information and perform classification prediction to obtain sample-level expert prediction results. This application further integrates the above results to combine the advantages of different expert classification networks in capturing features and uses multi-dimensional features to comprehensively predict classification results, thereby effectively improving the model's classification and recognition capabilities.
[0110] Figure 5 This is a schematic diagram illustrating the process of determining the sample network weights corresponding to the expert classification network, as provided in an embodiment of this application. Figure 5 As shown, determining the sample network weights corresponding to each expert classification network based on the differences between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels may include:
[0111] S501: Based on the difference between the expert prediction results of the target samples output by each expert classification network and the target classification label, determine the target sample classification accuracy of each expert classification network;
[0112] The target sample expert prediction result, the target classification label, and the target sample classification accuracy all correspond to the target category level in the preset category level.
[0113] The target category level is one of the preset category levels. The target sample classification accuracy of each expert classification network can be understood as the real-time loss value of each expert classification network for the target category level. For example, for a first-level category, assuming that the expert prediction result of the target sample output by expert classification network 1 is P0, then the target sample classification accuracy of expert classification network 1 can be the loss between p0 and the corresponding label of the first-level category; assuming that the expert prediction result of the target sample output by expert classification network 2 is P1, then the target sample classification accuracy of expert classification network 2 can be the loss between p1 and the corresponding label of the first-level category, and so on.
[0114] Specifically, the target sample classification accuracy of each expert classification network can be obtained by the following formula:
[0115]
[0116] Where p represents the target sample classification accuracy of the expert classification network, and y i Indicates the target category label, a i This represents the expert prediction results for the target sample in the expert classification network, where n represents the number of categories in the target category hierarchy.
[0117] S503: The target sample classification accuracy of each expert classification network is numerically summed to obtain the overall target sample classification accuracy.
[0118] S505: Determine the sample network weight of each expert classification network based on the proportion of the target sample classification accuracy of each expert classification network in the overall target sample classification accuracy.
[0119] For each target category level, the classification accuracy of the target samples can be determined through step S501 above, and the corresponding sample network weights can be obtained by the following formula:
[0120] p_i = pi / (p0 + ... + pm)
[0121] Where i is an integer from 0 to m, m is the number of expert classification networks, pi represents the target sample classification accuracy of the i-th expert classification network, and p_i represents the sample network weights corresponding to the i-th expert classification network.
[0122] It should be noted that the sample network weights of each expert classification network correspond to the target category level.
[0123] In this embodiment, the real-time classification accuracy is obtained by determining the difference between the output results of each expert classification network and the labels in real time. This accuracy is then used to determine the sample network weights, enabling the sample network weights to represent the classification accuracy of the expert classification networks in real time. This provides an accurate and reliable basis for the subsequent fusion of classification results. Compared to traditional methods that manually set weights, this method considers the accuracy of the output results during actual training, thus making it more accurate.
[0124] After obtaining the sample network weights of each expert classification network, this application further integrates the outputs of each expert classification network. Taking a preset category hierarchy including a first level and a second level as an example, the determination of the comprehensive classification result for the sample hierarchy based on the sample network weights corresponding to each expert classification network and the expert prediction results of each expert classification network at the sample level includes:
[0125] Obtain the first weight corresponding to the first layer and the second weight corresponding to the second layer in the sample network weights;
[0126] Obtain the first sample expert prediction result corresponding to the first level and the second sample expert prediction result corresponding to the second level from the expert prediction results of the sample level output by each expert classification network;
[0127] For the first level, the expert prediction results of the first sample output by each expert classification network are weighted and fused according to the first weight to obtain the comprehensive classification result of the first sample; the comprehensive classification result of the first sample can be obtained by the following formula:
[0128] d1 = sum(p_i*P i )
[0129] Where i is an integer from 0 to m, and m is the number of expert classification networks; p_i represents the sample network weight of the first level corresponding to the i-th expert classification network, i.e., the first weight, P i d1 represents the expert prediction result of the first sample corresponding to the first level output by the i-th expert classification network; d1 represents the comprehensive classification result of the first sample.
[0130] For the second level, the expert prediction results of the second sample output by each expert classification network are weighted and fused according to the second weight to obtain the comprehensive classification result of the second sample;
[0131] The calculation process for the comprehensive classification result of the second sample is similar to that of the comprehensive classification result of the first sample, and will not be repeated here.
[0132] The comprehensive classification results of the first sample and the comprehensive classification results of the second sample are determined as the comprehensive classification results of the sample hierarchy.
[0133] In this embodiment, by performing hierarchical differentiation and weighted fusion on the output of the expert classification network during fusion, the accuracy of the comprehensive classification result output by the model can be improved. On the other hand, the model does not need to consider features at different levels during feature extraction, thus greatly simplifying the model and improving classification efficiency.
[0134] Correspondingly, the multi-level classification labels include at least a first-level label and a second-level label; Figure 6 This is a schematic diagram illustrating the process of determining the target loss and training the model provided in an embodiment of this application. For example... Figure 6 As shown, training the multimedia resource classification model based on the differences between the sample hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchical associations among multiple sample comprehensive classification results, to obtain the target multimedia resource classification model, may include:
[0135] S601: Determine the first-level loss based on the difference between the comprehensive classification result of the first sample and the first-level label; and determine the second-level loss based on the difference between the comprehensive classification result of the second sample and the second-level label;
[0136] Both the first-level loss and the second-level loss can be calculated using the negative logarithmic loss function. The specific calculation process is as follows:
[0137]
[0138] Where, loss cls1 This represents the first-level loss, y. i Indicates the first-level tag, b i d1 represents the i-th item in the comprehensive classification result of the first sample, and n represents the number of categories in the first category level.
[0139] The calculation method for the second-level loss is similar to that for the first-level loss, and will not be repeated here.
[0140] S603: Determine the cascading loss based on the category hierarchy relationship between the comprehensive classification results of the first sample and the comprehensive classification results of the second sample;
[0141] When the item with the highest confidence in the comprehensive classification result of the first sample and the item with the highest confidence in the comprehensive classification result of the second sample do not have a membership relationship, i.e., there is no category hierarchical relationship, this application further introduces cascade loss. Cascade loss can be understood as a penalty coefficient, the value of which can be set according to actual training needs. For example, when there is no membership relationship, the penalty coefficient can be 1.1, while when there is a membership relationship, the penalty coefficient can be 1.
[0142] S605: Determine the target loss based on the first-level loss, the second-level loss, and the cascaded loss;
[0143] In this embodiment, the target loss is the sum of the first-level loss and the second-level loss multiplied by the cascaded loss. The formula for calculating the target loss is as follows:
[0144] Loss=(z)*loss_all=(z)*(loss cls1 +loss cls2 )
[0145] Where Loss represents the target loss, z represents the cascaded loss, and loss cls1 This represents the first-level loss. cls2 This indicates a second-level loss.
[0146] It should be understood that target losses can also include third-level losses, fourth-level losses, etc., which can be designed according to the actual preset category levels.
[0147] This application's embodiments utilize the correlation of hierarchical category systems and introduce the difference value between first-level and second-level classification results into the loss function to constrain the consistency of multi-level category results.
[0148] S607: Train the multiple expert classification networks in the multimedia resource classification model according to the target loss to obtain the target multimedia resource classification model.
[0149] Backpropagation is performed based on the target loss to update the network parameters of each expert classification network, thereby jointly training the aforementioned CNN, LSTM, FNN, and XGB networks until the training termination condition is met. The model at the end of training is then designated as the target multimedia resource classification model, and the sample network weights corresponding to each expert classification network are stored for subsequent applications. The training termination condition can be reaching a preset number of iterations or the target loss value being less than a preset loss threshold.
[0150] In this embodiment of the application, cascaded loss is further introduced when determining the loss. Cascaded loss is used to penalize the case where the hierarchical affiliation of the comprehensive classification results corresponding to different levels is inconsistent. This cascaded constraint on the model corrects the model and improves the accuracy of the model's hierarchical classification.
[0151] This application embodiment also provides a multimedia resource classification method, which includes: obtaining target description information corresponding to the target multimedia resource; inputting the target description information into a target multimedia resource classification model for multi-level label prediction processing to obtain a target-level comprehensive classification result; wherein, the target multimedia resource classification model is trained according to the training method of the multimedia resource classification model in the above embodiment.
[0152] It should be noted that the target multimedia resource classification model trained in the embodiments of this application is applicable to business scenarios that require extracting document categories, such as content classification in search systems and product title classification in e-commerce systems. It can be trained and applied according to actual needs.
[0153] Figure 7 This is a schematic diagram of the structure of the training device for the multimedia resource classification model provided in the embodiments of this application.
[0154] The multimedia resource classification model includes multiple expert classification networks. For example... Figure 7 As shown, the device 700 may include:
[0155] The sample acquisition module 701 is used to acquire sample description information corresponding to the sample multimedia resources; the sample description information is marked with multi-level classification tags of the preset category level corresponding to the sample multimedia resources.
[0156] The expert prediction module 702 is used to input the sample description information into the multimedia resource classification model, and use the multiple expert classification networks to perform multi-level label prediction processing on the sample description information to obtain the sample-level expert prediction results output by each expert classification network.
[0157] The network weight determination module 703 is used to determine the sample network weights corresponding to each expert classification network based on the differences between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels.
[0158] The comprehensive classification result determination module 704 is used to determine the comprehensive classification result at the sample level based on the sample network weights corresponding to each expert classification network and the sample level expert prediction results output by each expert classification network; the comprehensive classification result at the sample level includes the comprehensive classification results of samples that correspond one-to-one with the preset category level;
[0159] The model training module 705 is used to train the multimedia resource classification model based on the differences between the sample hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchical association relationship between multiple sample comprehensive classification results, to obtain the target multimedia resource classification model.
[0160] In some embodiments, the sample-level expert prediction results include sample expert prediction results that correspond one-to-one with the preset category levels; the network weight determination module may include:
[0161] The classification accuracy determination submodule is used to determine the target sample classification accuracy of each expert classification network based on the difference between the target sample expert prediction results output by each expert classification network and the target classification label; the target sample expert prediction results, the target classification label, and the target sample classification accuracy all correspond to the target category level in the preset category level;
[0162] The comprehensive classification accuracy determination submodule is used to sum the target sample classification accuracy of each expert classification network to obtain the comprehensive classification accuracy of the target sample.
[0163] The sample network weight determination submodule is used to determine the sample network weight of each expert classification network based on the proportion of the target sample classification accuracy of each expert classification network in the overall target sample classification accuracy.
[0164] In some embodiments, the preset category hierarchy includes a first level and a second level; the comprehensive classification result determination module may include:
[0165] The weight acquisition submodule is used to acquire the first weight corresponding to the first layer and the second weight corresponding to the second layer in the sample network weights.
[0166] The expert prediction result acquisition submodule is used to acquire the first sample expert prediction result corresponding to the first level and the second sample expert prediction result corresponding to the second level from the expert prediction results of the sample level output by each expert classification network.
[0167] The first classification result determination submodule is used to perform weighted fusion of the first sample expert prediction results output by each expert classification network according to the first weight for the first level, so as to obtain the first sample comprehensive classification result.
[0168] The second classification result determination submodule is used to perform weighted fusion of the second sample expert prediction results output by each expert classification network according to the second weight for the second level, so as to obtain the comprehensive classification result of the second sample.
[0169] The comprehensive classification result determination submodule is used to determine the comprehensive classification result of the first sample and the comprehensive classification result of the second sample as the comprehensive classification result of the sample level.
[0170] In some embodiments, the multi-level classification labels include at least a first-level label and a second-level label; the model training module may include:
[0171] The hierarchical loss determination submodule is used to determine a first hierarchical loss based on the difference between the comprehensive classification result of the first sample and the first hierarchical label; and to determine a second hierarchical loss based on the difference between the comprehensive classification result of the second sample and the second hierarchical label.
[0172] The cascade loss determination submodule is used to determine the cascade loss based on the category hierarchy relationship between the comprehensive classification results of the first sample and the comprehensive classification results of the second sample.
[0173] The target loss determination submodule is used to determine the target loss based on the first-level loss, the second-level loss, and the cascaded loss;
[0174] The training submodule is used to train the multiple expert classification networks in the multimedia resource classification model according to the target loss, so as to obtain the target multimedia resource classification model.
[0175] In some embodiments, the device 700 may further include:
[0176] The semantic feature extraction module is used to perform semantic feature extraction processing on the sample description information to obtain sample description semantic features;
[0177] The expert prediction module may include:
[0178] The feature prediction submodule is used to perform multi-level label prediction processing on the semantic features of the sample description using the multiple expert classification networks, and obtain the sample-level expert prediction results output by each expert classification network.
[0179] In some embodiments, the plurality of expert classification networks include at least one convolutional neural network, and the feature prediction submodule may include:
[0180] The kernel construction unit is used to determine the kernel size based on the dimension of the character feature corresponding to each character in the semantic features of the sample description and to construct the kernel according to the kernel size.
[0181] A convolutional unit is used to perform a sliding convolution operation on the semantic features describing the sample using the convolutional kernel to obtain a convolutional feature set.
[0182] A pooling unit is used to perform a pooling operation on each convolutional feature in the convolutional feature set to obtain a pooled feature set.
[0183] The dimensionality reduction unit is used to perform dimensionality reduction processing on each pooling feature in the pooling feature set to obtain the prediction feature set;
[0184] The prediction unit is used to perform multi-level label prediction processing on the prediction feature set to obtain the sample-level expert prediction results corresponding to the convolutional neural network.
[0185] In some embodiments, the plurality of expert classification networks include at least one long short-term memory network, and the feature prediction submodule may include:
[0186] The stacking layer number determination unit is used to determine the stacking layer number of the long short-term memory layer in the long short-term memory network based on the number of characters represented in the semantic features of the sample description.
[0187] The feature extraction unit is used to input the semantic features of the sample description into a long short-term memory layer with the number of stacked layers for feature extraction processing in sequence, and to determine the output of the last long short-term memory layer as the predicted feature set.
[0188] The prediction unit is used to perform multi-level label prediction processing on the prediction feature set to obtain the sample-level expert prediction results corresponding to the long short-term memory network.
[0189] In some embodiments, the plurality of expert classification networks includes at least one decision tree network, and the feature prediction submodule may include:
[0190] A character count acquisition unit is used to acquire the number of characters represented in the semantic features of the sample description;
[0191] The term frequency feature acquisition unit is used to acquire the term frequency features and inverse document frequency features of the characters represented in the semantic features of the sample description;
[0192] The character type proportion acquisition unit is used to determine the proportion of different types of characters in the characters represented by the semantic features of the sample description;
[0193] The prediction unit is used to input the quantity, the word frequency feature, the inverse document frequency feature, and the proportion of different types of characters as decision features into the decision tree network for multi-level label prediction processing, so as to obtain the sample-level expert prediction result corresponding to the decision tree network.
[0194] In some embodiments, the semantic feature extraction module may include:
[0195] The word segmentation submodule is used to segment the sample description information into words to obtain a character sequence.
[0196] The embedding encoding submodule is used to perform text embedding encoding on the character sequence to obtain the embedding feature corresponding to each character;
[0197] The semantic feature extraction submodule is used to extract semantic features from the embedded features corresponding to each character to obtain the character features corresponding to each character.
[0198] The feature concatenation submodule is used to concatenate the character features corresponding to each character to obtain the semantic features of the sample description.
[0199] Figure 8 This is a schematic diagram of the multimedia resource classification device provided in an embodiment of this application. Figure 8 As shown, the device 800 may include:
[0200] The acquisition module 801 is used to acquire the target description information corresponding to the target multimedia resource;
[0201] The classification module 802 is used to input the target description information into the target multimedia resource classification model for multi-level label prediction processing to obtain the target hierarchical comprehensive classification result; the target multimedia resource classification model is trained according to the training method of the multimedia resource classification model described above.
[0202] The apparatus and method embodiments described herein are based on the same inventive concept.
[0203] This application provides an electronic device including a processor and a memory. The memory stores at least one instruction or at least one program. The at least one instruction or at least one program is loaded and executed by the processor to implement the training method of the multimedia resource classification model or the multimedia resource classification method provided in the above method embodiments.
[0204] Embodiments of this application also provide a computer storage medium, which can be disposed in a terminal to store at least one instruction or at least one program related to the training method of the multimedia resource classification model in the method embodiment, or the multimedia resource classification method. The at least one instruction or at least one program is loaded and executed by the processor to implement the training method of the multimedia resource classification model or the multimedia resource classification method provided in the above method embodiment.
[0205] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a training method for the multimedia resource classification model provided in the above-described method embodiments, or a multimedia resource classification method.
[0206] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0207] The memory described in this application embodiment can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for the functions, etc.; the data storage area may store data created according to the use of the device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.
[0208] The training method for the multimedia resource classification model provided in this application can be executed on a mobile terminal, computer terminal, server, or similar computing device. Taking running on a server as an example... Figure 9 This is a hardware structure block diagram of a server for a training method of a multimedia resource classification model provided in an embodiment of this application. For example... Figure 9 As shown, server 900 can be considered as Figure 1 A specific example of the server shown can vary considerably due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 910 (CPUs 910 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 930 for storing data, and one or more storage media 920 (e.g., one or more mass storage devices) for storing application programs 923 or data 922. The memory 930 and storage media 920 may be temporary or persistent storage. The program stored in the storage media 920 may include one or more modules, each module including a series of instruction operations on the server. Furthermore, the CPU 910 may be configured to communicate with the storage media 920 and execute the series of instruction operations in the storage media 920 on the server 900. Server 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0209] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 900. In one example, the input / output interface 940 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 940 may be a radio frequency (RF) module for wireless communication with the Internet.
[0210] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 900 may also include... Figure 9 The more or fewer components shown, or having the same Figure 9 The different configurations shown.
[0211] As can be seen from the embodiments of the training method, classification method, apparatus, device, or storage medium of the multimedia resource classification model provided in this application, this application utilizes multiple expert classification networks in the multimedia resource classification model to perform multi-level label prediction processing on the sample description information, obtaining the sample-level expert prediction results output by each expert classification network. Each expert classification network can capture the features of the sample description information from different dimensions, thereby improving the accuracy of the classification results. Since the sample network weights change in real time with the expert prediction results during model training, the sample network weights can represent the classification accuracy of the expert classification networks in real time. Therefore, the fused sample-level comprehensive classification result can accurately integrate the outputs of multiple expert classification networks. Finally, based on the differences between the comprehensive classification results and the multi-level classification labels, as well as the category-level association relationships between the comprehensive classification results of multiple samples, the multimedia resource classification model is trained to obtain the target multimedia resource classification model. This model considers both the differences between the comprehensive classification results and the labels, and the consistency between the comprehensive classification results. On the one hand, it ensures the accuracy of the classification results; on the other hand, it eliminates the need to train the corresponding next-level classification model for each item in each category level, saving training and computational resources. Therefore, the method of this application can improve the accuracy and efficiency of multi-level classification of multimedia resources.
[0212] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0213] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0214] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer storage medium, such as a read-only memory, a disk, or an optical disk.
[0215] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A training method for a multimedia resource classification model, characterized in that, The multimedia resource classification model includes multiple expert classification networks, and the method includes: Obtain sample description information corresponding to the sample multimedia resources; the sample description information is labeled with multi-level classification tags of the preset category level corresponding to the sample multimedia resources; The sample description information is input into the multimedia resource classification model, and the sample description information is processed by multiple expert classification networks to perform multi-level label prediction, so as to obtain the sample-level expert prediction results output by each expert classification network. Based on the difference between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels, the sample network weights corresponding to each expert classification network are determined. Based on the sample network weights corresponding to each expert classification network and the sample-level expert prediction results output by each expert classification network, the comprehensive classification result at the sample level is determined; the comprehensive classification result at the sample level includes the comprehensive classification results corresponding one-to-one with the preset category level; The multimedia resource classification model is trained based on the differences between the sample hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchical associations among multiple sample comprehensive classification results, to obtain the target multimedia resource classification model.
2. The method according to claim 1, characterized in that, The sample-level expert prediction results include sample expert prediction results that correspond one-to-one with the preset category levels; the determination of the sample network weights corresponding to each expert classification network based on the differences between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels includes: Based on the difference between the expert prediction results of the target samples output by each expert classification network and the target classification label, the target sample classification accuracy of each expert classification network is determined; the expert prediction results of the target samples, the target classification label, and the target sample classification accuracy all correspond to the target category level in the preset category level; The overall classification accuracy of the target samples is obtained by summing the target sample classification accuracy of each expert classification network. The sample network weights of each expert classification network are determined based on the proportion of the target sample classification accuracy of each expert classification network in the overall target sample classification accuracy.
3. The method according to claim 1, characterized in that, The preset category hierarchy includes a first level and a second level; the determination of the comprehensive classification result of the sample hierarchy based on the sample network weights corresponding to each expert classification network and the expert prediction results of the sample hierarchy output by each expert classification network includes: Obtain the first weight corresponding to the first layer and the second weight corresponding to the second layer in the sample network weights; Obtain the first sample expert prediction result corresponding to the first level and the second sample expert prediction result corresponding to the second level from the expert prediction results of the sample level output by each expert classification network; For the first level, the first sample expert prediction results output by each expert classification network are weighted and fused according to the first weight to obtain the first sample comprehensive classification result; For the second level, the expert prediction results of the second sample output by each expert classification network are weighted and fused according to the second weight to obtain the comprehensive classification result of the second sample; The comprehensive classification results of the first sample and the comprehensive classification results of the second sample are determined as the comprehensive classification results of the sample hierarchy.
4. The method according to claim 3, characterized in that, The multi-level classification labels include at least a first-level label and a second-level label; the multimedia resource classification model is trained based on the differences between the sample hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchical associations among multiple sample comprehensive classification results, to obtain the target multimedia resource classification model, including: The first level loss is determined based on the difference between the comprehensive classification result of the first sample and the first level label; and the second level loss is determined based on the difference between the comprehensive classification result of the second sample and the second level label. The cascading loss is determined based on the category hierarchy relationship between the comprehensive classification results of the first sample and the comprehensive classification results of the second sample. The target loss is determined based on the first-level loss, the second-level loss, and the cascaded loss. The target multimedia resource classification model is obtained by training the multiple expert classification networks in the multimedia resource classification model based on the target loss.
5. The method according to claim 1, characterized in that, Before inputting the sample description information into the multimedia resource classification model, the method further includes: The sample description information is subjected to semantic feature extraction processing to obtain sample description semantic features; The process of using the multiple expert classification networks to perform multi-level label prediction on the sample description information, and obtaining the sample-level expert prediction results output by each expert classification network, includes: The sample description semantic features are processed by multiple expert classification networks to perform multi-level label prediction, and the sample-level expert prediction results output by each expert classification network are obtained.
6. The method according to claim 5, characterized in that, The plurality of expert classification networks includes at least one convolutional neural network. The process of using the plurality of expert classification networks to perform multi-level label prediction processing on the semantic features of the sample descriptions, and obtaining the sample-level expert prediction results output by each expert classification network, includes: Based on the dimension of the character feature corresponding to each character in the semantic features of the sample description, determine the convolution kernel size and construct the convolution kernel according to the convolution kernel size; The convolution kernel is used to perform a sliding convolution operation on the semantic features describing the sample to obtain a convolutional feature set. Perform pooling operations on each convolutional feature in the convolutional feature set to obtain a pooled feature set; Dimensionality reduction is performed on each pooling feature in the pooling feature set to obtain the prediction feature set; Multi-level label prediction processing is performed on the predicted feature set to obtain sample-level expert prediction results corresponding to the convolutional neural network.
7. The method according to claim 5, characterized in that, The plurality of expert classification networks includes at least one long short-term memory network. The process of using the plurality of expert classification networks to perform multi-level label prediction processing on the semantic features of the sample descriptions, respectively, yields the sample-level expert prediction results output by each expert classification network, including: The number of stacked layers of the long short-term memory (LSTM) in the LSM network is determined based on the number of characters represented in the semantic features of the sample description. The semantic features of the sample description are input into a long short-term memory layer with the number of stacked layers for feature extraction processing in sequence, and the output of the last long short-term memory layer is determined as the predicted feature set. Multi-level label prediction processing is performed on the predicted feature set to obtain the sample-level expert prediction results corresponding to the Long Short-Term Memory Network.
8. The method according to claim 5, characterized in that, The plurality of expert classification networks includes at least one decision tree network. The process of using the plurality of expert classification networks to perform multi-level label prediction processing on the semantic features of the sample descriptions, and obtaining the sample-level expert prediction results output by each expert classification network, includes: Obtain the number of characters represented in the semantic features of the sample description; Obtain the word frequency features and inverse document frequency features of the characters represented in the semantic features of the sample description; Determine the proportion of different types of characters among the characters represented in the semantic features of the sample description; The quantity, the word frequency feature, the inverse document frequency feature, and the proportion of different types of characters are used as decision features and input into the decision tree network for multi-level label prediction processing to obtain the sample-level expert prediction results corresponding to the decision tree network.
9. The method according to claim 5, characterized in that, The step of extracting semantic features from the sample description information to obtain sample description semantic features includes: The sample description information is segmented into words to obtain a character sequence; The character sequence is subjected to text embedding encoding to obtain the embedding feature corresponding to each character; Semantic features are extracted from the embedding features corresponding to each character to obtain the character features corresponding to each character; The character features corresponding to each character are concatenated to obtain the semantic features of the sample description.
10. A multimedia resource classification method, characterized in that, The method includes: Obtain the target description information corresponding to the target multimedia resource; The target description information is input into the target multimedia resource classification model for multi-level label prediction processing to obtain the target hierarchical comprehensive classification result; the target multimedia resource classification model is trained according to the training method of the multimedia resource classification model according to any one of claims 1-9.
11. A training device for a multimedia resource classification model, characterized in that, The multimedia resource classification model includes multiple expert classification networks, and the device includes: The sample acquisition module is used to acquire sample description information corresponding to sample multimedia resources; the sample description information is marked with multi-level classification tags of the preset category level corresponding to the sample multimedia resources; The expert prediction module is used to input the sample description information into the multimedia resource classification model, and use the multiple expert classification networks to perform multi-level label prediction processing on the sample description information to obtain the sample-level expert prediction results output by each expert classification network. The network weight determination module is used to determine the sample network weights corresponding to each expert classification network based on the differences between the sample-level expert prediction results output by each expert classification network and the multi-level classification labels. The comprehensive classification result determination module is used to determine the comprehensive classification result at the sample level based on the sample network weights corresponding to each expert classification network and the sample-level expert prediction results output by each expert classification network; the comprehensive classification result at the sample level includes the comprehensive classification results of samples that correspond one-to-one with the preset category level; The model training module is used to train the multimedia resource classification model based on the differences between the sample hierarchical comprehensive classification results and the multi-level classification labels, as well as the category hierarchical associations among multiple sample comprehensive classification results, to obtain the target multimedia resource classification model.
12. A multimedia resource classification device, characterized in that, The device includes: The acquisition module is used to acquire the target description information corresponding to the target multimedia resource; The classification module is used to input the target description information into the target multimedia resource classification model for multi-level label prediction processing to obtain the target hierarchical comprehensive classification result; the target multimedia resource classification model is trained by the training method of the multimedia resource classification model according to any one of claims 1-9.
13. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or at least one program being loaded and executed by the processor to implement the training method of the multimedia resource classification model as described in any one of claims 1-9, or the multimedia resource classification method as described in claim 10.
14. A computer storage medium, characterized in that, The computer storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the training method of the multimedia resource classification model as described in any one of claims 1-9, or the multimedia resource classification method as described in claim 10.
15. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the training method of the multimedia resource classification model according to any one of claims 1-9, or the multimedia resource classification method according to claim 10.