A data classification method, computer device, and readable storage medium
By fusing image features and text features in multimedia data classification and using data classification models to predict fusion features, the problem of low accuracy of multimedia data classification in the prior art is solved, and a more accurate and comprehensive multimedia data category reflection is achieved.
Patent Information
- Application Number
- CN202110011574.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-01-04
AI Technical Summary
The prior art relies on image information in multimedia data classification, resulting in inconsistent subjective judgment results and low accuracy.
By obtaining image data and text data in multimedia data, image features and text features are extracted, and feature fusion is performed, and the fusion features are predicted in combination with the data classification model to determine the category of multimedia data.
The accuracy of multimedia data classification is improved. By integrating image features and text features, taking into account the user's subjective emotions, it avoids inaccurate prediction results due to the differences in text features and multimedia data categories.
Smart Images

Figure CN113392236B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data classification method, a computer device, and a readable storage medium. Background Art
[0002] Multimedia data has been widely used in multiple industries. In some application scenarios, such as scenarios for classifying multimedia data, the prior art generally classifies multimedia data based on the images in the multimedia data to obtain the objective objects included in the multimedia data, and determines the category of the multimedia data based on the objective objects. However, when relying only on the information of the images themselves in the multimedia data to obtain the subjective classification results of the multimedia data, since subjective judgments are made for the same objective object, different judgment results may occur, resulting in low accuracy of data classification. Summary of the Invention
[0003] Embodiments of this application provide a data classification method, a computer device, and a readable storage medium, which can improve the accuracy of data classification.
[0004] Embodiments of this application on the one hand provide a data classification method, including:
[0005] Obtain image data and text data in the multimedia data;
[0006] Obtain the image features of the multimedia data according to the image data, obtain the text features of the multimedia data according to the text data, and perform feature fusion on the image features and the text features to obtain fusion features;
[0007] Based on a data classification model, predict the image features to obtain an object label, obtain at least two prediction labels associated with the object label and first probability values respectively corresponding to each prediction label, and based on the data classification model, predict the fusion features to obtain second probability values respectively corresponding to each prediction label; the at least two prediction labels include a prediction label p, and p is a positive integer;
[0008] Fuse the first probability value of the prediction label p and the second probability value of the prediction label p to obtain a third probability value of the prediction label p, until third probability values respectively corresponding to each prediction label are obtained, and based on the third probability values respectively corresponding to each prediction label and the object label, determine the media data category corresponding to the multimedia data.
[0009] Embodiments of this application on the one hand provide a data classification method, including:
[0010] Obtain sample image data and sample text data in the sample multimedia data, and obtain the sample label of the sample multimedia data;
[0011] Obtain the sample image features of the sample multimedia data according to the sample image data, obtain the sample text features of the sample multimedia data according to the sample text data, and perform feature fusion on the sample image features and the sample text features to obtain sample fusion features;
[0012] Based on the initial data classification model, predict the sample image features to obtain a sample object label, obtain at least two sample prediction labels associated with the sample object label and the first sample probability value corresponding to each sample prediction label, and based on the initial data classification model, predict the sample fusion features to obtain the second sample probability value corresponding to each sample prediction label; the at least two sample prediction labels include sample prediction label j, and j is a positive integer;
[0013] Fuse the first sample probability value of the sample prediction label j and the second sample probability value of the sample prediction label j to obtain the third sample probability value of the sample prediction label j until the third sample probability value corresponding to each sample prediction label is obtained. Determine the model output label corresponding to the sample multimedia data according to the third sample probability value corresponding to each sample prediction label and the sample object label; train the initial data classification model according to the loss function composed of the sample label and the model output label to obtain a data classification model.
[0014] An embodiment of the present application provides a data classification device on the one hand, including:
[0015] A data acquisition module for acquiring image data and text data in multimedia data;
[0016] A feature acquisition module for obtaining the image features of the multimedia data according to the image data, obtaining the text features of the multimedia data according to the text data, and performing feature fusion on the image features and the text features to obtain fusion features;
[0017] A probability prediction module for predicting the image features based on a data classification model to obtain an object label, obtaining at least two prediction labels associated with the object label and the first probability value corresponding to each prediction label, and predicting the fusion features based on the data classification model to obtain the second probability value corresponding to each prediction label; the at least two prediction labels include prediction label p, and p is a positive integer;
[0018] A category determination module for fusing the first probability value of the prediction label p and the second probability value of the prediction label p to obtain the third probability value of the prediction label p until the third probability value corresponding to each prediction label is obtained, and determining the media data category corresponding to the multimedia data based on the third probability value corresponding to each prediction label and the object label.
[0019] Optionally, the data acquisition module is configured to, if the multimedia data is video data, acquire at least two video frame images that make up the video data; acquire the image data from the at least two video frame images based on an image acquisition period; search for first text content associated with the multimedia data, and if the first text content is found, determine the first text content as the text data; if the first text content is not found, acquire voice data corresponding to the image data in the video data, perform voice conversion on the voice data to obtain second text content corresponding to the voice data, and determine the second text content as the text data.
[0020] Optionally, the feature acquisition module includes: a weight acquisition unit, a first feature determination unit, a second feature determination unit, and a feature fusion unit; the weight acquisition unit is configured to acquire a first weight matrix corresponding to the image feature and a second weight matrix corresponding to the text feature; the first feature determination unit is configured to perform a weighted operation on the image feature based on the first weight matrix to obtain an image weighted feature; the second feature determination unit is configured to perform a weighted operation on the text feature based on the second weight matrix to obtain a text weighted feature; the feature fusion unit is configured to perform feature splicing on the image weighted feature and the text weighted feature to obtain the fusion feature.
[0021] Optionally, the category determination module includes: a maximum probability determination unit and a label splicing unit; the maximum probability determination unit is configured to determine, among the at least two predicted labels, the predicted label with the largest third probability value as the target predicted label; the label splicing unit is configured to splice the target predicted label and the object label to obtain a media data label, and determine the data category corresponding to the media data label as the media data category of the multimedia data.
[0022] Optionally, the apparatus further includes: a data sending module, configured to, in response to an acquisition request for the multimedia data, acquire a media data acquisition label of a target user who sends the acquisition request; if the media data category matches the media data acquisition label, send the multimedia data to the target user; if the media data category does not match the media data acquisition label, send a media data exception message to the target user.
[0023] Optionally, the apparatus further includes: a data processing module, configured to acquire a label cluster to which the media data category belongs, and if the label cluster is a first label cluster, display the multimedia data on the home page of the application where the multimedia data is located; if the label cluster is a second label cluster, delete the multimedia data; the second label cluster includes labels that do not belong to the first label cluster.
[0024] One aspect of the embodiments of the present application provides a data classification device, including:
[0025] A sample data acquisition module, configured to acquire sample image data and sample text data in sample multimedia data, and acquire a sample label of the sample multimedia data;
[0026] A sample feature acquisition module, configured to acquire a sample image feature of the sample multimedia data according to the sample image data, acquire a sample text feature of the sample multimedia data according to the sample text data, and perform feature fusion on the sample image feature and the sample text feature to obtain a sample fusion feature;
[0027] A sample label determination module, configured to predict the sample image feature based on an initial data classification model to obtain a sample object label, acquire at least two sample prediction labels associated with the sample object label and a first sample probability value corresponding to each sample prediction label, and predict the sample fusion feature based on the initial data classification model to obtain a second sample probability value corresponding to each sample prediction label; the at least two sample prediction labels include a sample prediction label j, and j is a positive integer;
[0028] A label output module, configured to fuse the first sample probability value of the sample prediction label j with the second sample probability value of the sample prediction label j to obtain a third sample probability value of the sample prediction label j, until third sample probability values corresponding to each sample prediction label are obtained, and determine a model output label corresponding to the sample multimedia data according to the third sample probability values corresponding to each sample prediction label and the sample object label;
[0029] A model training module, configured to train the initial data classification model according to a loss function composed of the sample label and the model output label to obtain a data classification model.
[0030] Optionally, the sample label includes a reference sample label and a reference sample prediction label, and the loss function includes a first loss function and a second loss function; the model training module includes: a first training unit, a second training unit, and a model generation unit; the first training unit is configured to generate the first loss function according to the reference sample label and the sample object label; the second training unit is configured to splice the reference sample prediction label and the sample object label to generate a target sample label, and generate the second loss function based on the target sample label and the model output label; the model generation unit is configured to train the initial data classification model according to the first loss function and the second loss function to obtain the data classification model.
[0031] One aspect of the present application provides a computer device, including: a processor, a memory, and a network interface;
[0032] The above-mentioned processor is connected to a memory and a network interface. Among them, the network interface is used to provide data communication functions, the above-mentioned memory is used to store computer programs, and the above-mentioned processor is used to call the above-mentioned computer programs to execute the method in the above-mentioned one aspect in the embodiments of the present application.
[0033] On the one hand, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes a data classification method in the above-mentioned first aspect.
[0034] On the one hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative manners in one aspect of the embodiments of the present application.
[0035] In the embodiments of the present application, image data and text data in multimedia data are obtained; an image feature of the multimedia data is obtained according to the image data, a text feature of the multimedia data is obtained according to the text data, and the image feature and the text feature are feature - fused to obtain a fused feature; the image feature is predicted based on a data classification model to obtain an object label, at least two prediction labels associated with the object label, and a first probability value corresponding to each prediction label respectively, and the fused feature is predicted based on the data classification model to obtain a second probability value corresponding to each prediction label respectively. Since the image feature can reflect the image information in the multimedia data, the fused feature fuses the image feature and the text feature, and the text feature is a description of the image feature, the fused feature can reflect the user's subjective emotion towards the multimedia data. Moreover, by using the object label corresponding to the image feature to assist in judging the category of the multimedia data for the prediction label corresponding to the fused feature, it is possible to avoid inaccurate prediction results caused by too large a difference between the text feature and the category of the multimedia data (for example, the title of the multimedia data is exaggerated and does not match the video content). That is to say, by combining the image feature and the fused feature, when the prediction label corresponding to the fused feature is obtained, the object label corresponding to the image feature can be used for re - judgment. That is, the present application can add human emotion (because text features are generally added manually) through the fused feature of the text feature and the image feature, and predict the subjective emotion through the image feature to correct the subjective emotion corresponding to the fused feature, so that in the data classification of multimedia data, both human emotion can be considered and the content itself of the multimedia data (i.e., the image feature) will not be deviated from, thus achieving a more accurate and comprehensive reflection of the category of the multimedia data, and further improving the accuracy of data classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 is a schematic structural diagram of a data classification system provided by an embodiment of the present application;
[0038] Figure 2 is a schematic application scenario diagram of a data classification method provided by an embodiment of the present application;
[0039] Figure 3 is a schematic flowchart of a data classification method provided by an embodiment of the present application;
[0040] Figure 4It is a schematic diagram of obtaining image features and text features provided by an embodiment of the present application;
[0041] Figure 5 It is a schematic flowchart of a data classification method provided by an embodiment of the present application;
[0042] Figure 6 It is a schematic flowchart of a data classification method provided by an embodiment of the present application;
[0043] Figure 7 It is a schematic diagram of the composition structure of a data classification device provided by an embodiment of the present application;
[0044] Figure 8 It is a schematic diagram of the composition structure of a data classification device provided by an embodiment of the present application;
[0045] Figure 9 It is a schematic diagram of the composition structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0047] The technical solution of this application is applicable to scenarios of classifying multimedia data. For example, in scenarios of auditing multimedia data to determine whether to push the multimedia data to users. By obtaining image data and text data in the multimedia data; obtaining the image features of the multimedia data according to the image data, obtaining the text features of the multimedia data according to the text data, and performing feature fusion on the image features and the text features to obtain fusion features; predicting the image features based on a data classification model to obtain a first prediction label, obtaining at least two prediction labels associated with the object label and the first probability value corresponding to each prediction label respectively, predicting the fusion features based on the data classification model to obtain the second probability value corresponding to each prediction label respectively, where at least two prediction labels include prediction label p, p is a positive integer, fusing the first probability value of prediction label p and the second probability value of prediction label p to obtain the third probability value of prediction label p until the third probability value corresponding to each prediction label is obtained, and determining the media data category corresponding to the multimedia data based on the third probability value corresponding to each prediction label and the object label. Since the image features can reflect the image information in the multimedia data, and the text features can reflect the user's subjective emotion towards the multimedia data. Therefore, by using the object label corresponding to the image features to assist the prediction label corresponding to the fusion features to judge the category of the multimedia data, it is possible to avoid the situation where the text features are too different from the category of the multimedia data, and more accurately and comprehensively reflect the category of the multimedia data, thereby improving the accuracy of data classification.
[0048] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a data classification system provided by an embodiment of this application. As Figure 1 shown, the computer device can perform data interaction with the user terminal(s). The number of user terminals can be one or more. When the number of user terminals is multiple, the user terminals can include Figure 1 102a, 102b, and 102c in Figure 1 , and the computer device can be
[0049] Please refer to Figure 2 , Figure 2 which is a schematic diagram of an application scenario of a data classification method provided by an embodiment of the present application. As Figure 2 shown, after a computer device obtains image data and text data in multimedia data, first, the image data is input into an image feature extraction network to obtain image features; and the text data is input into a text feature extraction network to obtain text features. Secondly, the computer device inputs the extracted image features into a first image classifier in a data classification model for prediction to obtain an object label, and inputs the image features into a second image classifier in the data classification model for prediction to obtain at least two prediction labels associated with the object label and first probability values respectively corresponding to each prediction label. Among them, the first image classifier and the second image classifier may also belong to the same image classifier, which is not limited here. Further, the computer device can perform feature fusion on the image features and the text features to obtain fusion features, input the fusion features into a fusion classifier in the data classification model, and the fusion classifier predicts the fusion features to obtain second probability values respectively corresponding to each prediction label. The at least two prediction labels include a prediction label p; then, the computer device fuses the first probability value of the prediction label p and the second probability value of the prediction label p to obtain a third probability value of the prediction label p until third probability values respectively corresponding to each prediction label are obtained. Finally, the computer device determines the media data category corresponding to the multimedia data based on the third probability values respectively corresponding to each prediction label and the object label. For example, the computer device can determine the prediction label with the largest third probability value among the at least two prediction labels as the target prediction label, and then obtain a media data label according to the target prediction label and the object label, and determine the data category corresponding to the media data label as the media data category of the multimedia data. It can be understood that the above processes of processing the image data and the process of processing the text data can be carried out simultaneously, or the image data can be processed first and then the text data can be processed, which is not limited here.
[0050] It is understandable that the computer devices mentioned in the embodiments of the present application include, but are not limited to, terminal devices or servers. In other words, the computer device or user terminal can be a server or a terminal device, or a system composed of a server and a terminal device. Among them, the above-mentioned terminal device can be an electronic device, including but not limited to mobile phones, tablet computers, desktop computers, laptop computers, handheld computers, augmented reality / virtual reality (AR / VR) devices, head-mounted displays, wearable devices, smart speakers, digital cameras, cameras, and other mobile internet devices (MIDs) with network access capabilities. Among them, the client has a display function. Among them, the above-mentioned server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0051] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a data classification method provided by an embodiment of the present application. This method can be applied to a computer device, such as Figure 3 shown in the figure. The method includes:
[0052] S101, obtaining image data and text data in the multimedia data.
[0053] In the embodiments of the present application, the multimedia data may be video data or a single-frame picture, etc., which is not limited herein. Among them, if the multimedia data is video data, the video data includes at least two video frame images, and the image screen in each video frame image is the image data corresponding to the video frame image; if the multimedia data refers to a single-frame picture, the image screen in the single-frame picture is the image data corresponding to the single-frame picture. The text data may refer to the title of the multimedia data or the content introduction information of the multimedia data, etc. Optionally, the text data may also be the text information displayed in the video frame image. For example, when the multimedia data is video data, the text data may be the title corresponding to the video data, such as "The Origin of Dinosaurs in Man and Nature"; or, the text data may be the content introduction information corresponding to the multimedia data; or, the text data may be the text information in each video frame image of the multimedia data; or, the multimedia data may be the text data corresponding to the voice data in the multimedia data. Correspondingly, when the multimedia data is a single-frame picture, the text data may be the title corresponding to the single-frame picture, the text information included in the picture, etc.
[0054] In the embodiments of the present application, the manner in which the computer device obtains the image data and text data in the multimedia data may be: if the multimedia data is video data, obtain at least two video frame images that make up the video data; obtain the image data from at least two video frame images based on the image acquisition period; search for the first text content associated with the multimedia data, and if the first text content is found, determine the first text content as the text data; if the first text content is not found, obtain the voice data corresponding to the image data in the video data, perform voice conversion on the voice data to obtain the second text content corresponding to the voice data, and determine the second text content as the text data.
[0055] Specifically, the computer device may obtain the text display area of the image data, identify the text display area to obtain the text information corresponding to the image data, and use the text information as the text data in the multimedia data.
[0056] Among them, the image acquisition period can be a time acquisition period, a frame number acquisition period, etc., which is not limited here. For example, when the image acquisition period is a time acquisition period, the image acquisition period can be 0.1 second, 1 second, 2 seconds, etc.; when the image acquisition period is a frame number acquisition period, the image acquisition period can be "acquire a video frame image every n frames", where n is a positive integer and n is less than or equal to the number of video frame images included in at least two video frame images. The embodiments of the present application do not limit this. Specifically, taking the image acquisition period as 1 second as an example, the computer device can acquire image data from at least two video frame images of the video data every 1 second. Assuming that the total duration of the video data is 3 seconds, 3 image data are acquired. n can be determined according to the number of video frame images. For example, if the number of video frame images is 15 and 5 video frame images need to be acquired, then n is equal to 3, that is, the computer device acquires a video frame image from every 3 video frame images.
[0057] The computer device searches for the first text content associated with the multimedia data. The first text content can be the title of the multimedia data, the content summary of the multimedia data, or the text information included in the video frame images in the multimedia data, etc. If the first text content is found, the first text content is determined as the text data; for example, if the title of the multimedia data or the content summary of the multimedia data is found, etc., the title of the multimedia data or the content summary of the multimedia data is determined as the text data. If the first text content is not found, for example, the multimedia data does not contain a title or the title is a digital number, and the multimedia data does not contain a content summary, etc., the voice data corresponding to the image data in the video data is acquired. Here, the voice data can refer to the voice description of the user for the image data or the voice description of the user for the multimedia data. Perform voice conversion on the voice data. Specifically, voice conversion technology can be used to perform voice conversion on the voice data to obtain the second text content corresponding to the voice data, and the second text content is determined as the text data. Through the above method, the text data associated with each image data in the multimedia data can be acquired. Similarly, if multiple image data are acquired, the text data associated with each image data is determined according to the above method for acquiring text data.
[0058] The specific method for acquiring image features and text features from the multimedia data can be referred to Figure 4 , Figure 4It is a schematic diagram of obtaining image features and text features provided by an embodiment of the present application. Among them, the multimedia data is 30, and the multimedia data may refer to video data. The video data includes a title 3a. The computer device can obtain video frame images from the video data as image data. Specifically, the video frame images obtained by the computer device from the multimedia data may be as shown in 301, and the image screen in the video frame image 301 is used as the image data corresponding to the video frame. The computer device can obtain the title 3a in the multimedia data as text data, such as Figure 4 in Method ①; or, the computer device can obtain the text information 3b displayed in the video frame image as text data; or, the computer device can obtain the content summary in the video frame image as text data, such as Figure 4 in Method ②; or, the computer device can obtain the voice data 3c associated with the video frame image, perform voice conversion on the voice data 3c to obtain text content, and determine the text content as text data, such as Figure 4 in Method ③. After obtaining the image data, the computer device extracts features by inputting the image data into an image feature extraction network to obtain image features; and extracts features by inputting the text data into a text feature extraction network to obtain text features. Among them, Methods ①②③ are three different ways to obtain text data. In specific implementation, one of the methods can be used to obtain the text data associated with the image data, or at least two of the three methods can be combined to obtain the text data associated with the image data. In the embodiment of the present application, an example is given of obtaining an image data by obtaining a video frame image from video data, and obtaining image features and the text features corresponding to the image features according to the image data. The method of obtaining multiple video frame images and multiple video frame images to obtain image data, and obtaining multiple image features and the text features associated with each image feature can refer to the method of obtaining the image data, and will not be described in detail here.
[0059] In the embodiment of the present application, for example, during the process of auditing multimedia data, the computer device classifies the video data by obtaining the image data in the multimedia data (such as the image information in each frame of the image corresponding to the multimedia data) and obtaining the text data in the multimedia data (such as the title, content summary corresponding to the multimedia data, or the text data corresponding to the voice data), so as to determine whether to push the audited multimedia data to the user.
[0060] S102, obtain the image features of the multimedia data according to the image data, obtain the text features of the multimedia data according to the text data, and perform feature fusion on the image features and the text features to obtain fusion features.
[0061] Here, the computer device can extract features in the image data as image features. For example, the image features are used to reflect the image information in the image data, such as the target object included in the image data. Then, the image features can include the features of the target object. The computer device can extract features in the text data as text features. For example, the text features are used to reflect the text information in the text data. For example, the computer device can extract keyword information in the text data. The keyword information can include words representing human subjective emotions, such as "horrible", "nauseous", "dense", "lovely", "like", etc. The obtained keyword information is determined as the text feature. Further, feature fusion refers to the computer device fusing the image features and text features to obtain the fused features.
[0062] In specific implementation, the computer device can extract features from the image data through an image feature extraction network to obtain image features. The image feature extraction network can include, but is not limited to, Convolutional Neural Networks (CNN), Visual Geometry Group (VGG), or Residual Network (ResNet), etc. The computer device can extract features from the text data through a text feature extraction network to obtain text features. The text feature extraction network can refer to a Bidirectional Encoder Representations from Transformers (Bert) or other text feature extraction networks. The way for the computer device to fuse the image features and text features to obtain the fused features can be, for example, to fuse the image features and text features to obtain the fused features. Among them, the feature fusion can be to directly splice the image features and text features. For example, when the image features are composed of a 1*1024 matrix and the text features are composed of a 1*2048 matrix, the fused features obtained by splicing the features are a 1*(1024 + 2048) matrix. Or, the feature fusion can be to perform weighted splicing of the image features and text features, etc.
[0063] Specifically, the way for the computer device to perform weighted splicing of the image features and text features can be: the computer device obtains the first weight matrix corresponding to the image features and the second weight matrix corresponding to the text features; performs weighted operation on the image features based on the first weight matrix to obtain image weighted features; performs weighted operation on the text features based on the second weight matrix to obtain text weighted features; and performs feature splicing on the image weighted features and text weighted features to obtain the fused features.
[0064] Among them, the first weight matrix and the second weight matrix can be equal or not equal, and can be specifically set according to requirements. For example, when classifying multimedia data, if it is considered that the image features can more accurately reflect the category information of the multimedia data, the first weight matrix can be greater than the second weight matrix. If it is considered that the text features can more accurately reflect the category information of the multimedia data, the first weight matrix can be less than the second weight matrix. If it is considered that both the image features and the text features can accurately reflect the category information of the multimedia data, the first weight matrix can be equal to the second weight matrix. For example, the first weight matrix is A1, the second weight matrix is A2, the image feature is B1, and the text feature is B2. Then the weighted image feature obtained by weighted operation is A1*B1, and the weighted text feature obtained by weighted operation is A2*B2. The image weighted feature and the text weighted feature are concatenated to obtain a fusion feature of A1*B1 + A2*B2, where * is the matrix dot product algorithm. That is to say, A1*B1 means multiplying two matrices, and the result is a matrix. For example, if A1 is a 2*2 matrix and B1 is a 2*2 matrix, the weighted image feature A1*B1 obtained by weighted operation is a 2*2 matrix.
[0065] The computer device performs weighted operation on the image features by obtaining the first weight matrix corresponding to the image features, and performs weighted operation on the text features by obtaining the second weight matrix corresponding to the text features, and splices the weighted operation results of the image features and the weighted operation results of the text features to obtain a fusion feature. Since the weight matrix can reflect the classification result of the multimedia data, using the weight matrix to perform weighted operation on the image features and the text features can improve the accuracy of multimedia data classification.
[0066] S103. Based on the data classification model, predict the image features to obtain an object label, obtain at least two predicted labels associated with the object label and the first probability value corresponding to each predicted label respectively, and based on the data classification model, predict the fusion feature to obtain the second probability value corresponding to each predicted label respectively.
[0067] Here, the object label refers to the label obtained by predicting the image features through a data classification model. The object label can be used to indicate the category to which the target object in the image represented by the image features belongs. For example, it can indicate that the target object in the image is a dinosaur, a lizard, a frog, or other species categories. Or, the object label can also be used to indicate whether the image contains a dinosaur, a lizard, a frog, or other species categories, so as to determine the species category corresponding to the image features. The at least two prediction labels associated with the object label refer to the labels obtained by predicting the image features through a data classification model, or the labels obtained by predicting the fused features through a data classification model. The prediction labels can be used to indicate the subjective emotion categories of the multimedia data. For example, they can include categories such as thrilling, scary, dense, cute, silly, and like.
[0068] Specifically, the computer device can perform object recognition on the image features through a data classification model to obtain an object label, and predict the object label to obtain the probabilities of the image features being each of the prediction labels included in the data classification model. Denote this probability as the first probability value, that is, obtain the first probability values corresponding to at least two prediction labels respectively; or, the computer device can directly perform classification processing on the image features to obtain the first probability values of the image features being each of the prediction labels. Further, the computer device can also perform classification processing on the fused features through a data classification model to obtain the probabilities of the fused features being each of the prediction labels in the data classification model, that is, obtain the second probability values corresponding to each prediction label respectively. For example, the at least two prediction labels include thrilling, scary, dense, cute, silly, like, etc. Based on the above process, the first probability values and the second probability values corresponding to each prediction label can be obtained, such as the first probability value and the second probability value corresponding to thrilling, etc. Among them, the at least two prediction labels include a prediction label p, where p is a positive integer, that is, the prediction label p can refer to any one of the at least two prediction labels.
[0069] Optionally, the data classification model includes a first image classifier, a second image classifier, and a fusion classifier. The first image classifier is used to predict the object label from the image features, that is, to determine which one or more of the object categories in the image are dinosaurs, lizards, frogs, or other species. The second image classifier is used to predict at least two prediction labels associated with the object label and the first probability value corresponding to each prediction label from the image features, that is, to obtain the probability of each category in subjective emotion categories such as thrilling, scary, dense, cute, etc. For example, the first probability value corresponding to the prediction label "thrilling" is 0.4, the first probability value corresponding to the prediction label "scary" is 0.3, the first probability value corresponding to the prediction label "dense" is 0.2, the first probability value corresponding to the prediction label "cute" is 0.1, and so on. The fusion classifier is used to predict the second probability value corresponding to each prediction label from the fusion features, that is, to obtain the probability of each category in subjective emotion categories such as thrilling, scary, dense, cute, etc. For example, the second probability value of the prediction label "thrilling" is 0.2, the second probability value of the prediction label "scary" is 0.3, the second probability value of the prediction label "dense" is 0.4, the second probability value of the prediction label "cute" is 0.1, and so on. It can be understood that each prediction label corresponds to a first probability value and a second probability value respectively.
[0070] In a specific implementation, a computer device can input image features into a data classification model. The first image classifier in the data classification model can predict the image features and output the probabilities of the image features being various object labels in the data classification model. Among them, the object label can be a species category, such as dinosaur, lizard, gecko, etc. For example, the probability that the image features correspond to the object label "dinosaur" is 0.5, the probability that the image features correspond to the object label "lizard" is 0.35, the probability that the image features correspond to the object label "gecko" is 0.15, etc. The label with a probability greater than the image threshold can be determined as the object label corresponding to the image features. For example, if the image threshold is 0.5, then the dinosaur is determined as the object label. If the image threshold is 0.3, the object labels can include dinosaurs and lizards, that is, the image data includes multiple objects. The computer device can input the image features into the data classification model, and the second image classifier in the data classification model can classify the fusion features, and output the probabilities of the image features being various subjective emotion categories, obtaining at least two prediction labels associated with the object label and the first probability value corresponding to each prediction label respectively. For example, at least two prediction labels associated with the object label include thriller, dense, and cute. The computer device predicts the image features through the first image classifier, and obtains that the first probability value of the prediction label being thriller is 0.4, the first probability value of the prediction label being dense is 0.5, and the first probability value of the prediction label being cute is 0.1. The computer device can input the fusion features into the data classification model, and the fusion classifier in the data classification model can classify the fusion features, and output the probabilities of the fusion features being various subjective emotion categories, obtaining the second probability value corresponding to each prediction label respectively. For example, the computer device predicts the fusion features through the fusion classifier, and obtains that the second probability value of the prediction label being thriller is 0.6, the second probability value of the prediction label being dense is 0.35, and the second probability value of the prediction label being cute is 0.05. Thus, the computer device can obtain the object label, at least two prediction labels associated with the object label, the first probability value corresponding to each prediction label respectively, and the second probability value corresponding to each prediction label respectively. For example, if the prediction label p is dense, the first probability value corresponding to dense is 0.5, and the second probability value corresponding to dense is 0.35.
[0071] S104. Fuse the first probability value of the prediction label p and the second probability value of the prediction label p to obtain the third probability value of the prediction label p, until the third probability value corresponding to each prediction label is obtained. Based on the third probability value corresponding to each prediction label and the object label, determine the media data category corresponding to the multimedia data.
[0072] After obtaining the first probability value of the prediction label p and the second probability value of the prediction label p, the computer device fuses the first probability value of the prediction label p and the second probability value of the prediction label p to obtain the third probability value of the prediction label p until the third probability value corresponding to each prediction label is obtained; the computer device then determines the media data category corresponding to the multimedia data based on the third probability value corresponding to each prediction label and the object label. Among them, fusing the first probability value of the prediction label p and the second probability value of the prediction label p may refer to multiplying the first probability value of the prediction label p and the second probability value of the prediction label p to obtain the third probability value of the prediction label p. For example, if the prediction label p is "dense", the first probability value corresponding to "dense" is 0.5, and the second probability value corresponding to "dense" is 0.35, then the third probability value of the prediction label p is 0.5 * 0.35 = 0.175. Since the prediction label p is any one of at least two prediction labels, by splicing the first probability value of each prediction label and the second probability value of each prediction label among at least two prediction labels, the third probability value of each prediction label among at least two prediction labels can be obtained.
[0073] It can be understood that for each of at least two prediction labels, the computer device multiplies the first probability value and the second probability value of each prediction label bit by bit to obtain the third probability value of each prediction label among at least two prediction labels. Bit-by-bit multiplication means multiplying the first probability value and the second probability value of each prediction label. For example, if the at least two prediction labels include three labels: "thriller", "dense", and "lovely", the first probability value and the second probability value of "thriller" are 0.4 and 0.6 respectively, the first probability value and the second probability value of "dense" are 0.5 and 0.35 respectively, and the first probability value and the second probability value of "lovely" are 0.1 and 0.05 respectively, then the calculated third probability value of "thriller" is 0.4 * 0.6 = 0.24, the third probability value of "dense" is 0.5 * 0.35 = 0.175, and the third probability value of "lovely" is 0.1 * 0.05 = 0.005.
[0074] Optionally, the method for the computer device to determine the media data category corresponding to the multimedia data may be: determining the prediction label with the largest third probability value among at least two prediction labels as the target prediction label; splicing the target prediction label and the object label to obtain the media data label, and determining the data category corresponding to the media data label as the media data category of the multimedia data.
[0075] For example, the object label is "dinosaur", and at least two predicted labels include "thrilling", "dense", and "lovely". The computer device obtains the third probability value of the predicted label "thrilling" as 0.24 by concatenating the first probability value of the predicted label "thrilling" and the second probability value of the predicted label "thrilling", obtains the third probability value of the predicted label "dense" as 0.175 by concatenating the first probability value of the predicted label "dense" and the second probability value of the predicted label "dense", and obtains the third probability value of the predicted label "lovely" as 0.005 by concatenating the first probability value of the predicted label "lovely" and the second probability value of the predicted label "lovely". The computer device determines the predicted label with the largest third probability value, that is, determines the predicted label "thrilling" as the target predicted label; concatenates the target predicted label and the object label to obtain the media data label, that is, "dinosaur + horror", then the computer device determines the data category corresponding to "dinosaur + horror" as the media data category of the multimedia data, and the obtained media data category of the multimedia data is "dinosaur + horror".
[0076] Optionally, the computer device can also determine the predicted label with the third probability value greater than the category threshold as the target predicted label, concatenate the target predicted label and the object label to obtain the media data label, and determine the data category corresponding to the media data label as the media data category of the multimedia data. For example, if the category threshold is 0.15, then the third probability values of the predicted labels "thrilling" and "dense" are both greater than the category threshold, so the predicted labels "thrilling" and "dense" are determined as the target predicted labels. After concatenating the target predicted labels and the object label, the obtained media data label is "dinosaur + horror + dense", and the determined media data category of the multimedia data is "dinosaur + horror + dense".
[0077] Optionally, taking the example that the multimedia data only includes one image data and one text data, the number of object labels predicted by the computer device is one. The number of predicted labels can be determined according to the number of subjective emotion categories in the label library. For example, it is n, where n is a positive integer. By determining the target predicted label with the third largest probability value from the n predicted labels, the object label and the target predicted label are determined as the media data category corresponding to the multimedia data. If the multimedia data contains m image data and m text data, where m is a positive integer and one image data corresponds to one text data, then the target predicted label corresponding to each of the m object labels is determined, and the m object labels and the target predicted labels corresponding to each of the m object labels are determined as the media data category corresponding to the multimedia data. Alternatively, the m predicted labels can be matched with the key label library. If the key label library contains a reference label that matches any one of the m predicted labels, then the predicted label corresponding to the any one label and the reference label are determined as the media data category corresponding to the multimedia data. Among them, the key label library can include labels used to represent an atmosphere that makes people feel more depressed, worried, etc., which is relatively heavy. For example, the key label library can include labels such as thriller, fear, dense, horror, and depression.
[0078] It should be noted that the computer device can pre-train the data classification model, that is, train the first image classifier, the second image classifier, and the fusion classifier in the data classification model respectively. It can be understood that the three classifiers can be trained simultaneously, or the first image classifier and the second image classifier can be trained first and then the fusion classifier can be trained, etc., to obtain the trained data classification model. By using the trained data classification model to predict the image features and the fusion features, the obtained object labels, at least two predicted labels, and the first probability value and the second probability value respectively corresponding to each predicted label can more accurately reflect the category of the multimedia data. Specifically, the process of training the first image classifier, the second image classifier, and the fusion classifier in the data classification model can refer to Figure 5 the description in the corresponding embodiment, and no more description will be given here.
[0079] In the embodiments of the present application, image data and text data in multimedia data are obtained; image features of the multimedia data are obtained according to the image data, text features of the multimedia data are obtained according to the text data, and the image features and the text features are feature-fused to obtain fused features; the image features are predicted based on a data classification model to obtain an object label, at least two prediction labels associated with the object label, and first probability values respectively corresponding to each prediction label, and the fused features are predicted based on the data classification model to obtain second probability values respectively corresponding to each prediction label. Since the image features can reflect the image information in the multimedia data, the fused features fuse the image features and the text features, and the text features are descriptions of the image features, the fused features can reflect the subjective emotion of the user towards the multimedia data. Moreover, by using the object label corresponding to the image features to assist the prediction label corresponding to the fused features to determine the category of the multimedia data, it is possible to avoid inaccurate prediction results caused by too large a difference between the text features and the category of the multimedia data (for example, the title of the multimedia data is too exaggerated and does not match the video content). That is to say, by combining the image features and the fused features, when the prediction label corresponding to the fused features is obtained, the object label corresponding to the image features can be used for re-judgment. That is, the present application can add human emotion (because the text features are generally added manually) through the fused features of the text features and the image features, and predict the subjective emotion through the image features to correct the subjective emotion corresponding to the fused features, so that in the data classification of the multimedia data, both human emotion can be considered and the content itself of the multimedia data (i.e., the image features) will not be separated, thereby achieving a more accurate and comprehensive reflection of the category of the multimedia data, and further improving the accuracy of data classification.
[0080] Optionally, in order to improve the accuracy of the data classification model in predicting the image features and the fused features to obtain the object label, at least two prediction labels, and the first probability values and the second probability values respectively corresponding to each prediction label, so as to improve the accuracy of classifying the multimedia data, before using the data classification model to predict the image features and the text features, a large number of sample image features and fused features can be used to train and adjust the data classification model, so that the trained data classification model can more accurately predict the image features and the fused features, thereby improving the accuracy of data classification. For the specific method of training the data classification model, please refer to Figure 5 , Figure 5 is a schematic flowchart of a data classification method provided by an embodiment of the present application. This method can be applied to a computer device, such as Figure 5 shown, and this method includes:
[0081] S201. Obtain the sample image data and sample text data in the sample multimedia data, and obtain the sample label of the sample multimedia data.
[0082] Here, the sample multimedia data is the multimedia data prepared for training the data classification model. The sample multimedia data can be video data or a single-frame picture, etc. That is to say, the sample multimedia data is of the same type as the multimedia data. If the multimedia data is video data, then the sample multimedia data is also video data; if the multimedia data is a single-frame picture, then the sample multimedia data is also a single-frame picture. The specific method for obtaining the sample image data and sample text data in the sample multimedia data can refer to the method for obtaining the image data and text data in the multimedia data in step S101, which will not be elaborated here. The sample label of the sample multimedia data refers to the pre-set label. The purpose of training the data classification model is to make the predicted label obtained by using the model as similar as possible to the pre-set label, and the higher the accuracy of the corresponding model.
[0083] S202. Obtain the sample image feature of the sample multimedia data according to the sample image data, obtain the sample text feature of the sample multimedia data according to the sample text data, and perform feature fusion on the sample image feature and the sample text feature to obtain the sample fusion feature.
[0084] Here, the specific method for obtaining the sample image feature and sample text feature of the sample multimedia data, and for performing feature fusion on the sample image feature and sample text feature to obtain the sample fusion feature can refer to the method for obtaining the image feature and text feature of the multimedia data, and for performing feature fusion on the image feature and text feature to obtain the fusion feature in step S102, which will not be elaborated here. The way for the computer device to perform feature fusion on the sample image feature and the sample text feature to obtain the fusion feature can be, for example, to perform feature fusion on the sample image feature and the sample text feature to obtain the sample fusion feature. Among them, the sample feature fusion can be to directly splice the sample image feature and the sample text feature. For example, if the sample image feature is composed of a 1*1024 matrix and the sample text feature is composed of a 1*2048 matrix, then the sample fusion feature obtained by splicing the features is a 1*(1024 + 2048) matrix. Or, the feature fusion can be to perform weighted splicing on the sample image feature and the sample text feature, etc. The specific way of feature fusion can refer to the description in step S102.
[0085] S203. Predict the sample image features based on the initial data classification model to obtain the sample object label, acquire at least two sample prediction labels associated with the sample object label and the first sample probability value corresponding to each sample prediction label, and predict the sample fusion features based on the initial data classification model to obtain the second sample probability value corresponding to each sample prediction label respectively.
[0086] Here, the sample object label refers to the label obtained by predicting the sample image features through the initial data classification model. The sample object label can be used to indicate the category to which the target object in the image represented by the initial image features belongs. For example, it can indicate that the target object in the sample image is a dinosaur, a lizard, a frog, or other species categories. Or, the sample object label can also be used to indicate whether the sample image contains a dinosaur, a lizard, a frog, or other species categories, so as to determine the species category corresponding to the sample image features. The sample prediction label refers to the label obtained by predicting the sample fusion features through the initial data classification model. The sample prediction label can be used to indicate the subjective emotion category of the sample multimedia data. For example, it can include categories such as thrilling, scary, dense, cute, silly, and like.
[0087] Specifically, the computer device can perform object recognition on the sample image features through the initial data classification model to obtain the sample object label, predict the sample object label to obtain the probability that the sample image features are each sample prediction label included in the initial data classification model, and record this probability as the first sample probability value, that is, obtain the first sample probability values corresponding to at least two sample prediction labels respectively; or, the computer device can directly perform classification processing on the sample image features to obtain the first sample probability values of the sample image features for each sample prediction label. Further, the computer device can also perform classification processing on the sample fusion features through the initial data classification model to obtain the probability that the sample fusion features are each sample prediction label in the initial data classification model, that is, obtain the second sample probability values corresponding to each sample prediction label respectively. For example, at least two sample prediction labels include thrilling, scary, dense, cute, silly, and like. Based on the above process, the first sample probability value and the second sample probability value corresponding to each sample prediction label can be obtained, such as the first sample probability value and the second sample probability value corresponding to thrilling. Among them, at least two sample prediction labels include sample prediction label j, where j is a positive integer, and sample prediction label j refers to any one of at least two sample prediction labels. Sample prediction label j corresponds to the first sample probability value and the second sample probability value. The number and type of sample prediction labels are equal to the number and type of prediction labels. That is, if the prediction labels include thrilling, scary, dense, cute, silly, and like, then the sample prediction labels include thrilling, scary, dense, cute, silly, and like.
[0088] Optionally, the initial data classification model includes a first initial image classifier, a second initial image classifier, and an initial fusion classifier. The first initial image classifier is used to predict the sample image features to obtain sample object labels, that is, to obtain which one or several of the categories of dinosaurs, lizards, frogs, or other species are in the sample image. The second initial image classifier is used to predict the sample image features to obtain at least two sample prediction labels associated with the sample object labels and the first sample probability value corresponding to each sample prediction label, that is, to obtain the probability of each category in the subjective emotion categories such as thriller, fear, density, cuteness, etc. For example, the first sample probability value corresponding to the sample prediction label "thriller" is 0.4, the first sample probability value corresponding to the sample prediction label "fear" is 0.3, the first sample probability value corresponding to the sample prediction label "density" is 0.2, and the first sample probability value corresponding to the sample prediction label "cuteness" is 0.1, and so on. The initial fusion classifier is used to predict the sample fusion features to obtain the second sample probability value corresponding to each sample prediction label, that is, to obtain the probability of each category in the subjective emotion categories such as thriller, fear, density, cuteness, etc. For example, the second sample probability value of the sample prediction label "thriller" is 0.2, the second sample probability value of the sample prediction label "fear" is 0.3, the second sample probability value of the sample prediction label "density" is 0.4, and the second sample probability value of the sample prediction label "cuteness" is 0.1, and so on. It can be understood that each sample prediction label corresponds to a first sample probability value and a second sample probability value respectively.
[0089] In a specific implementation, the computer device inputs the sample image features into the initial data classification model. By using the initial data classification model to predict the sample image features, it can output the probabilities of the sample image features being various object labels in the initial data classification model. The computer device can determine the sample object prediction labels as all the labels with probabilities greater than the image threshold. For example, the probability of the sample image features corresponding to the sample object label "dinosaur" is 0.5, the probability of the sample image features corresponding to the sample object label "lizard" is 0.35, the probability of the sample image features corresponding to the sample object label "gecko" is 0.15, and so on. If the image threshold is 0.5, then "dinosaur" is determined as the sample object label. If the image threshold is 0.3, then the sample object labels can include "dinosaur" and "lizard", that is, the sample image data includes multiple sample objects. The computer device inputs the sample image features into the initial data classification model. By using the first initial image classifier in the initial data classification model to classify the sample fusion features, it can output the probabilities of the sample image features being various subjective emotion categories, and obtain at least two sample prediction labels associated with the sample object label and the first sample probability value corresponding to each sample prediction label. By inputting the sample fusion features into the initial data classification model and using the initial fusion classifier in the initial data classification model to classify the sample fusion features, it can output the probabilities of the sample fusion features being various subjective emotion categories, and obtain the second probability value corresponding to each sample prediction label. Thus, the computer device can obtain the sample object label, at least two sample prediction labels associated with the sample object label, the first sample probability value corresponding to each sample prediction label, and the second sample probability value corresponding to each sample prediction label. For example, if the prediction label j is "dense", the first sample probability value corresponding to "dense" is 0.2, and the second sample probability value corresponding to "dense" is 0.4.
[0090] S204. Fuse the first sample probability value of the sample prediction label j and the second sample probability value of the sample prediction label j to obtain the third sample probability value of the sample prediction label j, until the third sample probability value corresponding to each sample prediction label is obtained. Determine the model output label corresponding to the sample multimedia data according to the third sample probability value corresponding to each sample prediction label and the sample object label.
[0091] After obtaining the first sample probability value of the sample prediction label j and the second sample probability value of the sample prediction label j, the computer device fuses the first sample probability value of the sample prediction label j and the second sample probability value of the sample prediction label j to obtain the third sample probability value of the sample prediction label j until the third sample probability value corresponding to each sample prediction label is obtained; according to the third sample probability value corresponding to each sample prediction label and the sample object label, the model output label corresponding to the sample multimedia data is determined. Among them, fusing the first sample probability value of the sample prediction label j and the second sample probability value of the sample prediction label j may mean multiplying the first sample probability value of the sample prediction label j and the second sample probability value to obtain the third sample probability value of the sample prediction label j. Since the sample prediction label j is any one of at least two sample prediction labels, by splicing the first sample probability value of each sample prediction label in at least two sample prediction labels and the second sample probability value of each sample prediction label, the third sample probability value of each sample prediction label in at least two sample prediction labels can be obtained. The computer device determines the model output label corresponding to the sample multimedia data according to the third sample probability value corresponding to each sample prediction label and the sample object label. Among them, the model output label refers to the label composed of the sample prediction label with the largest third probability value among at least two sample prediction labels and the sample object label. For example, the computer device splices the sample prediction label with the largest third sample probability value and the sample object label to obtain the model output label corresponding to the sample multimedia data. For example, the sample prediction label with the largest third sample probability value is lizard, and the sample object label is dense, then the model output label is lizard + dense.
[0092] S205. Train the initial data classification model according to the loss function composed of the sample label and the model output label to obtain a data classification model.
[0093] Optionally, the sample label includes a reference sample label and a reference sample prediction label, and the loss function includes a first loss function and a second loss function. The specific method for training the initial data classification model to obtain a data classification model can be:
[0094] First, generate a first loss function according to the reference sample label and the sample object label.
[0095] Specifically, the computer device generates a first loss function according to the reference sample label and the sample object label, and trains the first initial image classifier according to the first loss function to obtain a first image classifier. The first loss function can be shown as formula (1-1):
[0096]
[0097] Among them, Limages refers to the loss value corresponding to the first initial image classifier, k refers to the total number of categories of the reference sample labels, and y i refers to the i-th element that makes up the reference sample label, and σ(L i ) is the predicted label of the reference sample. For example, if the i-th category element corresponding to this first initial image classifier is a crocodile, and if the reference sample label is a crocodile, then the corresponding y i = 1; if the reference sample label is not a crocodile, then y i = 0. Through this first loss function, the loss value corresponding to the first initial image classifier can be calculated. When the loss value corresponding to the first initial image classifier is greater than the first loss threshold, the parameters in the first initial image classifier are adjusted to implement the training of the first initial image classifier; when the loss value corresponding to the first initial image classifier is less than or equal to the first loss threshold, the trained first initial image classifier is saved to obtain the first image classifier. In a specific implementation, the loss value corresponding to the first initial image classifier can be determined based on the gradient descent method.
[0098] Secondly, the predicted label of the reference sample and the label of the sample object are concatenated to generate a target sample label, and a second loss function is generated based on the target sample label and the model output label.
[0099] Specifically, the computer device concatenates the predicted label of the reference sample and the label of the sample object to generate a target sample label, generates a second loss function based on the target sample label and the model output label, and trains the initial fusion classifier according to the second loss function to obtain the fusion classifier. The second loss function can be as shown in formula (1-2):
[0100]
[0101] where L fusion refers to the loss value corresponding to the initial fusion classifier, k refers to the total number of categories of the predicted labels of the reference samples, and y i refers to the i-th element that makes up the predicted label of the reference sample, and S(L i ) is the predicted label of the reference sample. Through this second loss function, the loss value corresponding to the initial fusion classifier can be calculated. When the loss value corresponding to the initial fusion classifier is greater than the second loss threshold, the parameters in the initial fusion classifier are adjusted to implement the training of the initial fusion classifier; when the loss value corresponding to the initial fusion classifier is less than or equal to the second loss threshold, the trained initial fusion classifier is saved to obtain the fusion classifier. In a specific implementation, the loss value corresponding to the initial fusion classifier can be determined based on the gradient descent method.
[0102] Finally, the initial data classification model is trained according to the first loss function and the second loss function to obtain the data classification model.
[0103] Here, the data classification model may include a first image classifier and a fusion classifier. That is, after saving the trained first image classifier and fusion classifier, the data classification model is obtained. Optionally, the second initial image classifier may also be trained, and when the loss value corresponding to the second initial image classifier is less than the third loss threshold, the trained second initial image classifier is saved to obtain the second image classifier, so as to obtain the data classification model according to the first image classifier, the second initial image classifier, and the fusion classifier.
[0104] In the embodiment of the present application, a large number of sample image data and sample text data are used to train the initial image classifier and the initial fusion classifier in the initial data classification model, obtain the loss value corresponding to the initial image classifier and the loss value corresponding to the initial fusion classifier, and train the initial image classifier according to the loss value corresponding to the initial image classifier, and train the initial fusion classifier according to the loss value corresponding to the initial fusion classifier. When the loss value is large, the initial image classifier and the initial fusion classifier are adjusted, and when the loss value corresponding to the initial image classifier and the loss value corresponding to the initial fusion classifier are small, the data classification model is obtained, so that the trained data classification model can more accurately predict the image features and the fusion features, thereby improving the accuracy of data classification.
[0105] Optionally, when determining the media data category corresponding to the multimedia data, the multimedia data may also be specifically applied. For the specific method, please refer to Figure 6 , Figure 6 is a schematic flowchart of a data classification method provided by an embodiment of the present application. This method can be applied to a computer device, such as Figure 6 shown. This method includes:
[0106] S301, obtaining the image data and text data in the multimedia data.
[0107] S302, obtaining the image features of the multimedia data according to the image data, obtaining the text features of the multimedia data according to the text data, and performing feature fusion on the image features and the text features to obtain fusion features.
[0108] S303, predicting the image features based on the data classification model to obtain an object label, obtaining at least two prediction labels associated with the object label and the first probability value corresponding to each prediction label respectively, and predicting the fusion features based on the data classification model to obtain the second probability value corresponding to each prediction label respectively.
[0109] S304. Fuse the first probability value of the predicted label p and the second probability value of the predicted label p to obtain the third probability value of the predicted label p until the third probability value corresponding to each predicted label is obtained. Based on the third probability value corresponding to each predicted label and the object label, determine the media data category corresponding to the multimedia data.
[0110] In the embodiments of the present invention, the specific implementation manners of steps S301 to S304 may refer to the descriptions of steps S101 to S104, which will not be elaborated here.
[0111] S305. In response to a request for obtaining multimedia data, obtain the media data acquisition label of the target user who sends the acquisition request.
[0112] Here, the acquisition request refers to a request for obtaining multimedia data, the target user refers to the user who needs to obtain multimedia data, and the media data acquisition label of the target user may refer to the user label set when the target user logs in to the application where the multimedia data is located. The media data acquisition label can reflect the preferences of the target user for different types of multimedia data. For example, if the media data acquisition label includes thriller, excitement, etc., it means that the target user has a preference for thriller and exciting multimedia data, and then more multimedia data of this category can be pushed to the target user to improve the user experience.
[0113] Specifically, the media data acquisition label of the target user may be carried in the acquisition request for multimedia data, or may be a media data acquisition label matched by the computer device for the target user from the media library according to the acquisition request for the multimedia data. For example, the computer device can obtain the acquisition records of the target user for multimedia data in a historical time period, and match the media data acquisition label for the user according to the media data category in the acquisition records. The target user can send an acquisition request for multimedia data to the computer device through the user terminal where the target user is located. After the computer device receives the acquisition request for the multimedia data, it responds to the acquisition request for the multimedia data and obtains the media data acquisition label of the target user in the user terminal.
[0114] S306. If the media data category matches the media data acquisition label, send the multimedia data to the target user.
[0115] Here, the media data category matching the media data acquisition label may mean that the media data category is the same as or similar to the media data acquisition label. For example, if the media data category is dinosaur + thriller and the media data acquisition label includes dinosaur, thriller, excitement, etc., it can be determined that the media data category matches the media data acquisition label, and then the multimedia data is sent to the target user to improve the user experience.
[0116] S307. If the media data category does not match the media data acquisition label, send a media data exception message to the target user.
[0117] Here, the non - matching of the media data category and the media data acquisition label may mean that the media data category is not the same or not similar to the media data acquisition label. For example, if the media data category is dinosaur + thriller, and the media data acquisition labels include lizard, cute, and dazed, etc., it can be determined that the media data category does not match the media data acquisition label, and then a media data exception message is sent to the target user. The media data exception message is used to indicate that the current multimedia data acquisition is abnormal, and it can also indicate the reason for the data acquisition exception. For example, it can be "The current video data does not match your preference type. Please confirm whether to continue viewing."
[0118] Optionally, the computer device can push the multimedia data to the corresponding user terminal based on the media data category of the multimedia data. Specifically, when the computer device acquires multimedia data, it determines the media data category of the multimedia data in the manner of the above - mentioned steps S101 - S104, and acquires the media data acquisition labels associated with each user. Then it matches each media data acquisition label with the media data category, determines the users whose media data acquisition labels match the media data category as the users to receive, and sends the multimedia data to the user terminal where the users to receive are located. By acquiring the media data acquisition labels associated with each user and matching them with the media data category corresponding to the multimedia data, the media data acquisition labels that match the multimedia data can be determined, so as to push the multimedia data to the users, improve the accuracy of data pushing, and thus enhance the user experience.
[0119] In the embodiment of the present application, after determining the media data category corresponding to the multimedia data, the media data acquisition label representing the preference of the target user is acquired. By matching the media data acquisition label with the media data category, targeted pushing of the multimedia data to the target user is realized, thereby enhancing the user experience.
[0120] Optionally, if the multimedia data is video data, at least two predicted labels corresponding to each frame of the picture can be obtained, and the first probability value corresponding to each predicted label and the second probability value corresponding to each predicted label can be obtained, so as to determine the third probability value corresponding to each predicted label, obtain the target predicted label, and use the target predicted label and the object label as the media data category of the video data. For example, the multimedia data includes 20 frames of pictures, the object label is dinosaur, and the target predicted label corresponding to the 1st - 17th frames of pictures is thriller, and the target predicted label corresponding to the 18th - 20th frames of pictures is cute. Then it can be considered that the media data category of the multimedia data is dinosaur + thriller. That is to say, at least two predicted labels can include positive labels and negative labels. Positive labels are used to represent labels corresponding to a relatively relaxed atmosphere, and negative labels are used to represent labels corresponding to a relatively heavy atmosphere such as depression and anxiety. For example, positive labels can include happy, like, cute, etc., and negative labels can include thriller, afraid, dense, fear, frustration, etc. When the multimedia data is video data, if the target predicted label corresponding to one or more frames of the images is a negative label, it means that the target predicted label corresponding to the multimedia data is a negative label. During the process of auditing the multimedia data, if it is determined that the target predicted label corresponding to the multimedia data is a negative label, operations such as deleting the multimedia data and not pushing it to the user can be performed. If it is determined that the target predicted label corresponding to the multimedia data is a positive label, operations such as pushing the multimedia data to the user or displaying it on the home page of the application can be performed to facilitate the user to view the multimedia data, etc.
[0121] Specifically, the label cluster to which the media data category belongs can be obtained. If the label cluster is the first label cluster, the multimedia data is displayed on the home page of the application where the multimedia data is located; if the label cluster is the second label cluster, the multimedia data is deleted; the second label cluster includes labels that do not belong to the first label cluster.
[0122] Among them, the first label cluster can be the positive labels as described above. For example, it can include happy, like, cute, etc. The second label cluster includes labels that do not belong to the first label cluster. For example, the second label cluster can be the negative labels as described above. For example, it can include thriller, afraid, dense, fear, frustration, etc.
[0123] In the embodiments of the present application, by obtaining the label cluster to which the media data category belongs, the type of operation to be performed on the multimedia data is determined. Since the second label cluster means that the content of the multimedia data is likely to make the user in negative emotions such as depression and fear, the multimedia data of this type can be deleted or not displayed on the home page of the application, which can improve the user's love for the application and thus improve the user experience.
[0124] The method of the embodiments of the present application is introduced above. Next, the device of the embodiments of the present application will be introduced.
[0125] See Figure 7 , Figure 7 FIG. is a schematic structural diagram of the composition of a data classification device provided by an embodiment of the present application. The above data classification device may be a computer program (including program code) running in a computer device. For example, the data classification device is an application software. The device can be used to execute the corresponding steps in the method provided by the embodiments of the present application. The device 70 includes:
[0126] A data acquisition module 71, configured to acquire image data and text data in multimedia data;
[0127] A feature acquisition module 72, configured to acquire the image feature of the multimedia data according to the image data, acquire the text feature of the multimedia data according to the text data, and perform feature fusion on the image feature and the text feature to obtain a fusion feature;
[0128] A probability prediction module 73, configured to predict the image feature based on a data classification model to obtain an object label, acquire at least two prediction labels associated with the object label and first probability values respectively corresponding to each prediction label, and predict the fusion feature based on the data classification model to obtain second probability values respectively corresponding to each prediction label. The at least two prediction labels include a prediction label p, and p is a positive integer;
[0129] A category determination module 74, configured to fuse the first probability value of the prediction label p with the second probability value of the prediction label p to obtain a third probability value of the prediction label p, until third probability values respectively corresponding to each prediction label are obtained, and determine the media data category corresponding to the multimedia data based on the third probability values respectively corresponding to each prediction label and the object label.
[0130] Optionally, the data acquisition module 71 is configured to:
[0131] If the multimedia data is video data, acquire at least two video frame images constituting the video data;
[0132] Acquire the image data from the at least two video frame images based on an image acquisition period;
[0133] Search for first text content associated with the multimedia data, and if the first text content is found, determine the first text content as the text data;
[0134] If the first text content is not found, obtain the voice data corresponding to the image data in the video data, perform voice conversion on the voice data to obtain the second text content corresponding to the voice data, and determine the second text content as the text data.
[0135] Optionally, the feature acquisition module 72 includes: a weight acquisition unit 721, a first feature determination unit 722, a second feature determination unit 723, and a feature fusion unit 724;
[0136] The weight acquisition unit 721 is configured to acquire a first weight matrix corresponding to the image feature and a second weight matrix corresponding to the text feature;
[0137] The first feature determination unit 722 is configured to perform a weighted operation on the image feature based on the first weight matrix to obtain an image weighted feature;
[0138] The second feature determination unit 723 is configured to perform a weighted operation on the text feature based on the second weight matrix to obtain a text weighted feature;
[0139] The feature fusion unit 724 is configured to perform feature splicing on the image weighted feature and the text weighted feature to obtain the fusion feature.
[0140] Optionally, the category determination module 74 includes: a maximum probability determination unit 741 and a label splicing unit 742;
[0141] The maximum probability determination unit 741 is configured to determine the prediction label with the largest third probability value among the at least two prediction labels as the target prediction label;
[0142] The label splicing unit 742 is configured to splice the target prediction label and the object label to obtain a media data label, and determine the data category corresponding to the media data label as the media data category of the multimedia data.
[0143] Optionally, the device 70 further includes: a data sending module 75, configured to:
[0144] In response to a request for obtaining the multimedia data, obtain a media data acquisition label of the target user who sends the acquisition request;
[0145] If the media data category matches the media data acquisition label, send the multimedia data to the target user;
[0146] If the media data category does not match the media data acquisition label, send a media data exception message to the target user.
[0147] Optionally, the device 70 further includes: a data processing module 76, configured to:
[0148] Obtain the label cluster to which the media data category belongs. If the label cluster is the first label cluster, display the multimedia data on the home page of the application where the multimedia data is located.
[0149] If the label cluster is the second label cluster, delete the multimedia data; the second label cluster includes labels that do not belong to the first label cluster.
[0150] It should be noted that Figure 7 For the content not mentioned in the corresponding embodiments, reference can be made to the description of the method embodiments, which will not be elaborated here.
[0151] In the embodiments of the present application, by obtaining the image data and text data in the multimedia data; obtaining the image features of the multimedia data according to the image data, obtaining the text features of the multimedia data according to the text data, and performing feature fusion on the image features and the text features to obtain fusion features; predicting the image features based on the data classification model to obtain an object label and at least two predicted labels associated with the object label and the first probability value corresponding to each predicted label respectively, and predicting the fusion features based on the data classification model to obtain the second probability value corresponding to each predicted label respectively. Since the image features can reflect the image information in the multimedia data, the fusion features fuse the image features and the text features, and the text features are the descriptions of the image features, so the fusion features can reflect the user's subjective emotion towards the multimedia data. Moreover, by using the object label corresponding to the image features to assist the predicted label corresponding to the fusion features to judge the category of the multimedia data, it is possible to avoid inaccurate prediction results caused by too large a difference between the text features and the category of the multimedia data (for example, the title of the multimedia data is exaggerated and does not match the video content). That is to say, by combining the image features and the fusion features, when obtaining the predicted label corresponding to the fusion features, the object label corresponding to the image features can be used for re-judgment. That is, the present application can add human emotion (because the text features are generally added manually) through the fusion features of the text features and the image features, and correct the subjective emotion corresponding to the fusion features by predicting the subjective emotion through the image features, so that in the data classification of the multimedia data, both human emotion can be considered and the content itself of the multimedia data (i.e., the image features) will not be separated, thereby achieving a more accurate and comprehensive reflection of the category of the multimedia data, and further improving the accuracy of data classification.
[0152] See Figure 8 , Figure 8It is a schematic structural diagram of a data classification device provided by an embodiment of the present application. The above data classification device may be a computer program (including program code) running on a computer device. For example, the data classification device is an application software. The device can be used to execute the corresponding steps in the method provided by the embodiment of the present application. The device 80 includes:
[0153] A sample data acquisition module 81, configured to acquire sample image data and sample text data in sample multimedia data, and acquire a sample label of the sample multimedia data;
[0154] A sample feature acquisition module 82, configured to acquire a sample image feature of the sample multimedia data according to the sample image data, acquire a sample text feature of the sample multimedia data according to the sample text data, and perform feature fusion on the sample image feature and the sample text feature to obtain a sample fusion feature;
[0155] A sample label determination module 83, configured to predict the sample image feature based on an initial data classification model to obtain a sample object label, acquire at least two sample prediction labels associated with the sample object label and a first sample probability value corresponding to each sample prediction label, and predict the sample fusion feature based on the initial data classification model to obtain a second sample probability value corresponding to each sample prediction label; the at least two sample prediction labels include a sample prediction label j, and j is a positive integer;
[0156] A label output module 84, configured to fuse the first sample probability value of the sample prediction label j with the second sample probability value of the sample prediction label j to obtain a third sample probability value of the sample prediction label j, until third sample probability values corresponding to each sample prediction label are obtained, and determine a model output label corresponding to the sample multimedia data according to the third sample probability values corresponding to each sample prediction label and the sample object label;
[0157] A model training module 85, configured to train the initial data classification model according to a loss function composed of the sample label and the model output label to obtain a data classification model.
[0158] Optionally, the sample label includes a reference sample label and a reference sample prediction label, and the loss function includes a first loss function and a second loss function; the model training module 85 includes: a first training unit 851, a second training unit 852, and a model generation unit 853;
[0159] The first training unit 851 is configured to generate the first loss function according to the reference sample label and the sample object label;
[0160] The second training unit 852 is configured to splice the reference sample prediction label and the sample object label to generate a target sample label, and generate a second loss function based on the target sample label and the model output label;
[0161] The model generation unit 853 is configured to train the initial data classification model according to the first loss function and the second loss function to obtain the data classification model.
[0162] It should be noted that Figure 8 For the content not mentioned in the corresponding embodiments, reference may be made to the description of the method embodiments, which will not be elaborated here.
[0163] In the embodiments of the present application, by using a large amount of sample image data and sample text data to train the initial image classifier and the initial fusion classifier in the initial data classification model, the loss value corresponding to the initial image classifier and the loss value corresponding to the initial fusion classifier are obtained, and the initial image classifier is trained according to the loss value corresponding to the initial image classifier, and the initial fusion classifier is trained according to the loss value corresponding to the initial fusion classifier. When the loss value is relatively large, the initial image classifier and the initial fusion classifier are adjusted, and when the loss value corresponding to the initial image classifier and the loss value corresponding to the initial fusion classifier are relatively small, the data classification model is obtained, so that the trained data classification model can more accurately predict the image features and the fusion features, thereby improving the accuracy of data classification.
[0164] According to another embodiment of the present application, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding methods shown in Figure 3 、 5 、6 on a general computer device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct a data classification device as shown in Figure 7 、 8 and to implement a data classification method of the embodiments of the present application. The above computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above computing device through the computer-readable recording medium and run therein.
[0165] See Figure 9 , Figure 9 is a schematic structural diagram of the composition of a computer device provided by the embodiments of the present application. As shown in Figure 9As shown in the figure, the above computer device 90 may include: a processor 901, a network interface 904, and a memory 905. In addition, the above computer device 90 may further include: a user interface 903 and at least one communication bus 902. Among them, the communication bus 902 is used to realize the connection and communication between these components. Among them, the user interface 903 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 903 may further include a standard wired interface and a wireless interface. The network interface 904 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 905 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 905 may also be at least one storage device located far from the aforementioned processor 901. As Figure 9 shown, the memory 905, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0166] In Figure 9 the computer device 90 shown in the figure, the network interface 904 can provide network communication functions; while the user interface 903 is mainly used to provide an input interface for users; and the processor 901 can be used to call the device control application program stored in the memory 905 to achieve:
[0167] Obtain the image data and text data in the multimedia data;
[0168] Obtain the image features of the multimedia data according to the image data, obtain the text features of the multimedia data according to the text data, and perform feature fusion on the image features and the text features to obtain fusion features;
[0169] Based on the data classification model, predict the image features to obtain object labels, obtain at least two prediction labels associated with the object labels and the first probability value corresponding to each prediction label respectively, and based on the data classification model, predict the fusion features to obtain the second probability value corresponding to each prediction label respectively; the at least two prediction labels include a prediction label p, and p is a positive integer;
[0170] Fuse the first probability value of the prediction label p and the second probability value of the prediction label p to obtain the third probability value of the prediction label p until the third probability value corresponding to each prediction label is obtained, and based on the third probability value corresponding to each prediction label and the object label, determine the media data category corresponding to the multimedia data
[0171] It should be understood that the computer device 90 described in the embodiments of the present application can execute the foregoing Figure 3, 5 , the description of the above data classification method in the corresponding embodiment of 6 can also be executed as described above Figure 7 , 8 , the description of the above data classification device in the corresponding embodiment of 6 will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated either.
[0172] In the embodiments of the present application, by obtaining image data and text data in multimedia data; obtaining the image features of the multimedia data according to the image data, obtaining the text features of the multimedia data according to the text data, and performing feature fusion on the image features and the text features to obtain fusion features; predicting the image features based on a data classification model to obtain an object label, at least two prediction labels associated with the object label, and a first probability value corresponding to each prediction label respectively, and predicting the fusion features based on the data classification model to obtain a second probability value corresponding to each prediction label respectively. Since the image features can reflect the image information in the multimedia data, the fusion features fuse the image features and the text features, and the text features are descriptions of the image features, the fusion features can reflect the user's subjective emotion towards the multimedia data. Moreover, by using the object label corresponding to the image features to assist in judging the category of the multimedia data for the prediction label corresponding to the fusion features, it is possible to avoid inaccurate prediction results caused by too large a difference between the text features and the category of the multimedia data (for example, the title of the multimedia data is exaggerated and does not match the video content). That is to say, by combining the image features and the fusion features, when obtaining the prediction label corresponding to the fusion features, the object label corresponding to the image features can be used for re-judgment. That is, the present application can add human emotion (because text features are generally added manually) through the fusion features of text features and image features, and correct the subjective emotion corresponding to the fusion features by predicting the subjective emotion through the image features, so that in the data classification of multimedia data, both human emotion can be considered and the content itself of the multimedia data (i.e., the image features) will not be deviated from, thereby achieving a more accurate and comprehensive reflection of the category of the multimedia data, and further improving the accuracy of data classification.
[0173] The embodiments of the present application further provide a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the method as described in the foregoing embodiments. The computer may be a part of the computer device mentioned above. For example, it may be the aforementioned processor 901. As an example, the program instructions may be deployed to be executed on a single computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. The multiple computer devices distributed at multiple locations and interconnected through a communication network may form a blockchain network.
[0174] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0175] The foregoing disclosure is only for the preferred embodiments of the present application, and of course cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data classification method, characterized in that, Including: Obtaining image data and text data in the multimedia data; Obtaining the image features of the multimedia data according to the image data, obtaining the text features of the multimedia data according to the text data, and performing feature fusion on the image features and the text features to obtain fusion features; Predicting the image features based on a data classification model to obtain an object label, obtaining at least two prediction labels associated with the object label and first probability values respectively corresponding to each prediction label, and predicting the fusion features based on the data classification model to obtain second probability values respectively corresponding to each prediction label; the object label is used to indicate the category to which the target object in the image represented by the image features belongs, and the at least two prediction labels are used to indicate the subjective emotion categories of the multimedia data; the number of the at least two prediction labels is determined according to the number of subjective emotion categories in the label library, and the at least two prediction labels include a prediction label p, where p is a positive integer; Fusing the first probability value of the prediction label p and the second probability value of the prediction label p to obtain a third probability value of the prediction label p, until the third probability values respectively corresponding to each prediction label are obtained, and determining the media data category corresponding to the multimedia data based on the third probability values respectively corresponding to each prediction label and the object label.
2. The method according to claim 1, characterized in that, The obtaining the image data and text data in the multimedia data includes: If the multimedia data is video data, obtaining at least two video frame images that make up the video data; Obtaining the image data from the at least two video frame images based on an image acquisition period; Searching for first text content associated with the multimedia data, and if the first text content is found, determining the first text content as the text data; If the first text content is not found, obtaining voice data corresponding to the image data in the video data, performing voice conversion on the voice data to obtain second text content corresponding to the voice data, and determining the second text content as the text data.
3. The method according to claim 1, characterized in that, The performing feature fusion on the image features and the text features to obtain fusion features includes: Obtaining a first weight matrix corresponding to the image features and a second weight matrix corresponding to the text features; Performing weighted operation on the image features based on the first weight matrix to obtain image weighted features; Performing weighted operation on the text features based on the second weight matrix to obtain text weighted features; Performing feature splicing on the image weighted features and the text weighted features to obtain the fusion features.
4. The method according to claim 1, characterized in that, The determining the media data category corresponding to the multimedia data based on the third probability values respectively corresponding to each prediction label and the object label includes: Determining the prediction label with the largest third probability value among the at least two prediction labels as the target prediction label; Splicing the target prediction label and the object label to obtain a media data label, and determining the data category corresponding to the media data label as the media data category of the multimedia data.
5. The method according to claim 1, characterized in that, The method further includes: In response to a request for obtaining the multimedia data, obtaining a media data acquisition label of a target user who sends the acquisition request; If the media data category matches the media data acquisition label, sending the multimedia data to the target user; If the media data category does not match the media data acquisition label, sending a media data exception message to the target user.
6. The method according to claim 1, characterized in that, The method further includes: Obtaining a label cluster to which the media data category belongs. If the label cluster is a first label cluster, displaying the multimedia data on the home page of the application where the multimedia data is located; If the label cluster is a second label cluster, deleting the multimedia data; the second label cluster includes labels that do not belong to the first label cluster.
7. A data classification method, characterized in that, Includes: Obtaining sample image data and sample text data in the sample multimedia data, and obtaining a sample label of the sample multimedia data; Obtaining sample image features of the sample multimedia data according to the sample image data, obtaining sample text features of the sample multimedia data according to the sample text data, and performing feature fusion on the sample image features and the sample text features to obtain sample fusion features; Predicting the sample image features based on an initial data classification model to obtain a sample object label, obtaining at least two sample prediction labels associated with the sample object label and a first sample probability value corresponding to each sample prediction label, and predicting the sample fusion features based on the initial data classification model to obtain a second sample probability value corresponding to each sample prediction label; the sample object label is used to indicate the category to which the target object in the image represented by the sample image features belongs, and the at least two sample prediction labels are used to indicate the subjective emotion types of the sample multimedia data; the number of the at least two sample prediction labels is determined according to the number of subjective emotion categories in the label library, and the at least two sample prediction labels include a sample prediction label j, where j is a positive integer; Fusing the first sample probability value of the sample prediction label j with the second sample probability value of the sample prediction label j to obtain a third sample probability value of the sample prediction label j until third sample probability values corresponding to each sample prediction label are obtained, and determining a model output label corresponding to the sample multimedia data according to the third sample probability values corresponding to each sample prediction label and the sample object label; Training the initial data classification model according to a loss function composed of the sample label and the model output label to obtain a data classification model.
8. The method according to claim 7, characterized in that, The sample label includes a reference sample label and a reference sample prediction label, and the loss function includes a first loss function and a second loss function; The training the initial data classification model according to a loss function composed of the sample label and the model output label to obtain a data classification model includes: Generating the first loss function according to the reference sample label and the sample object label; Concatenate the reference sample prediction label and the sample object label to generate a target sample label, and generate the second loss function based on the target sample label and the model output label; Train the initial data classification model according to the first loss function and the second loss function to obtain the data classification model.
9. A computer device, characterized in that, Including: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface. Among them, the network interface is used to provide data communication functions, the memory is used to store program codes, and the processor is used to call the program codes to execute the method according to any one of claims 1-6, or execute the method according to any one of claims 7-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by the processor, the processor executes the method according to any one of claims 1-6, or executes the method according to any one of claims 7-8.
Citation Information
Patent Citations
Multi-modal sentiment analysis method based on quantum theory
CN107832663A
Model generation method, video classification method and device, terminal and storage medium
CN109710800A