Multimedia resource classification model training method and multimedia resource recommendation method
By training the multimedia resource classification model, the problem of easy recommendation of low-quality resources in traditional multimedia resource recommendation methods is solved, more efficient resource recommendation is achieved, and user experience and computer resource utilization are improved.
Patent Information
- Application Number
- CN202110113770.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-01-27
AI Technical Summary
Traditional multimedia resource recommendation methods are easy to recommend low-quality resources, resulting in repeated searches and interface refreshes by users, reducing the effectiveness of recommendations.
The multimedia resource classification model training method is adopted, and the parameters of the multimedia resource classification model are adjusted until the convergence conditions are met, and the model used to classify the quality of the recommended multimedia resource is obtained.
It improves the effectiveness of multimedia resource recommendations, ensures that the resources recommended to users are of better quality, reduces user search and refresh behavior, and reduces the waste of computer resources.
Smart Images

Figure CN113590849B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a multimedia resource classification model training method, a multimedia resource recommendation method, an apparatus, a computer device and a storage medium. Background Art
[0002] With the development of computer technology, a variety of network applications have emerged. People can publish multimedia resources on network applications and browse multimedia resources on network applications.
[0003] In traditional technologies, multimedia resources are usually recommended to users randomly, which easily leads to the recommendation of low-quality multimedia resources to users. This may result in the multimedia resources recommended to users often not being of concern to or of interest to the users. In addition, low-quality multimedia resources will not only occupy storage resources, but also cause users to repeatedly search and refresh the interface, occupying a large amount of computer resources, ultimately resulting in low effectiveness of multimedia resource recommendations. Summary of the invention
[0004] Based on this, it is necessary to provide a multimedia resource classification model training method, multimedia resource recommendation method, device, computer equipment and storage medium that can improve the effectiveness of multimedia resource recommendation in response to the above technical problems.
[0005] A multimedia resource classification model training method, the method comprising:
[0006] Obtain a target attribute information set and a training label set of training multimedia resources; the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks;
[0007] Inputting the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained; the multimedia resource classification model includes multiple feature sub-networks and task sub-networks corresponding to each task;
[0008] Through each feature sub-network in the multimedia resource classification model, the target attribute information associated with the feature sub-network is vectorized to obtain the attribute feature vector output by each feature sub-network;
[0009] Input each attribute feature vector into each task sub-network to obtain the prediction label corresponding to each task;
[0010] The parameters of the corresponding task sub-network are adjusted based on the training labels and prediction labels corresponding to the same task, and the model parameters of each feature sub-network are adjusted based on the training labels and prediction labels corresponding to each task until the convergence conditions are met to obtain a trained multimedia resource classification model; the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended.
[0011] A multimedia resource classification model training device, the device comprising:
[0012] An information acquisition module is used to acquire a target attribute information set and a training label set of training multimedia resources; the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks;
[0013] The attribute information input module is used to input the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained; the multimedia resource classification model includes multiple feature sub-networks and task sub-networks corresponding to each task;
[0014] The attribute information processing module is used to vectorize the target attribute information associated with the feature sub-network through each feature sub-network in the multimedia resource classification model, and obtain the attribute feature vector output by each feature sub-network;
[0015] The label prediction module is used to input each attribute feature vector into each task sub-network to obtain the predicted label corresponding to each task;
[0016] The model adjustment module is used to adjust the parameters of the corresponding task sub-network according to the training labels and prediction labels corresponding to the same task, and adjust the model parameters of each feature sub-network based on the training labels and prediction labels corresponding to each task until the convergence conditions are met to obtain a trained multimedia resource classification model; the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended.
[0017] In one embodiment, the information acquisition module is also used to obtain recommended interaction information sets corresponding to multiple historical multimedia resources; the recommended interaction information sets include recommended interaction information corresponding to each task; the recommended interaction information corresponding to the same task is counted to obtain reference interaction information corresponding to each task; the quality of historical multimedia resources is classified based on the recommended interaction information and the corresponding reference interaction information to obtain a quality label set corresponding to each historical multimedia resource; training multimedia resources and a corresponding training label set are obtained based on the historical multimedia resources and the corresponding quality label set.
[0018] In one embodiment, the information acquisition module is also used to compare the recommended interactivity corresponding to the recommended interaction information of the same category and the reference interactivity corresponding to the reference interaction information in the recommended interaction information set corresponding to the same historical multimedia resource; determine the quality label of the task corresponding to the recommended interaction information whose recommended interactivity is greater than the reference interactivity as a positive label; and determine the quality label of the task corresponding to the recommended interaction information whose recommended interactivity is less than the reference interactivity as a negative label.
[0019] In one embodiment, the feature subnetwork includes a text feature subnetwork, the target attribute information associated with the text feature subnetwork includes a plurality of text attribute information, and the text feature subnetwork includes data processing channels corresponding to each piece of text attribute information. The attribute information processing module is further used to perform vectorization processing on the corresponding text attribute information through each data processing channel in the text feature subnetwork, and obtain the text feature vector output by each data processing channel; and obtain the attribute feature vector output by the text feature subnetwork based on each text feature vector.
[0020] In one embodiment, the feature subnetwork includes an atomic feature subnetwork. The attribute information processing module is further used to perform feature cross processing on target attribute information associated with the atomic feature subnetwork through the atomic feature subnetwork to obtain at least one cross feature vector; and obtain the attribute feature vector output by the atomic feature subnetwork based on each cross feature vector.
[0021] In one embodiment, the target attribute information associated with the atomic feature sub-network includes at least two of user attribute information, image attribute information, language attribute information, and text statistical attribute information.
[0022] In one embodiment, the feature subnetwork includes a text-image fusion feature subnetwork, the target attribute information associated with the text-image fusion feature subnetwork includes text attribute information and image attribute information, and the text-image fusion feature subnetwork includes a text data processing channel corresponding to the text attribute information and an image data processing channel corresponding to the image attribute information. The attribute information processing module is also used to encode the text attribute information through the text data processing channel to obtain an intermediate feature vector; encode the image attribute information through the image data processing channel to obtain an image feature vector; perform attention allocation processing on the intermediate feature vector based on the image feature vector to obtain a first text-image fusion feature vector; perform attention allocation processing on the image feature vector based on the intermediate feature vector to obtain a second text-image fusion feature vector; and obtain an attribute feature vector output by the text-image fusion feature subnetwork based on the first text-image fusion feature vector and the second text-image fusion feature vector.
[0023] In one embodiment, the attribute information processing module is further used to perform word encoding processing on the text attribute information to obtain a word feature vector; and perform sentence encoding processing on the word feature vector to obtain an intermediate feature vector.
[0024] In one embodiment, the feature subnetwork includes a style feature subnetwork, and the target attribute information associated with the style feature subnetwork includes style attribute information; the style feature subnetwork includes a first data processing channel and a second data processing channel. The attribute information processing module is further used to encode the style attribute information through the first data processing channel to obtain an initial feature vector, perform attention allocation processing on the initial feature vector to obtain a first feature vector; perform convolution processing on the style attribute information through the second data processing channel to obtain a second feature vector; and obtain an attribute feature vector output by the style feature subnetwork based on the first feature vector and the second feature vector.
[0025] In one embodiment, each task subnetwork includes an expert layer, a gating layer and a fusion layer; each task subnetwork shares the expert layer. The label prediction module is also used to perform feature processing on each attribute feature vector in the current task subnetwork through the expert layer to obtain a feature processing result, perform weighted processing on the feature processing result through the gating layer to obtain an intermediate processing result, and perform fusion processing on the intermediate processing result through the fusion layer to obtain a predicted label of the task corresponding to the current task subnetwork.
[0026] In one embodiment, the apparatus further comprises:
[0027] A model updating module is used to obtain a target attribute information set and a verification label set for verifying multimedia resources; verify that the multimedia resources are updated and recommended multimedia resources; input the target attribute information set for verifying multimedia resources into a trained multimedia resource classification model to obtain a predicted label set corresponding to the verified multimedia resources; calculate the classification accuracy based on the predicted label set and the verification label set corresponding to the verified multimedia resources; when the classification accuracy is less than an accuracy threshold, update the trained multimedia resource classification model based on the predicted label set and the training label set corresponding to the verified multimedia resources to obtain an updated multimedia resource classification model.
[0028] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0029] Obtain a target attribute information set and a training label set of training multimedia resources; the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks;
[0030] Inputting the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained; the multimedia resource classification model includes multiple feature sub-networks and task sub-networks corresponding to each task;
[0031] Through each feature sub-network in the multimedia resource classification model, the target attribute information associated with the feature sub-network is vectorized to obtain the attribute feature vector output by each feature sub-network;
[0032] Input each attribute feature vector into each task sub-network to obtain the prediction label corresponding to each task;
[0033] The parameters of the corresponding task sub-network are adjusted based on the training labels and prediction labels corresponding to the same task, and the model parameters of each feature sub-network are adjusted based on the training labels and prediction labels corresponding to each task until the convergence conditions are met to obtain a trained multimedia resource classification model; the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended.
[0034] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0035] Obtain a target attribute information set and a training label set of training multimedia resources; the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks;
[0036] Inputting the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained; the multimedia resource classification model includes multiple feature sub-networks and task sub-networks corresponding to each task;
[0037] Through each feature sub-network in the multimedia resource classification model, the target attribute information associated with the feature sub-network is vectorized to obtain the attribute feature vector output by each feature sub-network;
[0038] Input each attribute feature vector into each task sub-network to obtain the prediction label corresponding to each task;
[0039] The parameters of the corresponding task sub-network are adjusted based on the training labels and prediction labels corresponding to the same task, and the model parameters of each feature sub-network are adjusted based on the training labels and prediction labels corresponding to each task until the convergence conditions are met to obtain a trained multimedia resource classification model; the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended.
[0040] The above-mentioned multimedia resource classification model training method, device, computer equipment and storage medium obtain the target attribute information set and training label set of the training multimedia resource, the target attribute information set includes target attribute information of multiple dimensions, the training label set includes training labels corresponding to multiple tasks, the target attribute information set of the training multimedia resource is input into the multimedia resource classification model to be trained, the multimedia resource classification model includes multiple feature subnetworks and task subnetworks corresponding to each task, through each feature subnetwork in the multimedia resource classification model, the target attribute information associated with the feature subnetwork is vectorized, and the attribute feature vector output by each feature subnetwork is obtained, and each attribute feature vector is input into each task subnetwork to obtain the prediction label corresponding to each task, and the parameters of the corresponding task subnetwork are adjusted based on the training label and prediction label corresponding to the same task, and the model parameters of each feature subnetwork are adjusted based on the training label and prediction label corresponding to each task, until the convergence condition is met, and the trained multimedia resource classification model is obtained, and the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended. In this way, the multimedia resource classification model can be supervised based on the target attribute information set and the training label set of the training multimedia resource, and a multimedia resource classification model that can accurately classify the quality of the multimedia resources to be recommended is obtained. Among them, the target attribute information set of multimedia resources includes target attribute information of multiple dimensions. The target attribute information of different dimensions can reflect the content quality of multimedia resources from different angles. The target attribute information set is input into the multimedia resource classification model. By comprehensively considering the target attribute information of each dimension, the quality of multimedia resources can be accurately classified, and a prediction label that can accurately reflect the quality of multimedia resources can be obtained. In addition, the multimedia resource classification model includes multiple task sub-networks, which is a multi-task model that can predict the performance of multimedia resources on various tasks. During model training, multiple related tasks are learned in parallel at the same time, and gradients are back-propagated at the same time to learn the connection and difference between different tasks, thereby improving the learning efficiency and quality of each task. Finally, the trained multimedia resource classification model can be used to classify the quality of multimedia resources to be recommended, so that multimedia resources with better quality can be recommended to users, thereby improving the effectiveness of multimedia resource recommendation. Effective multimedia resource recommendation can avoid users from repeatedly searching or refreshing the interface due to low-quality and invalid multimedia resource recommendations. Repeated search or repeated refresh of the interface will occupy a large amount of computer equipment resources. Therefore, on the basis of improving the effectiveness of resource recommendation, it can also reduce the waste of resources of terminals or servers.
[0041] A multimedia resource recommendation method, the method comprising:
[0042] Acquire a target attribute information set of the multimedia resource to be recommended; the target attribute information set includes target attribute information of multiple dimensions;
[0043] Inputting the target attribute information set into the trained multimedia resource classification model; the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks;
[0044] Through each feature sub-network in the multimedia resource classification model, the target attribute information associated with the feature sub-network is vectorized to obtain the attribute feature vector output by each feature sub-network;
[0045] Input each attribute feature vector into each task sub-network to obtain the predicted label output by each task sub-network;
[0046] Based on each predicted label, a quality classification result corresponding to the multimedia resource to be recommended is obtained;
[0047] Recommend multimedia resources based on quality classification results.
[0048] A multimedia resource recommendation device, the device comprising:
[0049] The attribute information acquisition module is used to acquire a target attribute information set of the multimedia resource to be recommended; the target attribute information set includes target attribute information of multiple dimensions;
[0050] The attribute information input module is used to input the target attribute information set into the trained multimedia resource classification model; the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks;
[0051] The attribute information processing module is used to vectorize the target attribute information associated with the feature sub-network through each feature sub-network in the multimedia resource classification model, and obtain the attribute feature vector output by each feature sub-network;
[0052] The label prediction module is used to input each attribute feature vector into each task sub-network to obtain the predicted label output by each task sub-network;
[0053] A quality classification module, used to obtain quality classification results corresponding to the multimedia resources to be recommended based on each prediction label;
[0054] The resource recommendation module is used to recommend multimedia resources to be recommended based on the quality classification results.
[0055] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0056] Acquire a target attribute information set of the multimedia resource to be recommended; the target attribute information set includes target attribute information of multiple dimensions;
[0057] Inputting the target attribute information set into the trained multimedia resource classification model; the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks;
[0058] Through each feature sub-network in the multimedia resource classification model, the target attribute information associated with the feature sub-network is vectorized to obtain the attribute feature vector output by each feature sub-network;
[0059] Input each attribute feature vector into each task sub-network to obtain the predicted label output by each task sub-network;
[0060] Based on each predicted label, a quality classification result corresponding to the multimedia resource to be recommended is obtained;
[0061] Recommend multimedia resources based on quality classification results.
[0062] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0063] Acquire a target attribute information set of the multimedia resource to be recommended; the target attribute information set includes target attribute information of multiple dimensions;
[0064] Inputting the target attribute information set into the trained multimedia resource classification model; the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks;
[0065] Through each feature sub-network in the multimedia resource classification model, the target attribute information associated with the feature sub-network is vectorized to obtain the attribute feature vector output by each feature sub-network;
[0066] Input each attribute feature vector into each task sub-network to obtain the predicted label output by each task sub-network;
[0067] Based on each predicted label, a quality classification result corresponding to the multimedia resource to be recommended is obtained;
[0068] Recommend multimedia resources based on quality classification results.
[0069] The above-mentioned multimedia resource classification model training method, device, computer equipment and storage medium obtain the target attribute information set of the multimedia resource to be recommended, the target attribute information set includes target attribute information of multiple dimensions, input the target attribute information set into the trained multimedia resource classification model, the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks, through each feature sub-network in the multimedia resource classification model, respectively, the target attribute information associated with the feature sub-network is vectorized, and the attribute feature vector output by each feature sub-network is obtained, and each attribute feature vector is input into each task sub-network to obtain the prediction label output by each task sub-network, and the quality classification result corresponding to the multimedia resource to be recommended is obtained based on each prediction label, and the multimedia resource to be recommended is recommended based on the quality classification result. In this way, the target attribute information set of the multimedia resource includes target attribute information of multiple dimensions, and the target attribute information of different dimensions can reflect the content quality of the multimedia resource from different angles. The target attribute information set is input into the multimedia resource classification model, and the quality of the multimedia resource can be accurately classified by comprehensively considering the target attribute information of each dimension, and an accurate quality classification result is obtained. In addition, the multimedia resource classification model includes multiple task sub-networks and is a multi-task model that can predict the performance of multimedia resources on various tasks. The quality classification results of multimedia resources are obtained by integrating the performance of multimedia resources on various tasks, which can further improve the accuracy of multimedia resource quality classification. Finally, multimedia resources with better quality can be identified through the multimedia resource classification model, so that multimedia resources with better quality can be recommended to users, improving the effectiveness of multimedia resource recommendation. Effective multimedia resource recommendation can avoid users from repeatedly searching or refreshing the interface due to low-quality and invalid multimedia resource recommendations. Repeated searches or repeated refreshes of the interface will occupy a large amount of computer device resources. Therefore, on the basis of improving the effectiveness of resource recommendation, it can also reduce the waste of resources of terminals or servers. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 A diagram of an application environment of a multimedia resource classification model training method and a multimedia resource recommendation method in one embodiment;
[0071] Figure 2 A flowchart of a multimedia resource classification model training method in one embodiment;
[0072] Figure 3 A schematic diagram of a process of obtaining training labels for training multimedia resources in one embodiment;
[0073] Figure 4 A schematic diagram of a process for determining a quality label of a historical multimedia resource based on recommended interaction information and corresponding reference interaction information in one embodiment;
[0074] Figure 5 A schematic diagram of the structure of a text feature subnetwork in one embodiment;
[0075] Figure 6 A schematic diagram of the structure of an atomic feature subnetwork in one embodiment;
[0076] Figure 7 Schematic diagram of the structure of a graphic-text fusion feature sub-network in one embodiment;
[0077] Figure 8 A schematic diagram of the structure of a style feature sub-network in one embodiment;
[0078] Fig. 9A is a schematic diagram of the structure of a task subnetwork in one embodiment;
[0079] Fig. 9B A schematic diagram of the structure of a task subnetwork in another embodiment;
[0080] Fig. 9C is a schematic diagram of the structure of a task subnetwork in yet another embodiment;
[0081] Fig.10 A schematic diagram of a process for updating a multimedia resource classification model in one embodiment;
[0082] Fig.11 A schematic diagram of a process for training and updating a multimedia resource classification model in one embodiment;
[0083] Fig.12 A schematic diagram of a process of recommending multimedia resources in one embodiment;
[0084] Fig.13A A schematic diagram of a process for recommending high-quality content in one embodiment;
[0085] Fig. 13B A schematic diagram of the structure of a graphic content classification model in one embodiment;
[0086] Fig. 13C This is a schematic diagram of an interface for recommending graphic content in one embodiment;
[0087] Fig.14 is a structural block diagram of a multimedia resource classification model training device in one embodiment;
[0088] Fig.15 is a structural block diagram of a multimedia resource classification model training device in one embodiment;
[0089] Fig.16 is a structural block diagram of a multimedia resource recommendation device in one embodiment;
[0090] Fig.17 is an internal structure diagram of a computer device in one embodiment;
[0091] Fig.18 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0092] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0093] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0094] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0095] Computer Vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to the use of cameras and computers to replace human eyes to identify, track and measure targets, and further perform graphics processing so that the computer processes the images into images that are more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0096] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0097] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0098] The solution provided in the embodiments of the present application involves artificial intelligence computer vision, natural language processing, machine learning and other technologies, which are specifically described by the following embodiments:
[0099] The multimedia resource classification model training method and multimedia resource recommendation method provided in this application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 through a network. The terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented as an independent server or a server cluster consisting of multiple servers.
[0100] Both the terminal 102 and the server 104 can be used independently to execute the multimedia resource classification model training method and the multimedia resource recommendation method provided in the embodiments of the present application.
[0101] For example, the server 104 obtains a target attribute information set and a training label set for training multimedia resources, wherein the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks, and inputs the target attribute information set for training multimedia resources into a multimedia resource classification model to be trained, wherein the multimedia resource classification model includes multiple feature subnetworks and task subnetworks corresponding to each task. Through each feature subnetwork in the multimedia resource classification model, the target attribute information associated with the feature subnetwork is vectorized to obtain the attribute feature vector output by each feature subnetwork, and each attribute feature vector is input into each task subnetwork to obtain the prediction label corresponding to each task. The server 104 can adjust the parameters of the corresponding task subnetwork based on the training label and prediction label corresponding to the same task, and adjust the model parameters of each feature subnetwork based on the training label and prediction label corresponding to each task, until the convergence condition is met, and a multimedia resource classification model that has been trained is obtained, and the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended.
[0102] The terminal 102 obtains a target attribute information set of the multimedia resource to be recommended, the target attribute information set includes target attribute information of multiple dimensions, and inputs the target attribute information set into a trained multimedia resource classification model, the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks. Through each feature sub-network in the multimedia resource classification model, the target attribute information associated with the feature sub-network is vectorized to obtain the attribute feature vector output by each feature sub-network, and each attribute feature vector is input into each task sub-network to obtain the predicted label output by each task sub-network, and the quality classification result corresponding to the multimedia resource to be recommended is obtained based on each predicted label. The terminal 102 can recommend the multimedia resource to be recommended based on the quality classification result.
[0103] The terminal 102 and the server 104 can also be used together to execute the multimedia resource classification model training method and the multimedia resource recommendation method provided in the embodiments of the present application.
[0104] For example, the server 104 obtains the target attribute information set and the training label set of the training multimedia resources from the terminal 102, and the server 104 performs model training on the multimedia resource classification model based on the target attribute information set and the training label set of the training multimedia resources to obtain a trained multimedia resource classification model.
[0105] The terminal 102 obtains the trained multimedia resource classification model from the server 104. The terminal 102 classifies the quality of the multimedia resource to be recommended by the trained multimedia resource classification model, obtains the quality classification result of the multimedia resource to be recommended, and recommends the multimedia resource to be recommended based on the quality classification result.
[0106] In one embodiment, Figure 2 As shown, a multimedia resource classification model training method is provided, and the method is applied to Figure 1 The computer device in the example is used to illustrate, and the computer device can be the above Figure 1 The terminal 102 or the server 104 in FIG. Figure 2 ,The multimedia resource classification model training method includes the following steps:
[0107] Step S202, obtaining a target attribute information set and a training label set of training multimedia resources; the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks.
[0108] Among them, multimedia resources refer to resources containing at least two media, such as articles containing pictures, pictures containing text, videos containing subtitles, videos containing audio, etc. Users can publish multimedia resources on various resource service platforms, for example, publishing short videos of life on social applications and publishing news information on information applications. Training multimedia resources refer to multimedia resources used for model training. Target attribute information is information used to describe the attributes, characteristics, and functions of multimedia resources. The target attribute information can specifically be text-related information, image-related information, style-related information, user-related information of the publishing user of the multimedia resource, etc.
[0109] The training label refers to the quality label of the training multimedia resource. The quality label is used to measure the quality of the multimedia resource. For example, the quality label may include positive labels and negative labels. The positive label indicates that the multimedia resource performs better on the corresponding task and is a high-quality multimedia resource for the task. The negative label indicates that the multimedia resource performs generally or poorly on the corresponding task and is not a high-quality multimedia resource for the task. The quality label may also include a first label, a second label, a third label, etc. Different labels represent different quality levels. The quality level may reflect the quality of the multimedia resource. For example, the better the performance of the multimedia resource on the corresponding task, the higher the corresponding quality level on the task. A task corresponds to an interactive behavior between the multimedia resource and the user, such as a task related to click behavior, a task related to browsing behavior, and a task related to comment behavior. Then, the training label set may include a quality label corresponding to the click-through rate task, a quality label corresponding to the browsing time task, and a quality label corresponding to the comment rate task.
[0110] Specifically, the computer device can obtain the training multimedia resources from the multimedia resource database and determine the training label set corresponding to the training multimedia resources. Further, the computer device can perform content analysis on the training multimedia resources to obtain the target attribute information set of the training multimedia resources. For example, the text title, text body, image quality, and content layout of the multimedia resources are used as target attribute information to form the target attribute information set.
[0111] In one embodiment, the training multimedia resources may be multimedia resources that have been published and recommended. Then, a reasonable evaluation system for the quality of multimedia resources can be constructed from the perspective of user feedback, and a set of training labels for training multimedia resources can be determined based on the evaluation system. For example, the training labels for training multimedia resources can be determined based on a large number of users' feedback information and evaluation information on multimedia resources. For example, the quality label set of multimedia resources is determined based on the click-through rate, browsing time, and comment rate of multimedia resources. Further, a reasonable evaluation system for the quality of multimedia resources can also be constructed in combination with the content itself and user feedback. Of course, the training multimedia resources can also be multimedia resources that have not yet been published or recommended. Then, a reasonable evaluation system for the quality of multimedia resources can be constructed from the perspective of the content itself, and a set of training labels for training multimedia resources can be determined based on the evaluation system. For example, the training labels for training multimedia resources can be quality labels artificially determined based on the expert knowledge of experts. The training labels for training multimedia resources can also be determined based on the quality labels of published multimedia resources with similar content. For example, in the click-through rate task, it is known that the quality label of multimedia resource 1 is a positive label. When the content similarity between multimedia resource 1 and multimedia resource 2 is greater than a preset threshold, that is, the contents of multimedia resource 1 and multimedia resource 2 are relatively similar, it is determined that the quality label of multimedia resource 2 in the click-through rate task is also a positive label.
[0112] Step S204, inputting the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained; the multimedia resource classification model includes a plurality of feature sub-networks and task sub-networks corresponding to each task.
[0113] Among them, the multimedia resource classification model is a machine learning model used to classify the quality of multimedia resources, that is, to identify high-quality multimedia resources. The multimedia resource classification model includes multiple feature subnetworks and multiple task subnetworks. The feature subnetwork is used to convert target attribute information into attribute feature vectors, and different feature subnetworks are used to process different target attribute information. The task subnetwork is used to predict the performance of multimedia resources on specific tasks based on the attribute feature vectors, and different task subnetworks correspond to different tasks.
[0114] Step S206, through each feature sub-network in the multimedia resource classification model, vectorization processing is performed on the target attribute information associated with the feature sub-network respectively to obtain the attribute feature vector output by each feature sub-network.
[0115] Specifically, the computer device can input the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained, and each feature subnetwork in the multimedia resource classification model can receive the corresponding target attribute information, and each feature subnetwork respectively vectorizes the target attribute information associated with itself, thereby outputting the corresponding attribute feature vector. Among them, vectorization refers to representing the target attribute information with a dense vector, which helps the model to learn relevant information and knowledge more fully. The vectorization processing method of each feature subnetwork for the target attribute information can be the same or different.
[0116] In one embodiment, the computer device may first convert the target attribute information into an original feature vector, then input the original feature vector into a feature sub-network, and perform a densified representation of the original feature vector through the feature sub-network to obtain an attribute feature vector. The target attribute information may be converted into an original feature vector by One Hot encoding.
[0117] In one embodiment, the multimedia resource classification model includes at least two of a text feature subnetwork, an atomic feature subnetwork, a text-image fusion feature subnetwork, and a style feature subnetwork. The text feature subnetwork is used to process target attribute information related to the text corresponding to the multimedia resource. The atomic feature subnetwork is used to process atomic features corresponding to the multimedia resource, where an atomic feature refers to the smallest, indivisible feature. The text-image fusion feature subnetwork is used to process target attribute information related to the text and target attribute information related to the image corresponding to the multimedia resource, and to fuse the two types of target attribute information. The style feature subnetwork is used to process target attribute information related to the content style corresponding to the multimedia resource. Each feature subnetwork is used to process target attribute information of different dimensions, and each attribute feature vector finally obtained can represent highly refined content information of the multimedia resource from different angles.
[0118] Step S208: input each attribute feature vector into each task sub-network to obtain a prediction label corresponding to each task.
[0119] Specifically, after obtaining the attribute feature vectors output by each feature subnetwork, the computer device can input each attribute feature vector into each task subnetwork together, and each task subnetwork performs data processing on the input data to obtain the corresponding prediction label. For example, the multimedia resource classification model includes three feature subnetworks and two task subnetworks, feature subnetwork 1 outputs attribute feature vector 1, feature subnetwork 2 outputs attribute feature vector 2, and feature subnetwork 3 outputs attribute feature vector 3. Attribute feature vectors 1 to 3 can be input into task subnetwork 1 to obtain prediction label 1, and attribute feature vectors 1 to 3 can be input into task subnetwork 2 to obtain prediction label 2. Of course, in order to improve learning efficiency, the computer device can also input each attribute feature vector into the corresponding task subnetwork, and each task subnetwork performs data processing on the input data to obtain the corresponding task prediction result. For example, attribute feature vector 1 and attribute feature vector 2 can be input into task subnetwork 1 to obtain prediction label 1, and attribute feature vector 2 and attribute feature vector 3 can be input into task subnetwork 2 to obtain prediction label 2.
[0120] In one embodiment, each task subnetwork is used to predict user feedback information of different categories of multimedia resources. For example, a task subnetwork for predicting the click-through rate of multimedia resources, a task subnetwork for predicting the browsing time of multimedia resources, a task subnetwork for predicting the review rate of multimedia resources, a task subnetwork for predicting the forwarding rate of multimedia resources, etc.
[0121] Step S210, based on the training labels and prediction labels corresponding to the same task, the parameters of the corresponding task sub-network are adjusted, and based on the training labels and prediction labels corresponding to each task, the model parameters of each feature sub-network are adjusted until the convergence conditions are met, thereby obtaining a trained multimedia resource classification model, which is used to classify the quality of the multimedia resources to be recommended.
[0122] Specifically, the computer device can calculate the training loss value based on the training label and the prediction label corresponding to the same task, obtain the training loss value corresponding to each task, perform back propagation based on each training loss value at the same time, adjust the model parameters of the multimedia resource classification model, until the convergence condition is met, and obtain the multimedia resource classification model that has been trained. When adjusting the model parameters, the parameters of the corresponding task subnetwork are adjusted based on the training label and the prediction label corresponding to the same task, and the model parameters of each feature subnetwork are adjusted based on the training label and the prediction label corresponding to each task. For example, the multimedia resource classification model includes a task subnetwork corresponding to the click-through rate task and a task subnetwork corresponding to the browsing time task. The model parameters of the task subnetwork corresponding to the click-through rate task are adjusted based on the training label and the prediction label corresponding to the click-through rate task, the model parameters of the task subnetwork corresponding to the browsing time task are adjusted based on the training label and the prediction label corresponding to the browsing time task, and the model parameters of each feature subnetwork are adjusted based on the training label and the prediction label corresponding to the click-through rate task and the training label and the prediction label corresponding to the browsing time task. Among them, the convergence condition can be that the number of iterations of the model reaches the iteration threshold, and the training loss values corresponding to each task subnetwork are all less than the preset threshold.
[0123] In one embodiment, the multimedia resource classification model is a multi-task model. The multi-task model is a machine learning model based on multi-task learning. Multi-task learning is a machine learning method that puts multiple related tasks together for learning based on shared representation. Multi-task learning is also an inductive transfer learning method. The main task uses the domain-related information possessed by the training signal of the related tasks as an inductive bias to improve the generalization performance of the main task. Multi-task learning involves multiple related tasks being learned in parallel at the same time, and the gradients are back-propagated at the same time. Multiple tasks help each other learn through the underlying shared representation to improve the generalization effect. The essence of multi-task learning is an inductive transfer mechanism that uses additional sources of information to improve the learning performance of the current task, including improving the generalization accuracy, learning rate, and understandability of the learned model. The multi-task model can be a hard-parameter sharing model, a MOE model (Mixture-of-Experts), a MMOE model (Multi-gateMixture-of-Experts), etc. The main task of the multimedia resource classification model is the quality classification of multimedia resources, that is, to obtain the quality classification results of multimedia resources. The related tasks of the main task of the multimedia resource classification model are the tasks corresponding to each task sub-network, for example, predicting the click rate of multimedia resources, predicting the browsing time of multimedia resources, etc.
[0124] The trained multimedia resource classification model can be used to classify the quality of multimedia resources to be recommended and identify high-quality multimedia resources. When recommending multimedia resources, you can first filter out obviously low-quality multimedia resources, and then use the multimedia resource classification model to identify high-quality multimedia resources, and give priority to recommending the identified high-quality multimedia resources to users, thereby improving the effectiveness of multimedia resource recommendations. When identifying high-quality multimedia resources through the multimedia resource classification model, the target attribute information set of the multimedia resources to be recommended is input into the multimedia resource classification model, and each task subnetwork in the multimedia resource classification model can output the corresponding prediction label. The multimedia resource classification model outputs the final quality classification result based on each prediction label. For example, when all the prediction labels are positive labels, the output quality classification result is a positive label. It can also be that when most of the prediction labels are positive labels, the output quality classification result is a positive label.
[0125] In the above-mentioned multimedia resource classification model training method, the multimedia resource classification model can be supervisedly trained based on the target attribute information set of the training multimedia resources and the training label set, and a multimedia resource classification model that can accurately classify the quality of the multimedia resources to be recommended is obtained. Among them, the target attribute information set of the multimedia resources includes target attribute information of multiple dimensions, and the target attribute information of different dimensions can reflect the content quality of the multimedia resources from different angles. The target attribute information set is input into the multimedia resource classification model, and the quality of the multimedia resources can be accurately classified by comprehensively considering the target attribute information of each dimension, and a prediction label that can accurately reflect the quality of the multimedia resources is obtained. In addition, the multimedia resource classification model includes multiple task subnetworks, which is a multi-task model that can predict the performance of multimedia resources on various tasks. During model training, multiple related tasks are learned in parallel at the same time, and the gradients are back-propagated at the same time to learn the connection and difference between different tasks, thereby improving the learning efficiency and quality of each task. Finally, the trained multimedia resource classification model can be used to classify the quality of the multimedia resources to be recommended, so that multimedia resources with better quality can be recommended to users, and the effectiveness of multimedia resource recommendation can be improved. Effective multimedia resource recommendations can avoid users from repeatedly searching or refreshing the interface due to low-quality and invalid multimedia resource recommendations. Repeated searches or repeated refreshes of the interface will take up a lot of computer device resources. Therefore, on the basis of improving the effectiveness of resource recommendations, it can also reduce the waste of terminal or server resources.
[0126] In one embodiment, Figure 3 As shown, a training label set of training multimedia resources is obtained, including:
[0127] Step S302: obtaining recommendation interaction information sets corresponding to a plurality of historical multimedia resources respectively; the recommendation interaction information set includes recommendation interaction information corresponding to each task.
[0128] Among them, historical multimedia resources refer to multimedia resources that have been published and recommended. Recommended interactive information refers to the information generated by the interaction between ordinary users and the users who publish multimedia resources after the multimedia resources are published, that is, the feedback information of ordinary users on multimedia resources. Interactive behavior can specifically be ordinary users browsing, liking, commenting, forwarding, etc. on the multimedia resources published by the publishing users. Recommended interactive information can specifically include information such as the click-through rate, browsing time, comment rate, like rate, dislike rate, and forwarding rate of multimedia resources. It can be understood that the aforementioned recommended interactive information can be considered as recommended interactive information corresponding to different tasks.
[0129] Specifically, the computer device may obtain multiple historical multimedia resources and obtain recommended interaction information sets corresponding to each historical multimedia resource in the same time period. Each recommended interaction information set includes recommended interaction information corresponding to multiple tasks.
[0130] Step S304: Count each recommended interaction information corresponding to the same task to obtain reference interaction information corresponding to each task.
[0131] The reference interaction information refers to the statistical results of multiple recommended interaction information corresponding to the same task, which can reflect the average level of multiple recommended interaction information corresponding to the same task, for example, the average value of each recommended interaction information corresponding to the same task, or the median value of each recommended interaction information corresponding to the same task.
[0132] Specifically, the computer device can obtain the recommended interaction information corresponding to the same task from each set of recommended interaction information, and collect statistics on the recommended interaction information corresponding to the same task to obtain reference interaction information corresponding to each task. For example, the computer device obtains the click-through rate of each historical multimedia resource, calculates the average click-through rate, obtains the browsing time of each multimedia resource, and calculates the average browsing time.
[0133] Step S306 , classifying the quality of the historical multimedia resources based on the recommended interaction information and the corresponding reference interaction information, and obtaining a quality label set corresponding to each historical multimedia resource.
[0134] Specifically, in the recommended interaction information set corresponding to the same historical multimedia resource, the computer device can classify the quality of the historical multimedia resources under each task based on the comparison results of the recommended interaction information and the reference interaction information corresponding to the same task, and obtain the quality label set corresponding to the current historical multimedia resource, and then obtain the quality label set corresponding to each historical multimedia resource. For example, the recommended interaction information set includes click-through rate and browsing time. When the click-through rate of historical multimedia resource 1 is greater than the click-through rate mean, it is determined that the quality label corresponding to the click-through rate task of historical multimedia resource 1 is a positive label. When the browsing time of historical multimedia resource 1 is greater than the browsing time mean, it is determined that the quality label corresponding to the browsing time task of historical multimedia resource 1 is a positive label. The quality label set of historical multimedia resource 1 includes a positive label corresponding to the click-through rate task and a positive label corresponding to the browsing time task.
[0135] Step S308: obtaining training multimedia resources and corresponding training label sets based on historical multimedia resources and corresponding quality label sets.
[0136] Specifically, the computer device can determine the quality label sets corresponding to a large number of historical multimedia resources based on the recommended interaction information sets corresponding to a large number of historical multimedia resources. The computer device can select a part of the historical multimedia resources from the large number of historical multimedia resources as training multimedia resources, and use the training multimedia resources and the corresponding training label sets to train the multimedia resource classification model. Furthermore, the computer device can also select another part of the historical multimedia resources as verification multimedia resources, and use the verification multimedia resources to verify the classification accuracy of the trained multimedia resource classification model. If the classification accuracy is low, the computer device can update the multimedia resource classification model based on the latest relevant information of the multimedia resources.
[0137] In this embodiment, a statistical analysis is performed on a large number of historical multimedia resource recommendation interaction information sets, and a quality label set of historical multimedia resources is determined based on the statistical analysis results, thereby obtaining training data for a multimedia resource classification model. In this way, the quality of multimedia resources is classified based on user feedback information on multimedia resources, and multimedia resources that users are interested in and pay more attention to are determined as high-quality multimedia resources, so that the multimedia resource classification model finally trained can predict multimedia resources that users will be interested in and pay more attention to, and recommend these multimedia resources to users, which can improve the effectiveness of resource recommendation.
[0138] In one embodiment, Figure 4 As shown, the quality of historical multimedia resources is classified based on the recommended interaction information and the corresponding reference interaction information, and a set of quality labels corresponding to each historical multimedia resource is obtained, including:
[0139] Step S402 : comparing the recommended interactivity corresponding to the recommended interactive information of the same task with the reference interactivity corresponding to the reference interactive information in the recommended interactive information set corresponding to the same historical multimedia resource.
[0140] Among them, interaction refers to the normalized data that converts relevant interaction information into positive dimensions for data comparison. Recommended interaction refers to the interaction obtained by converting recommended interaction information. Reference interaction refers to the interaction obtained by converting reference interaction information.
[0141] Specifically, the recommended interactive information corresponding to each task is converted into the recommended interactive degree through a custom formula, and the reference interactive information corresponding to each task is converted into the reference interactive degree through a custom formula. For example, when the recommended interactive information is the click-through rate, it can be understood that the higher the click-through rate of the multimedia resource, the better the quality of the multimedia resource, and the more interested the user is in the multimedia resource. Then, the click-through rate can be directly converted into a percentage click score. In this way, the higher the click score of the multimedia resource, the better the quality of the multimedia resource, and the more interested the user is in the multimedia resource. Similarly, the average of the click-through rate can be converted into a percentage average click score. When the recommended interactive information is the click-down rate, it can be understood that the higher the click-down rate of the multimedia resource, the lower the quality of the multimedia resource, and the less interested the user is in the multimedia resource. Then, the click-down rate can be first converted into a percentage initial score, and the difference between 100 and the initial score is used as the click-down score. In this way, the higher the click-down score of the multimedia resource, the better the quality of the multimedia resource, and the more interested the user is in the multimedia resource. Furthermore, in the recommended interaction information set corresponding to the same historical multimedia resource, the computer device can directly compare the recommended interactivity corresponding to the recommended interaction information of the same task and the reference interactivity corresponding to the reference interaction information to obtain the quality label corresponding to each task.
[0142] Step S404: determine the quality label of the task corresponding to the recommended interaction information whose recommended interaction degree is greater than the reference interaction degree as a positive label.
[0143] Step S406 , determining the quality label of the task corresponding to the recommended interaction information whose recommended interaction degree is less than the reference interaction degree as a negative label.
[0144] Specifically, the computer device may determine the quality label of the task corresponding to the recommended interaction information whose recommended interaction degree is greater than the reference interaction degree as a positive label, and determine the quality label of the task corresponding to the recommended interaction information whose recommended interaction degree is less than the reference interaction degree as a negative label. For example, the recommended interaction information set of historical multimedia resources includes click-through rate, browsing time, and forwarding rate. When the click interactivity of historical multimedia resource 1 is greater than the reference click interactivity, the quality label corresponding to the click rate task of historical multimedia resource 1 is determined to be a positive label. When the browsing time interactivity of historical multimedia resource 1 is greater than the reference browsing time interactivity, the quality label corresponding to the browsing time task of historical multimedia resource 1 is determined to be a positive label. When the forwarding rate interactivity of historical multimedia resource 1 is less than the reference forwarding rate interactivity, the quality label corresponding to the forwarding rate task of historical multimedia resource 1 is determined to be a negative label. The quality label set of historical multimedia resource 1 includes a positive label corresponding to the click rate task, a positive label corresponding to the browsing time task, and a negative label corresponding to the forwarding rate task.
[0145] It is understandable that the statistical intervals corresponding to the labels of each quality level can also be set based on the reference interactivity. In a task, the label corresponding to the target statistical interval in which the recommended interactivity of the historical multimedia resource falls is used as the quality label corresponding to the historical multimedia resource under the task.
[0146] In this embodiment, by comparing the recommended interactivity corresponding to the recommended interaction information of the same task and the reference interactivity corresponding to the reference interaction information, the quality label corresponding to the historical multimedia resource under each task can be quickly determined.
[0147] In one embodiment, the feature subnetwork includes a text feature subnetwork, the target attribute information associated with the text feature subnetwork includes a plurality of text attribute information, and the text feature subnetwork includes data processing channels corresponding to each piece of text attribute information. Through each feature subnetwork in the multimedia resource classification model, the target attribute information associated with the feature subnetwork is vectorized to obtain the attribute feature vectors output by each feature subnetwork, including: through each data processing channel in the text feature subnetwork, the corresponding text attribute information is vectorized to obtain the text feature vectors output by each data processing channel; based on each text feature vector, the attribute feature vector output by the text feature subnetwork is obtained.
[0148] The text attribute information refers to the attribute information related to the text in the multimedia resource, such as the text title, text label, text body, etc. The text label can be the subject, category, keyword, etc. of the text.
[0149] Specifically, the multimedia resource classification model includes a text feature subnetwork, and the text feature subnetwork is used to process text attribute information. After the target attribute information set is input into the multimedia resource classification model, the text attribute information in the target attribute information set will be input into the text feature subnetwork. The text feature subnetwork includes data processing channels corresponding to each piece of text attribute information, and the text feature subnetwork can vectorize the corresponding text attribute information through each data processing channel to obtain a text feature vector. Furthermore, based on each text feature vector, the attribute feature vector output by the text feature subnetwork can be obtained, and specifically, each text feature vector can be spliced to obtain the corresponding attribute feature vector.
[0150] refer to Figure 5, the target attribute information associated with the text feature sub-network includes text title, text label and text body. The text feature sub-network can vectorize the text title through data processing channel one to obtain text feature vector one, vectorize the text label through data processing channel two to obtain text feature vector two, and vectorize the text body through data processing channel three to obtain text feature vector three, and then concatenate text feature vector one, text feature vector two and text feature vector three to obtain an attribute feature vector.
[0151] In this embodiment, different text attribute information is processed through different data processing channels, which can improve the specificity and accuracy of data processing, so that the attribute feature vectors obtained based on each accurate text feature vector can comprehensively and accurately represent the text information of the multimedia resource.
[0152] In one embodiment, the feature subnetwork includes an atomic feature subnetwork; through each feature subnetwork in the multimedia resource classification model, the target attribute information associated with the feature subnetwork is vectorized to obtain the attribute feature vector output by each feature subnetwork, including: through the atomic feature subnetwork, the target attribute information associated with the atomic feature subnetwork is subjected to feature cross processing to obtain at least one cross feature vector; based on each cross feature vector, the attribute feature vector output by the atomic feature subnetwork is obtained.
[0153] Specifically, the multimedia resource classification model includes an atomic feature subnetwork, which is used to process target attribute information of the atomic feature class. After the target attribute information set is input into the multimedia resource classification model, the target attribute information of the atomic feature class in the target attribute information set will be input into the atomic feature subnetwork. The atomic feature subnetwork can perform feature cross processing on each target attribute information to obtain at least one cross feature vector. Each target attribute information is cross-processed between each other to obtain a cross feature vector. The cross feature vector can reflect the correlation between the target attribute information.
[0154] In one embodiment, the target attribute information associated with the atomic feature sub-network includes at least two of user attribute information, image attribute information, language attribute information, and text statistical attribute information.
[0155] Specifically, the target attribute information associated with the atomic feature subnetwork includes various atomic features of multimedia resources, specifically including at least two of user attribute information, image attribute information, language attribute information and text statistical attribute information. User attribute information refers to the user attribute information of the user who publishes the multimedia resource, and can specifically be account-related information of the publishing user, such as account level, account verticality, account authority, etc. Image attribute information refers to image-related information in multimedia resources, such as the number of images, image clarity and beauty, etc. Language attribute information refers to information related to text language and wording in multimedia resources, such as article rhetoric (such as parallelism, metaphor number), ancient poetry citation, overall lexical diversity, syntactic diversity, etc. Text statistical attribute information refers to information obtained by statistically analyzing the text content in multimedia resources, such as text length and height, article title quality, article title and text matching degree, text typesetting, etc. Some atomic features can be directly obtained, and some atomic features can be obtained by statistics through corresponding software tools. It can be understood that multimedia resources published by users with advanced accounts are often more attractive and of higher quality. Multimedia resources with rich and exquisite pictures are often more attractive and of higher quality. Multimedia resources with gorgeous words are often more attractive and of higher quality. In addition to being able to characterize the quality of multimedia resources individually, various atomic features can also work together to maximize the quality of multimedia resources.
[0156] refer to Figure 6 , the target attribute information associated with the atomic feature sub-network includes user attribute information, image attribute information, language attribute information and text statistical attribute information. The atomic feature sub-network can perform feature cross processing between each target attribute information to obtain multiple cross feature vectors, and splice each cross feature vector to obtain an attribute feature vector.
[0157] In this embodiment, by performing feature cross processing on target attribute information pairwise, the combined features between target attribute information can be effectively learned, so that the attribute feature vectors obtained based on each cross feature vector can comprehensively and accurately characterize the atomic features of the multimedia resource.
[0158] In one embodiment, the feature subnetwork includes a text-image fusion feature subnetwork, the target attribute information associated with the text-image fusion feature subnetwork includes text attribute information and image attribute information, and the text-image fusion feature subnetwork includes a text data processing channel corresponding to the text attribute information and an image data processing channel corresponding to the image attribute information. Through each feature subnetwork in the multimedia resource classification model, the target attribute information associated with the feature subnetwork is vectorized to obtain the attribute feature vector output by each feature subnetwork, including: encoding the text attribute information through the text data processing channel to obtain an intermediate feature vector; encoding the image attribute information through the image data processing channel to obtain an image feature vector; performing attention allocation processing on the intermediate feature vector based on the image feature vector to obtain a first text-image fusion feature vector; performing attention allocation processing on the image feature vector based on the intermediate feature vector to obtain a second text-image fusion feature vector; and obtaining the attribute feature vector output by the text-image fusion feature subnetwork based on the first text-image fusion feature vector and the second text-image fusion feature vector.
[0159] Specifically, the multimedia resource classification model includes a text-image fusion feature subnetwork, which is used to process text attribute information and image attribute information. After the target attribute information set is input into the multimedia resource classification model, the text attribute information and image attribute information in the target attribute information set will be input into the text-image fusion feature subnetwork. The text-image fusion feature subnetwork includes a text data processing channel corresponding to the text attribute information and an image data processing channel corresponding to the image attribute information. The text data processing channel can encode the text attribute information to obtain an intermediate feature vector. The image data processing channel can encode the image attribute information to obtain an image feature vector.
[0160] Furthermore, images can affect text. Therefore, the computer device can perform attention allocation processing on the intermediate feature vector based on the image feature vector to obtain a first image-text fusion feature vector. The attention allocation processing on the intermediate feature vector is used to assign different attention weights to each sentence, and the attention weight of the sentence can reflect the importance of the sentence in the multimedia resource. The first image-text fusion feature vector is a feature vector obtained by weighted summing of the feature vectors corresponding to each sentence.
[0161] Text can also affect images. The computer device can perform attention allocation processing on the image feature vector based on the intermediate feature vector to obtain a second image-text fusion feature vector. Performing attention allocation processing on the image feature vector is used to assign different attention weights to each image. The attention weight of the image can reflect the importance of the image in the multimedia resource. The second image-text fusion feature vector is a feature vector obtained by weighted summing of the feature vectors corresponding to each image.
[0162] Finally, the computer device can obtain the attribute feature vector output by the image-text fusion feature subnetwork based on the first image-text fusion feature vector and the second image-text fusion feature vector. Specifically, the first image-text fusion feature vector and the second image-text fusion feature vector can be spliced to obtain the corresponding attribute feature vector.
[0163] The image-text fusion sub-network is a network based on image-text multimodal machine learning. Multimodal machine learning refers to the ability to process and understand multi-source modal information through machine learning methods. Single-modal representation learning is responsible for representing information as a numerical vector that can be processed by a computer or further abstracted into a higher-level feature vector, while multimodal representation learning refers to learning better feature representation by utilizing the complementarity between multiple modalities and eliminating the redundancy between modalities. Therefore, the image-text fusion sub-network can utilize the complementarity between images and texts, eliminate the redundancy between images and texts, and learn better feature representations.
[0164] In one embodiment, text attribute information is encoded through a text data processing channel to obtain an intermediate feature vector, including: performing word encoding on the text attribute information to obtain a word feature vector; performing sentence encoding on the word feature vector to obtain an intermediate feature vector.
[0165] Specifically, after receiving the text attribute information, the text data processing channel first performs word encoding processing on the text attribute information to obtain a word feature vector, that is, taking words as units, each word in the text of the multimedia resource text is encoded in turn to obtain a word feature vector. The word feature vector includes feature vectors corresponding to each word. Then, sentence encoding processing is performed on the sentence feature vector to obtain an intermediate feature vector, that is, taking sentences as units, each feature vector corresponding to each word is encoded in turn to obtain an intermediate feature vector. In this way, sentence expression can be obtained through word encoding processing, and text expression can be obtained through sentence encoding processing.
[0166] In one embodiment, sentence encoding is performed on a word feature vector to obtain an intermediate feature vector, including: performing attention allocation processing on the word feature vector to obtain a sentence feature vector, and performing sentence encoding processing on the sentence feature vector to obtain an intermediate feature vector. Specifically, for a sentence, the role and importance of each word in the sentence are different. Therefore, attention allocation processing can be performed on the word feature vector to obtain a sentence feature vector. Attention allocation processing on the word feature vector is used to allocate different attention weights to each word in a sentence, and the attention weight of the word can reflect the importance of the word in the sentence. Specifically, attention allocation processing can be performed on the word feature vector based on the target vector, and during model training, the parameters of the target vector are continuously adjusted to obtain the most suitable target vector. The sentence feature vector includes feature vectors corresponding to each sentence. Then, sentence encoding processing is performed on the sentence feature vector to obtain an intermediate feature vector, that is, the feature vectors corresponding to each word are encoded in sequence in units of sentences to obtain an intermediate feature vector.
[0167] In one embodiment, the text attribute information can be encoded by a transformer model to obtain an intermediate feature vector. The image attribute information can be encoded by a CNN model (convolutional neural network) to obtain an image feature vector.
[0168] refer to Figure 7 , the target attribute information associated with the image-text fusion feature subnetwork includes text attribute information and image attribute information. The image-text fusion feature subnetwork performs word encoding and sentence encoding on the text attribute information through transformer to obtain an intermediate feature vector. The image-text fusion feature subnetwork encodes the image attribute information through CNN to obtain an image feature vector. Next, based on the image feature vector, attention allocation is performed on the intermediate feature vector to obtain a first image-text fusion feature vector, and based on the intermediate feature vector, attention allocation is performed on the image feature vector to obtain a second image-text fusion feature vector. The first image-text fusion feature vector and the second image-text fusion feature vector are concatenated to obtain an attribute feature vector.
[0169] In this embodiment, the text attribute information and the image attribute information can be organically integrated through the text-image fusion feature sub-network, and the complementarity between text and image is utilized to eliminate the redundancy between text and image, so as to learn the feature representation of the text-image information that better characterizes the multimedia resources.
[0170] In one embodiment, the feature subnetwork includes a style feature subnetwork, and the target attribute information associated with the style feature subnetwork includes style attribute information; the style feature subnetwork includes a first data processing channel and a second data processing channel. Through each feature subnetwork in the multimedia resource classification model, the target attribute information associated with the feature subnetwork is vectorized to obtain the attribute feature vector output by each feature subnetwork, including: encoding the style attribute information through the first data processing channel to obtain an initial feature vector, performing attention allocation processing on the initial feature vector to obtain a first feature vector; performing convolution processing on the style attribute information through the second data processing channel to obtain a second feature vector; and obtaining the attribute feature vector output by the style feature subnetwork based on the first feature vector and the second feature vector.
[0171] The style attribute information refers to the attribute information related to the layout of information in the multimedia resource, that is, the layout information of the text and pictures of the multimedia resource. When the multimedia resource is an article containing pictures, the style attribute information can be composed of paragraphs and pictures arranged in order of appearance.
[0172] Specifically, the multimedia resource classification model includes a style feature subnetwork, and the style feature subnetwork is used to process style attribute information. After the target attribute information set is input into the multimedia resource classification model, the style attribute information in the target attribute information set is input into the style feature subnetwork. The style feature subnetwork includes a first data processing channel and a second data processing channel, and the first data processing channel and the second data processing channel are used to perform data processing on the style attribute information in different ways.
[0173] The style feature sub-network can encode the style attribute information through the first data processing channel to obtain an initial feature vector, and perform attention allocation processing on the initial feature vector to obtain a first feature vector. Specifically, each style sub-information in the style attribute information can be used as a unit, and each style sub-information can be encoded in turn to obtain an initial feature vector. The initial feature vector includes feature vectors corresponding to each style sub-information. Then, attention allocation is performed on the initial feature vector to obtain a first feature vector. Attention allocation processing is performed on the initial feature vector to allocate different attention weights to each style sub-information, and the attention weight of the style sub-information can reflect the importance of the style sub-information in the multimedia resource. For example, the style attribute information includes paragraph 1, paragraph 2, picture 1, and paragraph 3. The style attribute information is encoded to obtain an initial feature vector consisting of a feature vector corresponding to paragraph 1, a feature vector corresponding to paragraph 2, a feature vector corresponding to picture 1, and a feature vector corresponding to paragraph 3. Then, attention allocation processing is performed on each feature vector, and attention weights are allocated to each feature vector. Each feature vector and the corresponding attention weight are weighted and summed to obtain the first feature vector. The first data processing channel is mainly used to learn the features between each style sub-information in the style attribute information.
[0174] The style feature sub-network can perform convolution processing on the style attribute information through the second data processing channel to obtain a third feature vector. The second data processing channel is mainly used to learn the overall characteristics of the style attribute information.
[0175] Finally, the attribute feature vector output by the style feature subnetwork is obtained based on the first feature vector and the second feature vector. Specifically, the attribute feature vector can be obtained by concatenating the first feature vector and the second feature vector.
[0176] refer to Figure 8 , the target attribute information associated with the style feature sub-network includes style attribute information. The style feature sub-network encodes the style attribute information through LSTM (Long Short-Term Memory) to obtain an initial feature vector, performs attention allocation processing on the initial feature vector through attention to obtain a first feature vector, performs convolution processing on the style attribute information through CNN to obtain a second feature vector, and concatenates the first feature vector and the second feature vector to obtain an attribute feature vector.
[0177] In this embodiment, by performing different data processing on the style attribute information through different data processing channels, feature vectors representing the style attribute information from different angles can be obtained, thereby helping the model to learn knowledge related to the style.
[0178] In one embodiment, each task subnetwork includes an expert layer, a gating layer, and a fusion layer; each task subnetwork shares the expert layer. Each attribute feature vector is input into each task subnetwork to obtain a prediction label corresponding to each task, including: in the current task subnetwork, each attribute feature vector is processed by the expert layer to obtain a feature processing result, the feature processing result is weighted by the gating layer to obtain an intermediate processing result, and the intermediate processing result is fused by the fusion layer to obtain a prediction label of the task corresponding to the current task subnetwork.
[0179] Specifically, the multimedia resource classification model includes multiple task subnetworks, each of which includes an expert layer, a gating layer, and a fusion layer. After each feature subnetwork outputs an attribute feature vector, each attribute feature vector is input into each task subnetwork together. In the current task subnetwork, each attribute feature vector is feature processed by the expert layer to obtain a feature processing result, and the feature processing result is weighted by the gating layer to obtain an intermediate processing result. The intermediate processing result is fused by the fusion layer to obtain a predicted label of the task corresponding to the current task subnetwork. Each task subnetwork can output a predicted label.
[0180] In one embodiment, each task sub-network not only shares the expert layer but also shares the gating layer. Fig. 9A , the expert layer can be further divided into multiple expert sub-layers. This approach is more based on the idea of ensemble learning, that is, a single network of the same scale cannot effectively learn the common expressions between all tasks, but after dividing into multiple sub-networks, each sub-network can always learn some relevant and unique expressions in a certain task, and then the output of each expert sub-layer (Expert) is weighted by the output of the gating layer, and sent to each task sub-network's respective multi-layer full connection to learn the specific task better. In other words, the bottom multiple expert sub-layers learn different knowledge, and different expert sub-layers focus on different tasks. Some experts learn common patterns, and some experts learn independent patterns.
[0181] In one embodiment, each task sub-network only shares the expert layer. Fig. 9BCorresponding gating layers are set for different tasks. The advantage of this is that task-specific functions can be learned to balance shared expressions without adding a large number of new parameters, thereby modeling the relationship between tasks more clearly. The differences in tasks can be captured without significantly increasing the requirements of model parameters. On the one hand, the gating layer is lightweight, and the expert layer is shared by all tasks, so it has advantages in terms of computational complexity and parameter volume. On the other hand, compared with the gating layer shared by all tasks, each task uses a separate gating layer. The gating layer of each task achieves selective use of the expert sub-layer through different final output weights. The gating layers of different tasks can learn different patterns of combined feature processing results, so that the model takes into account the correlation and distinction between capturing different tasks.
[0182] The task prediction result can be expressed as: k =h k (f k (x)), g k (x) = softmax(W gk x) where y k represents the task prediction result of the kth task. k (x) represents the data processing of the fusion layer of the kth task. k (x) represents the data processing of the expert layer and the gating layer of the kth task. k (x) represents the data processing of the gated layer for the kth task. i represents the i-th expert sublayer. n represents the number of n expert sublayers. i (x) represents the data processing of the i-th expert sub-layer. k (x) i Represents the weight corresponding to the data processing result of the i-th expert sub-layer for the k-th task.
[0183] refer to Fig. 9C The multi-resource classification model can specifically include a task subnetwork corresponding to the click-through rate task and a task subnetwork corresponding to the browsing time task. The two task subnetworks share the expert layer, and the two task subnetworks respectively include their own gating layer and fusion layer. After processing the input data, the two task subnetworks can output the prediction label corresponding to the click-through rate task and the prediction label corresponding to the browsing time task respectively.
[0184] In one embodiment, Fig.10 As shown, the method also includes:
[0185] Step S1002, obtaining a target attribute information set and a verification tag set for verifying a multimedia resource; verifying that the multimedia resource is an updated recommended multimedia resource.
[0186] Step S1004 , inputting the target attribute information set of the verified multimedia resource into the trained multimedia resource classification model to obtain a predicted label set corresponding to the verified multimedia resource.
[0187] Step S1006 , calculating the classification accuracy based on the predicted label set and the verification label set corresponding to the verification multimedia resource.
[0188] Step S1008, when the classification accuracy is less than the accuracy threshold, the trained multimedia resource classification model is updated based on the prediction label set and the training label set corresponding to the verified multimedia resource to obtain an updated multimedia resource classification model.
[0189] The verification multimedia resource is a multimedia resource used to verify the accuracy of the multimedia resource classification model. The verification label refers to the quality label of the verification multimedia resource. The verification multimedia resource and the training multimedia resource are different multimedia resources. The verification multimedia resource is a newly released and newly recommended multimedia resource, that is, the verification multimedia resource is an updated recommended multimedia resource.
[0190] Specifically, the user's feedback evaluation of multimedia resources is subjective, and the user's subjective preference for multimedia resources will change in real time with the user group and the popularity of multimedia resources. Therefore, it is necessary to update the multimedia resource classification model to maintain the accuracy and adaptability of the multimedia resource classification model. The computer device can obtain the historical multimedia resources released recently as the verification multimedia resources, determine the verification label set of the verification multimedia resources based on the latest recommended interactive information set, and then perform a model test on the resource classification model based on the verification multimedia resources. If the test passes, there is no need to update the multimedia resource classification model. If the test fails, the multimedia resource classification model is updated based on the verification multimedia resources. When performing the model test, the computer device can obtain the target attribute information set of the verification multimedia resources, input the target attribute information set of the verification multimedia resources into the trained multimedia resource classification model, and each task subnetwork in the multimedia resource classification model respectively corresponds to the prediction label of each task. The computer device then calculates the classification accuracy of the multimedia resource classification model based on the prediction label set and the verification label set corresponding to the verification multimedia resources. When the classification accuracy is greater than the accuracy threshold, it indicates that the multimedia resource classification model still maintains a high accuracy and the model test passes. When the classification accuracy is less than the accuracy threshold, it indicates that the accuracy of the multimedia resource classification model is low and the model needs to be updated. The computer device can calculate the training loss value based on the predicted label set and the training label set corresponding to the verified multimedia resource, perform back propagation based on the training loss value, and update the model parameters of the trained multimedia resource classification model until the convergence condition is met to obtain the updated multimedia resource classification model. The updated multimedia resource classification model is suitable for the current recommendation environment and can identify the multimedia resources that the user is currently interested in.
[0191] After multimedia resources are released, their a posteriori consumption data is stored in the database of each resource service platform. The a posteriori consumption data is the recommended interaction information. Fig.11The computer device can obtain the a posteriori consumption data of historical multimedia resources from various databases, and select positive and negative samples based on the a posteriori consumption data to construct a training set and a verification set. A positive sample refers to a quality label corresponding to a historical multimedia resource in a certain task as a positive label, and a negative sample refers to a quality label corresponding to a historical multimedia resource in a certain task as a negative label. The computer device can select a part of the historical multimedia resources as a training set, train a multimedia resource classification model based on the training set, and obtain a trained multimedia resource classification model. The computer device can obtain another part of the historical multimedia resources as a verification set, verify the trained multimedia resource classification model based on the verification set, and calculate the classification accuracy of the trained multimedia resource classification model based on the verification set. When the classification accuracy is less than the accuracy threshold, the trained multimedia resource classification model can be updated based on the verification set. In addition, when determining the quality label of the multimedia resource, the computer device can determine the quality label of each historical multimedia resource in the training set based on the a posteriori consumption data of the training set, and determine the quality label of each historical multimedia resource in the verification set based on the a posteriori consumption data of the verification set, so as to avoid mutual interference between the training set and the verification set. It is understandable that the computer device may periodically update the latest multimedia resource classification model to ensure that the multimedia resource classification model always adapts to the current recommendation environment.
[0192] In one embodiment, when the overlap between the predicted label set and the verified label set of a multimedia resource is greater than the overlap threshold, it is determined that the multimedia resource classification model is accurate in classifying the quality of the multimedia resource and the model prediction is accurate. Then, the proportion of multimedia resources accurately predicted in the verification set can be counted to obtain the classification accuracy of the multimedia resource classification model.
[0193] In one embodiment, Fig.12 As shown, a multimedia resource recommendation method is provided, which is applied to Figure 1 The computer device in the example is used to illustrate, and the computer device can be the above Figure 1 The terminal 102 or the server 104 in FIG. Fig.12 ,The multimedia resource recommendation method includes the following steps:
[0194] Step S1202: obtaining a target attribute information set of the multimedia resource to be recommended; the target attribute information set includes target attribute information of multiple dimensions.
[0195] Step S1204, inputting the target attribute information set into the trained multimedia resource classification model; the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks.
[0196] Step S1206, through each feature sub-network in the multimedia resource classification model, vectorization processing is performed on the target attribute information associated with the feature sub-network respectively to obtain the attribute feature vector output by each feature sub-network.
[0197] Step S1208: input each attribute feature vector into each task sub-network to obtain the predicted label output by each task sub-network.
[0198] Step S1210: obtaining quality classification results corresponding to the multimedia resources to be recommended based on the prediction tags.
[0199] Step S1212: recommending multimedia resources to be recommended based on the quality classification result.
[0200] Specifically, in order to recommend better multimedia resources to users and improve the effectiveness of resource recommendations, before recommending multimedia resources, the computer device can use a trained multimedia resource classification model to identify high-quality multimedia resources from a large number of multimedia resources to be recommended, and then recommend the high-quality multimedia resources to the user.
[0201] The computer device can perform content analysis on the multimedia resources to be recommended, obtain a target attribute information set of the multimedia resources to be recommended, input the target attribute information set of the multimedia resources to be recommended into a trained multimedia resource classification model, and obtain a quality classification result corresponding to the multimedia resources to be recommended.
[0202] The trained multimedia resource classification model includes multiple feature subnetworks and multiple task subnetworks. The feature subnetwork is used to convert the target attribute information into an attribute feature vector, and different feature subnetworks are used to process different target attribute information. Through each feature subnetwork in the multimedia resource classification model, the target attribute information associated with the feature subnetwork is vectorized, and the attribute feature vector output by each feature subnetwork can be obtained.
[0203] The task subnetwork is used to predict the performance of multimedia resources on specific tasks based on the attribute feature vector. Different task subnetworks correspond to different tasks. After obtaining the attribute feature vectors output by each feature subnetwork, the computer device can input each attribute feature vector into each task subnetwork. After each task subnetwork processes the input data, it obtains the corresponding prediction label.
[0204] The computer device can predict the quality classification results corresponding to the multimedia resources to be recommended based on the prediction results of each task. Specifically, when the number of task prediction results that are all positive labels is greater than a preset threshold, the quality classification result is determined to be a positive label. For example, when all task prediction results are positive labels, the quality classification result is determined to be a positive label. When more than half of the task prediction results are positive labels, the quality classification result is determined to be a positive label. The multimedia resource classification model ultimately outputs the quality classification results corresponding to the multimedia resources to be recommended. It can be understood that the multimedia resource classification model can also output each prediction label, and determine the quality classification results based on each prediction label outside the model. It can also be that the multimedia resource classification model outputs each prediction label and the quality classification result.
[0205] After obtaining the quality classification results of the multimedia resources to be recommended, the computer device can recommend weighted high-quality multimedia resources that have been identified and downgrade the weights of low-quality multimedia resources. This recommendation method can effectively recommend high-quality and attractive multimedia resources to users first, providing users with a good reading experience and improving the effectiveness of resource recommendations.
[0206] It can be understood that the specific process of training and updating the multimedia resource classification model can refer to the methods described in the various relevant embodiments of the aforementioned multimedia resource classification model training method, and the model structure and data processing process of the multimedia resource classification model can also refer to the methods described in the various relevant embodiments of the aforementioned multimedia resource classification model training method, which will not be repeated here.
[0207] The above-mentioned multimedia resource classification model training method obtains a target attribute information set of the multimedia resource to be recommended, the target attribute information set includes target attribute information of multiple dimensions, inputs the target attribute information set into the trained multimedia resource classification model, the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks, and through each feature sub-network in the multimedia resource classification model, respectively performs vectorization processing on the target attribute information associated with the feature sub-network to obtain the attribute feature vector output by each feature sub-network, inputs each attribute feature vector into each task sub-network to obtain the prediction label output by each task sub-network, obtains the quality classification result corresponding to the multimedia resource to be recommended based on each prediction label, and recommends the multimedia resource to be recommended based on the quality classification result. In this way, the target attribute information set of the multimedia resource includes target attribute information of multiple dimensions, and the target attribute information of different dimensions can reflect the content quality of the multimedia resource from different angles. The target attribute information set is input into the multimedia resource classification model, and the quality of the multimedia resource can be accurately classified by comprehensively considering the target attribute information of each dimension, so as to obtain an accurate quality classification result. In addition, the multimedia resource classification model includes multiple task sub-networks and is a multi-task model that can predict the performance of multimedia resources on various tasks. The quality classification results of multimedia resources are obtained by integrating the performance of multimedia resources on various tasks, which can further improve the accuracy of multimedia resource quality classification. Finally, multimedia resources with better quality can be identified through the multimedia resource classification model, so that multimedia resources with better quality can be recommended to users, improving the effectiveness of multimedia resource recommendation. Effective multimedia resource recommendation can avoid users from repeatedly searching or refreshing the interface due to low-quality and invalid multimedia resource recommendations. Repeated searches or repeated refreshes of the interface will occupy a large amount of computer device resources. Therefore, on the basis of improving the effectiveness of resource recommendation, it can also reduce the waste of resources of terminals or servers.
[0208] The present application also provides an application scenario, which applies the above-mentioned multimedia resource classification model training and multimedia resource recommendation method. Specifically, the above-mentioned method is applied in the application scenario of graphic content recommendation as follows:
[0209] refer to Fig.13A When recommending graphic and text content, the resource service platform can first filter out low-quality content and then identify high-quality content. After the high-quality content is released from the library, it will make weighted recommendations for the high-quality content, thereby effectively giving priority to high-quality content with high tone and attractiveness to users, bringing users a good reading experience and improving the effectiveness of recommendations.
[0210] 1. Filter out low-quality content
[0211] 2. Identify quality content
[0212] The methods for identifying high-quality content include a priori quality and a posteriori quality. A priori quality refers to identifying high-quality content from objective perspectives such as text quality, image quality, and text and image layout based on the content of the image and text. A posteriori quality refers to identifying high-quality content from both objective and subjective perspectives, not only based on the content of the image and text, but also by further considering the user's evaluation of the image and text content.
[0213] The server can train models for high-quality image and text posterior recognition. First, construct features from various content dimensions such as image and text multimodality, article layout, account, linguistics, etc. to complete deep network modeling and establish a multi-task-based image and text content classification model. The model structure of the image and text content classification model can be referenced Fig. 13B The model includes text feature subnetwork, atomic feature subnetwork, image-text fusion feature subnetwork and style feature subnetwork. The data output by each feature subnetwork is then output to each task subnetwork. Each task subnetwork outputs the task prediction result. Based on each task prediction result, the quality classification result corresponding to the input data is obtained. Each feature subnetwork and task subnetwork can be connected through the MLP layer. Fig. 13B The model includes a task subnetwork for predicting click-through rate and a task subnetwork for predicting browsing time, that is, a task subnetwork corresponding to the click-through rate task and a task subnetwork corresponding to the browsing time task.
[0214] Next, we use the user's a posteriori consumption data to drive the understanding of high-quality content. That is, we use the a posteriori consumption data to screen out high-quality positive and negative samples of pictures and texts, and establish training sets and validation sets. Fig. 13C When users browse information in the News Highlights applet, they can like, dislike, comment, click to browse the text of the information article, and other operations. The background will compile statistics on these operation data to obtain a posteriori consumption data, such as click-through rate, like rate, browsing time, etc.
[0215] Furthermore, the model is trained based on the training set to obtain a trained graphic content classification model. The trained graphic content classification model can be used to classify the quality of the graphic content to be recommended. The target attribute information set of the graphic content to be recommended is input into the graphic content classification model to obtain the quality classification result corresponding to the graphic content to be recommended.
[0216] 3. Recommend high-quality content
[0217] When the quality classification result indicates that the graphic content to be recommended is high-quality content, the high-quality content is weightedly recommended.
[0218] After the model was launched, the accuracy rate reached 95%. After conducting a weighted recommendation experiment on the browser side for the identified high-quality text and picture content, it was achieved that high-quality content with good reading experience and attractiveness was recommended to users first, and good business results were achieved on the business side. The overall text and picture clicks on the browser side increased by 0.946%, the total browsing time of text and picture increased by 1.007%, the text and picture CTR increased by 0.729%, and the average number of comments per person in the interactive indicator data increased by 0.416%.
[0219] Furthermore, after the model is trained, it can be tested based on the validation set from time to time. If the test results show that the model accuracy is low, the model is updated to ensure the accuracy of the model. In addition, we also tested and analyzed the model update time, and the test results are shown in Table 1. According to the test results, the automatic update cycle of the model can be set to 5 days to maintain a high accuracy of the model.
[0220] Decay Time Quality ratio decrease AUC decay 5 days 0.452% 0.031% 7 days 0.831% 0.126% 10 days 1.856% 0.705%
[0221] Table 1
[0222] In this embodiment, the model automatic update scheme for posterior quality identification of pictures and texts based on multi-task and multi-modality of pictures and texts, article layout, account, linguistics, etc. is an innovation in algorithm and model structure based on specific business scenarios. The user's posterior consumption data is used to drive the understanding of high-quality content, that is, the posterior consumption data is used to screen high-quality positive and negative samples of pictures and texts, and features are constructed from various content dimensions such as multi-modality of pictures and texts, article layout, account, linguistics, etc. to complete deep network modeling, and the model automatic update scheme is used to continuously capture the latest consumption content rules, to a certain extent optimize the problem of real-time changes in the subjective consumption preferences of user groups over time, and improve the effectiveness of resource recommendations.
[0223] It should be understood that although Figure 2-13B The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-13B At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0224] In one embodiment, Fig.14As shown, a multimedia resource classification model training device is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an information acquisition module 1402, an attribute information input module 1404, an attribute information processing module 1406, a task prediction module 1408, a label prediction module 1410 and a model adjustment module 1412, wherein:
[0225] The information acquisition module 1402 is used to acquire a target attribute information set and a training label set of training multimedia resources; the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks;
[0226] The attribute information input module 1404 is used to input the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained; the multimedia resource classification model includes multiple feature sub-networks and task sub-networks corresponding to each task;
[0227] The attribute information processing module 1406 is used to perform vectorization processing on the target attribute information associated with each feature sub-network in the multimedia resource classification model, and obtain the attribute feature vector output by each feature sub-network;
[0228] The label prediction module 1408 is used to input each attribute feature vector into each task sub-network to obtain the predicted label corresponding to each task;
[0229] The model adjustment module 1410 is used to adjust the parameters of the corresponding task sub-network according to the training labels and prediction labels corresponding to the same task, and adjust the model parameters of each feature sub-network based on the training labels and prediction labels corresponding to each task until the convergence conditions are met, thereby obtaining a trained multimedia resource classification model; the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended.
[0230] In one embodiment, the information acquisition module is also used to obtain recommended interaction information sets corresponding to multiple historical multimedia resources; the recommended interaction information sets include recommended interaction information corresponding to each task; the recommended interaction information of the same task is counted to obtain reference interaction information corresponding to each task; the quality of historical multimedia resources is classified based on the recommended interaction information and the corresponding reference interaction information to obtain a quality label set corresponding to each historical multimedia resource; training multimedia resources and a corresponding training label set are obtained based on the historical multimedia resources and the corresponding quality label set.
[0231] In one embodiment, the information acquisition module is also used to compare the recommended interactivity corresponding to the recommended interaction information of the same task and the reference interactivity corresponding to the reference interaction information in the recommended interaction information set corresponding to the same historical multimedia resource; determine the quality label of the task corresponding to the recommended interaction information whose recommended interactivity is greater than the reference interactivity as a positive label; and determine the quality label of the task corresponding to the recommended interaction information whose recommended interactivity is less than the reference interactivity as a negative label.
[0232] In one embodiment, the feature subnetwork includes a text feature subnetwork, the target attribute information associated with the text feature subnetwork includes a plurality of text attribute information, and the text feature subnetwork includes data processing channels corresponding to each piece of text attribute information. The attribute information processing module is further used to perform vectorization processing on the corresponding text attribute information through each data processing channel in the text feature subnetwork, and obtain the text feature vector output by each data processing channel; and obtain the attribute feature vector output by the text feature subnetwork based on each text feature vector.
[0233] In one embodiment, the feature subnetwork includes an atomic feature subnetwork. The attribute information processing module is further used to perform feature cross processing on target attribute information associated with the atomic feature subnetwork through the atomic feature subnetwork to obtain at least one cross feature vector; and obtain the attribute feature vector output by the atomic feature subnetwork based on each cross feature vector.
[0234] In one embodiment, the target attribute information associated with the atomic feature sub-network includes at least two of user attribute information, image attribute information, language attribute information, and text statistical attribute information.
[0235] In one embodiment, the feature subnetwork includes a text-image fusion feature subnetwork, the target attribute information associated with the text-image fusion feature subnetwork includes text attribute information and image attribute information, and the text-image fusion feature subnetwork includes a text data processing channel corresponding to the text attribute information and an image data processing channel corresponding to the image attribute information. The attribute information processing module is also used to encode the text attribute information through the text data processing channel to obtain an intermediate feature vector; encode the image attribute information through the image data processing channel to obtain an image feature vector; perform attention allocation processing on the intermediate feature vector based on the image feature vector to obtain a first text-image fusion feature vector; perform attention allocation processing on the image feature vector based on the intermediate feature vector to obtain a second text-image fusion feature vector; and obtain an attribute feature vector output by the text-image fusion feature subnetwork based on the first text-image fusion feature vector and the second text-image fusion feature vector.
[0236] In one embodiment, the attribute information processing module is further used to perform word encoding processing on the text attribute information to obtain a word feature vector; and perform sentence encoding processing on the word feature vector to obtain an intermediate feature vector.
[0237] In one embodiment, the feature subnetwork includes a style feature subnetwork, and the target attribute information associated with the style feature subnetwork includes style attribute information; the style feature subnetwork includes a first data processing channel and a second data processing channel. The attribute information processing module is further used to encode the style attribute information through the first data processing channel to obtain an initial feature vector, perform attention allocation processing on the initial feature vector to obtain a first feature vector; perform convolution processing on the style attribute information through the second data processing channel to obtain a second feature vector; and obtain an attribute feature vector output by the style feature subnetwork based on the first feature vector and the second feature vector.
[0238] In one embodiment, each task subnetwork includes an expert layer, a gating layer and a fusion layer; each task subnetwork shares the expert layer. The label prediction module is also used to perform feature processing on each attribute feature vector in the current task subnetwork through the expert layer to obtain a feature processing result, perform weighted processing on the feature processing result through the gating layer to obtain an intermediate processing result, and perform fusion processing on the intermediate processing result through the fusion layer to obtain a predicted label of the task corresponding to the current task subnetwork.
[0239] In one embodiment, Fig.15 As shown, the device also includes:
[0240] The model updating module 1412 is used to obtain a target attribute information set and a verification label set for verifying multimedia resources; verify that the multimedia resources are updated and recommended multimedia resources; input the target attribute information set for verifying multimedia resources into a trained multimedia resource classification model to obtain a predicted label set corresponding to the verified multimedia resources; calculate the classification accuracy based on the predicted label set and the verification label set corresponding to the verified multimedia resources; when the classification accuracy is less than the accuracy threshold, update the trained multimedia resource classification model based on the predicted label set and the training label set corresponding to the verified multimedia resources to obtain an updated multimedia resource classification model.
[0241] In one embodiment, Fig.16 As shown, a multimedia resource recommendation device is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an attribute information acquisition module 1602, an attribute information input module 1604, an attribute information processing module 1606, a task prediction module 1608, a quality classification module 1610 and a resource recommendation module 1612, wherein:
[0242] The attribute information acquisition module 1602 is used to acquire a target attribute information set of the multimedia resource to be recommended; the target attribute information set includes target attribute information of multiple dimensions;
[0243] The attribute information input module 1604 is used to input the target attribute information set into the trained multimedia resource classification model; the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks;
[0244] The attribute information processing module 1606 is used to perform vectorization processing on the target attribute information associated with each feature sub-network in the multimedia resource classification model, and obtain the attribute feature vector output by each feature sub-network;
[0245] The label prediction module 1608 is used to input each attribute feature vector into each task sub-network to obtain the predicted label output by each task sub-network;
[0246] The quality classification module 1610 is used to obtain the quality classification results corresponding to the multimedia resources to be recommended based on each prediction tag;
[0247] The resource recommendation module 1612 is used to recommend multimedia resources to be recommended based on the quality classification results.
[0248] For the specific limitations on the multimedia resource classification model training device and the multimedia resource recommendation device, please refer to the limitations on the multimedia resource classification model training method and the multimedia resource recommendation method in the above text, which will not be repeated here. The various modules in the above-mentioned multimedia resource classification model training device and the multimedia resource recommendation device can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0249] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.17As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as a recommended interactive information set of multimedia resources, a target attribute information set, a quality label, a multimedia resource classification model, etc. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multimedia resource classification model training method and a multimedia resource recommendation method are implemented.
[0250] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig.18 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a multimedia resource classification model training method and a multimedia resource recommendation method are implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0251] Those skilled in the art will understand that Fig.17 , 18 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0252] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0253] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0254] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the above-mentioned method embodiments.
[0255] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0256] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0257] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A multimedia resource classification model training method, characterized in that: The method comprises: Acquire a target attribute information set and a training label set of training multimedia resources; the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks; Inputting the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained; the multimedia resource classification model includes a plurality of feature sub-networks and task sub-networks corresponding to each task; Through each feature sub-network in the multimedia resource classification model, target attribute information associated with the feature sub-network is vectorized to obtain attribute feature vectors output by each feature sub-network; Input each attribute feature vector into each task sub-network to obtain the prediction label corresponding to each task; Adjusting the parameters of the corresponding task subnetwork based on the training labels and prediction labels corresponding to the same task, and adjusting the model parameters of each feature subnetwork based on the training labels and prediction labels corresponding to each task, until the convergence condition is met, and obtaining a trained multimedia resource classification model; the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended; Wherein, each of the task sub-networks includes an expert layer, a gating layer and a fusion layer; each of the task sub-networks shares the expert layer; The method of inputting each attribute feature vector into each task sub-network to obtain the prediction label corresponding to each task includes: In the current task sub-network, the expert layer performs feature processing on each attribute feature vector to obtain a feature processing result, the gating layer performs weighted processing on the feature processing result to obtain an intermediate processing result, and the fusion layer performs fusion processing on the intermediate processing result to obtain a predicted label of the task corresponding to the current task sub-network.
2. The method according to claim 1, characterized in that: The step of obtaining a training label set of a training multimedia resource includes: Acquire a set of recommended interaction information corresponding to a plurality of historical multimedia resources respectively; the set of recommended interaction information includes recommended interaction information corresponding to each task; Count the recommended interaction information corresponding to the same task, and obtain the reference interaction information corresponding to each task; Classify the quality of historical multimedia resources based on the recommended interaction information and the corresponding reference interaction information to obtain a set of quality labels corresponding to each historical multimedia resource; The training multimedia resources and the corresponding training label set are obtained based on the historical multimedia resources and the corresponding quality label set.
3. The method according to claim 2, characterized in that The quality of the historical multimedia resources is classified based on the recommended interaction information and the corresponding reference interaction information to obtain a set of quality labels corresponding to each historical multimedia resource, including: In a set of recommended interaction information corresponding to the same historical multimedia resource, comparing the recommended interaction degree corresponding to the recommended interaction information of the same task with the reference interaction degree corresponding to the reference interaction information; The quality label of the task corresponding to the recommended interaction information whose recommended interaction degree is greater than the reference interaction degree is determined as a positive label; The quality label of the task corresponding to the recommended interaction information whose recommended interaction degree is less than the reference interaction degree is determined as a negative label.
4. The method according to claim 1, characterized in that: The feature subnetwork includes a text feature subnetwork, the target attribute information associated with the text feature subnetwork includes a plurality of text attribute information, and the text feature subnetwork includes data processing channels corresponding to each piece of text attribute information; The method of performing vectorization processing on target attribute information associated with the feature subnetworks through each feature subnetwork in the multimedia resource classification model to obtain attribute feature vectors output by each feature subnetwork includes: Through each data processing channel in the text feature subnetwork, the corresponding text attribute information is vectorized to obtain the text feature vector output by each data processing channel; The attribute feature vector output by the text feature subnetwork is obtained based on each text feature vector.
5. The method according to claim 1, characterized in that The feature subnetwork includes an atomic feature subnetwork; The method of performing vectorization processing on target attribute information associated with the feature subnetworks through each feature subnetwork in the multimedia resource classification model to obtain attribute feature vectors output by each feature subnetwork includes: Performing feature cross processing on target attribute information associated with the atomic feature subnetwork through the atomic feature subnetwork to obtain at least one cross feature vector; The attribute feature vector output by the atomic feature sub-network is obtained based on each cross feature vector.
6. The method according to claim 5, characterized in that The target attribute information associated with the atomic feature sub-network includes at least two of user attribute information, image attribute information, language attribute information and text statistical attribute information.
7. The method according to claim 1, characterized in that The feature subnetwork includes a picture-text fusion feature subnetwork, the target attribute information associated with the picture-text fusion feature subnetwork includes text attribute information and image attribute information, and the picture-text fusion feature subnetwork includes a text data processing channel corresponding to the text attribute information and an image data processing channel corresponding to the image attribute information; The method of performing vectorization processing on target attribute information associated with the feature subnetworks through each feature subnetwork in the multimedia resource classification model to obtain attribute feature vectors output by each feature subnetwork includes: Encoding the text attribute information through the text data processing channel to obtain an intermediate feature vector; Encoding the image attribute information through the image data processing channel to obtain an image feature vector; Performing attention allocation processing on the intermediate feature vector based on the image feature vector to obtain a first image-text fusion feature vector; Performing attention allocation processing on the image feature vector based on the intermediate feature vector to obtain a second image-text fusion feature vector; An attribute feature vector output by the image-text fusion feature subnetwork is obtained based on the first image-text fusion feature vector and the second image-text fusion feature vector.
8. The method according to claim 7, characterized in that The encoding process of the text attribute information through the text data processing channel to obtain an intermediate feature vector includes: Performing word encoding processing on the text attribute information to obtain a word feature vector; Sentence encoding is performed on the word feature vector to obtain the intermediate feature vector.
9. The method according to claim 1, characterized in that: The feature subnetwork includes a style feature subnetwork, and the target attribute information associated with the style feature subnetwork includes style attribute information; the style feature subnetwork includes a first data processing channel and a second data processing channel; The method of performing vectorization processing on target attribute information associated with the feature subnetworks through each feature subnetwork in the multimedia resource classification model to obtain attribute feature vectors output by each feature subnetwork includes: Performing encoding processing on the style attribute information through the first data processing channel to obtain an initial feature vector, and performing attention allocation processing on the initial feature vector to obtain a first feature vector; Performing convolution processing on the style attribute information through the second data processing channel to obtain a second feature vector; An attribute feature vector output by the style feature subnetwork is obtained based on the first feature vector and the second feature vector.
10. The method according to any one of claims 1 to 9, characterized in that: The method further comprises: Acquire a target attribute information set and a verification tag set of a verification multimedia resource; the verification multimedia resource is an update-recommended multimedia resource; Inputting the target attribute information set of the verification multimedia resource into the trained multimedia resource classification model to obtain a prediction label set corresponding to the verification multimedia resource; Calculating the classification accuracy based on the predicted label set and the verification label set corresponding to the verification multimedia resource; When the classification accuracy is less than the accuracy threshold, the trained multimedia resource classification model is updated based on the prediction label set and the training label set corresponding to the verification multimedia resource to obtain an updated multimedia resource classification model.
11. A multimedia resource recommendation method, characterized in that: The method comprises: Acquire a target attribute information set of the multimedia resource to be recommended; the target attribute information set includes target attribute information of multiple dimensions; Inputting the target attribute information set into a trained multimedia resource classification model; the multimedia resource classification model includes multiple feature sub-networks and multiple task sub-networks; Through each feature sub-network in the multimedia resource classification model, target attribute information associated with the feature sub-network is vectorized to obtain attribute feature vectors output by each feature sub-network; Input each attribute feature vector into each task sub-network to obtain the predicted label output by each task sub-network; Obtaining quality classification results corresponding to the multimedia resources to be recommended based on the prediction tags; recommending the multimedia resource to be recommended based on the quality classification result; Wherein, each of the task sub-networks includes an expert layer, a gating layer and a fusion layer; each of the task sub-networks shares the expert layer; The method of inputting each attribute feature vector into each task sub-network to obtain the prediction label corresponding to each task includes: In the current task sub-network, the expert layer performs feature processing on each attribute feature vector to obtain a feature processing result, the gating layer performs weighted processing on the feature processing result to obtain an intermediate processing result, and the fusion layer performs fusion processing on the intermediate processing result to obtain a predicted label of the task corresponding to the current task sub-network.
12. A multimedia resource classification model training device, characterized in that: The device comprises: An information acquisition module, used to acquire a target attribute information set and a training label set of training multimedia resources; the target attribute information set includes target attribute information of multiple dimensions, and the training label set includes training labels corresponding to multiple tasks; An attribute information input module, used to input the target attribute information set of the training multimedia resources into the multimedia resource classification model to be trained; the multimedia resource classification model includes a plurality of feature sub-networks and task sub-networks corresponding to each task; An attribute information processing module, used to perform vectorization processing on target attribute information associated with each feature sub-network in the multimedia resource classification model, respectively, to obtain attribute feature vectors output by each feature sub-network; The label prediction module is used to input each attribute feature vector into each task sub-network to obtain the predicted label corresponding to each task; A model adjustment module, used to adjust the parameters of the corresponding task sub-network based on the training labels and prediction labels corresponding to the same task, and to adjust the model parameters of each feature sub-network based on the training labels and prediction labels corresponding to each task, until the convergence condition is met, thereby obtaining a trained multimedia resource classification model; the multimedia resource classification model is used to classify the quality of the multimedia resources to be recommended; Wherein, each of the task sub-networks includes an expert layer, a gating layer and a fusion layer; each of the task sub-networks shares the expert layer; The method of inputting each attribute feature vector into each task sub-network to obtain the prediction label corresponding to each task includes: In the current task sub-network, the expert layer performs feature processing on each attribute feature vector to obtain a feature processing result, the gating layer performs weighted processing on the feature processing result to obtain an intermediate processing result, and the fusion layer performs fusion processing on the intermediate processing result to obtain a predicted label of the task corresponding to the current task sub-network.
13. A multimedia resource recommendation device, characterized in that: The device comprises: An attribute information acquisition module is used to acquire a target attribute information set of the multimedia resource to be recommended; the target attribute information set includes target attribute information of multiple dimensions; An attribute information input module, used to input the target attribute information set into a trained multimedia resource classification model; the multimedia resource classification model includes a plurality of feature sub-networks and a plurality of task sub-networks; An attribute information processing module, used to perform vectorization processing on target attribute information associated with each feature sub-network in the multimedia resource classification model, respectively, to obtain attribute feature vectors output by each feature sub-network; The label prediction module is used to input each attribute feature vector into each task sub-network to obtain the predicted label output by each task sub-network; A quality classification module, used for obtaining a quality classification result corresponding to the multimedia resource to be recommended based on each prediction label; A resource recommendation module, configured to recommend the multimedia resource to be recommended based on the quality classification result; Wherein, each of the task sub-networks includes an expert layer, a gating layer and a fusion layer; each of the task sub-networks shares the expert layer; The method of inputting each attribute feature vector into each task sub-network to obtain the prediction label corresponding to each task includes: In the current task sub-network, the expert layer performs feature processing on each attribute feature vector to obtain a feature processing result, the gating layer performs weighted processing on the feature processing result to obtain an intermediate processing result, and the fusion layer performs fusion processing on the intermediate processing result to obtain a predicted label of the task corresponding to the current task sub-network.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Image retrieval model training method, image retrieval method and computer equipment
CN109685121A
Multimedia resource recommendation method and device, electronic equipment and storage medium
CN112131411A