Multimedia Information Recommendation Model Training Method, Recommendation Method and Device

By building a graph neural network of multimedia information recommendation model and adjusting the recall strategy, the accuracy problem of traditional multimedia information recommendation systems in cold startup environment is solved, and higher recommendation accuracy and user experience are achieved.

CN116932784BActive Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210322890.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-07-25
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

The traditional multimedia information recommendation system is not accurate enough in a cold startup environment, and the existing tag representation method is single, resulting in poor recommendation results and poor user experience.

Method used

A graph neural network with multimedia information recommendation model is constructed, and the recommendation algorithm is optimized by extracting training sample sets, adjusting the recall strategy, and using nodes and edges in the graph neural network to match the title content text and labels.

Benefits of technology

It improves the accuracy and relevance of multimedia information recommendation, enhances the generalization ability of the model, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116932784B_ABST
    Figure CN116932784B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, and electronic device for training a multimedia information recommendation model. The method includes: extracting a pre-training information set based on the historical data of the target object; constructing a graph neural network of the multimedia information recommendation model; training a first recommendation network in the multimedia information recommendation model through a training sample set to determine the network parameters of the first recommendation network; calculating pre-training vectors for each label in the training sample set; and training a second recommendation network in the multimedia information recommendation model based on the pre-training vectors of each label and the training sample set to determine the network parameters of the second recommendation network. Thereby, the accuracy and relevance of multimedia information recommendation are enhanced, and the generalization of the multimedia information recommendation model is improved. Embodiments of the present invention can also be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to information processing technologies, and in particular to a method for training a multimedia information recommendation model, a method for recommending multimedia information, an apparatus, and an electronic device. Background Art

[0002] Artificial Intelligence (AI) is a comprehensive technology in computer science. By studying the design principles and implementation methods of various intelligent machines, machines are enabled to have functions of perception, reasoning, and decision-making. The AI technology is an interdisciplinary subject, covering a wide range of fields, such as natural language processing technology and machine learning / deep learning. It is believed that with the development of technology, the AI technology will be applied in more fields and play an increasingly important role.

[0003] In traditional technologies, in the process of recommending corresponding multimedia information to users by various multimedia information recommendation systems, a collaborative filtering-based recommendation method can be used. As an effective recommendation method, collaborative filtering is widely applied in various multimedia information recommendation systems. In the traditional collaborative filtering-based recommendation method, when calculating the similarity relationship of multimedia information, the algorithm ideas based on neighborhood or matrix factorization are mainly adopted. In order to ensure the accuracy of the multimedia information similarity calculation, these two algorithms often perform multi-dimensional filtering processing on the original data of user behavior, and at the same time require that the multimedia information to be calculated can obtain sufficient rich user behavior. Therefore, the processing of multimedia information recommendation in a cold start environment by such algorithms is not accurate enough and the coverage is low. For the recommendation based on user interests, the user's historical behavior is used to establish the user's interest scores in specific categories and tags. When recalling, if the tag information in the news multimedia information hits the corresponding user tag interest, then the multimedia information is recalled. This solution usually uses the tag information in a relatively single way. Through such tags, multimedia information can be roughly classified into entertainment multimedia information, sports multimedia information, or further subdivided into sports game highlights, movie and TV drama tidbits, etc. However, such a representation method is relatively rough, and the classification tag information needs to be set in advance and updated in a timely manner, and its content representation ability is limited. The classification tag information needs to be set in advance and updated in a timely manner, and its content representation ability is limited, resulting in poor recommendation effects of multimedia information, the user losing the freshness of viewing, and seriously affecting the user experience. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method for training a multimedia information recommendation model, an apparatus, an electronic device, and a storage medium. The technical solution of the embodiments of the present invention is implemented as follows:

[0005] Embodiments of the present invention provide a method for training a multimedia information recommendation model, including:

[0006] Obtain the historical data of the target object in the multimedia information recommendation environment;

[0007] Extract the pre-training information set based on the historical data of the target object;

[0008] Construct the graph neural network of the multimedia information recommendation model based on the pre-training information set;

[0009] Extract the training sample set through the graph neural network;

[0010] Train the first recommendation network in the multimedia information recommendation model through the training sample set to determine the network parameters of the first recommendation network;

[0011] Calculate the pre-training vector of each label in the training sample set based on the first recommendation network;

[0012] Train the second recommendation network in the multimedia information recommendation model according to the pre-training vector of each label and the training sample set to determine the network parameters of the second recommendation network, so as to realize the adjustment of the recall strategy of the multimedia information through the second recommendation network, and perform multimedia information recommendation through the adjusted recall strategy.

[0013] The embodiment of the present invention also provides a method for obtaining the multimedia information to be recommended in the multimedia information data source;

[0014] Process different multimedia information to be recommended through the multimedia information recommendation model to determine the similarity of different multimedia information to be recommended;

[0015] Adjust the recall strategy of the multimedia information according to the similarity of different multimedia information to be recommended, and perform multimedia information recommendation through the adjusted recall strategy.

[0016] The embodiment of the present invention also provides a multimedia information recommendation model training device, including:

[0017] An information transmission module, configured to obtain the historical data of the target object in the multimedia information recommendation environment;

[0018] An information processing module, configured to extract the pre-training information set based on the historical data of the target object;

[0019] The information processing module is configured to construct the graph neural network of the multimedia information recommendation model based on the pre-training information set;

[0020] The information processing module is configured to extract the training sample set through the graph neural network;

[0021] The information processing module is configured to train a first recommendation network in the multimedia information recommendation model through the training sample set to determine network parameters of the first recommendation network;

[0022] The information processing module is configured to calculate a pre-training vector of each label in the training sample set based on the first recommendation network;

[0023] The information processing module is configured to train a second recommendation network in the multimedia information recommendation model according to the pre-training vector of each label and the training sample set to determine network parameters of the second recommendation network, so as to adjust a recall strategy for multimedia information through the second recommendation network, and perform multimedia information recommendation through the adjusted recall strategy.

[0024] In the above solution,

[0025] The information processing module is configured to extract the title content text of multimedia information based on historical data of the target object;

[0026] The information processing module is configured to extract a text label, a video label, and a channel classification label corresponding to the title content text;

[0027] The information processing module is configured to combine the text label, the video label, the channel classification label, and the title content text into a piece of pre-training information;

[0028] The information processing module is configured to combine at least two pieces of pre-training information into a pre-training information set.

[0029] In the above solution,

[0030] The information processing module is configured to use each text label, video label, channel classification label, and title content text in the pre-training information set as a node of the graph neural network, where the nodes of the text label, video label, and channel classification label are used as label nodes;

[0031] The information processing module is configured to traverse the nodes corresponding to each title content text in the pre-training information set, and determine the edges of the graph neural network when the title content text matches any one of the text label, video label, and channel classification label;

[0032] The information processing module is configured to determine the graph neural network based on the nodes of the graph neural network and different edges of the graph neural network.

[0033] In the above solution,

[0034] The information processing module is used to randomly extract the title content text corresponding to a node in the graph neural network;

[0035] The information processing module is used to obtain the label nodes that match the title content text;

[0036] The information processing module is used to combine the title content text and the content in the label nodes into a positive example training sample;

[0037] The information processing module is used to obtain the label nodes that do not match the title content text;

[0038] The information processing module is used to combine the title content text and the content in the label nodes into a negative example training sample;

[0039] The information processing module is used to combine the positive example training samples and the negative example training samples into a training sample set.

[0040] In the above solution,

[0041] The information processing module is used to determine the first multi-task loss function that matches the first recommendation network;

[0042] The information processing module is used to train the first recommendation network in the multimedia information recommendation model through the training sample set, and adjust the network parameters of the first recommendation network based on the first multi-task loss function;

[0043] The information processing module is used to determine the network parameters of the first recommendation network until the first multi-task loss function corresponding to the first recommendation network reaches the corresponding convergence condition.

[0044] In the above solution,

[0045] The information processing module is used to process the title content text in the training sample set to determine the title content text vector;

[0046] The information processing module is used to perform vector splicing processing on the title content text vector and the pre-trained vector of each label to obtain a spliced training sample;

[0047] The information processing module is used to determine the second multi-task loss function that matches the second recommendation network;

[0048] The information processing module is used to train the second recommendation network in the multimedia information recommendation model through the spliced training sample, and adjust the network parameters of the second recommendation network based on the second multi-task loss function;

[0049] The information processing module is configured to determine the network parameters of the second recommendation network until the second multi-task loss function corresponding to the second recommendation network reaches the corresponding convergence condition.

[0050] In the above solution,

[0051] The information processing module is configured to, when the multimedia information is a short video,

[0052] The information processing module is configured to send the exposure parameters during the playback of the short video to the detection server, so that the detection server can obtain the exposure parameters of the short video;

[0053] The information processing module is configured to use the exposure parameters as evaluation parameters for the playback effect of the multimedia information, and search for target exposure parameters according to the adjustment result of the recall strategy.

[0054] In the above solution,

[0055] The information processing module is configured to obtain the historical browsing information of the target object;

[0056] The information processing module is configured to determine the multimedia information exposure history corresponding to the historical browsing information based on the historical browsing information of the target object;

[0057] The information processing module is configured to dynamically adjust the recall strategy of the multimedia information based on the multimedia information exposure history corresponding to the historical browsing information.

[0058] In the above solution,

[0059] The information processing module is configured to determine the category of the multimedia information to be played according to the multimedia information recommendation environment;

[0060] The information processing module is configured to trigger a matching multimedia information data source in response to the category of the multimedia information to be played, so as to adjust the multimedia information to be played through the multimedia information data source matching the category of the multimedia information to be played.

[0061] An embodiment of the present invention further provides a multimedia information recommendation model training device, including:

[0062] A data transmission module, configured to obtain the multimedia information to be recommended in the multimedia information data source;

[0063] A data processing module, configured to process different multimedia information to be recommended through a multimedia information recommendation model to determine the similarity of different multimedia information to be recommended;

[0064] The data processing module is used to adjust according to the recall strategy of similar multimedia information for different multimedia information to be recommended, and perform multimedia information recommendation through the adjusted recall strategy.

[0065] An embodiment of the present invention further provides an electronic device, which includes:

[0066] A memory for storing executable instructions;

[0067] A processor, when running the executable instructions stored in the memory, implements the foregoing multimedia information recommendation model training method or the foregoing multimedia information recommendation method.

[0068] An embodiment of the present invention further provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the foregoing multimedia information recommendation model training method or the foregoing multimedia information recommendation method.

[0069] The embodiments of the present invention have the following beneficial effects:

[0070] The present invention obtains historical data of a target object in a multimedia information recommendation environment; extracts a pre-training information set based on the historical data of the target object; constructs a graph neural network of a multimedia information recommendation model based on the pre-training information set; extracts a training sample set through the graph neural network; trains a first recommendation network in the multimedia information recommendation model through the training sample set to determine network parameters of the first recommendation network; calculates pre-training vectors of each label in the training sample set based on the first recommendation network; trains a second recommendation network in the multimedia information recommendation model according to the pre-training vectors of each label and the training sample set to determine network parameters of the second recommendation network, so as to adjust the recall strategy of multimedia information through the second recommendation network and perform multimedia information recommendation through the adjusted recall strategy. Thus, it can be realized that the multimedia information recommendation model can recommend multimedia information in the usage environment to different users, while enhancing the accuracy and relevance of multimedia information recommendation, effectively improving the quality of multimedia information recommendation, and also completing model training with fewer samples, improving the generalization of the multimedia information recommendation model and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is a schematic diagram of the usage scenario of the multimedia information recommendation model training method provided by the embodiment of the present invention;

[0072] Figure 2 It is a schematic diagram of the composition structure of the multimedia information recommendation model training device provided by the embodiment of the present invention;

[0073] Figure 3 It is an optional flowchart of the method for training a multimedia information recommendation model provided by an embodiment of the present invention;

[0074] Figure 4 It is a schematic structural diagram of a graph neural network in an embodiment of the present invention;

[0075] Figure 5 It is a schematic model structure diagram of the first recommendation network in an embodiment of the present invention;

[0076] Figure 6 It is an optional flowchart of the method for training a multimedia information recommendation model provided by an embodiment of the present invention;

[0077] Figure 7 It is a schematic model structure diagram of the first recommendation network in an embodiment of the present invention;

[0078] Figure 8 It is a schematic application environment diagram of the multimedia information recommendation method based on the multimedia information recommendation model in an embodiment of the present invention;

[0079] Figure 9 It is a schematic process diagram of the multimedia information recommendation method in an embodiment of the present invention;

[0080] Figure 10 It is a schematic diagram of an optional multimedia information recommendation in an embodiment of the present invention. Specific embodiments

[0081] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0082] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0083] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.

[0084] 1) Responsive to, which is used to represent the conditions or states upon which the executed operations depend. When the dependent conditions or states are met, one or more of the executed operations can be real-time or can have a set delay; without special specification, there is no restriction on the execution order of multiple executed operations.

[0085] 2) Based on, which is used to represent the conditions or states upon which the executed operations depend. When the dependent conditions or states are met, one or more of the executed operations can be real-time or can have a set delay; without special specification, there is no restriction on the execution order of multiple executed operations.

[0086] 3) Model training involves performing multi-classification learning on an image dataset. This model can be constructed using deep learning frameworks such as TensorFlow or torch, and a multi-classification model is formed by combining multiple layers of neural network layers such as CNN. The input to the model is a three-channel or original-channel matrix formed by reading an image using tools such as openCV, and the output of the model is multi-classification probabilities. The judgment of the similarity of multimedia information is finally output through algorithms such as softmax. During training, the model approaches the correct trend through objective functions such as cross-entropy.

[0087] 4) Neural Network (NN): Artificial Neural Network (ANN), simply referred to as neural network or neural-like network, is a mathematical model or computational model that mimics the structure and function of a biological neural network (the central nervous system of an animal, especially the brain) in the fields of machine learning and cognitive science, and is used to estimate or approximate a function.

[0088] 5) Graph Neural Network (GNN): A neural network that directly acts on graph structures, mainly for processing data with non-Euclidean space structures (graph structures). It has the characteristics of ignoring the input order of nodes; during the calculation process, the representation of a node is affected by its surrounding neighbor nodes, while the connections of the graph itself remain unchanged; the representation of the graph structure enables graph-based reasoning. Generally, a graph neural network consists of two modules: a Propagation Module and an Output Module. The Propagation Module is used to transmit information between nodes in the graph and update the state, and the Output Module is used to define an objective function according to different tasks based on the vector representations of the nodes and edges of the graph. Graph neural networks include: Graph Convolutional Networks (GCNs), Gated Graph Neural Networks (GGNNs), and Graph Attention Networks (GAT) based on the attention mechanism.

[0089] 6) Recommendation accuracy: The recommended multimedia information content has a certain effect within a certain period of time, and the effect is measured by the user's interest in the content of the multimedia information. Accuracy plays an important role in user retention, clicks, and CTR on the edge side.

[0090] 7) softmax: A very commonly used and important function in machine learning, especially widely used in multi-classification scenarios. It maps some inputs to real numbers between 0 and 1 and normalizes to ensure the sum is 1.

[0091] 8) tag: A label, which is a type of keyword marking, and a group of words representing the core content of the document are extracted from the article text and title.

[0092] 9) Meta-learning, also known as Learning to learn, refers to the process of learning how to learn. Traditional machine learning problems are about learning a mathematical model for prediction from scratch, which is quite different from the process of humans learning and accumulating historical experience (also known as meta-knowledge) to guide new learning tasks. Meta-learning is the learning and training process of different machine learning tasks, as well as learning how to train a model faster and better.

[0093] 9) Bidirectional Encoder Representations from Transformers (BERT) is a bidirectional attention neural network model.

[0094] 10) Multimedia information, various forms of information available on the Internet, such as advertisement information, video files, multimedia information to be recommended, news information, etc. presented in a client or intelligent device.

[0095] Among them, the embodiments of the present invention can be implemented in combination with cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing. It can also be understood as the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. Therefore, cloud technology needs to be supported by cloud computing.

[0096] It should be noted that cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool platform will be established, abbreviated as a cloud platform, generally referred to as Infrastructure as a Service (IaaS). Various types of virtual resources are deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (which can be virtual machines, including operating systems), storage devices, and network devices.

[0097] Before introducing the multimedia information recommendation method provided by this application, first briefly describe the defects of multimedia information recommendation in related technologies. When recommending multimedia information in related technologies, the methods that can be adopted include:

[0098] Directly train a text matching model for the title text content of multimedia information. Specifically, it includes: first collecting texts in the target field to construct a training set to be labeled. Then this part of the data will be given to annotators for annotation to mark whether each sample pair matches. After having the training data, train a matching model, use the matching model to determine the video similarity, and determine whether to recommend according to the similarity. However, since the title text content usually has a short text length and obvious text features, it is very difficult for the model to be trained, resulting in a relatively poor recommendation effect of the model. Changing this result depends on a large number of labeled training samples, which is time-consuming and laborious. At the same time, the generalization of the model is affected, which is not conducive to the large-scale use of multimedia information recommendation.

[0099] Figure 1 This is a schematic diagram of the usage scenario of the multimedia information recommendation model training method provided by an embodiment of the present invention. Refer to Figure 1 , corresponding clients capable of playing implanted multimedia information are set on terminals (including terminal 10-1 and terminal 10-2). The terminals are connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two, and uses a wireless link to implement data transmission. Among them, the multimedia information includes, but is not limited to, videos, pictures, GIF animations, and advertising information. Among them, the types of multimedia information obtained by the terminals (including terminal 10-1 and terminal 10-2) from the corresponding server 200 through the network 300 can be the same or different. For example: the terminals (including terminal 10-1 and terminal 10-2) can obtain video advertisements placed by advertisers from the corresponding server 200 through the network 300, and can also obtain image advertisements placed by advertisers from the corresponding server 200 through the network 300. The specific types are not limited in this application. Different multimedia information can be stored in the server 200. Among them, the multimedia information as advertisements can be content in different dynamic formats, such as gif, mp4, mov, etc.

[0100] During the process that the terminals (terminal 10-1 and / or terminal 10-2) obtain and present the corresponding services with implanted multimedia information from the server 200 through the network 300, users can perform different operations on the multimedia information presented in the multimedia information playback window through the terminals (terminal 10-1 and / or terminal 10-2), generating different user behaviors. For example, when the multimedia information is a video advertisement, users can share and / or like the exposed short video during the process of watching the information, and can also click. When the multimedia information is a dynamic GIF advertisement, during the exposure process of the advertisement through the terminals (terminal 10-1 and / or terminal 10-2), users can forward and / or comment on the advertisement, and can also jump to the corresponding product purchase link page through the GIF advertisement.

[0101] As an example, when the server 200 determines which multimedia information to recommend to the user's terminal 10-1 or 10-2 for playback, it is necessary to timely adjust the multimedia information to be played, such as replacing any multimedia information in the multimedia information set to be played, so as to adapt to the viewing needs of different target objects. Taking short video multimedia information as an example, the multimedia information recommendation model provided by the present invention can be applied to short video playback. In short video playback, different short video multimedia information from different data sources are usually processed, and finally the corresponding different multimedia information and the corresponding short video recommendation process corresponding to the recommended video are presented on the user interface UI (User Interface). The accuracy and timeliness of the characteristics of different multimedia information directly affect the user experience. The background database of video playback receives a large amount of video data from different sources every day. The different multimedia information obtained for multimedia information recommendation to the target object can also be called by other applications (for example, the recommendation result of the short video recommendation process is migrated to the long video recommendation process or the news recommendation process). Of course, the multimedia information recommendation model matching the corresponding target object can also be migrated to different video recommendation processes (for example, the web video recommendation process, the mini program video recommendation process or the video recommendation process of the long video client).

[0102] As an example, the server 200 is used to deploy the corresponding multimedia information recommendation model to implement the multimedia information recommendation model training method provided by the present invention, or deploy the multimedia information recommendation model training device to implement the multimedia information recommendation model training method. Specifically, by obtaining the historical data of the target object in the multimedia information recommendation environment; based on the historical data of the target object, extracting the pre-training information set; based on the pre-training information set, constructing the graph neural network of the multimedia information recommendation model; through the graph neural network, extracting the training sample set; through the training sample set, training the first recommendation network in the multimedia information recommendation model to determine the network parameters of the first recommendation network; based on the first recommendation network, calculating the pre-training vectors of each label in the training sample set; according to the pre-training vectors of each label and the training sample set, training the second recommendation network in the multimedia information recommendation model to determine the network parameters of the second recommendation network, so as to realize adjusting the recall strategy of the multimedia information through the second recommendation network, and performing multimedia information recommendation through the adjusted recall strategy, and displaying and outputting the to-be-recommended multimedia information matching the target object through the terminal (terminal 10-1 and / or terminal 10-2). Taking short multimedia information as an example, the multimedia information recommendation model provided by the present invention can be applied to short video playback. In short video playback, different short multimedia information from different data sources are usually processed, and finally the to-be-recommended multimedia information corresponding to different multimedia information and the corresponding short video recommendation process is presented on the user interface UI (User Interface). The accuracy and timeliness of the features of different multimedia information directly affect the user experience. The background database of video playback receives a large amount of multimedia information data from different sources every day. The different multimedia information obtained for multimedia information recommendation to the target object can also be called by other application programs (for example, the recommendation result of the short video recommendation process is migrated to the recommendation process in the instant messaging client or the news recommendation process). Of course, the multimedia information recommendation model matching the corresponding target object can also be migrated to different video recommendation processes (such as the web video recommendation process, the mini-program video recommendation process, or the video recommendation process in the instant messaging client). The recommended short videos can meet the viewing needs of users.

[0103] Among them, the multimedia information recommendation model training method provided in the embodiments of the present application is implemented based on artificial intelligence. Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0104] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0105] In the embodiments of the present application, the artificial intelligence software technologies mainly involved include the above-mentioned speech processing technology and machine learning and other directions. For example, it may involve the automatic speech recognition (ASR) technology in speech technology, which includes speech signal preprocessing, speech signal frequency domain analysis, speech signal feature extraction, speech signal feature matching / recognition, speech training, etc.

[0106] For example, it may involve machine learning (ML). Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning usually includes technologies such as deep learning. Deep learning includes artificial neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and deep neural networks (DNNs).

[0107] It can be understood that the multimedia information recommendation model training method and speech processing provided in this application can be applied to intelligent devices. An intelligent device can be any device with an information display function. For example, it can be a smart terminal, a smart home device (such as a smart speaker, a smart washing machine, etc.), a smart wearable device (such as a smart watch), an in-vehicle intelligent central control system (which displays multimedia information to users through small programs that perform different tasks), or an AI intelligent medical device (which displays treatment cases by showing multimedia information).

[0108] The structure of the multimedia information recommendation model training device according to the embodiments of the present invention will be described in detail below. The multimedia information recommendation model training device can be implemented in various forms, such as a dedicated terminal with multimedia information recommendation processing functions, or a server provided with multimedia information recommendation model training device processing functions. For example, the Figure 1 server 200 in the foregoing. Figure 2 FIG. is a schematic diagram of the composition structure of the multimedia information recommendation model training device provided by the embodiments of the present invention. It can be understood that Figure 2 only the exemplary structure of the multimedia information recommendation model training device is shown, rather than all structures. According to needs, Figure 2 part of the structures shown or all structures can be implemented.

[0109] The multimedia information recommendation model training device provided by the embodiments of the present invention includes: at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. Each component in the multimedia information recommendation model training device is coupled together through a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 205.

[0110] Among them, the user interface 203 may include a display, a keyboard, a mouse, a trackball, a click wheel, a button, a touchpad, or a touch screen, etc.

[0111] It can be understood that the memory 202 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The memory 202 in the embodiments of the present invention can store data to support the operation of the terminal (such as 10-1). Examples of these data include: any computer programs for operating on the terminal (such as 10-1), such as an operating system and application programs. Among them, the operating system contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs may include various application programs.

[0112] In some embodiments, the multimedia information recommendation model training device provided by the embodiments of the present invention can be implemented in a combination of software and hardware. As an example, the multimedia information recommendation model training device provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the multimedia information recommendation model provided by the embodiments of the present invention. For example, a processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs, Application Specific Integrated Circuits), DSPs, programmable logic devices (PLDs, Programmable Logic Devices), complex programmable logic devices (CPLDs, Complex Programmable Logic Devices), field-programmable gate arrays (FPGAs, Field-Programmable Gate Arrays), or other electronic components.

[0113] As an example of the multimedia information recommendation model training apparatus provided by the embodiments of the present invention implemented by combining software and hardware, the multimedia information recommendation model training apparatus provided by the embodiments of the present invention can be directly embodied as a combination of software modules executed by the processor 201. The software modules can be located in a storage medium, and the storage medium is located in the memory 202. The processor 201 reads the executable instructions included in the software modules in the memory 202 and combines the necessary hardware (for example, including the processor 201 and other components connected to the bus 205) to complete the training method of the multimedia information recommendation model provided by the embodiments of the present invention.

[0114] As an example, the processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0115] As an example of the multimedia information recommendation model training apparatus provided by the embodiments of the present invention implemented by hardware, the apparatus provided by the embodiments of the present invention can be directly implemented by a processor 201 in the form of a hardware decoding processor. For example, it is implemented by one or more application-specific integrated circuits (ASIC, Application Specific Integrated Circuit), DSP, programmable logic devices (PLD, Programmable Logic Device), complex programmable logic devices (CPLD, Complex Programmable Logic Device), field-programmable gate arrays (FPGA, Field-Programmable Gate Array) or other electronic components to complete the training method of the multimedia information recommendation model provided by the embodiments of the present invention.

[0116] The memory 202 in the embodiments of the present invention is used to store various types of data to support the operation of the multimedia information recommendation model training apparatus. Examples of these data include: any executable instructions for operating on the multimedia information recommendation model training apparatus, such as executable instructions. The program for implementing the training method of the multimedia information recommendation model according to the embodiments of the present invention can be included in the executable instructions.

[0117] In some other embodiments, the multimedia information recommendation model training apparatus provided by the embodiments of the present invention can be implemented in a software manner. Figure 2The figure shows a multimedia information recommendation model training device stored in the memory 202, which can be software in the form of a program and a plug-in, etc., and includes a series of modules. As an example of the program stored in the memory 202, it can include a multimedia information recommendation model training device, and the multimedia information recommendation model training device includes the following software modules:

[0118] An information transmission module 2081 and an information processing module 2082. When the software modules in the multimedia information recommendation model training device are read into the RAM by the processor 201 and executed, the training method of the multimedia information recommendation model provided by the embodiments of the present invention will be implemented. Among them, the functions of each software module in the multimedia information recommendation model training device include:

[0119] The information transmission module 2081 is used to obtain the historical data of the target object in the multimedia information recommendation environment.

[0120] The information processing module 2082 is used to extract a pre-training information set based on the historical data of the target object.

[0121] The information processing module 2082 is used to construct a graph neural network of the multimedia information recommendation model based on the pre-training information set.

[0122] The information processing module 2082 is used to extract a training sample set through the graph neural network.

[0123] The information processing module 2082 is used to train the first recommendation network in the multimedia information recommendation model through the training sample set, and determine the network parameters of the first recommendation network.

[0124] The information processing module 2082 is used to calculate the pre-training vector of each label in the training sample set based on the first recommendation network.

[0125] The information processing module 2082 is used to train the second recommendation network in the multimedia information recommendation model according to the pre-training vector of each label and the training sample set, and determine the network parameters of the second recommendation network, so as to realize the adjustment of the recall strategy of multimedia information through the second recommendation network, and perform multimedia information recommendation through the adjusted recall strategy.

[0126] When the multimedia information recommendation model training is completed, it can be deployed in an electronic device to execute the multimedia information recommendation method provided by this application, which can specifically include:

[0127] A data transmission module, which is used to obtain the multimedia information to be recommended in the multimedia information data source.

[0128] A data processing module, configured to process different multimedia information to be recommended through a multimedia information recommendation model, and determine the similarity of different multimedia information to be recommended.

[0129] The data processing module is configured to adjust the recall strategy of multimedia information according to the similarity of different multimedia information to be recommended, and perform multimedia information recommendation through the adjusted recall strategy.

[0130] According to Figure 2 In one aspect of the present application, the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes different embodiments and combinations of embodiments provided in various alternative implementations of the above multimedia information recommendation model training method.

[0131] In combination with Figure 2 The multimedia information recommendation model training method provided by the embodiments of the present invention is described with reference to the multimedia information recommendation model training device shown in Figure 3 , Figure 3 FIG. is an optional flowchart of the multimedia information recommendation model training method provided by the embodiments of the present invention. It can be understood that Figure 3 The steps shown in FIG. can be executed by various electronic devices running the multimedia information recommendation model training device, such as a dedicated terminal, a server or a server cluster with the multimedia information recommendation model training device. Among them, the dedicated terminal with the multimedia information recommendation model training device can be the previous Figure 2 The electronic device with the multimedia information recommendation model training device shown in the embodiment. The following is an explanation of the Figure 3 steps shown in FIG.

[0132] Step 301: The multimedia information recommendation model training device obtains the historical data of the target object in the multimedia information recommendation environment.

[0133] Step 302: The multimedia information recommendation model training device extracts a pre-training information set based on the historical data of the target object.

[0134] In some embodiments of the present invention, extracting a pre-training information set based on the historical data of the target object can be achieved in the following manner:

[0135] Extract the title content text of the multimedia information based on the historical data of the target object; extract the text tags, video tags, and channel classification tags corresponding to the title content text; combine the text tags, video tags, channel classification tags, and the title content text into a pre-training information; combine at least two pre-training information into a pre-training information set. Taking the multimedia information as a short video as an example, referring to Table 1, it specifically includes:

[0136] 1) Text tags: These text tags are obtained from the title text. There are two ways to obtain text tags, namely: using the hashtags in the title information, that is, the tags with the "#" symbol, which are provided by short video users; or calling an existing keyword extraction service to extract keywords from the title text, such as "funny" shown in Table 1

[0137] 2) Video tags: They are the tags obtained by classifying with a video classification model. For example, the video tags can be obtained by classifying with a deep residual resnet50 model. The pre-trained convolutional neural network of the deep residual resnet50 is used for feature extraction, and the image information of the video is extracted as a 128-dimensional feature vector.

[0138] 3) Channel tags: They can be obtained by a text classification BERT model. The input of the BERT model is the title text feature of the video. The bidirectional attention neural network model BERT (Bidirectional Encoder Representation from Transformers) is used to send the video title sentence into the model task to obtain a 64-dimensional (the dimension size can be customized) title feature vector. By using the BERT model, the generalization ability of the word vector model is further enhanced, and the sentence-level representation ability is realized.

[0139]

[0140] Step 303: The multimedia information recommendation model training device constructs a graph neural network of the multimedia information recommendation model based on the pre-training information set.

[0141] In some embodiments of the present invention, constructing a graph neural network of the multimedia information recommendation model based on the pre-training information set can be achieved in the following ways:

[0142] Each text label, video label, channel classification label, and title content text in the pre-trained information set is used as a node of the graph neural network, where the nodes of the text label, video label, and channel classification label are used as label nodes; traverse the nodes corresponding to each title content text in the pre-trained information set, and when the title content text matches any one of the text label, video label, and channel classification label, determine the edge line of the graph neural network; based on the nodes of the graph neural network and different edge lines of the graph neural network, determine the graph neural network. Refer to Figure 4 , Figure 4 is a schematic structural diagram of the graph neural network in an embodiment of the present invention. Among them, the graph neural network (Graph Neural Network, GNN) is a neural network that directly acts on the graph structure, mainly for processing data of non-Euclidean space structure (graph structure). It has the characteristics of ignoring the input order of nodes; during the calculation process, the representation of nodes is affected by their surrounding neighbor nodes, while the connection of the graph itself remains unchanged; the representation of the graph structure enables graph-based reasoning. Usually, the graph neural network consists of two modules: a propagation module (Propagation Module) and an output module (Output Module). The propagation module is used to transfer information between nodes in the graph and update the state, and the output module is used to define the objective function according to different tasks based on the vector representations of the nodes and edges of the graph. There are graph convolutional neural networks (Graph Convolutional Networks, GCNs), gated graph neural networks (Gated Graph Neural Networks, GGNNs), and graph attention neural networks (Graph Attention Networks, GAT) based on the attention mechanism in the graph neural network. The advantage of predicting the target stock through the graph neural network is that based on the constructed graph network, each node in the graph network can automatically transfer all the feature (trend) information of the node to adjacent neighbor nodes. Through multiple information propagations between neighbors, each node in the graph network can contain the attribute information of the nodes directly or indirectly related to itself. Since in the graph network, the nodes with direct or indirect connections mostly have similar trends. When constructing the graph neural network, as long as a text title contains this label tag, an edge line of the graph network is established between the title and the tag until all the edges are established between the titles and the tags.

[0143] Step 304: The multimedia information recommendation model training device extracts the training sample set through the graph neural network.

[0144] In some embodiments of the present invention, the extraction of the training sample set through the graph neural network can be achieved in the following manner:

[0145] Randomly extract the title content text corresponding to a node in the graph neural network; obtain a label node that matches the title content text; combine the title content text and the content in the label node as a positive training sample; obtain a label node that does not match the title content text; combine the title content text and the content in the label node as a negative training sample; combine the positive training sample and the negative training sample into a training sample set. Among them, the proportion of negative training samples can be flexibly adjusted according to the model accuracy of the multimedia information recommendation model. As shown in Table 1, the positive sample can be (these two people are really funny, funny), and (these two people are really funny, star B

[0146] ), the negative samples can be (骑风破浪, 搞笑) and (骑风破浪, 电影混剪).

[0147] Step 305: The multimedia information recommendation model training device trains the first recommendation network in the multimedia information recommendation model through the training sample set to determine the network parameters of the first recommendation network.

[0148] In some embodiments of the present invention, training the first recommendation network in the multimedia information recommendation model by using the training sample set and determining the network parameters of the first recommendation network can be achieved in the following manner:

[0149] Determine a first multi-task loss function that matches the first recommendation network; train the first recommendation network in the multimedia information recommendation model through the training sample set, and adjust the network parameters of the first recommendation network based on the first multi-task loss function; until the first multi-task loss function corresponding to the first recommendation network reaches a corresponding convergence condition, determine the network parameters of the first recommendation network. Figure 5 , Figure 5 This is a schematic diagram of the model structure of the first recommendation network in an embodiment of the present invention, where the title of the multimedia information can be input on the left, and the label text and the corresponding label identifier tagid can be input on the right. The label text and the label identifier are both passed through a BERT model to obtain their corresponding representation vectors, and finally the scores between them are obtained through cosine similarity, and the loss function can be calculated. The first multi-task loss function can be expressed as Formula 1:

[0150]

[0151] Among them, in Formula 1, V1 and V2 are the vectors of the first word unit token of the BERT structure. The score S uses cosine similarity, and the loss function is the hinge loss function. Since Formula 1 includes positive and negative example samples, S+ is the score of the title and the positive example label, and S- is the score of the title and the negative example label.

[0152] Step 306: The multimedia information recommendation model training device calculates the pre-training vectors of each label in the training sample set based on the first recommendation network.

[0153] Step 307: The multimedia information recommendation model training device trains the second recommendation network in the multimedia information recommendation model according to the pre-training vectors of each label and the training sample set, and determines the network parameters of the second recommendation network, so as to adjust the recall strategy of multimedia information through the second recommendation network, and perform multimedia information recommendation through the adjusted recall strategy.

[0154] Combined with Figure 2 The multimedia information recommendation model training device shown in this invention embodiment provides a multimedia information recommendation model training method. Refer to Figure 6 , Figure 6 This is an optional flowchart of the multimedia information recommendation model training method provided by the embodiments of the present invention. It can be understood that Figure 6 The steps shown can be executed by various electronic devices running the multimedia information recommendation model training device. For example, it can be a dedicated terminal, a server or a server cluster with the multimedia information recommendation model training device. Among them, the dedicated terminal with the multimedia information recommendation model training device can be the Figure 2 electronic device with the multimedia information recommendation model training device shown in the previous Figure 6 embodiment. The following will describe the

[0155] Step 601: The multimedia information recommendation model training device processes the title content text in the training sample set to determine the title content text vector.

[0156] Step 602: The multimedia information recommendation model training device performs vector splicing processing on the title content text vector and the pre-training vectors of each label to obtain spliced training samples.

[0157] Of course, for the short video processing environment, the feature extractor ResNet can also be directly used to extract the video frame sequence into frame-level features. For example, the video frame image features of a short video can be extracted using a pre-trained convolutional neural network based on the deep residual resnet50, and the video frame image information of the short video can be extracted into a 2048-dimensional feature vector. Resnet is beneficial to the representation of the video frame image information of short videos in image feature extraction. The video frame image information of short videos has great eye-catching attraction before the user watches. A reasonable and appropriate short video frame image can well improve the play click-through rate of the video.

[0158] In some embodiments of the present invention, netvlad (Vector of locally aggregated descriptors) can also be used for feature extraction to generate a 128-dimensional feature vector from the video frame image. During video viewing, the video frame information reflects the specific content and quality of the video, which is directly related to the user's viewing duration. Among them, when configuring the multimedia information recommendation model in the video server, the acquisition method of the frame-level feature vector can be flexibly configured according to different usage requirements.

[0159] Step 603: The multimedia information recommendation model training device determines a second multi-task loss function that matches the second recommendation network.

[0160] Reference Figure 7 , Figure 7 is a schematic diagram of the model structure of the second recommendation network in the embodiments of the present invention. Among them, the title text information passes through a BERT model to obtain the feature vector of the title. Then it will be concatenated with the pre-trained vector, and then passed through a feed-forward neural network to obtain the final representation vector of the video. The feed-forward neural network (FNN) processes to generate object embedding vectors such as user embedding vectors. During the modeling process, social attributes will also be introduced, which can help the recommended multimedia information from the user's social relationships for recommendation ranking. After obtaining the representation vector, the second multi-task loss function can be calculated through Formula 2:

[0161]

[0162] Step 604: The multimedia information recommendation model training device trains the second recommendation network in the multimedia information recommendation model through the concatenated training samples, and adjusts the network parameters of the second recommendation network based on the second multi-task loss function.

[0163] Step 605: The multimedia information recommendation model training device determines the network parameters of the second recommendation network until the second multi-task loss function corresponding to the second recommendation network reaches the corresponding convergence condition.

[0164] In some embodiments of the present invention, when the multimedia information recommendation model is solidified in a corresponding hardware mechanism (such as a news reading terminal, an e-book terminal, a financial news terminal), and the usage environment is to push different news multimedia information to users through the news reading terminal or the e-book terminal, by fixing the fixed noise threshold corresponding to the multimedia information recommendation model, the training speed of the multimedia information recommendation model can be effectively improved, and the waiting time of users can be reduced. Among them, in the usage environment where the noise is fixed, the training sample set can come from the historical data of the target object. The historical recommended multimedia information browsing data can be the recommended multimedia information viewing behavior data generated when the recommended multimedia information was recommended to the target object, and can be extracted from the historical browsing log. Here, the historical recommended multimedia information browsing data can be all the historical recommended multimedia information browsing data; it can also consider the timeliness of the behavior data and only include the historical recommended multimedia information browsing data within a preset time period, such as the historical recommended multimedia information browsing data within a week and other different historical data.

[0165] The following takes the video recommendation scenario in the short video playback interface as an example to illustrate the multimedia information recommendation method provided by the embodiments of the present invention. Among them, Figure 8 is a schematic diagram of the application environment of the multimedia information recommendation method based on the multimedia information recommendation model in the embodiments of the present invention. Among them, as Figure 8 shown, the short video playback interface can be presented in the corresponding APP, or can be triggered through the instant messaging client applet (the multimedia information recommendation model can be encapsulated in the corresponding APP after training or saved in the instant messaging client applet in the form of a plug-in). As the short video application products continue to develop and increase, the carrying capacity of video information is much larger than that of text information. Short videos can be continuously recommended to users through the corresponding application programs. Therefore, recommending fresh short videos to users and avoiding repeated recommendations can keep users fresh, and effective subsequent relevant video recommendations can effectively improve the user experience. Among them, Figure 9 is a schematic diagram of the process of the multimedia information recommendation method in the embodiments of the present invention, including the following steps:

[0166] Step 901: Obtain the multimedia information to be recommended in the multimedia information data source.

[0167] Step 902: Process different multimedia information to be recommended through the multimedia information recommendation model to determine the similarity of different multimedia information to be recommended.

[0168] Step 903: Adjust the recall strategy for similar multimedia information of different multimedia information to be recommended, and perform multimedia information recommendation through the adjusted recall strategy.

[0169] In some embodiments of the present invention, refer to Figure 10 , Figure 10 is a schematic diagram of an optional multimedia information recommendation in an embodiment of the present invention, in which the category of the multimedia information to be played can be determined; in response to the category of the multimedia information to be played, a matching multimedia information data source is triggered. For example, when in use, it is determined that the category of the multimedia information to be played is advertising information, the target resources include different advertisements of the same advertiser, and different short video advertisement information included in different resource groups can be played in sequence in the short video playback windows of different advertisement positions (for example, advertisement position 1, advertisement position 2, and advertisement position 3 play three different advertisements of the same advertiser respectively), or when all different short video playback areas of the display interface are contracted by the advertiser, the advertising information is cyclically displayed, and the advertising information of the same advertiser can be cyclically presented in the short video playback windows of different advertisement positions of the advertising information display interface. At the same time, when the different short video advertisements of the advertiser are video advertisements, the audio volume carried by the video can be adjusted to the maximum in sequence to prompt the user to watch the played video advertisement. Replace advertisement A with advertisement B to allocate more playback traffic for advertisement B and enable users to obtain a better viewing experience. Specifically, when dynamically adjusting the recall strategy of the advertising information based on the traffic parameters and iterative experiment parameters matched by the recall strategy of the advertising information, the advertising exposure rate can be increased. In some embodiments of the present invention, the exposure channel of advertisement A can also be adjusted from the current exposure in the short multimedia information playback client to the contact status information of the instant messaging client for advertising placement. Of course, when adjusting the exposure position of advertisement A, it can be adjusted from the circle of friends advertisement of the instant messaging client to the splash screen advertisement to conform to different dynamically adjusted recall strategies, so that the short videos of different advertisement positions can be recommended to different users in a short time to obtain a better video recommendation effect. Take Figure 10For example, when it is determined that an advertisement with feature B has been shared in the historical browsing information of male target object 1, other advertisement information containing feature B (such as advertisements or short videos containing feature B) can be used to replace the current advertisement by dynamically adjusting the recall strategy. When it is determined that an advertisement with feature X has been clicked and purchased in the historical browsing information of female target object 2, other advertisement information (such as advertisements or short video advertisement links containing feature X) can be used to replace advertisement A by dynamically adjusting the recall strategy, ensuring that users obtain more fresh advertisement information (recommending unviewed advertisement information to different types of users respectively), enabling users to have a better usage experience, and at the same time increasing the click-through rate of advertisements to obtain better advertisement delivery effects.

[0170] Beneficial technical effects:

[0171] The present invention obtains historical data of a target object in a multimedia information recommendation environment; extracts a pre-training information set based on the historical data of the target object; constructs a graph neural network of a multimedia information recommendation model based on the pre-training information set; extracts a training sample set through the graph neural network; trains a first recommendation network in the multimedia information recommendation model through the training sample set to determine network parameters of the first recommendation network; calculates pre-training vectors of each label in the training sample set based on the first recommendation network; trains a second recommendation network in the multimedia information recommendation model based on the pre-training vectors of each label and the training sample set to determine network parameters of the second recommendation network, so as to realize adjusting the recall strategy of multimedia information through the second recommendation network and performing multimedia information recommendation through the adjusted recall strategy. Thus, it can be realized that the multimedia information recommendation model can recommend multimedia information in the usage environment to different users, while enhancing the accuracy and relevance of multimedia information recommendation, effectively improving the quality of multimedia information recommendation, and also being able to complete model training with fewer samples, improving the generalization of the multimedia information recommendation model and enhancing the user's usage experience.

[0172] As described above, the above are only embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for training a multimedia information recommendation model, characterized in that The method includes: Obtaining historical data of a target object in a multimedia information recommendation environment; Extracting a pre-training information set based on the historical data of the target object; Constructing a graph neural network of a multimedia information recommendation model based on the pre-training information set; Extracting a training sample set through the graph neural network; Training a first recommendation network in the multimedia information recommendation model through the training sample set to determine network parameters of the first recommendation network; Calculating pre-training vectors of each label in the training sample set based on the first recommendation network; Training a second recommendation network in the multimedia information recommendation model according to the pre-training vectors of each label and the training sample set to determine network parameters of the second recommendation network, so as to adjust a recall strategy for multimedia information through the second recommendation network, and perform multimedia information recommendation through the adjusted recall strategy.

2. The method according to claim 1, wherein The extracting a pre-training information set based on the historical data of the target object includes: Extracting title content text of multimedia information based on the historical data of the target object; Extracting text labels, video labels, and channel classification labels corresponding to the title content text; Combining the text labels, video labels, channel classification labels, and the title content text into a piece of pre-training information; Combining at least two pieces of pre-training information into a pre-training information set.

3. The method according to claim 2, wherein The constructing a graph neural network of a multimedia information recommendation model based on the pre-training information set includes: Taking each text label, video label, channel classification label, and title content text in the pre-training information set as a node of the graph neural network, where nodes of the text labels, video labels, and channel classification labels are used as label nodes; Traversing nodes corresponding to each title content text in the pre-training information set, and determining edges of the graph neural network when the title content text matches any one of the text labels, video labels, and channel classification labels; Determining the graph neural network based on nodes of the graph neural network and different edges of the graph neural network.

4. The method according to claim 1, wherein The extracting a training sample set through the graph neural network includes: Randomly extracting title content text corresponding to a node in the graph neural network; Obtaining label nodes that match the title content text; Combining the title content text and contents in the label nodes into a positive example training sample; Obtaining label nodes that do not match the title content text; Combining the title content text and contents in the label nodes into a negative example training sample; Combining the positive example training samples and the negative example training samples into a training sample set.

5. The method according to claim 1, characterized in that, The training a first recommendation network in the multimedia information recommendation model through the training sample set to determine network parameters of the first recommendation network includes: Determining a first multi-task loss function that matches the first recommendation network; Train the first recommendation network in the multimedia information recommendation model through the training sample set, and adjust the network parameters of the first recommendation network based on the first multi-task loss function; When the first multi-task loss function corresponding to the first recommendation network reaches the corresponding convergence condition, determine the network parameters of the first recommendation network.

6. The method according to claim 1, characterized in that The training of the second recommendation network in the multimedia information recommendation model according to the pre-trained vector of each tag and the training sample set, and determining the network parameters of the second recommendation network includes: Process the title content text in the training sample set to determine the title content text vector; Perform vector splicing processing on the title content text vector and the pre-trained vector of each tag to obtain a spliced training sample; Determine a second multi-task loss function that matches the second recommendation network; Train the second recommendation network in the multimedia information recommendation model through the spliced training sample, and adjust the network parameters of the second recommendation network based on the second multi-task loss function; When the second multi-task loss function corresponding to the second recommendation network reaches the corresponding convergence condition, determine the network parameters of the second recommendation network.

7. The method according to claim 6, wherein The method further includes: When the multimedia information is a short video, Send the exposure parameters during the playback of the short video to the detection server to enable the detection server to obtain the exposure parameters of the short video; Use the exposure parameters as evaluation parameters for the playback effect of the multimedia information, and search for target exposure parameters according to the adjustment result of the recall strategy.

8. The method according to claim 1, wherein The method further includes: Obtain the historical browsing information of the target object; Based on the historical browsing information of the target object, determine the multimedia information exposure history corresponding to the historical browsing information; Dynamically adjust the recall strategy of the multimedia information based on the multimedia information exposure history corresponding to the historical browsing information.

9. The method according to claim 1, wherein The method further includes: Determine the category of the multimedia information to be played according to the multimedia information recommendation environment; In response to the category of the multimedia information to be played, trigger a matching multimedia information data source to enable the multimedia information to be played to be adjusted by the multimedia information data source matching the category of the multimedia information to be played.

10. A method for recommending multimedia information, characterized in that, The method includes: Obtain the multimedia information to be recommended in the multimedia information data source; Process different multimedia information to be recommended through the multimedia information recommendation model to determine the similarity of different multimedia information to be recommended; Adjust the recall strategy of the multimedia information according to the similarity of different multimedia information to be recommended, and perform multimedia information recommendation through the adjusted recall strategy, where the multimedia information recommendation model is trained based on any one of claims 1-9.

11. A multimedia information recommendation model training device, characterized in that, The device includes: An information transmission module for obtaining the historical data of the target object in the multimedia information recommendation environment; An information processing module for extracting a pre-trained information set based on the historical data of the target object; The information processing module is configured to construct a graph neural network of a multimedia information recommendation model based on the pre-training information set; The information processing module is configured to extract a training sample set through the graph neural network; The information processing module is configured to train a first recommendation network in the multimedia information recommendation model through the training sample set to determine network parameters of the first recommendation network; The information processing module is configured to calculate a pre-training vector of each label in the training sample set based on the first recommendation network; The information processing module is configured to train a second recommendation network in the multimedia information recommendation model according to the pre-training vector of each label and the training sample set to determine network parameters of the second recommendation network, so as to adjust a recall policy for multimedia information through the second recommendation network, and perform multimedia information recommendation through the adjusted recall policy.

12. A multimedia information recommendation model training device, characterized in that, The apparatus includes: A data transmission module, configured to obtain multimedia information to be recommended in a multimedia information data source; A data processing module, configured to process different multimedia information to be recommended through a multimedia information recommendation model to determine the similarity of different multimedia information to be recommended; The data processing module is configured to adjust a recall policy for multimedia information according to the similarity of different multimedia information to be recommended, and perform multimedia information recommendation through the adjusted recall policy, where the multimedia information recommendation model is trained based on any one of claims 1-9.

13. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by a processor, it implements the multimedia information recommendation model training method according to any one of claims 1 to 9, or implements the multimedia information recommendation method according to claim 10.

14. An electronic device, characterized in that, The electronic device includes: A memory, configured to store executable instructions; A processor, configured to implement the multimedia information recommendation model training method according to any one of claims 1 to 9, or implement the multimedia information recommendation method according to claim 10 when running the executable instructions stored in the memory.

15. A computer-readable storage medium stores executable instructions, characterized in that, When the executable instructions are executed by a processor, they implement the multimedia information recommendation model training method according to any one of claims 1 to 9, or implement the multimedia information recommendation method according to claim 10.

Citation Information

Patent Citations

  • Context recommendation method based on graph neural network and attention mechanism

    CN110879864A

  • Multimedia information recommendation method and device, electronic device and storage medium

    CN112989074A