Multimedia information recommendation method and device, electronic equipment and storage medium

By integrating features and adjusting recall strategies in the multimedia information recommendation model, the problem of inaccurate recommendations caused by the sparsity of historical data was solved, resulting in higher-quality multimedia information recommendations and improved user experience.

CN115482021BActive Publication Date: 2025-11-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110605178.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-31
Publication Date
2025-11-25
Estimated Expiration
2041-05-31

AI Technical Summary

Technical Problem

Traditional multimedia information recommendation systems struggle to accurately model data when historical data is sparse, failing to capture complex relationships between tasks. This leads to increased noise, loss of hidden features, and negatively impacts recommendation accuracy and user experience.

Method used

A multimedia information recommendation model is adopted. Feature fusion processing is used to obtain user playback and sharing feature vectors, and the recall strategy is adjusted. A hybrid expert network and feature weight adjustment network are used, combined with context features and noise threshold processing, to optimize multimedia information recommendation.

Benefits of technology

It improves the accuracy and relevance of multimedia information recommendations, enhances user experience, and improves recommendation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115482021B_ABST
    Figure CN115482021B_ABST
Patent Text Reader

Abstract

The application provides a multimedia information recommendation method and device and electronic equipment. The method comprises the following steps: determining a first input feature of a multimedia information recommendation model; obtaining a user playing feature vector and a user sharing feature vector matched with the target object; determining a second input feature of the multimedia information recommendation model; obtaining a multimedia information playing feature vector and a multimedia information sharing feature vector matched with the multimedia information to be recommended; and adjusting a recall strategy of the multimedia information based on the user playing feature vector, the user sharing feature vector, the multimedia information playing feature vector and the multimedia information sharing feature vector. Thus, the multimedia information recommendation model can recommend multimedia information in a use environment to different users, the accuracy and relevance of multimedia information recommendation are enhanced, the quality of multimedia information recommendation is effectively improved, and the use experience of users is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to information processing technology, and more particularly to multimedia information recommendation methods, apparatus, and electronic devices. Background Technology

[0002] Artificial Intelligence (AI) is a comprehensive technology within computer science that studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. AI technology is a multidisciplinary field, encompassing a wide range of areas, including natural language processing and machine learning / deep learning. It is believed that with technological advancements, AI will be applied in more fields and play an increasingly important role.

[0003] In traditional technologies, when historical data is sparse, it is difficult to accurately model different objectives. Furthermore, since the recommendation network is shared by all tasks, it may fail to capture more complex relationships between tasks, thus introducing noise into some tasks and losing hidden features in the input data. This makes it impossible to accurately recommend multimedia information, and high-quality multimedia information that meets user needs cannot be effectively disseminated. On the other hand, users receive multimedia information of varying quality, which affects their user experience. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a multimedia information recommendation method, apparatus, electronic device, and storage medium. The technical solution of the embodiments of the present invention is implemented as follows:

[0005] This invention provides a multimedia information recommendation method including:

[0006] Acquire historical data of target objects in a multimedia information recommendation environment;

[0007] Based on the historical data of the target object, determine the first input feature of the multimedia information recommendation model;

[0008] The multimedia information recommendation model is used to perform feature fusion processing on the first input features to obtain user playback feature vector and user sharing feature vector that match the target object.

[0009] Retrieve multimedia information to be recommended from multimedia information data sources;

[0010] Based on the multimedia information to be recommended, the second input feature of the multimedia information recommendation model is determined;

[0011] The multimedia information recommendation model is used to perform feature fusion processing on the second input features to obtain multimedia information playback feature vector and multimedia information sharing feature vector that match the multimedia information to be recommended.

[0012] Based on the user playback feature vector, the user sharing feature vector, the multimedia information playback feature vector, and the multimedia information sharing feature vector, the multimedia information recall strategy is adjusted.

[0013] This invention also provides a multimedia information recommendation device, comprising:

[0014] The information transmission module is used to acquire historical data of target objects in a multimedia information recommendation environment;

[0015] The information processing module is used to determine the first input feature of the multimedia information recommendation model based on the historical data of the target object;

[0016] The information processing module is used to perform feature fusion processing on the first input features through the multimedia information recommendation model to obtain user playback feature vector and user sharing feature vector that match the target object;

[0017] The information transmission module is used to acquire the multimedia information to be recommended from the multimedia information data source;

[0018] The information processing module is used to determine the second input feature of the multimedia information recommendation model based on the multimedia information to be recommended;

[0019] The information processing module is used to perform feature fusion processing on the second input feature through the multimedia information recommendation model to obtain a multimedia information playback feature vector and a multimedia information sharing feature vector that match the multimedia information to be recommended.

[0020] The information processing module is used to adjust the multimedia information recall strategy based on the user playback feature vector, the user sharing feature vector, the multimedia information playback feature vector, and the multimedia information sharing feature vector.

[0021] In the above scheme,

[0022] The information processing module is used to extract and process the playback type sub-information included in the historical data of the target object, and determine the first playback identifier feature, the first playback tag feature, and the first playback category feature that match the target object;

[0023] The information processing module is used to extract and process the sharing type sub-information included in the historical data of the target object, and determine the first sharing identifier feature, the first sharing tag feature, and the first sharing category feature that match the target object.

[0024] In the above scheme,

[0025] The information processing module is used to extract and process the playback type sub-information included in the historical data of the target object, and determine the first playback identifier feature, the first playback tag feature, and the first playback category feature that match the target object;

[0026] The information processing module is used to extract and process the sharing type sub-information included in the historical data of the target object, and determine the first sharing identifier feature, the first sharing tag feature, and the first sharing category feature that match the target object.

[0027] In the above scheme,

[0028] The information processing module is used to perform feature fusion processing on the first input features through the hybrid expert network in the multimedia information recommendation model to obtain the user playback high-order feature vector and the user sharing high-order feature vector.

[0029] The information processing module is used to perform weighted processing on the first input feature through the feature weight adjustment network in the multimedia information recommendation model to obtain the user playback low-order feature vector and the user sharing low-order feature vector corresponding to the target object.

[0030] The information processing module is used to concatenate the user playback high-order feature vector and the user playback low-order feature vector to obtain the user playback feature vector.

[0031] The information processing module is used to concatenate the user-shared high-order feature vector and the user-shared low-order feature vector to obtain the user-shared feature vector.

[0032] In the above scheme,

[0033] The information processing module is used to determine contextual features that match the target object based on the historical data of the target object;

[0034] The information processing module is used to adjust the user playback feature vector by utilizing the context features through the hybrid expert network in the multimedia information recommendation model;

[0035] The information processing module is used to adjust the user risk feature vector by utilizing the context features through the hybrid expert network in the multimedia information recommendation model.

[0036] In the above scheme,

[0037] The information processing module is used to determine the number of hybrid expert networks based on the multimedia information recommendation environment.

[0038] The information processing module is used to adjust the structure of the hybrid expert network in the multimedia information recommendation model according to the number of hybrid expert networks.

[0039] In the above scheme,

[0040] The information processing module is used to extract and process the playback type sub-information included in the multimedia information to be recommended, and determine the second playback identifier feature, the second playback tag feature, and the second playback category feature that match the media information to be recommended.

[0041] The information processing module is used to extract and process the sharing type sub-information included in the historical data of the target object, and determine the second sharing identifier feature, the second sharing tag feature, and the second sharing category feature that match the target object.

[0042] In the above scheme,

[0043] The information processing module is used to perform feature fusion processing on the second input features through the hybrid expert network in the multimedia information recommendation model to obtain high-order feature vectors for multimedia information playback and multimedia information sharing.

[0044] The information processing module is used to perform weighted processing on the second input feature through the feature weight adjustment network in the multimedia information recommendation model to obtain the low-order feature vector of multimedia information playback and the low-order feature vector of multimedia information sharing corresponding to the multimedia information, wherein all weights in the feature weight adjustment network are fixed values.

[0045] The information processing module is used to concatenate the high-order feature vector of multimedia information playback and the low-order feature vector of multimedia information playback to obtain the multimedia information playback feature vector.

[0046] The information processing module is used to concatenate the high-order feature vector and the low-order feature vector of multimedia information sharing to obtain the multimedia information sharing feature vector.

[0047] In the above scheme,

[0048] The information processing module is used to determine the first dot product value based on the user playback feature vector, the user playback high-order feature vector, the user playback low-order feature vector, the multimedia information playback high-order feature vector, and the multimedia information playback low-order feature vector in the multimedia information playback feature vector.

[0049] The information processing module is used to determine the second dot product value based on the user-shared high-order feature vector, the user-shared low-order feature vector, the multimedia information sharing high-order feature vector, and the multimedia information sharing low-order feature vector in the user-shared feature vector;

[0050] The information processing module is used to determine the multimedia information to be recalled based on the sum of the first dot product value and the second dot product value.

[0051] In the above scheme,

[0052] The information processing module is used to determine the third dot product value based on the user playback feature vector, the user playback low-order feature vector, the multimedia information playback feature vector, the multimedia information playback high-order feature vector and the multimedia information playback low-order feature vector, the user sharing feature vector, the user sharing high-order feature vector, the user sharing low-order feature vector, and the multimedia information sharing feature vector.

[0053] The information processing module is used to determine the multimedia information to be recalled based on the third dot product value.

[0054] In the above scheme,

[0055] The information processing module is used to perform data filtering processing on the multimedia information to be recommended, and to parse and obtain the title and tags of the multimedia information to be recommended;

[0056] The information processing module is used to trigger the target word segmentation library and perform word segmentation on the title and tags of the multimedia information to be recommended through the target word segmentation library to obtain word-level multimedia information to be recommended.

[0057] The information processing module is used to vectorize the word-level multimedia information to be recommended through the text information processing network in the multimedia information recommendation model, forming a multi-dimensional word-level title feature vector and a multi-dimensional word-level tag feature vector of the multimedia information to be recommended.

[0058] In the above scheme,

[0059] The information processing module is used to obtain the target user's historical browsing information;

[0060] The information processing module is used to determine the multimedia information exposure history corresponding to the historical browsing information based on the target user's historical browsing information;

[0061] The information processing module is used to dynamically adjust the playback strategy of the multimedia information based on the exposure history of the multimedia information corresponding to the historical browsing information.

[0062] In the above scheme,

[0063] The information processing module is used to determine the type of multimedia information recommendation environment;

[0064] The information processing module is used to determine the category of multimedia information to be played based on the type of the multimedia information recommendation environment.

[0065] The information processing module is used to respond to the category of the multimedia information to be played, trigger a matching multimedia information data source, so as to adjust the multimedia information to be played by a multimedia information data source that matches the category of the multimedia information to be played.

[0066] In the above scheme,

[0067] The information processing module is used to determine the historical parameters of the multimedia information to be recommended based on the type of multimedia information recommendation environment in which the multimedia information to be recommended is located.

[0068] The information processing module is used to determine a training sample set that matches the multimedia information recommendation model based on the historical parameters of the multimedia information to be recommended, wherein the training sample set includes at least one set of training samples.

[0069] The information processing module is used to extract a set of training samples that match the training samples by using a noise threshold matched by the multimedia information recommendation model.

[0070] The information processing module is used to train the multimedia information recommendation model based on a set of training samples that match the training samples.

[0071] In the above scheme,

[0072] The information processing module is used to determine a multi-task loss function that matches the multimedia information recommendation model;

[0073] The information processing module is used to adjust the parameters and feature weights of the hybrid expert network in the multimedia information recommendation model based on the multi-task loss function, until the loss functions of different dimensions corresponding to the multimedia information recommendation model reach the corresponding convergence conditions; so as to make the parameters of the multimedia information recommendation model compatible with the multimedia information recommendation environment.

[0074] In the above scheme,

[0075] The information processing module is used to determine a dynamic noise threshold that matches the usage environment of the multimedia information recommendation model when the multimedia information recommendation environment in which the multimedia information to be recommended is located is short video recommendation.

[0076] The information processing module is used to remove noise from the first training sample set according to the dynamic noise threshold, so as to form a second training sample set that matches the dynamic noise threshold.

[0077] The information processing module is used to determine a fixed noise threshold corresponding to the multimedia information recommendation model when the multimedia information recommendation environment in which the multimedia information to be recommended is played is an instant messaging client, and to remove noise from the first training sample set according to the fixed noise threshold to form a second training sample set that matches the fixed noise threshold.

[0078] This invention also provides an electronic device, the electronic device comprising:

[0079] Memory, used to store executable instructions;

[0080] The processor, when executing executable instructions stored in the memory, implements the aforementioned multimedia information recommendation method.

[0081] This invention also provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the aforementioned multimedia information recommendation method.

[0082] The embodiments of the present invention have the following beneficial effects:

[0083] This invention acquires historical data of target objects in a multimedia information recommendation environment; based on the historical data of the target objects, it determines a first input feature of the multimedia information recommendation model; through the multimedia information recommendation model, it performs feature fusion processing on the first input feature to obtain a user playback feature vector and a user sharing feature vector matching the target object; it acquires multimedia information to be recommended from a multimedia information data source; based on the multimedia information to be recommended, it determines a second input feature of the multimedia information recommendation model; through the multimedia information recommendation model, it performs feature fusion processing on the second input feature to obtain a multimedia information playback feature vector and a multimedia information sharing feature vector matching the multimedia information to be recommended; based on the user playback feature vector, the user sharing feature vector, and the multimedia information playback feature vector and multimedia information sharing feature vector, it adjusts the multimedia information recall strategy. Therefore, the multimedia information recommendation model can recommend multimedia information in the usage environment to different users, while enhancing the accuracy and relevance of multimedia information recommendations, effectively improving the quality of multimedia information recommendations, and enhancing the user experience. Attached Figure Description

[0084] Figure 1 This is a schematic diagram illustrating a usage scenario of the multimedia information recommendation method provided in an embodiment of the present invention;

[0085] Figure 2 This is a schematic diagram of the composition structure of the multimedia information recommendation device provided in an embodiment of the present invention;

[0086] Figure 3 A schematic flowchart of an optional multimedia information recommendation method provided in an embodiment of the present invention;

[0087] Figure 4 This is a schematic diagram of the structure of the hybrid expert network in an embodiment of the present invention;

[0088] Figure 5 This is a schematic diagram of the feature weight adjustment network in an embodiment of the present invention;

[0089] Figure 6 This is a schematic diagram illustrating the processing steps of user playback feature vectors and user sharing feature vectors in an embodiment of the present invention;

[0090] Figure 7 A schematic flowchart of an optional multimedia information recommendation method provided in an embodiment of the present invention;

[0091] Figure 8 This is a schematic diagram illustrating an optional multimedia information recommendation in an embodiment of the present invention;

[0092] Figure 9A schematic flowchart of an optional multimedia information recommendation method provided in an embodiment of the present invention;

[0093] Figure 10 This is a schematic diagram illustrating the application environment of the training method for the multimedia information recommendation model in this embodiment of the invention.

[0094] Figure 11 This is a schematic diagram illustrating the data processing of the multimedia information recommendation model provided in this embodiment of the invention applied to an instant messaging client. Detailed Implementation

[0095] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0096] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0097] Before providing a further detailed description of the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention will be explained, and the nouns and terms involved in the embodiments of the present invention shall be interpreted as follows.

[0098] 1) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0099] 2) Based on, used to indicate the conditions or states on which the operation is performed depends. When the conditions or states on which it depends are met, one or more operations can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0100] 3) Model training involves multi-class classification learning on the image dataset. This model can be built using deep learning frameworks such as TensorFlow and Torch, employing multiple layers of neural networks like CNNs to form a multi-class classification model. The model input is a three-channel or original-channel matrix generated from images read using tools like OpenCV. The model output is the multi-class probability, ultimately outputting the webpage category through algorithms such as softmax. During training, the model approximates the correct trend using objective functions such as cross-entropy.

[0101] 4) Neural Network (NN): Artificial Neural Network (ANN), also known as neural network or neural network-like network, is a mathematical or computational model in the fields of machine learning and cognitive science that imitates the structure and function of biological neural networks (the central nervous system of animals, especially the brain) and is used to estimate or approximate functions.

[0102] 5) Multi-task learning: In the field of machine learning, by simultaneously learning and optimizing multiple related tasks, better model accuracy can be achieved than that of a single task. Multiple tasks help each other by sharing a representation layer. This training method is called multi-task learning, also known as joint learning.

[0103] 6) Recommendation Accuracy: Recommended multimedia content has a certain effect over a period of time, and this effect is measured by the user's interest in the video content. Accuracy plays an important role in online user retention, clicks, and CTR on the client side.

[0104] 7) Contextual information: the time, location, and network status of the user's access to the recommendation system.

[0105] 8) MMoE: A model that applies the Mixture-of-Experts (MoE) structure to multi-task learning, explicitly models the relationships between tasks, and learns multiple gated recurrent unit networks to balance the shared expert representations in different tasks. It allows parameters to be automatically allocated without adding a large number of new parameters to capture common information between tasks and the model that distinguishes between tasks.

[0106] 9) ESMM: A model that borrows the idea of ​​multi-task learning (transfer learning), introduces two auxiliary tasks to fit pCTR and pCTCVR respectively, and treats pCVR as an intermediate variable, thereby reducing sample selection bias (SSB), sparse training data and delayed feedback problems.

[0107] 10) Attention: helps the model assign different weights to each part of the input, extracts more critical and important information, enables the model to make more accurate judgments, and does not bring greater overhead to the model's computation and storage.

[0108] 11) Multi-objective recall: This involves considering multiple objectives within a single recall model. In recommendation systems, it's often necessary to optimize multiple business objectives simultaneously to generate greater business revenue. For example, in e-commerce, the goal is to simultaneously optimize click-through rate and conversion rate, giving the platform more objectives; in news feed scenarios, the aim is to increase user engagement such as click-through rate, likes, and comments, creating a better community atmosphere and thus improving user retention.

[0109] 12) softmax: A very commonly used and important function in machine learning, especially in multi-class classification scenarios. It maps some inputs to real numbers between 0 and 1, and normalization ensures that the sum is 1.

[0110] 13) tag: A keyword marker, which is a set of words extracted from the article body and title to represent the core content of the document.

[0111] In this invention, embodiments can be implemented using cloud technology. Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. It can also be understood as a general term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies based on cloud computing business models. The backend services of network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites; therefore, cloud technology needs cloud computing as its support.

[0112] It's important to note that cloud computing is a computing model that distributes computing tasks across a resource pool comprised of numerous computers, enabling various application systems to access computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, resources in the "cloud" are infinitely scalable, readily available, and can be used on demand, expanded at any time, and paid for based on usage. As the foundational providers of cloud computing capabilities, they establish cloud resource pool platforms, often referred to as cloud platforms or Infrastructure as a Service (IaaS). These platforms deploy various types of virtual resources within the resource pool for external customers to choose from. The cloud resource pool primarily includes: computing devices (which can be virtualized machines containing operating systems), storage devices, and network devices.

[0113] Figure 1 This is a schematic diagram illustrating a usage scenario of the multimedia information processing method provided in an embodiment of the present invention. (See attached diagram.) Figure 1The terminals (including terminals 10-1 and 10-2) are equipped with corresponding clients capable of playing embedded multimedia information. The terminals connect to server 200 via network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both. Data transmission is achieved using a wireless link. The multimedia information includes, but is not limited to, videos, images, GIF animations, and advertising information. The types of multimedia information obtained by the terminals (including terminals 10-1 and 10-2) from the corresponding server 200 via network 300 can be the same or different. For example, the terminals (including terminals 10-1 and 10-2) can obtain video advertisements or image advertisements placed by advertisers from the corresponding server 200 via network 300. The specific types are not limited in this application. Server 200 can store different multimedia information, including advertising multimedia information in different dynamic formats, such as GIF, MP4, and MOV.

[0114] During the process of the terminal (terminal 10-1 and / or terminal 10-2) obtaining and displaying the corresponding service with embedded multimedia information from the server 200 via the network 300, the user can perform different operations on the multimedia information presented in the multimedia information playback window through the terminal (terminal 10-1 and / or terminal 10-2), resulting in different user behaviors. For example, when the multimedia information is a video advertisement, the user can share and / or like the exposed short video while watching the information, or click on it. When the multimedia information is a dynamic GIF advertisement, during the exposure of the advertisement through the terminal (terminal 10-1 and / or terminal 10-2), the user can forward and / or comment on the advertisement, or jump to the corresponding product purchase link page through the GIF advertisement.

[0115] As an example, when server 200 determines which multimedia information to recommend for playback to user's terminal 10-1 or 10-2, it needs to adjust the multimedia information to be played in a timely manner. For example, it may replace any multimedia information in the set of multimedia information to be played to adapt to the viewing needs of different target users. Taking short video information as an example, the multimedia information recommendation model provided by this invention can be applied to short video playback. In short video playback, different short video information from different data sources is usually processed, and finally, the corresponding different information and the corresponding recommended video to be recommended are presented on the user interface (UI). The accuracy and timeliness of the characteristics of different information directly affect the user experience. The background database of video playback receives a large amount of video data from different sources every day. The different information obtained for recommending information to target users can also be called by other applications (e.g., the recommendation results of the short video recommendation process are migrated to the long video recommendation process or the news recommendation process). Of course, the multimedia information recommendation model that matches the corresponding target user can also be migrated to different video recommendation processes (e.g., web video recommendation process, mini-program video recommendation process, or long video client video recommendation process).

[0116] As an example, server 200 is used to deploy a corresponding multimedia information recommendation model to implement the multimedia information recommendation method provided by this invention, or to deploy a multimedia information recommendation device to implement the multimedia information recommendation method. Specifically, it acquires historical data of target objects in the multimedia information recommendation environment; determines a first input feature of the multimedia information recommendation model based on the historical data of the target objects; performs feature fusion processing on the first input feature through the multimedia information recommendation model to obtain a user playback feature vector and a user sharing feature vector matching the target objects; acquires multimedia information to be recommended from the multimedia information data source; determines a second input feature of the multimedia information recommendation model based on the multimedia information to be recommended; performs feature fusion processing on the second input feature through the multimedia information recommendation model to obtain a multimedia information playback feature vector and a multimedia information sharing feature vector matching the multimedia information to be recommended; adjusts the multimedia information recall strategy based on the user playback feature vector, the user sharing feature vector, the multimedia information playback feature vector, and the multimedia information sharing feature vector, recommends the multimedia information to be recommended to the target users, and displays the multimedia information to be recommended matching the target users through terminals (terminal 10-1 and / or terminal 10-2). Taking short multimedia information as an example, the multimedia information recommendation model provided by this invention can be applied to short video playback. During short video playback, different short multimedia information from different data sources is typically processed, and finally, the corresponding multimedia information and the multimedia information to be recommended corresponding to the short video recommendation process are presented on the user interface (UI). The accuracy and timeliness of the characteristics of different multimedia information directly affect the user experience. The background database of video playback receives a large amount of multimedia information data from different sources every day. The different multimedia information obtained for recommending multimedia information to the target user can also be called by other applications (e.g., the recommendation results of the short video recommendation process can be migrated to the recommendation process in an instant messaging client or a news recommendation process). Of course, the multimedia information recommendation model matching the corresponding target user can also be migrated to different video recommendation processes (e.g., webpage video recommendation process, mini-program video recommendation process, or video recommendation process in an instant messaging client).

[0117] The multimedia information recommendation method provided in this application is based on artificial intelligence (AI). AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making functions.

[0118] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0119] In the embodiments of this application, the main artificial intelligence software technologies involved include the aforementioned speech processing technologies and machine learning. For example, it may involve Automatic Speech Recognition (ASR) technology in speech technology, including speech signal preprocessing, speech signal frequency analyzing, speech signal feature extraction, speech signal feature matching / recognition, and speech training.

[0120] For example, this could involve machine learning (ML), a multidisciplinary field encompassing probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning typically includes techniques such as deep learning, which includes artificial neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and deep neural networks (DNNs).

[0121] It is understood that the multimedia information recommendation method and voice processing provided in this application can be applied to intelligent devices. Intelligent devices can be any device with information display functions, such as smart terminals, smart home devices (such as smart speakers, smart washing machines, etc.), smart wearable devices (such as smartwatches), in-vehicle intelligent central control systems (which display multimedia information to users through applets that perform different tasks), or AI intelligent medical devices (which display treatment cases by showing multimedia information), etc.

[0122] The structure of the multimedia information recommendation device according to an embodiment of the present invention will be described in detail below. The multimedia information recommendation device can be implemented in various forms, such as a dedicated terminal with multimedia information recommendation processing function, or a server equipped with multimedia information recommendation processing function, for example, the preceding... Figure 1 Server 200 in the middle. Figure 2 This is a schematic diagram of the composition structure of the multimedia information recommendation device provided in the embodiments of the present invention. It can be understood that... Figure 2 The diagram shows only an exemplary structure of the multimedia information recommendation device, not the entire structure; it can be implemented as needed. Figure 2 The structure shown may be part or all of the structure.

[0123] The multimedia information recommendation device provided in this embodiment of the invention includes at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. The various components in the multimedia information recommendation device are coupled together via a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 205.

[0124] The user interface 203 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.

[0125] It is understood that memory 202 can be volatile memory or non-volatile memory, or both. In this embodiment of the invention, memory 202 is capable of storing data to support the operation of the terminal (e.g., 10-1). Examples of this data include any computer programs used to operate on the terminal (e.g., 10-1), such as operating systems and applications. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications.

[0126] In some embodiments, the multimedia information recommendation device provided in this invention can be implemented using a combination of hardware and software. For example, the multimedia information recommendation device provided in this invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the multimedia information recommendation model provided in this invention. For instance, the processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0127] As an example of the multimedia information recommendation device provided in this embodiment of the invention, which is implemented using a combination of hardware and software, the multimedia information recommendation device provided in this embodiment of the invention can be directly embodied as a combination of software modules executed by processor 201. The software modules can be located in a storage medium, which is located in memory 202. Processor 201 reads the executable instructions included in the software modules in memory 202 and combines them with necessary hardware (e.g., including processor 201 and other components connected to bus 205) to complete the training method of the multimedia information recommendation model provided in this embodiment of the invention.

[0128] As an example, processor 201 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0129] As an example of the hardware implementation of the multimedia information recommendation device provided in this embodiment of the invention, the device provided in this embodiment of the invention can be directly executed by a processor 201 in the form of a hardware decoding processor. For example, it can be executed by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components to implement the training method of the multimedia information recommendation model provided in this embodiment of the invention.

[0130] In this embodiment of the invention, the memory 202 is used to store various types of data to support the operation of the multimedia information recommendation device. Examples of such data include: any executable instructions for operation on the multimedia information recommendation device, such as executable instructions that implement the training method of the multimedia information recommendation model in this embodiment of the invention, which may be included in the executable instructions.

[0131] In other embodiments, the multimedia information recommendation device provided in this invention can be implemented in software. Figure 2A multimedia information recommendation device stored in memory 202 is shown. This device can be software in the form of programs and plugins, and includes a series of modules. As an example of a program stored in memory 202, it may include the multimedia information recommendation device, which includes the following software modules:

[0132] Information transmission module 2081 and information processing module 2082. When the software modules in the multimedia information recommendation device are read into RAM and executed by processor 201, the training method of the multimedia information recommendation model provided in this embodiment of the invention will be implemented. The functions of each software module in the multimedia information recommendation device include:

[0133] The information transmission module 2081 is used to acquire historical data of target objects in a multimedia information recommendation environment.

[0134] The information processing module 2082 is used to determine the first input feature of the multimedia information recommendation model based on the historical data of the target object.

[0135] The information processing module 2082 is used to perform feature fusion processing on the first input features through the multimedia information recommendation model to obtain user playback feature vector and user sharing feature vector that match the target object.

[0136] The information transmission module 2082 is used to acquire multimedia information to be recommended from the multimedia information data source.

[0137] The information processing module 2082 is used to determine the second input feature of the multimedia information recommendation model based on the multimedia information to be recommended.

[0138] The information processing module 2082 is used to perform feature fusion processing on the second input features through the multimedia information recommendation model to obtain multimedia information playback feature vector and multimedia information sharing feature vector that match the multimedia information to be recommended.

[0139] The information processing module 2082 is used to adjust the multimedia information recall strategy based on the user playback feature vector, the user sharing feature vector, the multimedia information playback feature vector, and the multimedia information sharing feature vector.

[0140] The information processing module 2082 is used to determine the target user that matches the multimedia information to be recommended in the multimedia information data source based on the feature classification processing result, and recommend the multimedia information to be recommended to the target user.

[0141] according to Figure 2In one aspect of this application, the electronic device also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform various embodiments and combinations of embodiments provided in the various optional implementations of the multimedia information recommendation method described above.

[0142] Combination Figure 2 The multimedia information recommendation device shown illustrates the multimedia information recommendation method provided in this embodiment of the invention. See also: Figure 3 , Figure 3 This is an optional flowchart illustrating the multimedia information recommendation method provided in an embodiment of the present invention. It can be understood that... Figure 3 The steps shown can be performed by various electronic devices that run the multimedia information recommendation device, such as a dedicated terminal with a multimedia information recommendation device, a server, or a server cluster, wherein the dedicated terminal with the multimedia information recommendation device can be a preceding step. Figure 2 The illustrated embodiment is an electronic device with a multimedia information recommendation device. The following section addresses... Figure 3 The steps shown are explained.

[0143] Step 301: The multimedia information recommendation device receives a multimedia information recommendation request sent by the terminal.

[0144] Step 302: The multimedia information recommendation device responds to the multimedia information recommendation request by obtaining historical data of the target object in the media information recommendation environment.

[0145] In some embodiments of the present invention, various user behaviors matched with corresponding clients can be collected through different program components. This involves effectively extracting raw logs of user behavior data, such as the user's device ID (user account), multimedia information type, multimedia information browsing duration, and multimedia information browsing completeness parameters. The user's historical click behavior and corresponding information browsing duration are recorded through a subscription service and stored in Redis. When a user request arrives, the online recommendation system retrieves the corresponding user's historical click behavior to determine the historical data of the target object.

[0146] Step 303: The multimedia information recommendation device determines the first input feature of the multimedia information recommendation model based on the historical data of the target object.

[0147] In some embodiments of the present invention, the first input feature of the multimedia information recommendation model is determined based on the historical data of the target object, which can be achieved in the following ways:

[0148] The playback type sub-information included in the historical data of the target object is extracted and processed to determine the first playback identifier feature, first playback tag feature, and first playback category feature that match the target object; the sharing type sub-information included in the historical data of the target object is extracted and processed to determine the first sharing identifier feature, first sharing tag feature, and first sharing category feature that match the target object. For multimedia information recommendation models, recommendations can be made based on implicit user feedback. User satisfaction with the recommendation results usually depends on many indicators. For example, in e-commerce, the satisfaction evaluation of item recommendations is based on indicators related to behaviors such as clicks, browsing depth (dwell time), adding to cart, favorites, purchases, repeat purchases, and positive reviews. Different recommendation systems, different periods, and different product forms often have different indicators. In the use of multimedia information recommendation models, when the playback rate and sharing rate are the objectives, the playback ID, playback tag, playback category, sharing ID, sharing tag, and sharing category can be used as network inputs, and the feature vector can be adjusted through contextual information.

[0149] Step 304: The multimedia information recommendation device performs feature fusion processing on the first input features through the multimedia information recommendation model to obtain user playback feature vector and user sharing feature vector that match the target object.

[0150] In some embodiments of the present invention, the first input features are fused using the multimedia information recommendation model to obtain user playback feature vectors and user sharing feature vectors that match the target object. This can be achieved in the following ways:

[0151] The first input features are fused using a hybrid expert network in the multimedia information recommendation model to obtain a high-order feature vector for user playback and a high-order feature vector for user sharing. The first input features are then weighted using a feature weight adjustment network in the multimedia information recommendation model to obtain low-order feature vectors for user playback and user sharing corresponding to the target object. These high-order and low-order feature vectors are then concatenated to obtain a user playback feature vector. Finally, these high-order and low-order feature vectors are concatenated to obtain a user sharing feature vector. (Reference) Figure 4 , Figure 4This is a schematic diagram of the hybrid expert network structure in an embodiment of the present invention. In the multimedia information recommendation model, the hybrid expert network can establish multiple different expert sub-networks on the bottom pooling layer embedding, depending on usage requirements. Different targets can select expert networks according to their needs. The input for each target is obtained by weighted summation of the outputs of the expert networks through a gated recurrent unit network (where each target is a separate MLP tower). The input of the gated recurrent unit network is the same as that of the expert networks. When processing multimedia information, to fully utilize the contextual information of the target user, additional contextual features can be added to the gated recurrent unit network, and then the weights of each expert corresponding to the task are obtained through linear transformation and softmax. This input, which integrates target and contextual features, can enhance the network's learning of the target and help alleviate conflicts between different targets. Specifically, contextual features matching the target object can be determined based on the historical data of the target object; the user playback feature vector can be adjusted using the contextual features through the hybrid expert network in the multimedia information recommendation model; and the user risk feature vector can be adjusted using the contextual features through the hybrid expert network in the multimedia information recommendation model.

[0152] refer to Figure 5 , Figure 5 This is a schematic diagram of the feature weight adjustment network in an embodiment of the present invention. When the feature weight adjustment network is working, it can solve the feature combination problem under sparse data, and its prediction complexity is linear. It has good versatility for both continuous and discrete features.

[0153] In typical linear models, features are considered independently without taking into account the relationships between them. However, in multimedia information recommendation environments, features may exhibit certain correlations. For example, in news recommendation, male users tend to view more military news, while female users prefer emotional news. Similarly, male users browse sports products more often, while female users prefer clothing products. This demonstrates a correlation between gender and news channels, and the types of products are also related to the target user's gender.

[0154] To optimize the multimedia information recommendation process, this application only considers the second-order crossover case. The specific model can be found in Formula 1:

[0155]

[0156] Where n represents the number of features in the sample, x i It is the value of the i-th feature, w0, w i w ij These are model parameters; only when x...i With x j Crossover is only meaningful when all values ​​are zero. However, in cases of sparse data, there will be very few samples that satisfy the condition that the crossover terms are not zero. When there are insufficient training samples, it is easy to lead to insufficient and inaccurate parameter training, which will ultimately affect the performance of the model.

[0157] Therefore, the training problem of the cross-term parameters can be approximated by matrix factorization, as shown in Formula 2.

[0158]

[0159] The model needs to consider the parameters w0∈R, w∈R n V∈R n×k , <·, ·> are the inner product of two k-dimensional vectors, see Formula 3:

[0160]

[0161] For any positive definite matrix W, as long as k is large enough, there exists a matrix W such that W = VV T However, in cases of sparse data, a smaller k should be chosen because there is not enough data to estimate w. ij Limiting the size of k improves the model's generalization ability. The time complexity of directly calculating formula (2) is O(kn). 2 This is because all cross features need to be calculated. However, by changing the formula, the complexity can be reduced to linear, as shown in Formula 4:

[0162]

[0163] like Figure 4 As shown, the target context is the embedding of the context features. The inner product is calculated between this embedding and the features of each user. The mathematical expression for the inner product is shown in Formula 5.

[0164]

[0165] It can determine the importance of a feature in the current target under the current context. The larger the inner product value, the more important the current feature is. Therefore, the introduction of contextual features can add strong priors to the training of the target.

[0166] refer to Figure 6 , Figure 6This diagram illustrates the processing steps for user playback and user sharing feature vectors in this embodiment of the invention. The user features consist of two parts: 1) Higher-order features obtained from MMoE: Higher-order features output by different expert networks are combined with weights from a gated recurrent unit network that incorporates context features to form different target features, with dimension d1; 2) Lower-order features obtained from FM: The original FM features are weighted by context features to obtain features with dimension d2. Directly concatenating the higher-order and lower-order features minimizes loss, ultimately yielding user sharing and user playback feature vectors with dimensions d1+d2.

[0167] Step 305: The multimedia information recommendation device acquires the multimedia information to be recommended from the multimedia information data source.

[0168] Step 306: The multimedia information recommendation device determines the second input feature of the multimedia information recommendation model based on the multimedia information to be recommended.

[0169] In some embodiments of the present invention, the second input feature of the multimedia information recommendation model is determined based on the historical data of the target object, which can be achieved in the following ways:

[0170] The playback type sub-information included in the multimedia information to be recommended is extracted and processed to determine the second playback identifier feature, the second playback tag feature, and the second playback category feature that match the media information to be recommended; the sharing type sub-information included in the historical data of the target object is extracted and processed to determine the second sharing identifier feature, the second sharing tag feature, and the second sharing category feature that match the target object.

[0171] Step 307: The multimedia information recommendation device performs feature fusion processing on the second input feature through the multimedia information recommendation model to obtain a multimedia information playback feature vector and a multimedia information sharing feature vector that match the multimedia information to be recommended.

[0172] Combination Figure 2 The multimedia information recommendation device shown illustrates the multimedia information recommendation method provided in this embodiment of the invention. See also: Figure 7 , Figure 7 This is an optional flowchart illustrating the multimedia information recommendation method provided in an embodiment of the present invention. It can be understood that... Figure 7 The steps shown can be performed by various electronic devices that run the multimedia information recommendation device, such as a dedicated terminal with a multimedia information recommendation device, a server, or a server cluster, wherein the dedicated terminal with the multimedia information recommendation device can be a preceding step. Figure 2 The illustrated embodiment is an electronic device with a multimedia information recommendation device. The following section addresses... Figure 7The steps shown are explained.

[0173] Step 701: Through the hybrid expert network in the multimedia information recommendation model, the second input features are fused to obtain the high-order feature vector of multimedia information playback and the high-order feature vector of multimedia information sharing.

[0174] Step 702: The second input feature is weighted through the feature weight adjustment network in the multimedia information recommendation model to obtain the low-order feature vector of multimedia information playback and the low-order feature vector of multimedia information sharing corresponding to the multimedia information, wherein all weights in the feature weight adjustment network are fixed values.

[0175] Step 703: Concatenate the high-order feature vector of multimedia information playback and the low-order feature vector of multimedia information playback to obtain the multimedia information playback feature vector.

[0176] Step 704: Concatenate the high-order feature vector and the low-order feature vector of multimedia information sharing to obtain the multimedia information sharing feature vector.

[0177] It should be noted that the difference between obtaining the feature vector corresponding to multimedia information (Item) and obtaining the feature vector corresponding to the target user lies in the fact that, since items lack contextual features, the gated recurrent unit network of the Hybrid Expert Network (MMoE) reduces the need for contextual features. Figure 2 In the extra input, all feature weights are set to 1 during FM feature weight adjustment, ultimately yielding item sharing and item playback features of dimension d1+d2. In this process, the Gated Recurrent Unit (GRU) network, with fewer parameters than LSTM, is a model that can handle sequence information well. Next, the feature input feedforward neural network will be fused to process effective information from other features. Predicting cheating behavior is treated as a probability prediction problem, using the sigmoid function (logistic function) as the output layer, and the loss function is the standard cross-entropy loss, as shown in Equation 6.

[0178]

[0179] The GRU layer is used for deep feature extraction. Alternatively, the GRU layer can be omitted and replaced with several feedforward neural network layers, which can also effectively process and fuse features.

[0180] Step 308: The multimedia information recommendation device adjusts the multimedia information recall strategy based on the user playback feature vector, the user sharing feature vector, the multimedia information playback feature vector, and the multimedia information sharing feature vector, and recommends multimedia information through the adjusted recall strategy.

[0181] In some embodiments of the present invention, the adjustment of the multimedia information recall strategy can be achieved in the following ways:

[0182] Based on the user playback feature vector (including the higher-order and lower-order user playback feature vectors) and the multimedia information playback feature vector (including both higher-order and lower-order multimedia information playback feature vectors), a first dot product value is determined. Based on the user sharing feature vector (including the higher-order and lower-order user sharing feature vectors) and the multimedia information sharing feature vector (including both higher-order and lower-order multimedia information sharing feature vectors), a second dot product value is determined. The multimedia information to be recalled is determined by summing the first and second dot products. Specifically, the dot product of <user playback higher-order feature vector, user playback lower-order feature vector> and <multimedia information playback higher-order feature vector, multimedia information playback lower-order feature vector> is taken, and the top k1 multimedia information items are selected as the multimedia information to be recalled. The dot product of <user sharing higher-order feature vector, user sharing lower-order feature vector> and <multimedia information sharing higher-order feature vector, multimedia information sharing lower-order feature vector> is taken, and the top k2 items are selected as the recall candidates. The inner product of the two recall candidates is added together, and the top k with the highest score is taken as the final recall result.

[0183] In some embodiments of the present invention, the adjustment of the multimedia information recall strategy can also be achieved in the following ways:

[0184] Based on the user playback feature vector (including higher-order and lower-order user playback feature vectors), the multimedia information playback feature vector (including both higher-order and lower-order multimedia information playback feature vectors), the user sharing feature vector (including both higher-order and lower-order user sharing feature vectors), and the multimedia information sharing feature vector (including both higher-order and lower-order multimedia information sharing feature vectors), a third dot product value is determined. Based on this third dot product value, the top k values ​​are used as candidates for recall to determine the multimedia information to be recalled.

[0185] In some embodiments of the present invention, see Figure 8 , Figure 8This is a schematic diagram of an optional multimedia information recommendation in an embodiment of the present invention. All tasks share a single network structure, but the features extracted from different network layers correspond to different tasks. Generally, the lower layers of the model structure correspond to less complex NLP tasks. In use, when the target resource includes different advertisements from the same advertiser, the time-sensitive short video advertisements included in different resource groups can be played sequentially in the time-sensitive short video playback window. When all time-sensitive short video playback areas in the display interface are occupied by the advertiser, after the advertisement playback ends, the advertisement information of the same advertiser can be displayed cyclically in the time-sensitive short video playback window of the advertisement information display interface. Simultaneously, when the advertiser's time-sensitive short video is a video advertisement, the video advertisement information of the same advertiser can be displayed cyclically. The audio volume carried by the video is adjusted to the maximum to prompt the user to watch the played video advertisement. Replacing advertisement A with advertisement B allocates more playback bandwidth to advertisement B, resulting in a better viewing experience for the user. Specifically, based on the traffic parameters and iterative experimental parameters matched to the ad playback strategy, the ad playback strategy can be dynamically adjusted to increase ad exposure. In some embodiments of this invention, the exposure channel of ad A can be changed from the current short video playback client to the contact status information of the instant messaging client. Of course, when adjusting the exposure position of ad A, it can be changed from a Moments ad in the instant messaging client to a splash screen ad, to conform to different dynamically adjusted playback strategies. This allows time-sensitive short videos to be recommended to different users in a short period of time, achieving better video recommendation results. Figure 8 For example, if it is determined that male target user 1 has shared an ad with feature B in their browsing history, the playback strategy can be dynamically adjusted to replace the current ad with other ad information containing feature B (such as an ad or short video containing feature B). Similarly, if it is determined that female target user 2 has clicked on an ad with feature X in their browsing history, the playback strategy can be dynamically adjusted to replace ad A with other ad information (such as an ad or short video ad link containing feature X) to match the user's usage habits and provide a better user experience.

[0186] Combination Figure 2 The multimedia information recommendation device shown illustrates the multimedia information recommendation method provided in this embodiment of the invention. See also: Figure 9 , Figure 9 This is an optional flowchart illustrating the multimedia information recommendation method provided in an embodiment of the present invention. It can be understood that... Figure 9The steps shown can be performed by various electronic devices that run the multimedia information recommendation device, such as a dedicated terminal with a multimedia information recommendation device, a server, or a server cluster, wherein the dedicated terminal with the multimedia information recommendation device can be a preceding step. Figure 2 The illustrated embodiment is an electronic device with a multimedia information recommendation device. The following section addresses... Figure 9 The steps shown are explained.

[0187] Step 901: Determine the historical parameters of the multimedia information to be recommended based on the type of multimedia information recommendation environment in which the multimedia information to be recommended is located.

[0188] Step 902: Based on the historical parameters of the multimedia information to be recommended, determine the training sample set that matches the multimedia information recommendation model.

[0189] Step 903: Extract a set of training samples that match the training samples by using the noise threshold matched by the multimedia information recommendation model.

[0190] Step 904: Determine the multi-task loss function that matches the multimedia information recommendation model.

[0191] Step 905: Based on the multi-task loss function, adjust the parameters and feature weights of the hybrid expert network in the multimedia information recommendation model to adjust the network parameters.

[0192] Therefore, during the training process, until the loss functions of different dimensions corresponding to the multimedia information recommendation model reach the corresponding convergence conditions, the parameters of the multimedia information recommendation model can be adapted to the multimedia information recommendation environment. For example, when the multimedia information recommendation model is used in a short video recommendation environment, during the process of recommending different short videos to users, the short video playback interface can be displayed in the corresponding APP or triggered by an instant messaging client mini-program (the multimedia information recommendation model can be encapsulated in the corresponding APP or stored as a plugin in the instant messaging client mini-program after training). As short video application products continue to develop and increase, the amount of multimedia information carried is far greater than that of text information. Different types of short videos in the short video server can be continuously recommended to users through corresponding applications. During this training process, in the usage environment of short video recommendations triggered by the instant messaging client mini-program, the dynamic noise threshold matching the usage environment of the multimedia information recommendation model needs to be lower than the dynamic noise threshold for recommending short videos directly to users in the short multimedia information playback client.

[0193] In some embodiments of the present invention, when the multimedia information recommendation model is applied to the news information recommendation process, a fixed noise threshold corresponding to the news information recommendation process is determined, and the first training sample set is denoised according to the fixed noise threshold to form a second training sample set that matches the fixed noise threshold. Wherein, when the multimedia information recommendation model is embedded in a corresponding hardware device (e.g., a news reading terminal, an e-book terminal, a financial news terminal), and the usage environment involves pushing different news information to users through a news reading terminal or e-book terminal, fixing the fixed noise threshold corresponding to the multimedia information recommendation model can effectively improve the training speed of the multimedia information recommendation model and reduce the user's waiting time. Wherein, in a usage environment with fixed noise, the training sample set can come from the target user's historical data. Historical recommended multimedia information browsing data can be recommended multimedia information viewing behavior data generated when recommending multimedia information to the target user, which can be extracted from historical browsing logs. Here, historical recommended multimedia information browsing data can be all historical recommended multimedia information browsing data; or, considering the timeliness of the behavior data, it can only include historical recommended multimedia information browsing data within a preset time period, such as historical recommended multimedia information browsing data within a week, etc.

[0194] The following example illustrates the multimedia information recommendation method provided in this embodiment of the invention, using a product advertisement recommendation scenario within a short video playback interface as an example. Figure 10 This is a schematic diagram illustrating the application environment of the training method for the multimedia information recommendation model in this embodiment of the invention, wherein, as shown... Figure 10 As shown, the product insert advertising information playback interface can be displayed in the corresponding APP or triggered through an instant messaging client mini-program (the multimedia information recommendation model can be trained and encapsulated in the corresponding APP or stored as a plugin in the instant messaging client mini-program, with the usage environment being news and information recommendation). With the continuous development and increase of short video application products, the carrying capacity of video news multimedia information far exceeds that of text information, and product insert advertising information can be continuously recommended to users through the corresponding application. For example, the "Take a Look" entry in the discovery page of an instant messaging client application, or the audio recommendation entry in an audio application, or the video recommendation entry in a video application, or the live recommendation entry in a live streaming application, etc. When the target terminal runs the target application according to the user's operation and controls the target application to display the application page including the trigger entry for triggering the opening of the recommended content display page, it can detect the trigger operation of the trigger entry. When a trigger operation corresponding to the trigger entry is generated, a recommendation request is sent to the server, and after receiving the recommended content fed back by the server in response to the recommendation request, the recommended content is displayed on the recommended content display page in the recommended content display order.

[0195] refer to Figure 11 , Figure 11 This is a schematic diagram illustrating the data processing of the multimedia information recommendation model provided in this embodiment of the invention applied to an instant messaging client, specifically including the following steps:

[0196] Step 1101: Obtain historical data of the target object in the instant messaging client.

[0197] Step 1102: Obtain the multimedia information to be recommended from the multimedia information data source of the instant messaging client.

[0198] Step 1103: Obtain user playback feature vectors and user sharing feature vectors that match the target object through the multimedia information recommendation model.

[0199] Step 1104: Using the multimedia information recommendation model, obtain the multimedia information playback feature vector and multimedia information sharing feature vector that match the multimedia information to be recommended.

[0200] Step 1105: Adjust the multimedia information recall strategy based on the user playback feature vector, the user sharing feature vector, the multimedia information playback feature vector, and the multimedia information sharing feature vector.

[0201] Step 1106: Recommend multimedia information using the adjusted recall strategy.

[0202] Multimedia information recommendation mainly includes: 1) Recall logic: Context-aware multimedia information recommendation models are used to help recall candidate items for each user who has read or clicked. This involves taking the user's most recent playback tag, playback ID, playback category, sharing tag, sharing ID, sharing category, and contextual features, and then calculating the highest-scoring candidate items according to the strategy described in the principle section for recall and recommendation. 2) Coarse ranking logic: Context-aware multimedia information recommendation models can not only be used to help with candidate selection but also in the ranking stage. Each user's embedding vector and the candidate item's embedding have already been calculated in the recall stage. In the coarse ranking stage, the cosine similarity between the candidate item's embedding vector and the user's embedding vector is calculated, and the top item is taken as the recall result.

[0203] Beneficial technical effects:

[0204] This invention acquires historical data of target objects in a multimedia information recommendation environment; based on the historical data of the target objects, it determines a first input feature of the multimedia information recommendation model; through the multimedia information recommendation model, it performs feature fusion processing on the first input feature to obtain a user playback feature vector and a user sharing feature vector matching the target object; it acquires multimedia information to be recommended from a multimedia information data source; based on the multimedia information to be recommended, it determines a second input feature of the multimedia information recommendation model; through the multimedia information recommendation model, it performs feature fusion processing on the second input feature to obtain a multimedia information playback feature vector and a multimedia information sharing feature vector matching the multimedia information to be recommended; based on the user playback feature vector, the user sharing feature vector, and the multimedia information playback feature vector and multimedia information sharing feature vector, it adjusts the multimedia information recall strategy. Therefore, the multimedia information recommendation model can recommend multimedia information in the usage environment to different users, while enhancing the accuracy and relevance of multimedia information recommendations, effectively improving the quality of multimedia information recommendations, and enhancing the user experience.

[0205] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multimedia information recommendation method, characterized in that, The method includes: Acquire historical data of target objects in a multimedia information recommendation environment; Based on the historical data of the target object, determine the first input feature of the multimedia information recommendation model; The multimedia information recommendation model performs feature fusion processing on the first input features, including the first playback identifier feature, the first playback tag feature, the first playback category feature, the first sharing identifier feature, the first sharing tag feature, and the first sharing category feature, to obtain a user playback feature vector and a user sharing feature vector that match the target object. Retrieve multimedia information to be recommended from multimedia information data sources; Based on the multimedia information to be recommended, the second input feature of the multimedia information recommendation model is determined; The multimedia information recommendation model performs feature fusion processing on the second input features, including the second playback identifier feature, the second playback tag feature, the second playback category feature, the second sharing identifier feature, the second sharing tag feature, and the second sharing category feature, to obtain a multimedia information playback feature vector and a multimedia information sharing feature vector that match the multimedia information to be recommended. Based on the user playback feature vector and the multimedia information playback feature vector, a first dot product value is determined; based on the user sharing feature vector and the multimedia information sharing feature vector, a second dot product value is determined; the multimedia information to be recalled is determined by summing the first dot product value and the second dot product value; or, based on the user playback feature vector, the multimedia information playback feature vector, the user sharing feature vector, and the multimedia information sharing feature vector, a third dot product value is determined; the multimedia information to be recalled is determined by the third dot product value. Based on the multimedia information to be recalled, multimedia information recommendations are made.

2. The method according to claim 1, characterized in that, The determination of the first input feature of the multimedia information recommendation model based on the historical data of the target object includes: The playback type sub-information included in the historical data of the target object is extracted and processed to determine the first playback identifier feature, the first playback tag feature, and the first playback category feature that match the target object; The sharing type sub-information included in the historical data of the target object is extracted and processed to determine the first sharing identifier feature, the first sharing tag feature, and the first sharing category feature that match the target object.

3. The method according to claim 1, characterized in that, The multimedia information recommendation model performs feature fusion processing on the first input features, including the first playback identifier feature, the first playback tag feature, the first playback category feature, the first sharing identifier feature, the first sharing tag feature, and the first sharing category feature, to obtain a user playback feature vector and a user sharing feature vector that match the target object, including: The first input features are fused using the hybrid expert network in the multimedia information recommendation model to obtain the user playback high-order feature vector and the user sharing high-order feature vector. The first input feature is weighted by the feature weight adjustment network in the multimedia information recommendation model to obtain the user playback low-order feature vector and the user sharing low-order feature vector corresponding to the target object. The user playback high-order feature vector and the user playback low-order feature vector are concatenated to obtain the user playback feature vector. The user-shared high-order feature vector and the user-shared low-order feature vector are concatenated to obtain the user-shared feature vector.

4. The method according to claim 3, characterized in that, The method further includes: Based on the historical data of the target object, determine the contextual features that match the target object; By using the hybrid expert network in the multimedia information recommendation model and the contextual features, the user playback feature vector is adjusted. The user-shared feature vector is adjusted using the contextual features through the hybrid expert network in the multimedia information recommendation model.

5. The method according to claim 3, characterized in that, The method further includes: The number of hybrid expert networks is determined based on the multimedia information recommendation environment. The structure of the hybrid expert network in the multimedia information recommendation model is adjusted according to the number of hybrid expert networks.

6. The method according to claim 1, characterized in that, The step of determining the second input feature of the multimedia information recommendation model based on the multimedia information to be recommended includes: The playback type sub-information included in the multimedia information to be recommended is extracted and processed to determine the second playback identifier feature, the second playback tag feature, and the second playback category feature that match the media information to be recommended. The sharing type sub-information included in the historical data of the target object is extracted and processed to determine the second sharing identifier feature, the second sharing tag feature, and the second sharing category feature that match the target object.

7. The method according to claim 1, characterized in that, The multimedia information recommendation model performs feature fusion processing on the second input features, including the second playback identifier feature, the second playback tag feature, the second playback category feature, the second sharing identifier feature, the second sharing tag feature, and the second sharing category feature, to obtain a multimedia information playback feature vector and a multimedia information sharing feature vector that match the multimedia information to be recommended, including: The hybrid expert network in the multimedia information recommendation model performs feature fusion processing on the second input features to obtain high-order feature vectors for multimedia information playback and multimedia information sharing. The second input feature is weighted by the feature weight adjustment network in the multimedia information recommendation model to obtain the low-order feature vector of multimedia information playback and the low-order feature vector of multimedia information sharing corresponding to the multimedia information. All weights in the feature weight adjustment network are fixed values. The high-order feature vector of multimedia information playback and the low-order feature vector of multimedia information playback are concatenated to obtain the multimedia information playback feature vector. The high-order feature vector and the low-order feature vector of multimedia information sharing are concatenated to obtain the multimedia information sharing feature vector.

8. The method according to claim 1, characterized in that, The method further includes: The multimedia information to be recommended is subjected to data filtering processing, and the title and tags of the multimedia information to be recommended are parsed and obtained; The target word segmentation library is triggered, and the title and tags of the multimedia information to be recommended are segmented into words using the target word segmentation library to obtain word-level multimedia information to be recommended. The text information processing network in the multimedia information recommendation model is used to vectorize the word-level multimedia information to be recommended, forming a multi-dimensional word-level title feature vector and a multi-dimensional word-level tag feature vector for the multimedia information to be recommended.

9. The method according to claim 1, characterized in that, The method further includes: Obtain the target user's browsing history; Based on the target user's historical browsing information, determine the multimedia information exposure history corresponding to the historical browsing information; Based on the exposure history of the multimedia information corresponding to the historical browsing information, the playback strategy of the multimedia information is dynamically adjusted.

10. The method according to claim 1, characterized in that, The method further includes: Determine the type of multimedia information recommendation environment; Based on the type of the multimedia information recommendation environment, the category of the multimedia information to be played is determined; In response to the category of the multimedia information to be played, a matching multimedia information data source is triggered to adjust the multimedia information to be played by using a multimedia information data source that matches the category of the multimedia information to be played.

11. A multimedia information recommendation device, characterized in that, The device includes: The information transmission module is used to acquire historical data of target objects in a multimedia information recommendation environment; The information processing module is used to determine the first input feature of the multimedia information recommendation model based on the historical data of the target object; The information processing module is used to perform feature fusion processing on the first input features, including the first playback identifier feature, the first playback tag feature, the first playback category feature, the first sharing identifier feature, the first sharing tag feature, and the first sharing category feature, through the multimedia information recommendation model, to obtain a user playback feature vector and a user sharing feature vector that match the target object. The information transmission module is used to acquire the multimedia information to be recommended from the multimedia information data source; The information processing module is used to determine the second input feature of the multimedia information recommendation model based on the multimedia information to be recommended; The information processing module is used to perform feature fusion processing on the second input features, including the second playback identifier feature, the second playback tag feature, the second playback category feature, the second sharing identifier feature, the second sharing tag feature, and the second sharing category feature, through the multimedia information recommendation model, to obtain a multimedia information playback feature vector and a multimedia information sharing feature vector that match the multimedia information to be recommended. The information processing module is configured to: determine a first dot product value based on the user playback feature vector and the multimedia information playback feature vector; determine a second dot product value based on the user sharing feature vector and the multimedia information sharing feature vector; determine multimedia information to be recalled based on the sum of the first and second dot products; or, determine a third dot product value based on the user playback feature vector, the multimedia information playback feature vector, the user sharing feature vector, and the multimedia information sharing feature vector; determine multimedia information to be recalled based on the third dot product value; and recommend multimedia information based on the multimedia information to be recalled.

12. The apparatus according to claim 11, characterized in that, The information processing module is also used to extract and process the playback type sub-information included in the historical data of the target object, and determine the first playback identifier feature, the first playback tag feature, and the first playback category feature that match the target object; The sharing type sub-information included in the historical data of the target object is extracted and processed to determine the first sharing identifier feature, the first sharing tag feature, and the first sharing category feature that match the target object.

13. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the multimedia information recommendation method according to any one of claims 1 to 10.

14. A computer-readable storage medium storing executable instructions, characterized in that, When the executable instructions are executed by the processor, they implement the multimedia information recommendation method according to any one of claims 1-10.

15. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement the multimedia information recommendation method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Method for pushing and acquiring multimedia playing information between strangers

    CN110427503A

  • Video information processing method and device based on video information processing model

    CN111191078A