Recommendation Method, Device, and Computer-Readable Storage Medium for Multimedia Data

By using multi-feature extraction networks and gated networks in the personalized recommendation system to process multimedia data features and generate differentiated recommendation features and splicing features, the problem of feature signal loss in the rough arrangement stage is solved, and the effect and applicability of personalized recommendations are improved.

CN115114461BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210422793.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-06-27
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

In the prior art, the screening of push media content in the rough arrangement stage leads to serious loss of original feature signal, limiting the ability of personalized recommendation models to learn differentiated representations, resulting in poor recommendation effect and poor applicability.

Method used

By acquiring the characteristics of multimedia data, using multiple feature extraction networks and gated networks for feature extraction and weighting, differentiated recommended features and splicing features are generated, and the business target prediction model is input to obtain the target recommended media.

Benefits of technology

It improves the effectiveness of selecting multimedia data to be recommended, enhances the personalized recommendation experience of multimedia data, is highly applicable, solves the problem of original feature signal loss, and improves the learning ability of the recommendation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115114461B_ABST
    Figure CN115114461B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method, device, and computer-readable storage medium for recommending multimedia data. The method includes: obtaining multimedia data features, where the multimedia data features include first multimedia data features and second multimedia data features. Feature extraction is performed on the first multimedia data features through multiple first feature extraction networks to obtain multiple recommended features, and multiple weighted recommended features are obtained through multiple first gating networks. Multiple concatenated vectors obtained by concatenating each of the weighted recommended features in the multiple weighted recommended features with the second multimedia data features are obtained, and multiple concatenated features are obtained based on the second feature extraction network. The multiple concatenated vectors and the concatenated features of each concatenated vector are input into a business objective prediction model for obtaining recommended media to obtain multiple target recommended media. By adopting the present application, the selection effectiveness of multimedia data to be recommended can be improved, the personalized recommendation experience of multimedia data can be enhanced, and the applicability is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, and computer-readable storage medium for recommending multimedia data. Background Art

[0002] As a product of the development of the Internet and artificial intelligence (AI), personalized recommendation is a technology that provides personalized information services and decision-making support to customers based on massive data mining. In the actual application process, personalized recommendation generally can be divided into stages such as recall, rough ranking, and fine ranking from the process. The recall stage generally includes multiple models or strategies, and can quickly screen out some (such as thousands) of media content to be pushed from a large number (such as millions) of items to be recommended (such as media content to be pushed) and provide them to the rough ranking stage for unified pre-ranking. The rough ranking includes further screening out some (such as hundreds) of media content to be pushed from the part of the media content to be pushed recalled and outputting them to the fine ranking stage for more accurate ranking, and pushing the above fine ranking results to the user. However, in the prior art, in the rough ranking stage, the media content to be pushed is usually screened and then output to the fine ranking. The original feature signals contained in the screened media content to be pushed are severely lost, which limits the ability of the personalized recommendation model to learn differential representations, resulting in poor personalized recommendation effects and poor applicability. Summary of the Invention

[0003] The embodiments of the present application provide a method, device, and computer-readable storage medium for recommending multimedia data, which can improve the effectiveness of selecting multimedia data to be recommended, enhance the personalized recommendation experience of multimedia data, and have high applicability.

[0004] In a first aspect, the embodiments of the present application provide a method for recommending multimedia data, and the method includes:

[0005] Obtain multimedia data features corresponding to the multimedia data, where the multimedia data features include a first multimedia data feature and a second multimedia data feature. The first multimedia data feature is a data feature of the recommended object, and the second multimedia data feature is a data feature of the multimedia data to be recommended, or the first multimedia data feature is a data feature of the multimedia data to be recommended, and the second multimedia data feature is a data feature of the recommended object;

[0006] Extract features from the above-mentioned first multimedia data features through multiple first feature extraction networks to obtain multiple recommended features corresponding to the above-mentioned first multimedia data features, and obtain multiple weighted recommended features corresponding to the above-mentioned multiple recommended features through multiple first gating networks, where one of the above-mentioned first feature extraction networks is used to obtain one recommended feature corresponding to the above-mentioned first multimedia data features, and one of the above-mentioned first gating networks is used to obtain one weighted recommended feature corresponding to the above-mentioned multiple recommended features;

[0007] Obtain multiple concatenated vectors obtained by concatenating each weighted recommended feature among the above-mentioned multiple weighted recommended features with the above-mentioned second multimedia data features respectively, and obtain concatenated features of each concatenated vector based on the second feature extraction network corresponding to each first gating network that obtains each weighted recommended feature to obtain multiple concatenated features;

[0008] Input the above-mentioned multiple concatenated vectors and the concatenated features of each concatenated vector into a business objective prediction model for obtaining recommended media, so as to obtain multiple target recommended media based on the above-mentioned business objective prediction model.

[0009] In a possible implementation manner, the above-mentioned obtaining multiple weighted recommended features corresponding to the above-mentioned multiple recommended features through multiple first gating networks includes:

[0010] Obtain feature combination weights corresponding to each recommended feature used by each of the above-mentioned first gating networks to obtain weighted recommended features based on the above-mentioned first multimedia data features;

[0011] Perform weighted summation on the above-mentioned multiple recommended features through any one of the above-mentioned first gating networks based on the feature combination weights corresponding to each recommended feature used by any one of the above-mentioned first gating networks, so as to obtain one of the above-mentioned weighted recommended features obtained by any one of the above-mentioned first gating networks;

[0012] Obtain each of the above-mentioned weighted recommended features obtained by each of the above-mentioned first gating networks to obtain the above-mentioned multiple weighted recommended features;

[0013] Wherein, the above-mentioned first gating network includes at least one of a gating network based on linear transformation or a gating network based on normalized weighting.

[0014] In a possible implementation manner, after the above-mentioned obtaining multiple concatenated vectors obtained by concatenating each weighted recommended feature among the above-mentioned multiple weighted recommended features with the above-mentioned second multimedia data features respectively, the method further includes:

[0015] Obtain cross features, and concatenate the cross features with each of the above-mentioned multiple concatenated vectors respectively to obtain updated multiple concatenated vectors.

[0016] In a possible implementation, the above-mentioned business objective prediction model includes multiple second gating networks and prediction networks corresponding to each of the above-mentioned second gating networks;

[0017] After inputting the above-mentioned multiple concatenated vectors and the concatenation features of each of the above-mentioned concatenated vectors into the business objective prediction model for obtaining recommended media, the method further includes:

[0018] Based on the above-mentioned multiple concatenated vectors, obtain the feature combination weights corresponding to the concatenation features used by each of the above-mentioned second gating networks to obtain target recommendation features. Through any one of the above-mentioned second gating networks, perform weighted summation on the above-mentioned multiple concatenation features based on the feature combination weights corresponding to the concatenation features used by any one of the above-mentioned second gating networks to obtain a target recommendation feature obtained by any one of the above-mentioned second gating networks, and obtain the target recommendation features obtained by each of the above-mentioned second gating networks to obtain the above-mentioned multiple target recommendation features;

[0019] Based on the prediction networks corresponding to each of the above-mentioned second gating networks that obtain the above-mentioned target recommendation features, obtain the business objective prediction values corresponding to the above-mentioned target recommendation features to obtain multiple business objective prediction values, and determine multiple target recommended media from the above-mentioned multimedia data based on the above-mentioned multiple business objective prediction values;

[0020] Wherein, the above-mentioned second gating network includes at least one of a gating network based on linear transformation or a gating network based on normalized weighting.

[0021] In a possible implementation, the above-mentioned first multimedia data feature is a recommended object data feature and the above-mentioned second multimedia data feature is a multimedia data feature to be recommended;

[0022] The above-mentioned obtaining, through multiple first feature extraction networks, multiple recommendation features corresponding to the above-mentioned first multimedia data feature, and obtaining multiple weighted recommendation features corresponding to the above-mentioned multiple recommendation features through multiple first gating networks includes:

[0023] Through multiple first feature extraction networks, perform feature extraction on the above-mentioned recommended object data feature to obtain multiple recommended object features corresponding to the above-mentioned recommended object data feature, and obtain multiple weighted recommended object features corresponding to the above-mentioned multiple recommended object features through multiple first gating networks as multiple weighted recommendation features;

[0024] The above-mentioned obtaining multiple concatenated vectors obtained by concatenating each of the above-mentioned weighted recommendation features in the above-mentioned multiple weighted recommendation features with the above-mentioned second multimedia data feature respectively includes:

[0025] Obtain multiple concatenated vectors obtained by concatenating each of the above-mentioned weighted recommended object features in the above-mentioned multiple weighted recommended object features with the above-mentioned multimedia data feature to be recommended respectively.

[0026] In a possible implementation, the above first multimedia data feature is the feature of the multimedia data to be recommended, and the above second multimedia data feature is the feature of the recommended object data;

[0027] The above-mentioned extraction of the first multimedia data feature by multiple first feature extraction networks to obtain multiple recommended features corresponding to the first multimedia data feature, and the obtaining of multiple weighted recommended features corresponding to the multiple recommended features by multiple first gating networks includes:

[0028] Extracting the features of the multimedia data to be recommended by multiple first feature extraction networks to obtain multiple recommended multimedia features corresponding to the multimedia data to be recommended, and obtaining multiple weighted recommended multimedia features corresponding to the multiple recommended multimedia features by multiple first gating networks as multiple weighted recommended features;

[0029] The above-mentioned obtaining of multiple concatenated vectors obtained by concatenating each weighted recommended feature in the multiple weighted recommended features with the above second multimedia data feature respectively includes:

[0030] Obtaining multiple concatenated vectors obtained by concatenating each weighted recommended multimedia feature in the multiple weighted recommended multimedia features with the above recommended object data feature respectively.

[0031] In a possible implementation, the above-mentioned obtaining of the multimedia data feature corresponding to the multimedia data includes:

[0032] Obtaining multimedia data, where the multimedia data includes recommended object data and multimedia data to be recommended, and the multimedia data to be recommended includes at least one of graphic media data, audio data, and video data;

[0033] Obtaining the recommended object data feature corresponding to the recommended object data in the multimedia data through vectorization processing and embedding compression processing, and obtaining the multimedia data feature to be recommended corresponding to the multimedia data to be recommended.

[0034] In a second aspect, an embodiment of the present application provides a recommendation device for multimedia data, and the device includes:

[0035] An obtaining module, configured to obtain a multimedia data feature corresponding to multimedia data, where the multimedia data feature includes a first multimedia data feature and a second multimedia data feature, the first multimedia data feature is a recommended object data feature and the second multimedia data feature is a multimedia data feature to be recommended, or the first multimedia data feature is a multimedia data feature to be recommended and the second multimedia data feature is a recommended object data feature;

[0036] The weighted recommendation feature generation module is used to extract features from the above-mentioned first multimedia data features through multiple first feature extraction networks to obtain multiple recommendation features corresponding to the above-mentioned first multimedia data features, and obtain multiple weighted recommendation features corresponding to the above-mentioned multiple recommendation features through multiple first gating networks. Among them, one of the above-mentioned first feature extraction networks is used to obtain one recommendation feature corresponding to the above-mentioned first multimedia data features, and one of the above-mentioned first gating networks is used to obtain one weighted recommendation feature corresponding to the above-mentioned multiple recommendation features;

[0037] The concatenated feature generation module is used to obtain multiple concatenated vectors by concatenating each of the above-mentioned weighted recommendation features with the above-mentioned second multimedia data features respectively, and obtain concatenated features of each of the above-mentioned concatenated vectors based on the second feature extraction network corresponding to each of the first gating networks that obtain the above-mentioned weighted recommendation features to obtain multiple concatenated features;

[0038] The target recommended media generation module is used to input the above-mentioned multiple concatenated vectors and the concatenated features of each of the above-mentioned concatenated vectors into a service target prediction model for obtaining recommended media, so as to obtain multiple target recommended media based on the above-mentioned service target prediction model.

[0039] In a possible implementation manner, the above-mentioned weighted recommendation feature generation module is used for:

[0040] Based on the above-mentioned first multimedia data features, obtain the feature combination weights corresponding to the respective recommendation features used by each of the above-mentioned first gating networks to obtain weighted recommendation features;

[0041] Through any one of the above-mentioned first gating networks, perform weighted summation on the above-mentioned multiple recommendation features based on the feature combination weights corresponding to the respective recommendation features used by any one of the above-mentioned first gating networks, so as to obtain one of the above-mentioned weighted recommendation features obtained by any one of the above-mentioned first gating networks;

[0042] Obtain the respective weighted recommendation features obtained by each of the above-mentioned first gating networks to obtain the above-mentioned multiple weighted recommendation features;

[0043] Among them, the above-mentioned first gating network includes at least one of a gating network based on linear transformation or a gating network based on normalized weighting.

[0044] In a possible implementation manner, after obtaining the multiple concatenated vectors obtained by concatenating each of the above-mentioned weighted recommendation features with the above-mentioned second multimedia data features respectively, the above-mentioned concatenated feature generation module is used for:

[0045] Obtain cross features, and concatenate the cross features with each of the above-mentioned multiple concatenated vectors respectively to obtain updated multiple concatenated vectors.

[0046] In a possible implementation manner, the above business objective prediction model includes a plurality of second gating networks and prediction networks corresponding to each of the above second gating networks;

[0047] After inputting the above plurality of concatenated vectors and the concatenation features of each of the above concatenated vectors into the business objective prediction model for obtaining recommended media, the above target recommended media generation module is configured to:

[0048] Based on the above plurality of concatenated vectors, obtain the feature combination weights corresponding to the concatenation features used by each of the above second gating networks to obtain target recommendation features, and through any one of the above second gating networks, perform weighted summation on the above plurality of concatenation features based on the feature combination weights corresponding to the concatenation features used by any one of the above second gating networks, so as to obtain a target recommendation feature obtained by any one of the above second gating networks, and obtain the target recommendation features obtained by each of the above second gating networks to obtain the above plurality of target recommendation features;

[0049] Based on the prediction networks corresponding to each of the above second gating networks that obtain the above respective target recommendation features, obtain the business objective prediction values corresponding to the above respective target recommendation features to obtain a plurality of business objective prediction values, and determine a plurality of target recommended media from the above multimedia data based on the above plurality of business objective prediction values;

[0050] Wherein, the above second gating network includes at least one of a gating network based on linear transformation or a gating network based on normalization weighting.

[0051] In a possible implementation manner, the above acquisition module is configured to:

[0052] Acquire multimedia data, where the multimedia data includes recommendation object data and multimedia data to be recommended, and the multimedia data to be recommended includes at least one of graphic media data, audio data, and video data;

[0053] Obtain the recommendation object data features corresponding to the above recommendation object data in the above multimedia data through vectorization processing and embedding compression processing, and obtain the multimedia data to be recommended features corresponding to the above multimedia data to be recommended.

[0054] In a third aspect, an embodiment of the present application provides a computer device, where the computer device includes: a processor, a memory, and a network interface;

[0055] The above processor is connected to the memory and the network interface, wherein the network interface is used to provide a data communication function, the above memory is used to store program code, and the above processor is used to call the above program code to execute the method in the first aspect of the embodiment of the present application.

[0056] Fourthly, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the processor executes the program instructions, the method in the first aspect of the embodiment of the present application is executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0058] Figure 1 is a schematic diagram of a multi-gated mixture-of-experts network architecture;

[0059] Figure 2 is a schematic diagram of the system architecture provided by the embodiment of the present application;

[0060] Figure 3 is a schematic diagram of the scenario of the multimedia data recommendation method provided by the embodiment of the present application;

[0061] Figure 4 is a schematic diagram of the flow of the multimedia data recommendation method provided by the embodiment of the present application;

[0062] Figure 5 is a schematic diagram of the personalized recommendation model structure provided by the embodiment of the present application;

[0063] Figure 6 is a schematic diagram of the experimental effect comparison of the multimedia data recommendation method provided by the embodiment of the present application;

[0064] Figure 7 is a schematic diagram of the structure of the multimedia data recommendation device provided by the embodiment of the present application;

[0065] Figure 8 is a schematic diagram of the structure of the computer device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0067] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0068] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0069] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.

[0070] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0071] The solution provided in the embodiments of this application involves natural language processing and machine learning technologies in the field of artificial intelligence, and will be specifically described through the following embodiments:

[0072] The recommendation method for multimedia data provided by the embodiments of the present application (or simply referred to as the method provided by the embodiments of the present application) is applicable to screening and then recommending and outputting various types of multimedia data (such as video and graphic content, etc.) in application programs with functions of viewing and / or uploading video and graphic content (such as video applications, news applications, short video applications, and social applications, etc.). Specifically, the method provided by the embodiments of the present application is applicable to the rough ranking stage of personalized recommendation, screening the multimedia data selected in the recall stage to obtain some multimedia data (for convenience of description, the target recommended media can be used as an example) from the multimedia data selected in the recall stage and sending it to the fine ranking stage for further screening. The above-mentioned various types of video and graphic content can be professional generated content (PGC) published by self-media and institutions, or user generated content (UGC) from user creation, which can be specifically determined according to the actual application scenario, and the present application does not limit this here. To achieve better personalized recommendation effects, generally, a network model based on multi-objective learning is used for screening multimedia data in each stage of personalized recommendation. For example, in the rough ranking stage, a personalized recommendation model based on a multi-gate mixture-of-experts (MMoE) can be used to screen multimedia data. The multi-gate mixture-of-experts is a commonly used network structure for multi-objective learning. Refer to Figure 1 , Figure 1 which is a schematic diagram of the multi-gate mixture-of-experts network architecture. As Figure 1 shown, the multi-gate mixture-of-experts network includes multiple expert networks (expert network 1, expert network 2,..., expert network N) and multiple gating networks (gating network 1,..., gating network K). Among them, the expert networks are used to extract different features, and the gating networks are used to assign weights to each expert network. The multi-gate mixture-of-experts network can perform multi-objective prediction based on the input multimedia data to obtain multiple business objective prediction values corresponding to each multimedia data (the business objectives can be click-through rate, conversion rate, and browsing duration, etc.), so that the multimedia data can be screened based on multiple business objective prediction values. In addition, to enhance the personalization of multimedia data screening, the screening can be combined with the to-be-recommended multimedia data and the recommended object data (which can be various attribute information of the recommended object). And since in the rough ranking stage, the multimedia data with a certain quantity (usually several thousand or several tens of thousands of data) selected in the recall stage (for convenience of description, the to-be-recommended multimedia data can be used as an example) is processed. From the perspective of screening efficiency, the input recommended object data and the to-be-recommended multimedia data can be dimensionally reduced and then screened through the multi-gate mixture-of-experts network. For example, Figure 1The recommended object data and the multimedia data to be recommended in [[]] are respectively subjected to dimensionality reduction processing through tower structure 1 and tower structure 2. At the same time, the recommended object data and the multimedia data to be recommended after the above dimensionality reduction processing are concatenated with the cross data (which can introduce the internal correlation information between the recommended object and the multimedia data) and then input into each expert network for feature extraction. However, compared with directly inputting the multimedia data to be recommended and the recommended object data, the original feature signals contained in the dimensionality-reduced recommended object data and the multimedia data to be recommended are all lost, which limits the ability of different expert networks in the multi-gate mixture of experts network to learn differential representations end-to-end, resulting in poor multi-objective learning effects, poor screening effects of the multimedia data to be recommended, and low applicability. Therefore, before screening the multimedia data through a network model based on multi-objective learning (such as a multi-gate mixture of experts network), differential representations can be extracted from the original data (such as the multimedia data to be recommended and the recommended object data), so as to enhance the ability of the expert network in the multi-gate mixture of experts network to learn differential representations, enhance the effect of multi-objective prediction through the multi-gate mixture of experts network, and thus improve the screening effectiveness of the multimedia data to be recommended in the rough ranking stage, with high applicability.

[0073] In the method provided in the embodiments of the present application, during the recommendation process of multimedia data, the multimedia data features corresponding to the multimedia data can be obtained, and the recommended object data features or the to-be-recommended multimedia data features (which can be the first multimedia data features) in the multimedia data features are input into the first feature extraction network in the personalized recommendation model. Each first feature extraction network performs feature extraction based on the first multimedia data features to obtain multiple differentiated recommendation features, providing recommendation features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. Multiple first gating networks perform weighted summation on the multiple recommendation features output by the above-mentioned first feature extraction network to obtain multiple weighted recommendation features. Since different first gating networks assign different feature combination weights to each recommendation feature, the weighted recommendation features generated by each first gating network based on the multiple recommendation features have obvious differences in emphasis corresponding to the above-mentioned multiple first feature extraction networks. Therefore, multiple differentiated weighted recommendation features can be obtained through each first gating network, providing weighted recommendation features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. By extracting differentiated representations (the above-mentioned first multimedia data features) from the original data (such as the to-be-recommended multimedia data and the recommended object data), the ability of the personalized recommendation model to learn differentiated representations is enhanced, so as to enhance the effect of multi-object prediction through the personalized recommendation model. The multiple concatenated vectors obtained by concatenating each weighted recommendation feature with the second multimedia data features (the recommended object data features or the to-be-recommended multimedia data features, and different from the first multimedia data features) and the concatenation features of the above-mentioned concatenated vectors are input into the service objective prediction model (such as the multi-gating mixture of experts network) in the personalized recommendation model for obtaining the recommended media, with strong adaptability.

[0074] See Figure 2 , Figure 2 is the schematic diagram of the system architecture provided by the embodiments of the present application. As Figure 2As shown in the figure, the system architecture may include a business server 100 and a terminal cluster. The terminal cluster may include terminal devices such as terminal device 200a, terminal device 200b, terminal device 200c, ……, terminal device 200n. Among them, the above-mentioned business server 100 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal devices (including terminal device 200a, terminal device 200b, terminal device 200c, ……, terminal device 200n) may be intelligent terminals such as smart phones, tablet computers, laptop computers, desktop computers, handheld computers, mobile internet devices (MIDs), wearable devices (such as smart watches, smart bracelets, etc.), intelligent computers, and intelligent vehicles. Among them, the business server 100 may establish communication connections with each terminal device in the terminal cluster, and communication connections may also be established between the terminal devices in the terminal cluster. In other words, the business server 100 may establish communication connections with each terminal device among terminal device 200a, terminal device 200b, terminal device 200c, ……, terminal device 200n. For example, a communication connection may be established between terminal device 200a and the business server 100. A communication connection may be established between terminal device 200a and terminal device 200b, and a communication connection may also be established between terminal device 200a and terminal device 200c. Among them, the above-mentioned communication connection is not limited to the connection method, and may be directly or indirectly connected by a wired communication method, or may be directly or indirectly connected by a wireless communication method, etc., which can be specifically determined according to the actual application scenario, and this application does not make any restrictions here.

[0075] It should be understood that, as Figure 2 shown, each terminal device in the terminal cluster may be installed with an application client. When the application client runs on each terminal device, it may respectively communicate with the above-mentioned Figure 2Data interaction is carried out between the business servers 100 shown, so that the business servers 100 can receive business data from each terminal device (such as official accounts, multimedia data to be recommended uploaded by users through terminal devices). Among them, the application client can be a social application, an instant messaging application, a live broadcast application, a news application, a short video application, a video application, a music application, a shopping application, a novel application, a payment application, etc., which have the function of displaying data information such as text, images, and videos. Specifically, it can be determined according to the actual application scenario requirements and is not limited here. Among them, the application client can be an independent client or an embedded sub-client integrated in a certain client (such as an instant messaging client, a social client, etc.). Specifically, it can be determined according to the actual application scenario and is not limited here. Taking the news application as an example, when a user uses the news application through a terminal device, the user can upload relevant user-generated content (such as news texts written and uploaded by the user himself, news videos made by himself, etc.). The business server 100, as the server of the news application, can be a collection of multiple servers including the background server corresponding to the application client, the data processing server, etc. The business server 100 can receive the multimedia data to be recommended from the terminal device (for example, user-generated content uploaded by the user through the terminal device). In addition, the business server 100 can also receive recommendation object data for personalized recommendation of specific objects. The above recommendation object data can include the attribute information of the recommendation object. For example, the above attribute information can include the portrait information of the recommendation object (such as object level, etc.) and the behavior information of the recommendation object (such as the click-through rate, download rate, etc. of the recommendation object based on media data). The business server 100 can screen out some multimedia data from the above multimedia data to be recommended as the target recommended media based on the above recommendation object data. Or, the above multimedia data to be recommended can also be the multimedia data screened and generated in the recall stage, that is, it can be part of the multimedia data to be recommended obtained by quickly screening multiple multimedia data to be recommended input or made by users, official accounts, and institutions through the application client (such as a news application) installed in the terminal device. Specifically, it can be determined according to the actual application scenario and is not limited here. The business server 100 screens the above multimedia data to be recommended to obtain the target recommended media. Here, the business server 100 can return the above target recommended media to the above terminal device to recommend and display it to the user through the installed news application. Or, the business server 100 can further screen the above target recommended media again and return the screened target recommended media to the above terminal device to recommend and display it to the user through the installed news application. Please refer to Figure 3 , Figure 3 which is a schematic diagram of the scenario of the method for recommending multimedia data provided by an embodiment of the present application. As Figure 3As shown, the user can use the above news application as Figure 3 shown in Interface 1 in Figure 3 . Specifically, in Interface 1, multiple target recommended media recommended and presented to the user can be displayed, which may include Target Recommended Media 10, Target Recommended Media 20, and Target Recommended Media 20. Among them, each target recommended media may include corresponding visual elements. For example, Target Recommended Media 10 may include a corresponding cover image 10a, a title text 10b (which may include a multimedia data title and account information, not shown in the figure), and an abstract text 10c. Target Recommended Media 20 may include a corresponding cover image 20a, a title text 20b, and an abstract text 20c. Target Recommended Media 30 may include a corresponding cover image 30a, a title text 30b, and an abstract text 30c. The user can browse the visual elements of each displayed target recommended media and select a favorite target recommended media to further view the corresponding text, video, etc. (which may be clicking on the corresponding visual element).

[0076] The method provided by the embodiments of this application can be executed by a service server 100 as Figure 2 shown, or can be executed by a terminal device (such as Figure 2 any one of the terminal devices 200a, 200b,..., 200n shown), or can be jointly executed by the terminal device and the service server, which can be specifically determined according to the actual application scenario and is not limited here.

[0077] In some feasible embodiments, the terminal device 200a can be used as the provider of the multimedia data to be recommended. The service server 100 filters out some multimedia data from the multimedia data to be recommended obtained from the terminal device 200a (which can be the multimedia data to be recommended generated in the recall stage) as the target recommended media. A personalized recommendation model can be deployed in the service server 100. The service server 100 can obtain the multimedia data features corresponding to the multimedia data (including the recommended object data features and the multimedia data features to be recommended), and input the recommended object data features or the multimedia data features to be recommended (which can be the first multimedia data features) in the multimedia data features into the first feature extraction network in the personalized recommendation model, so as to respectively perform feature extraction on the above-mentioned first multimedia data features through each first feature extraction network to obtain multiple recommendation features. The network structures of the above-mentioned multiple first feature extraction networks can be the same (the network parameters of each first feature extraction network are different) or different. Through each first feature extraction network, feature extraction is performed based on the first multimedia data features to obtain multiple differentiated recommendation features, providing more diverse recommendation features for the subsequent personalized recommendation model to recommend multimedia data. The service server 100 can input the first multimedia data features into each first gating network. Each first gating network can determine the feature combination weights corresponding to each recommended feature input by the first gating network according to the currently input first multimedia data features. Thus, each first gating network can respectively perform weighted summation on the above-mentioned multiple recommendation features based on the feature combination weights corresponding to each recommended feature to obtain multiple weighted recommendation features. Since the feature combination weights assigned to each recommended feature by different first gating networks are different, the weighted recommendation features generated by each first gating network based on the multiple recommendation features have obvious differences in emphasis corresponding to the above-mentioned multiple first feature extraction networks. Therefore, multiple differentiated weighted recommendation features can be obtained through each first gating network, providing more diverse weighted recommendation features for the subsequent personalized recommendation model to recommend multimedia data. The service server 100 can also splice each weighted recommendation feature in the above-mentioned multiple weighted recommendation features with the second multimedia data features respectively to obtain multiple spliced vectors (here, cross features can also be obtained, and the cross features and each spliced vector in the above-mentioned multiple spliced vectors are respectively spliced to obtain updated multiple spliced vectors), and input the spliced vectors obtained by splicing each weighted recommendation feature and the second multimedia data features into the second feature extraction network corresponding to the first gating network that outputs each weighted recommendation feature for feature extraction, so as to obtain multiple differentiated spliced features, providing more diverse spliced features for the subsequent personalized recommendation model to recommend multimedia data.The service server 100 may input the above-mentioned multiple spliced vectors into each second gating network. Each of the above-mentioned second gating networks may determine the feature combination weights of each spliced feature input thereto according to the currently input spliced vector. Thus, each second gating network may respectively perform weighted summation on the above-mentioned multiple spliced features based on the feature combination weights corresponding to its respective spliced features to obtain multiple target recommendation features. The service server 100 may input the target recommendation features output by each second gating network into the prediction network corresponding to each second gating network, so as to obtain the service target prediction values corresponding to the respective target recommendation features to obtain multiple service target prediction values, and determine multiple target recommendation media from the multimedia data based on the multiple service target prediction values. The above-mentioned multiple target recommendation media are screened from multiple to-be-recommended multimedia data by integrating multiple service targets (such as duration and click-through rate, etc.), solving the problem of obtaining target recommendation media that are too biased towards partial service targets by only optimizing a single service target, and enhancing the personalized recommendation experience of multimedia data.

[0078] In some feasible embodiments, it may be that the terminal device 200a obtains a plurality of multimedia data to be recommended through the application client (such as a news application) installed thereon. A multi-modal vector generation model may be deployed in the terminal device 200a. The terminal device 200a may obtain the multimedia data features corresponding to the multimedia data (including the recommended object data features and the multimedia data features to be recommended), and input the recommended object data features or the multimedia data features to be recommended (which may be the first multimedia data features) in the multimedia data features into the first feature extraction network in the personalized recommendation model, so as to respectively perform feature extraction on the above-mentioned first multimedia data features through each first feature extraction network to obtain a plurality of recommended features. The network structures of the above-mentioned plurality of first feature extraction networks may be the same (the network parameters of each first feature extraction network are different) or different. Through each first feature extraction network, feature extraction is performed based on the first multimedia data features to obtain a plurality of differentiated recommended features, providing recommended features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. The terminal device 200a may input the first multimedia data features into each first gating network. Each first gating network may determine the feature combination weights corresponding to each recommended feature input thereto according to the currently input first multimedia data features. Thus, each first gating network may respectively perform weighted summation on the above-mentioned plurality of recommended features based on the feature combination weights corresponding to each recommended feature thereof to obtain a plurality of weighted recommended features. Since the feature combination weights assigned to each recommended feature by different first gating networks are different, the weighted recommended features generated by each first gating network based on the plurality of recommended features have obvious differences in emphasis corresponding to the above-mentioned plurality of first feature extraction networks. Therefore, through each first gating network, a plurality of differentiated weighted recommended features can be obtained, providing weighted recommended features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. The terminal device 200a may also splice each weighted recommended feature in the above-mentioned plurality of weighted recommended features with the second multimedia data features respectively to obtain a plurality of spliced vectors (here, cross features may also be obtained, and the above-mentioned cross features are respectively spliced with each spliced vector in the above-mentioned plurality of spliced vectors to obtain updated plurality of spliced vectors), and input the spliced vectors obtained by splicing each weighted recommended feature and the second multimedia data features into the second feature extraction network corresponding to the first gating network that outputs the above-mentioned weighted recommended features for feature extraction, so as to obtain a plurality of differentiated spliced features, providing spliced features with better diversity for the subsequent personalized recommendation model to recommend multimedia data.The terminal device 200a may input the above-mentioned multiple splicing vectors into each second gating network. Each of the above second gating networks may determine the feature combination weights of each second gating network for the input splicing features according to the currently input splicing vector. Thus, each second gating network may perform weighted summation on the above-mentioned multiple splicing features respectively based on the feature combination weights corresponding to its respective splicing features to obtain multiple target recommendation features. The terminal device 200a may input the target recommendation features output by each second gating network into the prediction network corresponding to each second gating network, so as to obtain the service target prediction values corresponding to each target recommendation feature to obtain multiple service target prediction values, and determine multiple target recommendation media from the multimedia data based on the multiple service target prediction values. The above-mentioned multiple target recommendation media are screened from multiple multimedia data to be recommended by integrating multiple service targets (such as duration and click-through rate, etc.), which solves the problem of obtaining target recommendation media that are too biased towards some service targets by only optimizing a single service target, and enhances the personalized recommendation experience of multimedia data.

[0079] For ease of description, hereinafter, the terminal device will be used as the execution subject of the method provided in the embodiments of the present application, and a specific embodiment will be used to specifically illustrate the method for recommending multimedia data through the terminal device.

[0080] See Figure 4 , Figure 4 is a schematic flowchart of the method for recommending multimedia data provided in the embodiments of the present application. As Figure 4 shown, the method includes the following steps:

[0081] S101, obtain multimedia data features corresponding to the multimedia data.

[0082] In some feasible embodiments, a terminal device (such as terminal device 200a) may obtain multimedia data, where the multimedia data may include recommended object data and multimedia data to be recommended. Specifically, the recommended object data may include attribute information of the recommended object. For example, the attribute information may include portrait information of the recommended object (such as object level, etc.) and behavior information of the recommended object (such as click-through rate, download rate, etc. of the recommended object based on media data). The multimedia data to be recommended may include multiple multimedia data to be recommended (which may include graphic media data, audio data, video data, etc.) input or produced by users, public accounts, and institutions through application clients (such as news applications) installed in the terminal device, and their attribute information (such as media type, theme, duration, etc.). The terminal device may obtain the multimedia data to be recommended through the news application. Among them, the news application may be an independent client, or may be an embedded sub-client integrated in a certain client (such as an instant messaging client, a social client, etc.), or may also be a web application accessed through a browser, which can be specifically determined according to the actual application scenario and is not limited herein. Alternatively, the multimedia data to be recommended may also be multimedia data screened out in the recall stage, that is, it may be part of the multimedia data to be recommended obtained by quickly screening multiple multimedia data to be recommended input or produced by users, public accounts, and institutions through application clients (such as news applications) installed in the terminal device. Taking the multimedia data to be recommended as the multimedia data screened out in the recall stage as an example, the terminal device may receive the multimedia data to be recommended obtained by screening in the recall stage.

[0083] Optionally, after vectorizing the above recommended object data and the multimedia data to be recommended, a plurality of feature vectors are obtained. The above plurality of feature vectors may include discrete feature vectors (for example, the feature vectors obtained by vectorizing the media type in the multimedia data to be recommended) and continuous feature vectors (for example, the feature vectors obtained by vectorizing the duration in the multimedia data to be recommended). Since the dimension of the discrete feature vectors is often very large, if directly input into the personalized recommendation model, the number of network parameters of the entire personalized recommendation model will be extremely large, and at the same time, the discrete feature vectors will cause the convergence of the entire personalized recommendation model to be very slow. Therefore, usually, it is necessary to first perform Embedding (embedding) compression processing on the discrete feature vectors to make them dense (eliminating useless features) to generate low-dimensional dense feature vectors, and then input the low-dimensional dense feature vectors into the personalized recommendation model for processing. Specifically, the continuous feature vectors (which can be the continuous feature vectors obtained by vectorizing the recommended object data or the multimedia data to be recommended) can be discretized to obtain the discrete feature vectors corresponding to the continuous feature vectors, and then the Embedding processing is performed on the discrete feature vectors corresponding to the continuous feature vectors to obtain the encoded features of the continuous feature vectors (which can be the recommended object data features or the multimedia data to be recommended features). The Embedding processing is directly performed on the discrete feature vectors (which can be the discrete feature vectors obtained by vectorizing the recommended object data or the multimedia data to be recommended) to obtain the encoded features of the discrete feature vectors (which can be the recommended object data features or the multimedia data to be recommended features). The above recommended object data features and the multimedia data to be recommended features constitute multimedia data features (which can include the first multimedia data features and the second multimedia data features). Here, it can be that the first multimedia data features are the recommended object data features and the second multimedia data features are the multimedia data to be recommended features, or it can also be that the first multimedia data features are the multimedia data to be recommended features and the second multimedia data features are the recommended object data features. The following will take the first multimedia data features being the recommended object data features and the second multimedia data features being the multimedia data to be recommended features as an example for illustration, and will not be elaborated below.

[0084] In some feasible embodiments, the above-mentioned recommended object data may include recommended object data corresponding to multiple objects (for example, it may include portrait information and behavior information respectively corresponding to multiple objects), and the above-mentioned multimedia data to be recommended may include multimedia data to be recommended corresponding to multiple multimedia data (for example, it may include multiple multimedia data and their respective corresponding attribute information). The above-mentioned recommended object data features may include recommended object data features corresponding to multiple objects (for example, the recommended object data features of object 1, the recommended object data features of object 2, and the recommended object data features of object 3, etc.), and the above-mentioned multimedia data to be recommended features may include multimedia data to be recommended features corresponding to multiple multimedia data (for example, the multimedia data to be recommended features of multimedia data 1, the multimedia data to be recommended features of multimedia data 2, and the multimedia data to be recommended features of multimedia data 3, etc.). That is, the terminal device can determine multiple target recommended media corresponding to one object, and the terminal device can also simultaneously determine multiple target recommended media corresponding to each object among multiple objects. Specifically, it can be determined according to the actual application scenario, and the embodiments of the present application do not limit this here.

[0085] S102, perform feature extraction on the first multimedia data feature through multiple first feature extraction networks to obtain multiple recommended features corresponding to the first multimedia data feature, and obtain multiple weighted recommended features corresponding to the multiple recommended features through multiple first gating networks.

[0086] In some feasible embodiments, the terminal device may input the first multimedia data feature among the above-mentioned multimedia data features into multiple expert networks (for the convenience of description, the first feature extraction network may be taken as an example for illustration) to respectively perform feature extraction on the first multimedia data feature through each first feature extraction network to obtain multiple recommended features. Specifically, each of the above-mentioned first feature extraction networks is used to obtain a recommended feature corresponding to the first multimedia data feature. Here, the first multimedia data feature may be a recommended object data feature, and the recommended object data feature may include recommended object data features corresponding to one or more objects. It can be understood that when the recommended object data feature includes recommended object data features corresponding to multiple objects (that is, when the terminal device simultaneously determines multiple target recommended media corresponding to each object among multiple objects), a recommended feature corresponding to the first multimedia data feature may include recommended features corresponding to each recommended object data feature obtained through the first feature extraction network based on each recommended object data feature among the multiple recommended object data features. The network structures of the multiple first feature extraction networks may be the same (the network parameters of each first feature extraction network are different). For example, the network structures of the multiple first feature extraction networks may all be Deep Neural Networks (DNN) structures, or may all be Factorization Machine (FM) structures, or may all be Deep&Cross Network (DCN) structures. Or, the network structures of the multiple first feature extraction networks may be a combination of at least two network structures among the deep neural network structure, the factorization machine structure, and the deep cross network structure, which can be specifically determined according to the actual application scenario, and the embodiments of the present application do not limit this here. By performing feature extraction on the first multimedia data feature through each first feature extraction network, multiple differentiated recommended features are obtained, providing recommended features with better diversity for the subsequent personalized recommendation model to perform multimedia data recommendation.

[0087] Further, the terminal device can obtain multiple weighted recommendation features corresponding to the above-mentioned multiple recommendation features through multiple first gating networks. Specifically, the terminal device can input the first multimedia data feature into each first gating network. Each first gating network can determine the feature combination weights corresponding to the recommendation features input thereto based on the currently input first multimedia data feature. Thus, each first gating network can respectively perform weighted summation on the above-mentioned multiple recommendation features based on the feature combination weights corresponding to the recommendation features thereof to obtain multiple weighted recommendation features. Since the feature combination weights assigned by different first gating networks to the recommendation features are different, the weighted recommendation features generated by each first gating network based on the multiple recommendation features have obvious differences in emphasis corresponding to the above-mentioned multiple first feature extraction networks (for example, the first gating network 1 assigns more weights to the recommendation features output by the first feature extraction network 1, and the first gating network 2 assigns more weights to the recommendation features output by the first feature extraction network 2). Therefore, multiple differentiated weighted recommendation features can be obtained through each first gating network, providing weighted recommendation features with better diversity for the subsequent personalized recommendation model to perform multimedia data recommendation. Optionally, the above-mentioned first gating network can include at least one of a gating network based on linear transformation or a gating network based on normalized weighting. It can be specifically determined according to the actual application scenario, and the embodiments of the present application do not limit this here. For example, if the first multimedia data feature is U (the dimension can be L) and the number of first gating networks is N and the number of first feature extraction networks is M, the feature combination weights corresponding to the recommendation features determined by the nth first gating network can be expressed as:

[0088]

[0089] Among them, is a parameter matrix with a dimension of M×L, that is, G n (U) can have a dimension of M, and the values of each dimension respectively represent the feature combination weights of the recommendation features generated by the corresponding first feature extraction network.

[0090] S103. Obtain multiple concatenated vectors obtained by concatenating each of the weighted recommendation features in the multiple weighted recommendation features with the second multimedia data feature respectively, and based on the second feature extraction network corresponding to each first gating network that obtains each weighted recommendation feature, obtain the concatenated features of each concatenated vector obtained based on each weighted recommendation feature to obtain multiple concatenated features.

[0091] In some feasible embodiments, the terminal device may splice each of the multiple weighted recommendation features with the second multimedia data feature respectively to obtain multiple spliced vectors, and input each of the spliced vectors obtained by splicing each weighted recommendation feature and the second multimedia data feature into a second feature extraction network corresponding to the first gating network that outputs each of the weighted recommendation features. Specifically, one first gating network may correspond to one second feature extraction network, that is, the number of the first gating networks may be the same as the number of the second feature extraction networks. The spliced vector obtained by splicing the weighted recommendation feature output by any one of the first gating networks and the second multimedia data feature may be input into the second feature extraction network corresponding to the first gating network for feature extraction to obtain the spliced feature corresponding to the spliced vector. Here, the network structures of the multiple second feature extraction networks may all be deep neural network structures, may all be factorization machine structures, or may all be deep cross network structures. Or, the network structures of the multiple second feature extraction networks may be a combination of at least two network structures among deep neural network structures, factorization machine structures, and deep cross network structures, which can be specifically determined according to the actual application scenario, and the embodiments of the present application do not limit this here. By each second feature extraction network performing feature extraction based on the spliced vector, multiple differentiated spliced features are obtained, providing spliced features with better diversity for the subsequent personalized recommendation model to perform multimedia data recommendation.

[0092] In some feasible embodiments, after the terminal device obtains multiple spliced vectors obtained by splicing each of the multiple weighted recommendation features with the second multimedia data feature respectively, the terminal device may further obtain cross features, and splice the cross features with each of the multiple spliced vectors respectively to obtain updated multiple spliced vectors. Specifically, the cross features may include cross features obtained based on cross data generated by cross-correlating the portrait information of the sample recommended object and the sample recommended multimedia data, and the cross features may further include cross features obtained based on cross data generated by cross-correlating the behavior information of the sample recommended object and the sample recommended multimedia data. Taking the cross data generated by cross-correlating the portrait information of the sample recommended object and the sample recommended multimedia data as an example, the portrait information of the sample recommended object may be the object level, and the sample recommended multimedia data (such as the attribute information of the multimedia data to be recommended) may be the theme. Then, distribution data of different object levels and different themes may be generated based on the object level and the theme as cross data (for example, it may be the statistical distribution of sample recommended objects of different levels corresponding to each theme under theme A and theme B). By adding cross features to obtain updated multiple spliced vectors, the internal correlation information between the recommended object and the multimedia data can be introduced into the spliced vector, so as to further enhance the personalized recommendation effect of the personalized recommendation model.

[0093] It can be understood that the above-mentioned multiple first feature extraction networks and multiple first gating networks can form a multi-gated mixture-of-experts network (which can be called the lower-layer multi-gated mixture-of-experts network here). When the above-mentioned first multimedia data feature is the recommended object data feature, that is, the lower-layer multi-gated mixture-of-experts network only involves processing the recommended object data features of the recommended object (such as a user) (including feature extraction, weighted summation, etc.). Therefore, in the process of personalized recommendation of a recommended object based on a large amount of multimedia data (such as thousands or tens of thousands) through a personalized recommendation model, the lower-layer multi-gated mixture-of-experts network can process a small number of recommended object data features to obtain corresponding weighted recommendation features, and copy the weighted recommendation features to obtain weighted recommendation features with the same quantity as the multimedia data features (i.e., the second multimedia data features), so as to splice the weighted recommendation features and the second multimedia data features to obtain multiple spliced vectors, and then the personalized recommendation model performs multi-objective estimation based on the multiple spliced vectors to generate target recommended media (which can be multi-objective estimation through the upper-layer multi-gated mixture-of-experts network), further improving the multi-objective estimation efficiency and enhancing the multi-objective estimation performance.

[0094] S104, input the multiple spliced vectors and the splicing features of each spliced vector into a business objective prediction model for obtaining recommended media, so as to obtain multiple target recommended media based on the business objective prediction model.

[0095] In some feasible embodiments, the above business objective prediction model may include multiple second gating networks and multiple prediction networks. Among them, one prediction network in the multiple prediction networks corresponds to one second gating network. The terminal device may input the multiple concatenated vectors into each second gating network. Each second gating network may determine the feature combination weights of each concatenated feature input to the second gating network according to the currently input concatenated vector. Thus, each second gating network may respectively perform weighted summation on the multiple concatenated features based on the feature combination weights corresponding to the respective concatenated features thereof to obtain multiple target recommendation features. Since different second gating networks assign different feature combination weights to each concatenated feature, the target recommendation features generated by each second gating network based on the multiple concatenated features have obvious differences in emphasis corresponding to the multiple second feature extraction networks (for example, the second gating network 1 assigns more weights to the concatenated features output by the second feature extraction network 1, and the second gating network 2 assigns more weights to the concatenated features output by the second feature extraction network 2). Therefore, multiple differentiated weighted concatenated features can be obtained through each second gating network, providing weighted concatenated features with better diversity for the subsequent personalized recommendation model to perform multimedia data recommendation. Optionally, the above second gating network may include at least one of a gating network based on linear transformation or a gating network based on normalized weighting. It can be specifically determined according to the actual application scenario, and the embodiments of the present application do not limit this here. For example, if the number of the first gating networks is N, the number of the second feature extraction networks is also N, the number of the first feature extraction networks is M, and the output of the m-th first feature extraction network can be expressed as u m (U), then the weighted recommendation feature output by the n-th first gating network can be expressed as:

[0096]

[0097] Among them, represents the feature combination weight determined by the n-th first gating network for the recommendation feature output by the m-th first feature extraction network. Concatenating each weighted recommendation feature in the above N weighted recommendation features with the second multimedia data feature (which can be expressed as i) and the cross feature (which can be expressed as c) respectively to obtain N concatenated vectors. The n-th concatenated vector can be expressed as:

[0098] v n = concat(u n , i, c)

[0099] If the number of the first feature extraction networks is K, the feature combination weights corresponding to each concatenated feature determined by the k-th second gating network can be expressed as:

[0100]

[0101] Among them is a parameter matrix with dimensions N×D, and V represents a D-dimensional vector formed by concatenating the outputs of all first gating networks, the second multimedia data features, and the cross features, that is

[0102] V = concat(u 1 , …, u n , i, c)

[0103] Thus, the target recommendation feature output by the k-th second gating network can be expressed as:

[0104]

[0105] Among them represents the feature combination weight determined by the k-th second gating network for the concatenated feature output by the n-th second feature extraction network, and f n (v n ) represents the concatenated feature output by the n-th second feature extraction network. The terminal device can input the target recommendation features output by each second gating network into the prediction network corresponding to each second gating network, so as to obtain the service target prediction values corresponding to each target recommendation feature to obtain multiple service target prediction values, and determine multiple target recommendation media from the multimedia data based on the multiple service target prediction values. Specifically, each prediction network in the above multiple prediction networks corresponds one-to-one with a second gating network, and each prediction network can carry the service target prediction values for generating different service targets.

[0106] In some feasible implementation manners, the service target may include click-through rate, conversion rate, click conversion rate, and duration, etc., which can be specifically determined according to the actual application scenario and are not limited herein. Among them, the click-through rate can be the ratio of the number of clicks to the number of exposures, which can be used to indicate the popularity of the multimedia data to be recommended. The conversion rate is the ratio of the number of conversions to the number of clicks, which can be used to indicate the popularity of the multimedia data to be recommended. The click conversion rate is the ratio of the number of conversions to the number of exposures, that is, the product of the click-through rate and the conversion rate, which can be used to indicate the popularity of the multimedia data to be recommended. The duration refers to the browsing duration of the multimedia data to be recommended, such as reading duration, stay duration, etc., which can be used to indicate the degree of interest of the recommended object in the multimedia data to be recommended.

[0107] In some feasible embodiments, the predicted value of the business objective is used to predict the operation behavior of the recommended object for the to-be-recommended multimedia data. Optionally, the predicted value of the business objective may be the predicted probability value of the business objective or the predicted value of the business objective, which can be specifically determined according to the actual application scenario and is not limited in the embodiments of the present application. Taking the predicted value of the business objective being the predicted probability value of the business objective as an example, if the business objective is duration, if the prediction result is 1, it can indicate that the browsing duration of the recommended object for the to-be-recommended multimedia data will exceed the threshold duration; if the prediction result is 0, it can indicate that the browsing duration of the recommended object for the to-be-recommended multimedia data will not exceed the threshold duration; if the prediction result is a value between 0 and 1 (such as 0.5), it can indicate that there is a 50% probability that the browsing duration of the recommended object for the to-be-recommended multimedia data will exceed the threshold duration. For example, if the prediction network includes prediction network 1 and prediction network 2, and the business objectives corresponding to prediction network 1 and prediction network 2 are duration and click-through rate respectively, taking the terminal device to determine multiple target recommended media corresponding to an object for an object as an example, the terminal device can obtain the predicted values of the business objectives of each to-be-recommended multimedia data in the multiple to-be-recommended multimedia data through the above prediction network, that is, the duration prediction result obtained by each to-be-recommended multimedia data through prediction network 1 and the click-through rate prediction result obtained by prediction network 2, so that each to-be-recommended multimedia data has corresponding multiple predicted values of business objectives. Assume the number of prediction networks is K, the scoring hidden layer of the k-th prediction network is h k , and the target recommended feature output by the k-th second gating network is f k , then the output of the k-th prediction network can be expressed as:

[0108] p k =h k (f k )

[0109] The terminal device can determine multiple target recommended media from the multiple to-be-recommended multimedia data based on the multiple to-be-recommended multimedia data and the multiple predicted values of the business objectives corresponding to each to-be-recommended multimedia data. The above multiple target recommended media are screened from the multiple to-be-recommended multimedia data by comprehensively considering multiple business objectives (such as duration and click-through rate), which solves the problem of obtaining target recommended media that are too biased towards some business objectives by only optimizing a single business objective, and enhances the personalized recommendation experience of multimedia data.

[0110] See Figure 5 , Figure 5It is a schematic diagram of the personalized recommendation model structure provided by the embodiments of the present application. In the method provided by the embodiments of the present application, the terminal device can obtain the multimedia data features corresponding to the multimedia data, and input the recommended object data features or the to-be-recommended multimedia data features (which may be the first multimedia data features) in the multimedia data features into the first feature extraction network in the personalized recommendation model, so as to respectively perform feature extraction on the above-mentioned first multimedia data features through each first feature extraction network to obtain multiple recommendation features. As Figure 5 shown, the first multimedia data features are input into the first feature extraction network 1, the first feature extraction network 2,..., the first feature extraction network M, and the recommendation features are respectively output through the first feature extraction network 1, the first feature extraction network 2,..., the first feature extraction network M. The network structures of the above-mentioned multiple first feature extraction networks can be the same (the network parameters of each first feature extraction network are different) or different. By each first feature extraction network performing feature extraction based on the first multimedia data features, multiple differentiated recommendation features are obtained, providing recommendation features with better diversity for the subsequent personalized recommendation model to perform multimedia data recommendation. The terminal device can input the first multimedia data features into each first gating network, such as Figure 5 the first gating network 1, the first gating network 2,..., the first gating network N in. Each first gating network can determine the feature combination weights corresponding to each of the recommendation features input to it according to the currently input first multimedia data features, so that each first gating network can respectively perform weighted summation on the above-mentioned multiple recommendation features based on the feature combination weights corresponding to each of the recommendation features input to it to obtain multiple weighted recommendation features. Since the feature combination weights assigned by different first gating networks to each recommendation feature are different, the weighted recommendation features generated by each first gating network based on the multiple recommendation features have obvious differences in emphasis corresponding to the above-mentioned multiple first feature extraction networks. Therefore, multiple differentiated weighted recommendation features can be obtained through each first gating network, providing weighted recommendation features with better diversity for the subsequent personalized recommendation model to perform multimedia data recommendation. The above-mentioned first gating network can include at least one of a gating network based on linear transformation or a gating network based on normalized weighting. The terminal device can also splice each of the weighted recommendation features in the above-mentioned multiple weighted recommendation features with the second multimedia data features to obtain multiple spliced vectors (here, cross features can also be obtained, and the cross features are respectively spliced with each of the above-mentioned multiple spliced vectors to obtain updated multiple spliced vectors), and input the spliced vectors obtained by splicing each weighted recommendation feature and the second multimedia data features into the second feature extraction network corresponding to the first gating network that outputs each of the weighted recommendation features, such as Figure 5The second feature extraction networks 1, 2, …, N in [description] respectively perform feature extraction based on the concatenated vectors through the second feature extraction networks 1, 2, …, N to obtain multiple differentiated concatenated features, providing concatenated features with better diversity for subsequent multimedia data recommendation by the personalized recommendation model. The terminal device can input the above multiple concatenated vectors into each second gating network, such as Figure 5 the second gating networks 1, …, K in [description]. Each of the above second gating networks can determine the feature combination weights of each concatenated feature input to it according to the currently input concatenated vector. Thus, each second gating network can respectively perform weighted summation on the above multiple concatenated features based on the feature combination weights corresponding to its respective concatenated features to obtain multiple target recommendation features. The terminal device can input the target recommendation features output by each second gating network into the prediction networks corresponding to each second gating network, such as Figure 5 the prediction networks 1, …, K in [description] (the second gating networks 1, …, K and the prediction networks 1, …, K are in one-to-one correspondence respectively), so as to obtain the service target prediction values corresponding to each target recommendation feature to obtain multiple service target prediction values, and determine multiple target recommendation media from the multimedia data based on the multiple service target prediction values. The above multiple target recommendation media are screened from multiple multimedia data to be recommended by integrating multiple service targets (such as duration and click-through rate, etc.), solving the problem of obtaining target recommendation media that are too biased towards some service targets by only optimizing a single service target, and enhancing the personalized recommendation experience of multimedia data.

[0111] In some feasible embodiments, the above multiple first feature extraction networks and multiple first gating networks can form a multi-gating mixture of experts network (here, it can be called the lower multi-gating mixture of experts network). In addition, the above multiple second feature extraction networks and multiple second gating networks can also form a multi-gating mixture of experts network (here, it can be called the upper multi-gating mixture of experts network). It can be understood that the personalized recommendation model provided by the embodiments of the present application can be a personalized recommendation model obtained by improving a multi-gating mixture of experts network (for example, it can refer to Figure 1 the multi-gating mixture of experts network in [description]), and it can be the tower structure in the multi-gating mixture of experts network (for example, Figure 1The tower structures 1 and 2) in are replaced with another multi-gated mixture-of-experts network that includes multiple feature extraction networks (which can also be called expert networks) and multiple gating networks. That is, the tower structures in the upper multi-gated mixture-of-experts network above are replaced with the lower multi-gated mixture-of-experts network above to obtain the personalized recommendation model provided by the embodiments of the present application. Thus, the problem that after the tower structure reduces the dimensionality of the recommended object data and the to-be-recommended multimedia data, the original feature signals contained in the recommended object data and the to-be-recommended multimedia data are all lost and the ability of different expert networks in the multi-gated mixture-of-experts network to learn differential representations end-to-end is restricted is solved. By improving the tower structure to the lower multi-gated mixture-of-experts network associated with the upper multi-gated mixture-of-experts network, the lower multi-gated mixture-of-experts network can provide a differential low-dimensional user vector representation (such as differential recommendation features) for the upper multi-gated mixture-of-experts network, improving the effect of multi-object prediction while ensuring the multi-object prediction performance of the multi-gated mixture-of-experts network, thereby improving the screening effectiveness of the to-be-recommended multimedia data in the rough ranking stage and having high applicability.

[0112] See Figure 6 , Figure 6 is a schematic diagram for comparing the experimental effects of the method for recommending multimedia data provided by the embodiments of the present application. As Figure 6 shown, Figure 6 Chart 601 in the figure is a comparison result chart generated by the experimental group (adopting the technical solution of the present application) and the control group (adopting the technical solution provided by other related technologies) based on the duration service target (such as the browsing duration of graphic content) during the idle running period. Chart 602 is a comparison result chart generated by the experimental group and the control group based on the duration service target during the experimental period. Chart 603 is a bar chart of the result difference comparison between the experimental group and the control group based on the duration service target during the idle running period and the experimental period (a positive value represents that the experimental group is better than the control group, and a negative value represents that the experimental group lags behind the control group). It can be seen that during the experimental period, the experimental group adopting the method for recommending multimedia data proposed in the present application can achieve better duration results compared to the control group (during the experimental period, the duration results of the experimental group are always better than those of the control group, and the leading percentages are: 0.64%, 0.67%, 0.84%, 1.05%, 1.17%, and 0.99% respectively). That is, by using the method for recommending multimedia data proposed in the present application, the target recommended media that can better meet the user's needs can be screened out from the to-be-recommended multimedia data, and the personalized recommendation effect of the multimedia data is good.

[0113] In the method provided by the embodiments of the present application, the terminal device can obtain the multimedia data features corresponding to the multimedia data (including the recommended object data features and the multimedia data features to be recommended), and input the recommended object data features or the multimedia data features to be recommended (which can be the first multimedia data features) in the multimedia data features into the first feature extraction network in the personalized recommendation model, so as to respectively perform feature extraction on the above-mentioned first multimedia data features through each first feature extraction network to obtain multiple recommended features. By each first feature extraction network performing feature extraction based on the first multimedia data features, multiple differentiated recommended features are obtained, providing recommended features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. The terminal device can input the first multimedia data features into each first gating network, and each first gating network can determine the feature combination weights corresponding to the recommended features input by it according to the currently input first multimedia data features. Thus, each first gating network can respectively perform weighted summation on the above-mentioned multiple recommended features based on the feature combination weights corresponding to the recommended features corresponding to it to obtain multiple weighted recommended features. Since the feature combination weights assigned to the recommended features by different first gating networks are different, the weighted recommended features generated by each first gating network based on the multiple recommended features have obvious differences in emphasis corresponding to the above-mentioned multiple first feature extraction networks. Therefore, multiple differentiated weighted recommended features can be obtained through each first gating network, providing weighted recommended features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. The terminal device can also splice each weighted recommended feature in the above-mentioned multiple weighted recommended features with the second multimedia data features respectively to obtain multiple spliced vectors (here, cross features can also be obtained, and the above-mentioned cross features and each spliced vector in the above-mentioned multiple spliced vectors are respectively spliced to obtain updated multiple spliced vectors), and input the spliced vectors obtained by splicing each weighted recommended feature and the second multimedia data features into the second feature extraction network corresponding to the first gating network outputting the above-mentioned weighted recommended features for feature extraction to obtain multiple differentiated spliced features, providing spliced features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. The terminal device can input the above-mentioned multiple spliced vectors into each second gating network, and each of the above-mentioned second gating networks can determine the feature combination weights of the spliced features input by it according to the currently input spliced vectors. Thus, each second gating network can respectively perform weighted summation on the above-mentioned multiple spliced features based on the feature combination weights corresponding to the spliced features corresponding to it to obtain multiple target recommended features. The terminal device can input the target recommended features output by each second gating network into the prediction network corresponding to each second gating network, so as to obtain the service target prediction values corresponding to the target recommended features to obtain multiple service target prediction values, and determine multiple target recommended media from the multimedia data based on the multiple service target prediction values.The above-mentioned multiple target recommended media are selected from multiple multimedia data to be recommended by integrating multiple service objectives (such as duration and click-through rate, etc.), which solves the problem that the target recommended media obtained by only optimizing a single service objective is too biased towards some service objectives, and enhances the personalized recommendation experience of multimedia data.

[0114] Based on the description of the embodiment of the recommendation method for multimedia data above, an embodiment of the present application also discloses a recommendation device for multimedia data. This recommendation device for multimedia data can be applied to Figures 4 to 5 the recommendation method for multimedia data in the shown embodiment to execute the steps in the recommendation method for multimedia data. Here, the recommendation device for multimedia data can be the Figures 4 to 5 service server or terminal device in the shown embodiment above, that is, the recommendation device for multimedia data can be the Figures 4 to 5 execution subject of the recommendation method for multimedia data in the shown embodiment above. Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the recommendation device for multimedia data provided by an embodiment of the present application. In the embodiment of the present application, the device can operate the following modules:

[0115] An acquisition module 41, configured to acquire multimedia data features corresponding to multimedia data. The multimedia data features include a first multimedia data feature and a second multimedia data feature. The first multimedia data feature is a recommended object data feature and the second multimedia data feature is a multimedia data feature to be recommended, or the first multimedia data feature is a multimedia data feature to be recommended and the second multimedia data feature is a recommended object data feature;

[0116] A weighted recommendation feature generation module 42, configured to perform feature extraction on the first multimedia data feature through multiple first feature extraction networks to obtain multiple recommendation features corresponding to the first multimedia data feature, and obtain multiple weighted recommendation features corresponding to the multiple recommendation features through multiple first gating networks, where one of the first feature extraction networks is used to obtain one recommendation feature corresponding to the first multimedia data feature, and one of the first gating networks is used to obtain one weighted recommendation feature corresponding to the multiple recommendation features;

[0117] A splicing feature generation module 43, configured to obtain multiple splicing vectors by splicing each weighted recommendation feature in the multiple weighted recommendation features with the second multimedia data feature respectively, and obtain splicing features based on each splicing vector through a second feature extraction network corresponding to each first gating network for obtaining each weighted recommendation feature to obtain multiple splicing features;

[0118] A target recommended media generation module 44, configured to input the multiple concatenated vectors and the concatenation features of the respective concatenated vectors into a business objective prediction model for obtaining recommended media, so as to obtain multiple target recommended media based on the business objective prediction model.

[0119] In some feasible implementation manners, the weighted recommended feature generation module 42 is further configured to:

[0120] Obtain the feature combination weights corresponding to the respective recommended features used by the respective first gating networks to obtain weighted recommended features based on the first multimedia data feature;

[0121] Perform weighted summation on the multiple recommended features through any one of the first gating networks based on the feature combination weights corresponding to the respective recommended features used by any one of the first gating networks, so as to obtain one weighted recommended feature obtained by any one of the first gating networks;

[0122] Obtain the respective weighted recommended features obtained by the respective first gating networks to obtain the multiple weighted recommended features;

[0123] Wherein, the first gating network includes at least one of a gating network based on linear transformation or a gating network based on normalized weighting.

[0124] In some feasible implementation manners, after the splicing feature generation module 43 obtains the multiple concatenated vectors obtained by respectively splicing each of the weighted recommended features in the multiple weighted recommended features and the second multimedia data feature, the splicing feature generation module 43 is further configured to:

[0125] Obtain cross features, and respectively splice the cross features with each of the multiple concatenated vectors to obtain updated multiple concatenated vectors.

[0126] In some feasible implementation manners, the business objective prediction model includes multiple second gating networks and prediction networks corresponding to the respective second gating networks;

[0127] After inputting the multiple concatenated vectors and the concatenation features of the respective concatenated vectors into the business objective prediction model for obtaining recommended media, the target recommended media generation module 44 is further configured to:

[0128] Based on the above-mentioned multiple concatenated vectors, obtain the feature combination weights corresponding to the respective concatenated features used by each of the above-mentioned second gating networks to obtain the target recommendation features. Through any one of the above-mentioned second gating networks, perform weighted summation on the above-mentioned multiple concatenated features based on the feature combination weights corresponding to the respective concatenated features used by any one of the above-mentioned second gating networks to obtain a target recommendation feature obtained by any one of the above-mentioned second gating networks, and obtain the respective target recommendation features obtained by each of the above-mentioned second gating networks to obtain the above-mentioned multiple target recommendation features;

[0129] Based on the prediction networks corresponding to the respective second gating networks that obtain the above-mentioned respective target recommendation features, obtain the service target prediction values corresponding to the above-mentioned respective target recommendation features to obtain multiple service target prediction values, and determine multiple target recommendation media from the above-mentioned multimedia data based on the above-mentioned multiple service target prediction values;

[0130] Wherein, the above-mentioned second gating network includes at least one of a gating network based on linear transformation or a gating network based on normalized weighting.

[0131] In some feasible implementation manners, the above-mentioned first multimedia data feature is a recommended object data feature and the above-mentioned second multimedia data feature is a multimedia data feature to be recommended;

[0132] The above-mentioned concatenated feature generation module 43 is further configured to:

[0133] Extract features from the above-mentioned recommended object data features through multiple first feature extraction networks to obtain multiple recommended object features corresponding to the above-mentioned recommended object data features, and obtain multiple weighted recommended object features corresponding to the above-mentioned multiple recommended object features through multiple first gating networks as multiple weighted recommended features;

[0134] The above-mentioned obtaining multiple concatenated vectors obtained by concatenating each of the above-mentioned weighted recommended features in the above-mentioned multiple weighted recommended features with the above-mentioned second multimedia data feature respectively includes:

[0135] Obtain multiple concatenated vectors obtained by concatenating each of the above-mentioned weighted recommended object features in the above-mentioned multiple weighted recommended object features with the above-mentioned multimedia data feature to be recommended respectively.

[0136] In some feasible implementation manners, the above-mentioned first multimedia data feature is a multimedia data feature to be recommended and the above-mentioned second multimedia data feature is a recommended object data feature;

[0137] The above-mentioned concatenated feature generation module 43 is further configured to:

[0138] Extracting features of the multimedia data to be recommended through a plurality of first feature extraction networks to obtain a plurality of recommended multimedia features corresponding to the multimedia data to be recommended, and obtaining a plurality of weighted recommended multimedia features corresponding to the plurality of recommended multimedia features as a plurality of weighted recommendation features through a plurality of first gating networks;

[0139] The multiple splicing vectors obtained by splicing each weighted recommendation feature in the multiple weighted recommendation features and the second multimedia data feature include:

[0140] A plurality of concatenated vectors are obtained by concatenating each of the plurality of weighted recommended multimedia features and the recommended object data feature.

[0141] In some feasible implementations, the acquisition module 41 is further used for:

[0142] Acquire multimedia data, the multimedia data including recommendation object data and multimedia data to be recommended, the multimedia data to be recommended including at least one of graphic media data, audio data and video data;

[0143] The recommended object data features corresponding to the recommended object data in the multimedia data are obtained through vectorization processing and embedded compression processing, and the multimedia data features to be recommended corresponding to the multimedia data to be recommended are obtained.

[0144] According to the above Figure 4 The corresponding embodiment, Figure 4 The implementation method described in steps S101 to S104 in the multimedia data recommendation method shown in FIG. Figure 7 Each module of the device shown in the figure is executed. For example, the above Figure 4 The implementation method described in step S101 of the multimedia data recommendation method shown in FIG. Figure 7 The implementation method described in step S102 can be performed by the weighted recommendation feature generation module 42, the implementation method described in step S103 can be performed by the splicing feature generation module 43, and the implementation method described in step S104 can be performed by the target recommended media generation module 44. The implementation methods performed by the acquisition module 41, the weighted recommendation feature generation module 42, the splicing feature generation module 43, and the target recommended media generation module 44 can be referred to in the above Figure 4 The implementation methods provided in each step of the corresponding embodiment will not be repeated here.

[0145] In the embodiments of the present application, a multimedia data recommendation device can obtain multimedia data features corresponding to multimedia data (including recommended object data features and multimedia data features to be recommended), and input the recommended object data features or the multimedia data features to be recommended (which may be the first multimedia data features) in the multimedia data features into the first feature extraction network in the personalized recommendation model, so as to respectively perform feature extraction on the above-mentioned first multimedia data features through each first feature extraction network to obtain multiple recommended features. The network structures of the above-mentioned multiple first feature extraction networks may be the same (the network parameters of each first feature extraction network are different) or different. By each first feature extraction network performing feature extraction based on the first multimedia data features, multiple differentiated recommended features are obtained, providing recommended features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. The multimedia data recommendation device can input the first multimedia data features into each first gating network, and each first gating network can determine the feature combination weights corresponding to each recommended feature input thereto according to the currently input first multimedia data features. Thus, each first gating network can respectively perform weighted summation on the above-mentioned multiple recommended features based on the feature combination weights corresponding to each recommended feature input thereto to obtain multiple weighted recommended features. Since the feature combination weights assigned to each recommended feature by different first gating networks are different, the weighted recommended features generated by each first gating network based on the multiple recommended features have obvious differences in emphasis corresponding to the above-mentioned multiple first feature extraction networks. Therefore, multiple differentiated weighted recommended features can be obtained through each first gating network, providing weighted recommended features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. The multimedia data recommendation device can also splice each weighted recommended feature in the above-mentioned multiple weighted recommended features with the second multimedia data features respectively to obtain multiple spliced vectors (here, cross features can also be obtained, and the above-mentioned cross features and each spliced vector in the above-mentioned multiple spliced vectors are respectively spliced to obtain updated multiple spliced vectors), and input the spliced vectors obtained by splicing each weighted recommended feature and the second multimedia data features into the second feature extraction network corresponding to the first gating network outputting each weighted recommended feature for feature extraction to obtain multiple differentiated spliced features, providing spliced features with better diversity for the subsequent personalized recommendation model to recommend multimedia data. The multimedia data recommendation device can input the above-mentioned multiple spliced vectors into each second gating network, and each of the above-mentioned second gating networks can determine the feature combination weights of each spliced feature input thereto according to the currently input spliced vector. Thus, each second gating network can respectively perform weighted summation on the above-mentioned multiple spliced features based on the feature combination weights corresponding to each spliced feature input thereto to obtain multiple target recommended features.The recommendation device for multimedia data may input the target recommendation features output by each second gating network into the prediction network corresponding to each second gating network, so as to obtain the service target prediction values corresponding to each target recommendation feature to obtain multiple service target prediction values, and determine multiple target recommendation media from the multimedia data based on the multiple service target prediction values. The above-mentioned multiple target recommendation media are screened from multiple to-be-recommended multimedia data by integrating multiple service targets (such as duration and click-through rate, etc.), which solves the problem of obtaining target recommendation media that are too biased towards some service targets by only optimizing a single service target, and enhances the personalized recommendation experience of multimedia data.

[0146] In the embodiments of the present application, each of the above Figure 7 modules in the shown device may be separately or all combined into one or several other modules to form, or some of them may be further split into multiple smaller modules in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module may also be implemented by multiple modules, or the functions of multiple modules may be implemented by one module. In other feasible implementation manners of the present application, the above device may also include other modules. In practical applications, these functions may also be assisted by other modules and may be implemented by multiple modules collaborating, which is not limited herein.

[0147] Please refer to Figure 8 , Figure 8 which is the structural schematic diagram of the computer device provided by the embodiments of the present application. As Figure 8 shown, the computer device 1000 may be the terminal device in the corresponding embodiment of the above Figures 4 to 7 . The computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 8As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0148] Among them, the network interface 1004 in the computer device 1000 may also be network-connected to the terminal 200a in the corresponding embodiment described above. Figure 4 And the optional user interface 1003 may further include a display screen (Display) and a keyboard (Keyboard). In the Figure 8 computer device 1000 shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for users (or developers); and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement the Figure 4 recommendation method for multimedia data in the corresponding embodiment described above.

[0149] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the recommendation method for multimedia data in the corresponding embodiment described above, which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Figure 4

[0150] In addition, it should be noted here that the embodiments of the present application also provide a computer-readable storage medium, and the computer program executed by the recommendation device for multimedia data mentioned above is stored in the computer-readable storage medium. The computer program includes program instructions. When the processor executes the program instructions, it can execute the Figure 4 description of the recommendation method for multimedia data in the corresponding embodiment described above. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.

[0151] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0152] ​The above disclosure is only for the preferred embodiments of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A method for recommending multimedia data, characterized in that, The method includes: Obtaining multimedia data features corresponding to multimedia data, where the multimedia data features include first multimedia data features and second multimedia data features. The first multimedia data features are recommended object data features and the second multimedia data features are multimedia data features to be recommended, or the first multimedia data features are multimedia data features to be recommended and the second multimedia data features are recommended object data features; Performing feature extraction on the first multimedia data features through a plurality of first feature extraction networks to obtain a plurality of recommended features corresponding to the first multimedia data features, and obtaining a plurality of weighted recommended features corresponding to the plurality of recommended features through a plurality of first gating networks. Among them, one of the first feature extraction networks is used to obtain one recommended feature corresponding to the first multimedia data features, and one of the first gating networks is used to obtain one weighted recommended feature corresponding to the plurality of recommended features; Obtaining a plurality of concatenated vectors obtained by concatenating each of the weighted recommended features in the plurality of weighted recommended features with the second multimedia data features respectively, and obtaining concatenated features of each of the concatenated vectors in the plurality of concatenated vectors based on a second feature extraction network corresponding to each of the first gating networks that obtain the respective weighted recommended features to obtain a plurality of concatenated features; Inputting the plurality of concatenated vectors and the concatenated features of each of the concatenated vectors into a business objective prediction model for obtaining recommended media, so as to obtain a plurality of target recommended media based on the business objective prediction model; Among them, the business objective prediction model includes a plurality of second gating networks and prediction networks corresponding to each of the second gating networks; after inputting the plurality of concatenated vectors and the concatenated features of each of the concatenated vectors into the business objective prediction model for obtaining recommended media, the method further includes: Based on the plurality of concatenated vectors, obtaining feature combination weights corresponding to the concatenated features used by each of the second gating networks to obtain target recommended features, and performing weighted summation on the plurality of concatenated features through any one of the second gating networks based on the feature combination weights corresponding to the concatenated features used by the any one of the second gating networks to obtain one target recommended feature obtained by the any one of the second gating networks, and obtaining the target recommended features obtained by each of the second gating networks to obtain the plurality of target recommended features; Based on the prediction networks corresponding to each of the second gating networks that obtain the respective target recommended features, obtaining business objective prediction values corresponding to the respective target recommended features to obtain a plurality of business objective prediction values, and determining a plurality of target recommended media from the multimedia data based on the plurality of business objective prediction values.

2. The method according to claim 1, wherein The obtaining the plurality of weighted recommended features corresponding to the plurality of recommended features through a plurality of first gating networks includes: Based on the first multimedia data features, obtaining feature combination weights corresponding to the recommended features used by each of the first gating networks to obtain weighted recommended features; Performing weighted summation on the multiple recommended features by any one of the first gating networks based on the feature combination weights corresponding to the respective recommended features adopted by any one of the first gating networks, to obtain one weighted recommended feature obtained by any one of the first gating networks; Obtaining the respective weighted recommended features obtained by the first gating networks to obtain the multiple weighted recommended features; Wherein, the first gating network includes at least one of a gating network based on linear transformation or a gating network based on normalized weighting.

3. The method according to claim 2, wherein After obtaining the multiple concatenated vectors obtained by concatenating each of the weighted recommended features in the multiple weighted recommended features with the second multimedia data feature respectively, the method further includes: Obtaining cross features, and concatenating the cross features with each of the concatenated vectors in the multiple concatenated vectors respectively to obtain updated multiple concatenated vectors; Wherein, the cross features include cross features obtained based on cross data generated by cross - matching the portrait information of the sample recommended object and the sample recommended multimedia data, or cross features obtained based on cross data generated by cross - matching the behavior information of the sample recommended object and the sample recommended multimedia data.

4. The method according to claim 3, characterized in that, The second gating network includes at least one of a gating network based on linear transformation or a gating network based on normalized weighting.

5. The method according to claim 4, wherein The first multimedia data feature is the recommended object data feature and the second multimedia data feature is the multimedia data feature to be recommended; The step of performing feature extraction on the first multimedia data feature by multiple first feature extraction networks to obtain multiple recommended features corresponding to the first multimedia data feature, and obtaining multiple weighted recommended features corresponding to the multiple recommended features by multiple first gating networks includes: Performing feature extraction on the recommended object data feature by multiple first feature extraction networks to obtain multiple recommended object features corresponding to the recommended object data feature, and obtaining multiple weighted recommended object features corresponding to the multiple recommended object features by multiple first gating networks as the multiple weighted recommended features; The step of obtaining the multiple concatenated vectors obtained by concatenating each of the weighted recommended features in the multiple weighted recommended features with the second multimedia data feature respectively includes: Obtaining the multiple concatenated vectors obtained by concatenating each of the weighted recommended object features in the multiple weighted recommended object features with the multimedia data feature to be recommended respectively.

6. The method according to claim 4, characterized in that, The first multimedia data feature is the multimedia data feature to be recommended and the second multimedia data feature is the recommended object data feature; The step of performing feature extraction on the first multimedia data feature by multiple first feature extraction networks to obtain multiple recommended features corresponding to the first multimedia data feature, and obtaining multiple weighted recommended features corresponding to the multiple recommended features by multiple first gating networks includes: Performing feature extraction on the multimedia data feature to be recommended by multiple first feature extraction networks to obtain multiple recommended multimedia features corresponding to the multimedia data feature to be recommended, and obtaining multiple weighted recommended multimedia features corresponding to the multiple recommended multimedia features by multiple first gating networks as the multiple weighted recommended features; The obtaining of multiple concatenated vectors obtained by concatenating each weighted recommendation feature in the multiple weighted recommendation features with the second multimedia data feature respectively includes: Obtaining multiple concatenated vectors obtained by concatenating each weighted recommendation multimedia feature in the multiple weighted recommendation multimedia features with the recommended object data feature respectively.

7. The method according to any one of claims 1-6, wherein the obtaining of the multimedia data feature corresponding to the multimedia data includes: Obtaining multimedia data, where the multimedia data includes recommended object data and multimedia data to be recommended, and the multimedia data to be recommended includes at least one of graphic media data, audio data, and video data; Obtaining the recommended object data feature corresponding to the recommended object data in the multimedia data through vectorization processing and embedding compression processing, and obtaining the multimedia data feature to be recommended corresponding to the multimedia data to be recommended.

8. A recommendation device for multimedia data, characterized in that, including: An obtaining module, configured to obtain the multimedia data feature corresponding to the multimedia data, where the multimedia data feature includes a first multimedia data feature and a second multimedia data feature, the first multimedia data feature is the recommended object data feature and the second multimedia data feature is the multimedia data feature to be recommended, or the first multimedia data feature is the multimedia data feature to be recommended and the second multimedia data feature is the recommended object data feature; A weighted recommendation feature generation module, configured to perform feature extraction on the first multimedia data feature through multiple first feature extraction networks to obtain multiple recommendation features corresponding to the first multimedia data feature, and obtain multiple weighted recommendation features corresponding to the multiple recommendation features through multiple first gating networks, where one of the first feature extraction networks is used to obtain one recommendation feature corresponding to the first multimedia data feature, and one of the first gating networks is used to obtain one weighted recommendation feature corresponding to the multiple recommendation features; A concatenated feature generation module, configured to obtain multiple concatenated vectors by concatenating each weighted recommendation feature in the multiple weighted recommendation features with the second multimedia data feature respectively, and obtain concatenated features based on each of the concatenated vectors through a second feature extraction network corresponding to each of the first gating networks that obtain the respective weighted recommendation features, so as to obtain multiple concatenated features; A target recommended media generation module, configured to input the multiple concatenated vectors and the concatenated features of each of the concatenated vectors into a service target prediction model for obtaining recommended media, so as to obtain multiple target recommended media based on the service target prediction model; wherein, the service target prediction model includes multiple second gating networks and prediction networks corresponding to each of the second gating networks; after inputting the multiple concatenated vectors and the concatenated features of each of the concatenated vectors into the service target prediction model for obtaining recommended media, the target recommended media generation module is configured to: Based on the multiple splicing vectors, obtain the feature combination weights corresponding to the respective splicing features used by each of the second gating networks to obtain the target recommendation features. Through any one of the second gating networks, perform weighted summation on the multiple splicing features based on the feature combination weights corresponding to the respective splicing features used by the any one of the second gating networks, so as to obtain one target recommendation feature obtained by the any one of the second gating networks, and obtain the respective target recommendation features obtained by each of the second gating networks to obtain the multiple target recommendation features; Based on the prediction networks corresponding to the respective second gating networks that obtain the respective target recommendation features, obtain the service target prediction values corresponding to the respective target recommendation features to obtain multiple service target prediction values, and determine multiple target recommendation media from the multimedia data based on the multiple service target prediction values.

9. A computer device, characterized in that, Comprising: a processor, a memory, and a network interface; The processor is connected to the memory and the network interface. Among them, the network interface is used to provide data communication functions, the memory is used to store program codes, and the processor is used to call the program codes to execute the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is adapted to be loaded and executed by the processor to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-target recommendation method and device, computing equipment and medium

    CN112699305A

  • Information recommendation method and device, equipment, storage medium and computer program product

    CN113626719A