METHOD, DEVICE, COMPUTER DEVICE, AND COMPUTER PROGRAM FOR RECOMMENDING MEDIA DATA

By extracting and fusing media and text representation vectors with entity information from a knowledge graph, the method enhances media recommendation relevance and diversity, addressing the limitations of uniform similarity-based recommendations.

JP2026501509APending Publication Date: 2026-01-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025530789
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-18
Filing Date
2023-11-29
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing media recommendation systems often provide unified recommendations based on high similarity to viewed media, leading to a lack of diversity and failing to capture the user's true interests.

Method used

A method and device that extract media and text representation vectors, perform knowledge search in a graph to determine entity subgraphs, and fuse these vectors to obtain a knowledge enrichment vector, recommending target media data based on this vector.

Benefits of technology

Improves the relevance and diversity of media recommendations by incorporating entity information, ensuring the recommended media aligns with user interests without manual searches, thus reducing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501509000001_ABST
    Figure 2026501509000001_ABST
Patent Text Reader

Abstract

The present application relates to a media data recommendation method, apparatus, computer device, storage medium, and computer program product. The method can be applied to the field of artificial intelligence, for example, to a scenario in which a smart terminal determines target media data that a target subject is interested in. The method includes: extracting a media representation vector and a text representation vector from media data and corresponding description text (step 202); performing knowledge search in a knowledge graph based on the media representation vector to obtain an entity subgraph corresponding to the media data and determine an entity representation vector corresponding to the entity subgraph (step 204); performing feature fusion processing on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge enrichment vector (step 206); obtaining target media data based on the knowledge enrichment vector and recommending the target media data to the target subject (step 208).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to a Chinese patent application filed with the China Patent Office on July 18, 2023, bearing application number 2023108802404 and entitled "Media data recommendation method, device, computer device, and storage medium," the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the field of computer technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for recommending media data. [Background technology]

[0003] With the development of Internet technology, media viewing is increasingly popular among a growing number of subjects. In related technologies, a recommendation system can determine other media that may be of interest to a subject based on the media content that the subject has viewed. Most of the determined other media that may be of interest to the subject are media that have a relatively high similarity to the viewed media content, which tends to result in unified recommendations. Summary of the Invention [Problem to be solved by the invention]

[0004] According to various embodiments provided herein, a method, apparatus, computer device, computer-readable storage medium, and computer program product for recommending media data are provided. [Means for solving the problem]

[0005] In a first aspect, the present application provides a method for recommending media data, the method being executed by a server, comprising: The method includes the steps of extracting a media representation vector and a text representation vector from the media data and the description text of the media data; performing a knowledge search in a knowledge graph based on the media representation vector to obtain an entity subgraph of the media data and determine the entity representation vector of the entity subgraph; performing a feature fusion process on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge enrichment vector; and obtaining target media data based on the knowledge enrichment vector and recommending the target media data to a target subject.

[0006] In a second aspect, the present application further provides a media data recommendation device, the device comprising: a vector extraction module used for extracting media representation vectors and text representation vectors from the media data and the description text of the media data; a first knowledge retrieval module used for performing knowledge retrieval in the knowledge graph based on the media representation vector to obtain an entity subgraph of the media data and determine the entity representation vector of the entity subgraph; a first fusion module for performing feature fusion processing on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge enriched vector; a recommendation module used for obtaining target media data based on the knowledge enrichment vector and recommending the target media data to the target subject.

[0007] In a third aspect, the present application further provides a computer device including a memory and a processor, wherein a computer program is stored in the memory, and the processor, when executing the computer program, realizes the media data recommendation method provided in the first aspect.

[0008] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored therein, the computer program realizing the media data recommendation method provided in the first aspect when executed by a processor.

[0009] In a fifth aspect, the present application further provides a computer program product, the computer program product including a computer program that, when executed by a processor, implements the media data recommendation method provided in the first aspect.

[0010] In a sixth aspect, the present application provides a method for processing a recommendation model, the method being executed by a server, comprising: According to the feature extraction model, extracting a first media training vector and a first text training vector from the first sample media data and the corresponding first sample text; according to the knowledge retrieval model, performing a knowledge retrieval process on the first media training vector and the knowledge graph to obtain a training subgraph of the first sample media data and determine an entity training vector of the training subgraph; according to the knowledge enrichment model, performing a feature fusion process on the first media training vector, the first text training vector, and the entity training vector to obtain a knowledge enrichment training vector; and determining a visual loss value and a linguistic loss value based on the knowledge enriched training vector and the training subgraph; adjusting parameters of the feature extraction model, the knowledge retrieval model and the knowledge enrichment model based on the visual loss value, the linguistic loss value and the knowledge retrieval loss value to obtain an enriched vector extraction model; and determining a recommendation model based on the enriched vector extraction model and the classification model, wherein the recommendation model is used to extract a knowledge enrichment vector based on the media data, the description text and the knowledge graph, and determine an interest type based on the knowledge enrichment vector, thereby obtaining target media data based on the interest type and recommending the target media data to the target subject.

[0011] In a seventh aspect, the present application further provides a processing device for a recommendation model, said device comprising: a training vector extraction module used to extract a first media training vector and a first text training vector from the first sample media data and the corresponding first sample text according to the feature extraction model; a second knowledge retrieval module, which is used to perform a knowledge retrieval process on the first media training vector and the knowledge graph according to the knowledge retrieval model, to obtain a training subgraph of the first sample media data, and to determine an entity training vector of the training subgraph; a second fusion module, which is used to perform feature fusion processing on the first media training vector, the first text training vector, and the entity training vector according to the knowledge enrichment model, to obtain a knowledge enrichment training vector; a first loss value determination module used to determine a visual loss value and a linguistic loss value according to the knowledge-enhanced training vector and the sample label; a second loss value determination module, used for determining a knowledge retrieval loss value according to the knowledge enrichment training vector and the training subgraph; a parameter adjustment module used for adjusting parameters of the feature extraction model, the knowledge retrieval model and the knowledge enrichment model according to the visual loss value, the language loss value and the knowledge retrieval loss value to obtain an enriched vector extraction model; and a recommendation model determination module used to determine a recommendation model based on the reinforcement vector extraction model and the classification model, wherein the recommendation model is used to extract a knowledge reinforcement vector based on media data, description text, and knowledge graph, determine an interest type based on the knowledge reinforcement vector, obtain target media data based on the interest type, and recommend the target media data to a target subject.

[0012] In an eighth aspect, the present application further provides a computer device, the computer device including a memory and a processor, a computer program stored in the memory, and the processor, when executing the computer program, realizes the recommendation model processing method provided in the sixth aspect.

[0013] In a ninth aspect, the present application further provides a computer-readable storage medium having a computer program stored therein, the computer program realizing the recommendation model processing method provided in the sixth aspect when executed by a processor.

[0014] In a tenth aspect, the present application further provides a computer program product, the computer program product including a computer program that, when executed by a processor, implements the recommendation model processing method provided in the sixth aspect.

[0015] The details of one or more embodiments of the application are set forth in the drawings and description below. Other features, objects, and advantages of the application will become apparent from the description, drawings, and claims.

[0016] In order to more clearly explain the technical solutions of the embodiments of the present application, the following briefly introduces the drawings that need to be used in the description of the embodiments. It is obvious that the drawings in the following description are only exemplary embodiments of the present application, and those skilled in the art can further obtain other drawings based on these drawings without any creative work. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is an application environment diagram of a media data recommendation method in one embodiment; [Figure 2] 1 is a flow diagram of a media data recommendation method according to one embodiment; [Figure 3] FIG. 2 is a schematic diagram of extracting media representation vectors in one embodiment. [Figure 4] FIG. 2 is a schematic diagram of extracting text representation vectors in one embodiment. [Figure 5] FIG. 2 is a schematic diagram of determining target media data based on media data, description text, and a knowledge graph in one embodiment. [Figure 6]FIG. 2 is a schematic diagram of determining entity representation vectors in one embodiment. [Figure 7] FIG. 1 is a schematic diagram of determining a knowledge reinforcement vector in one embodiment. [Figure 8] FIG. 1 is a structural schematic diagram of a recommendation model in one embodiment. [Figure 9] FIG. 10 is a schematic diagram of a media data recommendation method according to another embodiment; [Figure 10] FIG. 1 is a schematic diagram of a processing method of a recommendation model in one embodiment. [Figure 11] FIG. 10 is a schematic diagram of determining a first media training vector in one embodiment; [Figure 12] FIG. 2 is a schematic diagram of determining a first text training vector in one embodiment. [Figure 13] FIG. 1 is a schematic diagram of determining knowledge-enhanced training vectors in the training process of an enhanced vector extraction model in one embodiment; [Figure 14] FIG. 10 is a schematic diagram of a processing method for a recommendation model in another embodiment. [Figure 15] FIG. 1 is a structural block diagram of a media data recommendation device in one embodiment. [Figure 16] FIG. 2 is a structural block diagram of a processing unit of a recommendation model in one embodiment. [Figure 17] 1 is a diagram illustrating the internal structure of a computer device according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0018] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be described in more detail below in conjunction with drawings and examples. It should be understood that the specific examples described herein are merely for the purpose of interpreting the present application, and are not intended to limit the present application.

[0019] In this specification and drawings, substantially the same or similar steps and elements are represented by the same or similar reference numerals, and repeated descriptions of these steps and elements are omitted. At the same time, in the description of this application, the terms "first", "second", etc. are used only to distinguish between descriptions, and should not be understood as indicating or implying relative importance or order.

[0020] The media data recommendation method provided by the embodiments of the present application can be applied in the application environment shown in Fig. 1 , where a terminal 102 communicates with a server 104 via a network. A data storage system can store data that the server 104 needs to process. The data storage system may be integrated into the server 104, or may be located in a cloud or another network server. The media data recommendation method may be executed by the terminal 102, by the server 104, or even by both the terminal 102 and the server 104.

[0021] For example, taking the media data recommendation method as being performed by the server 104, the server 104 can extract a media representation vector and a text representation vector from the media data and the description text of the media data, the server 104 can perform a knowledge search in a knowledge graph based on the media representation vector to obtain an entity subgraph of the media data and determine an entity representation vector of the entity subgraph, the server 104 can perform a feature fusion process on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge enriched vector, and the server 104 can further obtain target media data based on the knowledge enriched vector and recommend the target media data to a target subject.

[0022] Here, the terminal 102 may be a smartphone, a tablet PC, a laptop, a desktop computer, a smart speaker, a smart watch, an Internet of Things device, or a portable wearable device, and the Internet of Things device may be a smart speaker, a smart TV, a smart air conditioner, a smart in-vehicle device, etc. The portable wearable device may be a smart watch, a smart band, a head-mounted device, etc.

[0023] The server 104 may be an independent physical server or a service node in a blockchain system, and each service node in the blockchain system forms a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol that operates on the Transmission Control Protocol (TCP) protocol.

[0024] In addition, the server 104 may also be a server cluster consisting of multiple physical servers, or may be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0025] The terminal 102 and the server 104 may be connected by a communication connection method such as Bluetooth, USB (Universal Serial Bus), or a network, and the present application is not limited thereto.

[0026] In some embodiments, as shown in FIG. 2, a method for recommending media data is provided, which may be performed by the server or the terminal in FIG. 1, or may be performed by both the server and the terminal in FIG. 1. The method will be described taking the example of being performed by the server in FIG. 1, and includes the following steps:

[0027] Step 202: Extract a media representation vector and a text representation vector from the media data and the description text of the media data.

[0028] Here, the media data may be media data being viewed by the target subject or media data viewed by the target subject, and the media data may specifically be videos, images, or live rooms, and the target subject refers to a user. It should be noted that in recommending media data, the media data being viewed by the user or the media data viewed by the user may be media data that the user is interested in, and by recommending media based on media data that the user may be interested in, the matching between the media data determined to be recommended and the user's preferences can be improved.

[0029] The descriptive text is used to describe the content of the media data, exemplarily, the media data is a video, for example, the content of the video is a cat eating fish, and the descriptive text of the video may be "The newly purchased dried fish has arrived and the cat is eating it with gusto", exemplarily, the media data is a picture, for example, the content of the picture is a pitcher pitching in a baseball game, and the descriptive text of the picture may be "The baseball player is pitching".

[0030] Media representation vectors are obtained by performing feature extraction on media data and are used to reflect the content of the media data, while text representation vectors are obtained by performing feature extraction on descriptive text and are used to reflect the content of the descriptive text.

[0031] In some embodiments, the server can obtain media data that the target subject is viewing and descriptive text of the media data, and the server can also obtain media data that the target subject has viewed and descriptive text of the media data, and the server can extract media representation vectors from the media data and text representation vectors from the descriptive text through a feature extraction model.

[0032] Illustratively, the server inputs the media data and the description text into a feature extraction model, and extracts a media representation vector of the media data and a text representation vector of the description text through the feature extraction model.

[0033] In some embodiments, step 202 includes performing feature extraction on the media data using an image feature extraction model to obtain media representation vectors, and extracting text representation vectors from descriptive text of the media data using a text feature extraction model.

[0034] Here, the image feature extraction model includes a first self-attention layer and a visual feedforward layer, and the text feature extraction model includes a second self-attention layer and a text feedforward layer.

[0035] As shown in Figure 3, media data is input into the image feature extraction model, the first self-attention layer outputs the initial representation vector of the media data, and the visual feedforward layer processes the initial representation vector of the media data to obtain the media representation vector.

[0036] As shown in Figure 4, the description text is input into the text feature extraction model, the initial representation vector of the description text is output by the second self-attention layer, and the initial representation vector of the description text is processed by the text feedforward layer to obtain the text representation vector.

[0037] In the above embodiment, the image feature extraction model is used to extract the media representation vector of the media data, and the text feature extraction model is used to extract the text representation vector of the description text, so that the media representation vector can reflect the content of the media data, and the text representation vector can reflect the content of the description text, thereby improving the quality of the media representation vector and the text representation vector.

[0038] In some embodiments, performing feature extraction on the media data using the image feature extraction model to obtain a media representation vector includes, when the media data is a video, extracting features from a plurality of image frames in the video using the image feature extraction model to obtain a media representation vector; and, when the media data is an image, performing feature extraction on a plurality of image blocks of the image using the image feature extraction model to obtain a media representation vector.

[0039] Here, the plurality of image frames may be some image frames in a video, the number of the plurality of image frames may be a first preset number, and the size of the image frame may be a preset size. The first preset number and the preset size can both be set based on actual needs, and the embodiments of the present application do not limit the first preset number and the preset size.

[0040] Here, the multiple image blocks may be obtained by dividing an image, the number of the multiple image blocks may be a first predetermined number, and the size of the image blocks may be a predetermined size, i.e., the size of the image block and the size of the image frame are the same, and the number of the multiple image blocks and the number of the multiple image frames are the same.

[0041] When the media data is a video, the server samples the video to obtain a first predetermined number of image frames, and performs padding or cropping on the first predetermined number of image frames, so that the sizes of the first predetermined number of image frames are all predetermined sizes; the server inputs the first predetermined number of image frames into an image representation extraction model, and outputs a media representation vector through the image representation extraction model, where the media representation vector includes image sub-representation vectors of each image frame, i.e., the media representation vector includes the first predetermined number of image sub-representation vectors.

[0042] When the media data is an image, the server divides the image into a first predetermined number of image blocks, performs padding or cropping on the first predetermined number of image blocks, and the sizes of the first predetermined number of image blocks can all be preset sizes; the server inputs the first predetermined number of image blocks into an image representation extraction model, and outputs a media representation vector through the image representation extraction model, where the media representation vector includes image sub-representation vectors of each image block, i.e., the media representation vector includes the first predetermined number of image sub-representation vectors.

[0043] In some embodiments, the media data may also be a live room, and when the media data is a live room, the image features are used to extract features of multiple live screen frames in the live room to obtain a media expression vector.

[0044] Here, the plurality of live screen frames may be a portion of screen frames in a screen that has already been played in the live room, the number of the plurality of live screen frames may be a first preset number, and the size of the live screen frame may be a preset size.

[0045] In the above embodiment, the media data may be a video or an image, so that the media data recommendation method can be applied to scenes where target media data is recommended by viewing a video or viewing an image, thereby improving the applicability of the media data recommendation method.

[0046] Step 204: Perform knowledge search in the knowledge graph based on the media expression vector to obtain an entity subgraph of the media data, and determine the entity expression vector of the entity subgraph.

[0047] Here, the knowledge graph includes multiple entities and relationships between multiple entities, and the knowledge graph belongs to a node-connection diagram, where nodes are used to represent entities and connections between nodes are used to represent relationships between nodes, and multiple sets of entity relationships can be obtained through the knowledge graph. For example, an entity relationship: {E1, r1, E2} can be obtained from the knowledge graph, where E1 is the first entity, E2 is the second entity, and r1 is the entity relationship between the first entity and the second entity, for example, the first entity is a ball, and the second entity is table tennis, and the relationship is a belonging relationship. In the node-connection diagram, the first entity is represented by node E1, and the second entity is represented by node E2, and the connection r1 between node E1 and node E2 represents the entity relationship r1 between the first entity and the second entity.

[0048] In practical applications, the knowledge graph may be constructed in the background of an application that views media data; for example, a target subject views media data in an instant communication application, and the knowledge graph is constructed in the background of the instant communication application.

[0049] Here, an entity subgraph is a graph that is a part of a knowledge graph, and the entity subgraph relates to multiple entities, and the multiple entities are a part of all entities included in the knowledge graph, and the entity subgraph can be used to reflect the relationships between multiple entities in the knowledge graph.

[0050] For example, all entities included in the knowledge graph are E1, E2, E3, ..., and En, respectively, and multiple entities included in the entity subgraph are E1, E2, ..., and Eu, respectively, and the entity subgraph is used to reflect the relationships between E1, E2, ..., and Eu in the knowledge graph.

[0051] The entity representation vector is used to reflect the entities and the relationships between the entities in the entity subgraph.

[0052] In some embodiments, the server determines representation vectors of multiple entities in the knowledge graph, selects each entity in the multiple entities in the knowledge graph that is related to media data based on each entity representation vector and the media representation vector, determines an entity subgraph of the media data based on the multiple entities related to the media data and the relationships of the multiple entities in the knowledge graph, and performs feature extraction on the entity subgraph to obtain the entity representation vector.

[0053] In some embodiments, the server's selection of multiple entities related to the media data from the multiple entities in the knowledge graph based on the multiple entity representation vectors and the media representation vector may involve the server determining respective similarities between the media representation vector and the multiple entity representation vectors, sorting the multiple entity representation vectors according to order of similarity to obtain an entity representation vector sequence, selecting the first of a second predetermined number of multiple target entity representation vectors in the ordered sequence from the entity representation vector sequence, and determining that the entities represented by the multiple target entity representation vectors are the multiple entities related to the media data.

[0054] In some embodiments, after the server determines the multiple entities associated with the media data, it can obtain neighboring entities of each entity in the knowledge graph, and determine an entity subgraph of the media data based on the multiple entities, the neighboring entities of each entity, and the relationships between the multiple entities and the multiple neighboring entities in the knowledge graph.

[0055] Step 206: Perform feature fusion processing on the media expression vector, the text expression vector, and the entity expression vector to obtain a knowledge-enhanced vector.

[0056] Here, the knowledge enrichment vector is obtained by fusing the knowledge information of the entity in the media expression vector and the text expression vector.

[0057] In some embodiments, the server obtains the preset weights of the media representation vector, the text representation vector, and the entity representation vector, respectively, and performs weighted addition based on the media representation vector, the text representation vector, the entity representation vector, the preset weights of the media representation vector, the preset weights of the text representation vector, and the preset weights of the entity representation vector to obtain a knowledge enrichment vector.

[0058] It is necessary to explain that the sum of the preset weights of the media representation vectors, the preset weights of the text representation vectors, and the preset weights of the entity representation vectors is 1, which is equivalent to averaging the media representation vectors, the text representation vectors, and the entity representation vectors when the preset weights of the media representation vectors, the text representation vectors, and the entity representation vectors are all the same.

[0059] In some embodiments, the server can combine the media representation vector and the text representation vector to obtain a first combined vector, and the server performs feature extraction on the first combined vector using a self-attention network to obtain a first fusion vector. The server obtains weights for the first fusion vector and the text representation vector, and performs weighting on the first fusion vector and the text representation vector based on the weights of the first fusion vector and the text representation vector to obtain a knowledge-enhanced vector.

[0060] In some embodiments, the server may combine the media representation vector and the entity representation vector to obtain a second combined vector, combine the text representation vector and the entity representation vector to obtain a third combined vector, perform feature extraction on the second combined vector using a self-attention network to obtain a second fusion vector, perform feature extraction on the third combined vector using a self-attention network to obtain a third fusion vector, obtain weights for the second fusion vector and the third fusion vector, and perform weighting on the second fusion vector and the third fusion vector based on the weights of the second fusion vector and the third fusion vector to obtain a knowledge-enhanced vector.

[0061] In some embodiments, the server combines the media representation vector, the text representation vector, and the entity representation vector to obtain a combined vector, and performs feature extraction on the combined vector by a self-attention network to obtain a knowledge-enhanced vector.

[0062] In some embodiments, the server performs weighted addition on the media representation vector, the text representation vector, and the entity representation vector based on the preset weights of the media representation vector, the preset weights of the text representation vector, and the preset weights of the entity representation vector to obtain a fusion vector, and performs feature extraction on the fusion vector using a self-attention network to obtain a knowledge-enhanced vector.

[0063] Step 208: Obtain target media data according to the knowledge enrichment vector, and recommend the target media data to the target subject.

[0064] Here, the target media data is media data that is recommended to a target subject.

[0065] In some embodiments, the knowledge-enhanced vector is obtained by fusing entity knowledge information in the media representation vector and the text representation vector. The knowledge-enhanced vector can reflect the content of the media data and the descriptive text, and the entity information related to the media data. Therefore, by obtaining the target media data based on the knowledge-enhanced vector, the target media data whose content is related to the media data and the descriptive text and whose entity is related to the media data can be obtained.

[0066] In some embodiments, step 208 includes performing a classification process on the knowledge-enhanced vector to obtain an interest type of the target subject; obtaining target media data based on the interest type; and recommending the target media data to the target subject.

[0067] Here, the interest type is a type that the target subject may be interested in, and the number of interest types may be one or more, and the number of target media data may be one or more.

[0068] In some embodiments, the server inputs the knowledge-enhanced vector into a classification model, and obtains a predicted probability that each knowledge-enhanced vector belongs to a predetermined type through the classification model. Then, the server selects a target probability from the plurality of predicted probabilities, and the selected target probability is greater than the unselected predicted probabilities. The predicted type corresponding to the selected target probability is the interest type of the target subject. It is necessary to explain that when there is one selected target probability, there is one determined interest type, and when there are multiple selected target probabilities, there are multiple determined interest types.

[0069] When there is one interest type, the server can select target media data from multiple candidate media data belonging to the interest type. The server obtains the popularity values ​​of the multiple candidate media data and may select the candidate media data with the highest popularity value from the multiple candidate media data as the target media data, or may select multiple candidate media data with relatively high popularity values ​​from the multiple candidate media data as the target media data, where the popularity value of the target media data is greater than the popularity value of the candidate media data that are not selected. In practical applications, the popularity value of the candidate media data may be determined based on the number of views, comments, and likes of the candidate media data.

[0070] When there are multiple interest types, the server can select target media data belonging to each one of the multiple candidate media data belonging to the multiple interest types, thereby obtaining multiple target media data.

[0071] When there is one target media data, the server can transmit the target media data to the terminal used by the target subject to view the media data, and when the target subject triggers an operation to switch to the next media data in the process of viewing the media data, the terminal can play the target media data.

[0072] When there are multiple target media data, the server can sort the multiple target media data based on the order of popularity values ​​to obtain a target media data list and send the target media data list to the terminal used by the target subject to view the media data, and the terminal can display the target media data list in the recommendation area of ​​the media data viewing page, and in response to a trigger operation on any one of the target media data in the target media data list, the terminal can play the target media data that received the trigger operation, and when the target subject triggers an operation to switch to the next media data in the process of viewing the media data, the terminal can play the target media data that is listed first in the target media data list.

[0073] In the above embodiment, by determining the interest type of the target subject for the knowledge enrichment vector, target media data whose content is related to the media data and the description text and whose entities are related to the media data can be obtained, the relevance between the target media data and the media data can be improved, and further, the target media data may be media data that the target subject is interested in, thereby improving the media recommendation effect. In this way, the target subject can obtain and view the media data that interests them without needing to perform manual searches, thereby reducing resource consumption caused by manual searches and improving resource utilization.

[0074] For example, as shown in FIG. 5, the server extracts a media representation vector of the media data, extracts a text representation vector of the description text of the media data, searches in the knowledge graph based on the media representation vector to obtain an entity subgraph, and determines the entity representation vector of the entity subgraph; the server merges the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge enrichment vector; obtains target media data based on the knowledge enrichment vector, and recommends the target media data to the target subject.

[0075] In the above media data recommendation method, a media representation vector and a text representation vector are extracted from the media data and the description text of the media data, and an entity subgraph is searched for and obtained in a knowledge graph based on the media representation vector, and an entity representation vector of the entity subgraph is determined. A feature fusion process is performed on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge enrichment vector. A target media data to be recommended to the target subject is obtained based on the knowledge enrichment vector. An entity subgraph related to the content of the media data is searched for and obtained in the knowledge graph based on the media representation vector. Further, an entity representation vector related to the content of the media data can be obtained based on the entity subgraph. The media representation vector, the text representation vector, and the entity representation vector are then combined. The text expression vector and the entity expression vector are fused to obtain a knowledge-enhanced vector, so that the knowledge-enhanced vector can reflect the content of the media data and the descriptive text, and the entity information related to the content of the media data. Therefore, based on the knowledge-enhanced vector, target media data whose content is similar to that of the media data and whose entities are related to the entities of the media data can be obtained, thereby improving the relevance between the target media data and the media data. Furthermore, the target media data may be media data that the target subject is interested in, thereby improving the media recommendation effect. In this way, the target subject can obtain and view the media data that interests them without needing to perform a manual search, thereby reducing resource consumption caused by manual searches and improving resource utilization.

[0076] In some embodiments, the step of performing knowledge search in the knowledge graph based on the media representation vector, obtaining an entity subgraph of the media data, and determining an entity representation vector of the entity subgraph includes the steps of searching for a target entity related to the media data in the knowledge graph based on the media representation vector, determining an entity subgraph of the media data based on the target entity and the knowledge graph, and extracting features of multiple entities in the entity subgraph to obtain an entity representation vector.

[0077] Here, the target entity is an entity that is part of the knowledge graph and is related to the content of the media data, and may be understood as an entity of the media data. For example, if the media data is an image and the content of the image is a baseball player pitching in a stadium, the target entity may include, but is not limited to, a baseball, a player, a pitch, and a stadium.

[0078] Here, the entity representation vector includes the entity sub-representation vectors of multiple entities in the entity sub-graph.

[0079] In some embodiments, the server obtains a plurality of initial entity vectors in the knowledge graph, determines respective relevance degrees between the media representation vector and the plurality of initial entity vectors, selects candidate relevance degrees for each relevance degree between the media representation vector and the plurality of initial entity vectors, and determines the entity represented by the initial entity vector used to calculate the candidate relevance degrees as the target entity related to the media data.

[0080] For each relevance degree between the media representation vector and the multiple initial entity vectors, selecting a candidate relevance degree may involve arranging the respective relevance degrees between the media representation vector and the multiple initial entity vectors in order of magnitude to obtain an initial relevance degree sequence, and selecting the first third predetermined number of candidate relevance degrees arranged in the initial relevance degree sequence.

[0081] In some embodiments, determining an entity subgraph of the media data based on the target entities and the knowledge graph may include determining relationships between the target entities based on the knowledge graph, and determining the entity subgraph based on the target entities and the relationships between the target entities.

[0082] The server performs feature extraction on the entity subgraph to obtain entity sub-representation vectors of the multiple entities in the entity subgraph, and determines an entity feature representation vector based on the entity sub-representation vectors of the multiple entities in the entity subgraph.

[0083] In the above embodiment, the media representation vector is used to search for target entities related to media data in the knowledge graph, and then the entity subgraph of the media entities is determined based on the target entities. Since the entity subgraph includes the target entities and the relationships between the target entities, by determining the entity representation vector based on the entity subgraph, the entity representation vector can more accurately reflect the entities in the media data, and the accuracy of the entity representation vector is improved.

[0084] In some embodiments, the media representation vector includes at least two image sub-representation vectors, and the step of searching for a target entity related to the media data in the knowledge graph based on the media representation vector includes the steps of obtaining a plurality of initial entity vectors in the knowledge graph, searching for candidate entities in the knowledge graph based on the plurality of initial entity vectors and the at least two image sub-representation vectors, and selecting a target entity related to the media data from the candidate entities.

[0085] Here, the initial entity vector is a representation vector of multiple entities in the knowledge graph, and the initial entity vector may be predetermined by the encoder.

[0086] Here, when the media data is an image, the media representation vector is obtained by performing feature extraction on at least two image blocks obtained by dividing the image, and at least two image sub-representation vectors are further used to represent the at least two image blocks; when the media data is a video, the media representation vector is obtained by performing feature extraction on at least two image frames in the video, and at least two image sub-representation vectors are further used to represent the at least two image frames.

[0087] In some embodiments, for each initial entity vector, a degree of association between the initial entity vector and each of the image sub-representation vectors is determined, and candidate entities are determined based on the degree of association between the initial entity vector and each of the image sub-representation vectors.

[0088] Determining a candidate entity based on the respective degrees of association between the initial entity vector and each image sub-representation vector includes determining whether or not there is at least one degree of association between the initial entity vector and each image sub-representation vector that belongs to a predetermined interval, and if there is, determining that the entity represented by the initial entity vector is a candidate entity, and if there is no such degree of association, determining that the entity represented by the initial entity vector is not a candidate entity.

[0089] Here, the relevance of a section that belongs to a predetermined section is greater than the relevance of a section that does not belong to a predetermined section, and the predetermined section can be set based on actual needs, and the embodiments of the present application do not limit the specific range of the predetermined section.

[0090] In some embodiments, for each image sub-representation vector, a degree of association between the image sub-representation vector and each of the initial entity vectors is determined, and a candidate entity is determined based on the degree of association between the image sub-representation vector and each of the initial entity vectors.

[0091] Determining candidate entities based on the respective degrees of relevance between the image sub-representation vector and each initial entity vector may involve determining a set of relevance degrees for the image sub-representation vector based on the respective degrees of relevance between the image sub-representation vector and each initial entity vector, selecting a plurality of relatively high relevance degrees from the set of relevance degrees for the image sub-representation vector, setting the initial entity vector used to calculate the relatively high relevance degrees as the initial entity vector associated with the image sub-representation vector, and further setting the entity represented by the initial entity vector associated with the image sub-representation vector as the candidate entity associated with the image sub-representation vector.

[0092] Selecting a relatively high relevance in the relevance set of the image sub-representation vector may be selecting a fourth preset number of relevance in the relevance set according to the order of relevance, where the fourth preset number can be set based on actual needs, and the embodiment of the present application does not limit the specific value of the fourth preset number.

[0093] In some embodiments, the candidate entities are obtained by searching based on at least two image sub-representation vectors, and selecting a target entity related to the media data from the candidate entities may involve obtaining at least one target entity from the candidate entities obtained by searching based on each image sub-representation vector, and obtaining a target entity related to the media data; and obtaining at least one target entity from the candidate entities obtained by searching based on each image sub-representation vector may involve determining a candidate association degree between the candidate entity and the image sub-representation vector, and obtaining at least one target entity from the candidate entity based on the candidate association degree, and the candidate association degree between the target entity and the image sub-representation vector is greater than the candidate association degree between other candidate entities and the image sub-representation vector.

[0094] In some embodiments, the candidate entities are obtained by searching based on at least two image sub-representation vectors, and selecting a target entity related to the media data from the candidate entities may involve selecting at least one target entity from all candidate entities obtained by searching based on at least two image sub-representation vectors. For example, for each candidate entity obtained by searching based on each image sub-representation vector, the relevance between the image sub-representation vector and the candidate entity is taken as the candidate relevance of the candidate entity. The server sorts all candidate entities obtained by searching based on at least two image sub-representation vectors according to the order of the candidate relevance to obtain a candidate entity sequence, and selects the first fifth predetermined number of candidate entities in the candidate entity sequence, and determines the fifth predetermined number of candidate entities as the target entities related to the media data. Here, the fifth predetermined number can be set based on actual needs, and the embodiments of the present application do not limit the specific value of the fifth predetermined number.

[0095] In the above embodiment, based on the initial entity vector in the knowledge graph and at least two image sub-representation vectors, candidate entities associated with the image sub-representation vectors are searched for in the knowledge graph, and a target entity associated with the media data is selected from the candidate entities associated with the image sub-representation vectors, so that the determined target entity is associated with multiple image sub-representation vectors of the media data, and further, the target entity can reflect the content of the media data, thereby improving the accuracy of searching for target entities associated with the media data.

[0096] In some embodiments, the step of searching for candidate entities in the knowledge graph based on the plurality of initial entity vectors and the at least two image sub-representation vectors includes the steps of determining a relevance set of the at least two image sub-representation vectors based on the plurality of initial entity vectors and the at least two image sub-representation vectors, wherein the relevance set includes respective relevances between the image sub-representation vectors and the initial entity vectors; and selecting candidate entities related to the at least two image sub-representation vectors in the knowledge graph based on the relevance set.

[0097] Here, the candidate entities associated with at least two image sub-representation vectors include candidate entities associated with each of the image sub-representation vectors.

[0098] In some embodiments, for each image sub-representation vector, the server determines a relevance degree between the image sub-representation vector and a plurality of initial entity vectors, and determines a relevance set for the image sub-representation vector based on the relevance degrees between the image sub-representation vector and a plurality of initial entity vectors.

[0099] Illustratively, the determination of the degree of association between the image sub-representation vector and the initial entity vector is shown in Equation (1).

[0100]

number

number

number

number

number

[0101] The server sorts the multiple relevance degrees in the relevance set in order of magnitude, selects the first six predetermined number of target relevance degrees from the sorted relevance set, obtains initial entity representation vectors for the sixth predetermined number of target relevance degrees, and sets the entities represented by the obtained multiple initial entity representation vectors as candidate entities related to the image sub-representation vector.

[0102] Illustratively, if the media representation vector includes s image sub-representation vectors and the sixth preset number is t, then s*t candidate entities can be obtained.

[0103] In the above embodiment, a relatively high relevance is obtained in the relevance set of each image sub-representation vector, and a candidate entity related to the image sub-representation vector is selected based on the relatively high relevance, so that the candidate entity can reflect the content represented in the image sub-representation vector, thereby improving the accuracy of the selected candidate entity.

[0104] In some embodiments, the step of determining an entity subgraph of the media data based on the target entities and the knowledge graph includes the steps of determining adjacent nodes of each target entity in the knowledge graph, determining extended entities based on multiple target entities and adjacent nodes of each target entity and determining relationships between the extended entities in the knowledge graph, and determining an entity subgraph of the media data based on the extended entities and the relationships between the extended entities.

[0105] Here, the adjacent nodes of the target entity may be the first-order adjacent nodes of the target entity, or may include the first-order adjacent nodes and the second-order adjacent nodes of the target entity, and the extended entity includes the target entity and the entities represented by the adjacent nodes.

[0106] For example, the entities included in the knowledge graph may be represented as {E1, E2, ..., En}, and based on the media representation vector V, search for multiple target entities {E1, E2, ..., Eq} related to media in {E1, E2, ..., En}, determine the first-order neighbor nodes of the multiple target entities in the knowledge graph, realize extensions to the target entities, obtain extended entities {E1, E2, ..., Eu}, and construct an entity subgraph G based on the extended entities {E1, E2, ..., Eu} in the knowledge graph.

[0107] In the above embodiment, the adjacent nodes of the target entity in the knowledge graph are obtained to realize the expansion to the target entity, which makes the entities contained in the entity subgraph richer and further improves the quality of the entity representation vector of the entity subgraph.

[0108] In some embodiments, the step of searching for multiple target entities related to the media data in the knowledge graph based on the media representation vector includes searching for multiple target entities related to the media data in the knowledge graph based on the media representation vector by a search sub-model of the knowledge retrieval model; the step of determining an entity sub-graph of the media data based on the target entities and the knowledge graph includes determining an entity sub-graph of the media data based on the multiple target entities and the knowledge graph by a sub-graph construction network of the knowledge retrieval model; and the step of extracting features of the multiple entities in the entity sub-graph and obtaining an entity representation vector includes extracting features of the multiple entities in the entity sub-graph and obtaining an entity representation vector by a graph neural network of the knowledge retrieval model.

[0109] Here, the knowledge retrieval model includes a retrieval sub-model, a sub-graph construction network, and a graph neural network. As shown in FIG. 6, the server inputs the media representation vector and the knowledge graph into the retrieval sub-model, searches through the retrieval sub-model to obtain a target entity related to the media data, inputs the target entity and the knowledge graph into the sub-graph construction network, outputs an entity sub-graph of the media data through the sub-graph construction network, inputs the entity sub-graph into the graph neural network, and extracts the features of multiple entities in the entity sub-graph through the graph neural network to obtain an entity representation vector.

[0110] In some embodiments, the search sub-model encodes multiple entities in the knowledge graph, obtains initial entity vectors for the multiple entities, and determines a relevance set of the multiple image sub-representation vectors based on the initial entity vector and multiple image sub-representation vectors in the media representation vector, selects candidate entities related to the multiple image sub-representation vectors in the knowledge graph based on the relevance set, and selects multiple target entities related to the media data from the multiple candidate entities.

[0111] A target entity and a knowledge graph are input into a subgraph construction network, and the subgraph construction network can obtain adjacent nodes of the target entity in the knowledge graph, determine extended entities based on the target entity and the adjacent nodes, determine relationships between the extended entities in the knowledge graph, and determine an entity subgraph based on the extended entities and the relationships between the extended entities.

[0112] For each entity in the entity subgraph, a node trajectory graph of the entity is obtained based on a random walk in the entity subgraph of the entity, and the node trajectory graph is input into a graph neural network, and the graph neural network outputs a representation vector of the entity. The representation vector of each entity in the entity subgraph is determined in a similar manner, and the entity representation vector is obtained based on the representation vector of each entity. In practical applications, the graph neural network may be a GNN (Graph Neural Networks).

[0113] In the above embodiment, the search sub-model, sub-graph construction network, and graph neural network in the knowledge search model are used to search for target entities related to media data, and an entity sub-graph of the media data is constructed, and entity representation vectors of multiple entities in the entity sub-graph are extracted and obtained, so that the entity representation vectors can more accurately reflect the entities in the media data, and the accuracy of the entity representation vectors is improved.

[0114] In some embodiments, the step of performing a feature fusion process on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge-enhanced vector includes the steps of combining the media representation vector, the text representation vector, and the entity representation vector, and, during the combining process, adding a separator between the media representation vector and the text representation vector, and adding a separator between the text representation vector and the entity representation vector, to obtain a combined vector; and performing a feature fusion process on the combined vector using a knowledge-enhanced model to obtain a knowledge-enhanced vector.

[0115] Here, the knowledge-augmented model includes a regularization layer, an encoder, and a feedforward network layer.

[0116] The separator in the combined vector can be used to distinguish between the media representation vector, the text representation vector, and the entity representation vector in the combined vector.

[0117] The knowledge enriched vector includes a media enriched vector, a text enriched vector, and an entity enriched vector, and in a situation where the combined vector includes a separator, the knowledge enriched vector also includes a separator used to distinguish between the media enriched vector, the text enriched vector, and the entity enriched vector.

[0118] In some embodiments, the media representation vectors are {v1, v2, ..., vn}, the text representation vectors are {t1, t2, ..., tn}, and the entity representation vectors are {e1, e2, ..., en}, and the server combines the media representation vectors, the text representation vectors, and the entity representation vectors, and in the process of combining, adds a separator [sep] between the media representation vectors and the text representation vectors, and between the text representation vectors and the entity representation vectors, to obtain a combined vector {v1, v2, ..., vn [sep] t1, t2, ..., tn [sep] e1, e2, ..., en}.

[0119] In some embodiments, as shown in Figure 7, the server inputs the combined vector into a knowledge-enhanced model and randomly drops the combined vector through a regularization layer to reduce the amount of data processed and obtain a regularized vector. For example, the regularization layer processes the combined vector {v1,v2,...,vn [sep] t1,t2,...,tn [sep] e1,e2,...,en} to obtain the regularized vector {v1,0,...,vn [sep] 0,t2,...,tn [sep] e1,e2,...,0}.

[0120] The regularized vector is processed by the encoder to realize multimodal fusion of the media representation vector, the text representation vector, and the entity representation vector in the regularized vector, thereby obtaining a fused vector. In practical applications, the encoder can be realized by a multi-head attention network. For example, the encoder performs multimodal fusion on the regularized vector {v1,0,...,vn [sep] 0,t2,...,tn [sep] e1,e2,...,0} to obtain a fused vector {a1,a2,...,an [sep] b1,b2,...,bn [sep] c1,c2,...,cn}.

[0121] The feedforward network layer performs activation processing on the fusion vector to obtain a knowledge-enhanced vector, which has higher expression for media data, descriptive text, and entities than the fusion vector. For example, the feedforward network layer activates and fuses the fusion vector {a1,a2,...,an [sep] b1,b2,...,bn [sep] c1,c2,...,cn} to obtain a knowledge-enhanced vector {x1,x2,...,xn [sep] y1,y2,...,yn [sep] z1,z2,...,zn}.

[0122] It should be noted that the knowledge enrichment vectors include media enrichment vectors {x1, x2, ..., xn}, text enrichment vectors {y1, y2, ..., yn}, and entity enrichment vectors {z1, z2, ..., zn}.

[0123] In the above embodiment, the media representation vector, the text representation vector, and the entity representation vector are combined to obtain a combined vector, and feature fusion is performed on the combined vector according to the instruction enrichment model to obtain a knowledge enrichment vector. The multi-modal representation vectors are fused, so that the knowledge enrichment vector can reflect the media data, the description text, and their contents, as well as the entity information related to the content of the media data. Furthermore, based on the knowledge enrichment vector, target media data whose contents are similar to the media data and whose entities are related to the entities of the media data can be obtained, thereby improving the media recommendation effect. In this way, the target subject can obtain and view the media data they are interested in without having to perform a manual search, thereby reducing the resource consumption caused by manual searches and improving resource utilization.

[0124] In some embodiments, the media data recommendation method includes: extracting a first media training vector and a first text training vector from first sample media data and a first sample text of the first sample media data based on a feature extraction model; performing a knowledge retrieval process on the first media training vector and a knowledge graph based on a knowledge retrieval model to obtain a training subgraph of the first sample media data and determine an entity training vector of the training subgraph; performing a feature fusion process on the first media training vector, the first text training vector, and the entity training vector based on a knowledge enrichment model to obtain a knowledge enrichment training vector; and and the sample label of the first sample media data, determining a visual loss value and a linguistic loss value; determining a knowledge retrieval loss value based on the knowledge enrichment training vector and the training sub-graph; adjusting parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model based on the visual loss value, the linguistic loss value, and the knowledge retrieval loss value to obtain an enrichment vector extraction model; and determining a recommendation model based on the enrichment vector extraction model and the classification model, wherein the recommendation model is used to extract a knowledge enrichment vector based on the media data, the description text, and the knowledge graph, and determine an interest type based on the knowledge enrichment vector, thereby obtaining target media data based on the interest type, and recommending the target media data to the target subject.

[0125] In some embodiments, the media data recommendation method can be applied to a recommendation model, and as shown in FIG. 8 , the recommendation model includes an enhanced vector extraction model and a classification model, and the enhanced vector extraction model includes an image feature extraction model, a text feature extraction model, a knowledge retrieval model, and a knowledge enrichment model, and the enhanced vector extraction model is obtained by performing parameter adjustment on the pre-training feature extraction model, the knowledge retrieval model, and the knowledge enrichment model, where the pre-training feature extraction model includes the pre-training image feature extraction model and the text feature extraction model.

[0126] In practical application, the media data, description text, and knowledge graph are processed by the reinforcement vector model in the recommendation model to obtain a knowledge reinforcement vector, and the knowledge reinforcement vector is classified by the classification model in the recommendation model to obtain the interest type of the target subject, so as to obtain the target media data based on the interest type, and recommend the target media data to the target subject.

[0127] In some embodiments, the knowledge-enhanced training vectors include media-enhanced training vectors and text-enhanced training vectors, and the sample labels include hidden sub-image labels and hidden word labels. The step of determining a visual loss value and a linguistic loss value based on the knowledge-enhanced training vectors and the sample labels of the first sample media data includes the steps of: obtaining hidden sub-image training vectors in the media-enhanced training vectors, and determining a visual loss value based on the hidden sub-image training vectors and the hidden sub-image labels; and performing a classification process on the text-enhanced training vectors to obtain hidden word prediction probabilities, and determining a linguistic loss value based on the hidden word prediction probabilities and the hidden word labels.

[0128] In some embodiments, the knowledge-enhanced training vectors further include entity-enhanced training vectors, and the step of determining a knowledge search loss value based on the knowledge-enhanced training vectors and the training sub-graph includes the steps of selecting an entity-enhanced training vector pair in the entity-enhanced training vectors and determining a first score for the entity-enhanced training vector pair; obtaining an entity-negative sample pair in the training sub-graph and determining a second score for the entity-negative sample pair, where the entity-negative sample pair includes two training entities that have no entity relationship in the training sub-graph; and determining a knowledge search loss value based on the first score and the second score.

[0129] In some embodiments, the first sample media data and the first sample text belong to a sample set, and the sample set further includes second sample media data and the second sample text. The method further includes: performing feature extraction on the second sample media data and the second sample text to obtain second media training vectors and second text training vectors; and determining an image-text contrast loss value based on the second media training vectors, the second text training vector, the first media training vector, and the first text training vector. The method further includes adjusting parameters of the pre-training feature extraction model, the pre-training knowledge retrieval model, and the pre-training knowledge enrichment model based on the visual loss value, the language loss value, and the knowledge retrieval loss value to obtain a post-training recommendation model.

[0130] In some embodiments, determining an image-text contrast loss value based on the second media training vector, the second text training vector, the first media training vector, and the first text training vector includes determining a first similarity based on the first media training vector and the second text training vector; determining a second similarity based on the first text training vector and the second media training vector; and determining an image-text contrast loss value based on the first similarity, the second similarity, the first similarity label, and the second similarity label.

[0131] In some embodiments, as shown in FIG. 9, a method for recommending media data includes: Step 901: when the media data is a video, extracting features of a plurality of image frames in the video using an image feature extraction model to obtain a media representation vector; when the media data is an image, performing feature extraction on a plurality of image blocks of the image using the image feature extraction model to obtain a media representation vector, wherein the media representation vector includes at least two image sub-representation vectors; Step 902 of extracting text representation vectors from the description text of the media data by a text feature extraction model; Step 903: obtain initial entity vectors of multiple entities in the knowledge graph through a search sub-model of the knowledge search model; determine a relevance set of at least two image sub-representation vectors based on the initial entity vectors and the at least two image sub-representation vectors, where the relevance set includes respective relevances between the image sub-representation vectors and the initial entity vector; select candidate entities related to the at least two image sub-representation vectors in the knowledge graph based on the relevance set; and select multiple target entities related to the media data from the candidate entities; Step 904: determine the adjacent nodes of each target entity in the knowledge graph through the subgraph construction network of the knowledge retrieval model; determine extended entities according to the target entities and the adjacent nodes of each target entity, and determine the relationships between the extended entities in the knowledge graph; and determine an entity subgraph of the media data according to the extended entities and the relationships between the extended entities; Step 905: extracting features of multiple entities in the entity subgraph by a graph neural network of the knowledge retrieval model to obtain entity representation vectors; Step 906: combine the media representation vector, the text representation vector, and the entity representation vector, and in the process of combining, add a separator between the media representation vector and the text representation vector, and add a separator between the text representation vector and the entity representation vector to obtain a combined vector, and perform feature fusion processing on the combined vector through a knowledge enrichment model to obtain a knowledge enrichment vector; Step 907 includes performing a classification process on the knowledge-enhanced vector to obtain the interest type of the target subject, obtaining target media data based on the interest type, and recommending the target media data to the target subject.

[0132] In the above media data recommendation method, a media expression vector and a text expression vector are extracted from the media data and the description text of the media data, and an entity subgraph is searched for and obtained in the knowledge graph based on the media expression vector, and an entity expression vector of the entity subgraph is determined, and a feature fusion process is performed on the media expression vector, the text expression vector, and the entity expression vector to obtain a knowledge enrichment vector, and a target media data to be recommended to the target subject is obtained based on the knowledge enrichment vector, and an entity subgraph related to the content of the media data is searched for and obtained in the knowledge graph based on the media expression vector, and an entity expression vector related to the content of the media data can be further obtained based on the entity subgraph, and the media expression vector, the text expression vector, and the entity expression vector are then combined. The text representation vector and the entity representation vector are fused to obtain a knowledge-enhanced vector, so that the knowledge-enhanced vector can reflect the content of the media data and the description text, and the entity information related to the content of the media data. Therefore, based on the knowledge-enhanced vector, target media data whose content is similar to the media data and whose entities are related to the entities of the media data can be obtained, thereby improving the relevance between the target media data and the media data. Furthermore, the target media data may be media data that the target subject is interested in, thereby improving the media recommendation effect. In this way, the target subject can obtain and view the media data that interests them without needing to perform a manual search, thereby reducing resource consumption caused by manual searches and improving resource utilization.

[0133] In some embodiments, as shown in FIG. 10 , a method for processing a recommendation model is provided, which can be performed by a server or a terminal. In the following description, the method is performed by a server, and includes the following steps:

[0134] Step 1002: Extract a first media training vector and a first text training vector from the first sample media data and a first sample text of the first sample media data according to the feature extraction model.

[0135] Here, the feature extraction model includes an image feature extraction model before training and a text feature extraction model.

[0136] In some embodiments, feature extraction is performed on first sample media data using a pre-trained image feature extraction model to obtain a first media training vector, and feature extraction is performed on first sample text using a pre-trained text feature extraction model to obtain a first text training vector.

[0137] In some embodiments, the pre-trained image feature extraction model can be realized by a first bidirectional encoding model (Transformer), and the first bidirectional encoding model includes a plurality of image encoders. The first sample media data can include a plurality of sample images and concealment sub-images, and can be obtained by dividing an initial sample image to obtain a plurality of sample images, and concealing some sample images in the plurality of sample images to obtain the first media sample data including the plurality of sample images and the concealment sub-images, or by sampling image frames in a sample video to obtain a plurality of sample images, and concealing some sample images in the plurality of sample images to obtain the first media sample data.

[0138] For example, as shown in FIG. 11 , an initial sample image is divided into N sample images, and a visual mask model is used to perform mask processing on the N sample images to hide some of the sample images in the N sample images, thereby obtaining first sample media data. The first bidirectional encoding model includes L image encoders, and the L image encoders process the first sample media data to obtain a first media training vector.

[0139] In some embodiments, the pre-trained image feature extraction model can be realized by a second bidirectional encoding model, which includes a plurality of text encoders. The first sample text includes a plurality of words and a concealment word, and a word segmentation process is performed on the initial sample text to obtain the plurality of words, and some words in the plurality of words are concealed to obtain the first sample text including the plurality of words and the concealment word.

[0140] For example, as shown in FIG. 12, a word segmentation process is performed on an initial sample to obtain individual words. For example, the initial sample is “A baseball player throwing a ball in a game”, and the individual words are “A”, “baseball”, “player”, “throwing”, “a”, “ball”, “in”, “a”, and “game”, respectively. A masking process is performed on the individual words using a text masking model to conceal some of the words in the individual words. A start mark is added before the sample text after the masking process to obtain a first sample text. For example, the first sample text includes “[cls]”, “A”, “[MASK]”, “[MASK]”, “throwing”, “a”, “[MASK]”, “in”, “a”, and “game”. The second bidirectional coding model includes L text encoders, and the L text encoders process the first sample text to obtain a first text training vector.

[0141] Step 1004: Based on the knowledge retrieval model, perform knowledge retrieval processing on the first media training vector and the knowledge graph to obtain a training subgraph of the first sample media data, and determine the entity training vector of the training subgraph.

[0142] Here, the knowledge retrieval model in this step is a pre-training knowledge retrieval model, which includes a pre-training retrieval sub-model, a pre-training sub-graph construction network, and a pre-training graph neural network.

[0143] In some embodiments, the first media training vector and the knowledge graph are processed by a pre-training search sub-model to search for training entities related to the first sample media data, the training entities and the knowledge graph are processed by a pre-training sub-graph construction network to construct a training sub-graph of the first sample media data, and the pre-training graph neural network performs feature extraction on the training sub-graph to obtain entity training vectors of the training sub-graph.

[0144] Step 1006: Based on the knowledge-enhanced model, perform feature fusion processing on the first media training vector, the first text training vector, and the entity training vector to obtain a knowledge-enhanced training vector.

[0145] Here, the knowledge-enhanced model in this step is a pre-training knowledge-enhanced model, which includes a pre-training regularization layer, a pre-training encoder, and a pre-training feedforward layer.

[0146] In some embodiments, the server combines the first media training vector, the first text training vector, and the entity training vector, and during the combining process, adds separators between the first media training vector and the first text training vector and between the first text training vector and the entity training vector to obtain a training combined vector, performs a random drop process on the training combined vector using a pre-training regularization layer to obtain a training regularized vector, and performs a fusion process on the training regularized vector using a pre-training encoder to achieve multimodal fusion of the first media training vector, the first text training vector, and the entity training vector in the training regularized vector to obtain a training fusion vector, and performs an activation process on the training fusion vector using a pre-training feedforward layer to obtain a knowledge-enhanced training vector.

[0147] It should be noted that knowledge-enhanced training vectors include media-enhanced training vectors, text-enhanced training vectors, and entity-enhanced training vectors.

[0148] In some embodiments, the pre-training encoder includes a self-attention layer, a first normalization layer, a feedforward layer, and a second normalization layer, and the process by which the self-attention layer processes the training regularization vector is shown in equation (2).

[0149]

number

number

number

number

number

number

number

number

[0150] The process of processing the representation vectors output by the self-attention layer and the training regularization vectors by the first normalization layer is shown in Equation (3).

[0151]

number

number

number

number

[0152] The process of processing the representation vector output by the first normalization layer by the feedforward layer is shown in equation (4).

[0153]

number

number

number

number

number

[0154] The process of processing the representation vectors output by the feedforward layer and the first normalization layer by the second normalization layer is shown in equation (5).

[0155]

number

number

number

number

[0156] Step 1008: Determine visual loss values ​​and linguistic loss values ​​based on the knowledge-enhanced training vectors and sample labels.

[0157] Here, the sample labels include hidden sub-image labels and hidden word labels, the visual loss value is used to reflect the difference between the media-enhanced training vectors and the hidden sub-image labels, and the linguistic loss value is used to reflect the difference between the predicted probabilities corresponding to the text-enhanced training vectors and the hidden word labels.

[0158] In some embodiments, step 1008 includes obtaining a hidden sub-image enhanced vector in the media-enhanced training vector, and determining a visual loss value based on the hidden sub-image enhanced vector and the hidden sub-image label; and performing a classification process on the text-enhanced training vector to obtain hidden word prediction probabilities, and determining a language loss value based on the hidden word prediction probabilities and the hidden word labels.

[0159] Here, the media-enhanced training vector includes a hidden sub-image enhanced vector of the hidden sub-image, and the hidden sub-image enhanced vector is a feature vector obtained by reconstructing the hidden sub-image using the first sample media data, the first sample text, and the knowledge graph.

[0160] The text-enhanced training vectors include hidden word-enhanced vectors of hidden words, and the hidden word-enhanced vectors are feature vectors obtained by predicting the hidden words using the first sample media data, the first sample text, and the knowledge graph.

[0161] In some embodiments, the server obtains a hidden subimage enhancement vector in the media-enhanced training vector, obtains a hidden subimage label of the hidden subimage enhancement vector, and calculates a visual loss value based on the hidden subimage enhancement vector and the hidden subimage label of the hidden subimage enhancement vector. It should be noted that the hidden subimage enhancement vector and the hidden subimage label of the hidden subimage enhancement vector correspond to the same hidden subimage.

[0162] Exemplarily, the visual loss value can be determined by a cross-entropy loss function, as shown in Equation (6).

[0163]

number

number

number

number

number

[0164] In some embodiments, when the number of hidden sub-image enhancement vectors is multiple, a loss value of each hidden sub-image enhancement vector is determined based on the hidden sub-image enhancement vector and the corresponding hidden sub-image label, and an average value is calculated based on the different loss values ​​of the multiple hidden sub-image enhancement vectors to obtain a visual loss value.

[0165] In some embodiments, the server can use a classifier to perform classification processing on the text-enhanced training vectors to obtain hidden word prediction probabilities, and the server obtains hidden word reinforcement vectors from the text-enhanced training vectors, and calculates a language loss value according to the hidden word prediction probabilities and hidden word labels. It should be noted that the hidden word prediction probabilities and hidden word labels correspond to the same hidden words.

[0166] Illustratively, the language loss value can be determined by the cross-entropy loss function, shown in Equation (7).

[0167]

number

number

number

number

number

[0168] In the above embodiment, the visual loss value is determined by predicting the expression vector of the hidden sub-image, and the linguistic loss value is determined by predicting the expression vector of the hidden word, so that the accuracy of the visual loss value and the linguistic loss value is improved, and then it is easy to adjust the parameters of the model according to the visual loss value and the linguistic loss value.

[0169] Step 1010: Determine a knowledge retrieval loss value based on the knowledge enriched training vector and the training subgraph.

[0170] Here, the knowledge-enriched training vector includes an entity-enriched training vector, which includes a plurality of entity-enriched sub-vectors.

[0171] In some embodiments, for each entity enrichment subvector, the server can determine a target entity enrichment subvector for the entity enrichment subvector in other entity enrichment subvectors, and determine entity positive sample pairs based on the entity enrichment subvector and the target entity enrichment subvector; the server obtains entity negative sample pairs in the training subgraph; and the server determines a knowledge search loss value based on the entity positive sample pairs and the entity negative sample pairs.

[0172] In some embodiments, step 1010 includes obtaining entity-positive sample pairs in the training subgraph and determining a first score for the entity-positive sample pairs based on the entity-enriched training vectors; obtaining entity-negative sample pairs in the training subgraph and determining a second score for the entity-negative sample pairs based on the entity-enriched training vectors, where the entity-negative sample pairs include two training entities that do not have an entity relationship in the training subgraph; and determining a knowledge search loss value based on the first score and the second score.

[0173] Here, the two entities included in the entity-positive sample pair have an entity relationship in the training subgraph, and the two entities included in the entity-negative sample pair do not have an entity relationship in the training subgraph.

[0174] In some embodiments, the server obtains each entity-positive sample pair in which an entity relationship exists in the training subgraph, and for each entity-positive sample pair, obtains, in the entity-enriched training vector, separate entity-enriched subvectors for the two entities in the entity-positive sample pair, and determines a first score for the entity-positive sample pair based on the separate entity-enriched subvectors for the two entities in the entity-positive sample pair.

[0175] The server obtains each entity-negative sample pair in which no entity relationship exists in the training subgraph, and for each entity-negative sample pair, obtains separate entity-enriched subvectors for the two entities in the entity-negative sample pair in the entity-enriched training vector, and determines a second score for the entity-positive sample pair based on the separate entity-enriched subvectors for the two entities in the entity-negative sample pair.

[0176] For example, the entities included in the training subgraph are E1, E2, E3, E4, and E5, respectively, where there is no entity relationship between E2 and E3, and there is no entity relationship between E4 and E5; further, the entity negative sample pairs include {E2, E3} and {E4, E5}; a second score for {E2, E3} is determined based on the separate entity enrichment subvectors of E2 and E3; and a second score for {E4, E5} is determined based on the separate entity enrichment subvectors of E4 and E5.

[0177] In some embodiments, a knowledge retrieval loss value can be determined based on the first score and the second score, and can refer to equation (8).

[0178]

number

number

number

number

number

number

number

number

number

[0179] In some embodiments, the step of determining a knowledge search loss value based on a knowledge-enhanced training vector and a training subgraph includes the steps of determining an entity sample pair based on the training subgraph, and determining a score for the entity sample pair based on the entity-enhanced training vector, and determining the entity sample pair as an entity positive sample pair when the score belongs to a positive sample interval, and determining the entity sample pair as an entity negative sample pair when the score does not belong to the positive sample interval.

[0180] Determining entity sample pairs based on the training subgraph may involve combining all entities included in the training subgraph two by two to obtain each entity sample pair.

[0181] In the above embodiment, entity-positive sample pairs and entity-negative sample pairs are determined based on the training subgraph, and a first score for the entity-positive sample pairs and a second score for the entity-negative sample pairs are determined based on the entity-enriched training vector. Furthermore, the knowledge retrieval loss value can be used to reflect the difference between training entities with entity relationships and training entities without entity relationships, thereby improving the accuracy of the knowledge retrieval loss value.

[0182] Step 1012: According to the visual loss value, the language loss value, and the knowledge retrieval loss value, adjust the parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model to obtain an enriched vector extraction model.

[0183] Here, the enhanced vector extraction model includes a post-training feature extraction model, a post-training knowledge retrieval model, and a post-training knowledge enhancement model, and the post-training feature extraction model includes a post-training image feature extraction model and a post-training text feature extraction model.

[0184] In some embodiments, the server combines the visual loss value, the language loss value, and the knowledge retrieval loss value to obtain a total loss value, and adjusts parameters of the pre-training feature extraction model, the pre-training knowledge retrieval model, and the pre-training knowledge enrichment model according to the total loss value to obtain an enriched vector extraction model until the pre-training feature extraction model, the pre-training knowledge retrieval model, and the pre-training knowledge enrichment model converge.

[0185] In practical applications, the AdamW optimizer can adjust the parameters of the pre-trained feature extraction model, the pre-trained knowledge retrieval model, and the pre-trained knowledge reinforcement model according to a preset learning rate and a preset weight decay. The AdamW optimizer is used to update the parameters of the neural network based on the gradient and minimize the total loss value. The preset learning rate can be set based on actual needs, for example, the preset learning rate can be 5e-5, and the preset weight decay can be set based on actual needs, for example, the preset weight decay can be 0.02.

[0186] Step 1014: Determine a recommendation model based on the enrichment vector extraction model and the classification model. The recommendation model extracts a knowledge enrichment vector based on the media data, the description text, and the knowledge graph, and determines an interest type based on the knowledge enrichment vector, thereby obtaining target media data based on the interest type and recommending the target media data to the target subject.

[0187] Here, the recommendation model includes an enhanced vector extraction model and a classification model.

[0188] In some embodiments, a recommendation model is obtained by connecting a previously trained classification model after the reinforcement vector extraction model. In practical application, the media data viewed by the target subject, the description text of the media data, and the knowledge graph are input into the recommendation model, the knowledge reinforcement vector is determined by the reinforcement vector extraction model of the recommendation model, and the interest type of the knowledge reinforcement vector is output by the classification model of the recommendation model, so as to obtain the target media data according to the interest type, and recommend the target media data to the target subject.

[0189] It needs to be explained that the enhanced vector extraction model may be a kind of pre-training model. After obtaining the enhanced vector extraction model through pre-training, the enhanced vector extraction model can be used for the downstream task of media data recommendation, by fixing the parameters of the enhanced vector extraction model and adjusting the parameters of the initial classification model to obtain a post-training classification model, and determining a recommendation model based on the enhanced vector extraction model and the post-training classification model.

[0190] In the processing method of the recommendation model, a first media training vector and a first text training vector are extracted by a feature extraction model, an entity training vector is searched in a knowledge graph based on the first media training vector, and a feature fusion process is performed on the first media training vector, the first text training vector, and the entity training vector to obtain a knowledge-enhanced training vector, that is, an entity related to the first media sample data is searched in the knowledge graph, and the entity training vector, the first media training vector, and the first text amount of the related entity are fused to realize multi-modal data exchange, and the expression of the first media sample data, the first text sample, and the related entity is enhanced, and the quality of the knowledge-enhanced training vector is improved. is improved, and the parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model are adjusted in conjunction with the visual loss value, the language loss value, and the knowledge retrieval loss value, so that in the process of parameter adjustment, content information of the first media sample data and the first text sample can be learned, and entity information related to the first media sample data can be learned, thereby improving the quality of the enriched vector extraction model obtained by training, and further improving the quality of the recommendation model including the enriched vector extraction model, so that target media data to be recommended to the target subject can be determined based on the recommendation model, and the media recommendation effect can be improved. In this way, the target subject can obtain and view media data that interests them without needing to perform manual searches, thereby reducing resource consumption caused by manual searches and improving resource utilization.

[0191] In some embodiments, the first sample media data and the first sample text belong to a sample set, and the sample set further includes second sample media data and second sample text. The recommendation model processing method further includes: performing feature extraction on the second sample media data and the second sample text to obtain second media training vectors and second text training vectors; and determining an image-text contrast loss value based on the second media training vectors, the second text training vectors, the first media training vectors, and the first text training vectors. The recommendation model processing method further includes adjusting parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model based on the visual loss value, the language loss value, and the knowledge retrieval loss value to obtain an enriched vector extraction model.

[0192] Here, the image-text contrast loss value can reflect the difference between the similarity between the sample media data and the sample text of the sample media data and the similarity between the sample media data and the sample text of other sample media data.

[0193] In some embodiments, the server may perform feature extraction on the second sample media data and the second sample text using a feature extraction model to obtain a second media training vector and a second text training vector, and the server may determine a first candidate similarity based on the first media training vector and the first text training vector, determine a second candidate similarity based on the first media training vector and the second text training vector, determine a third candidate similarity based on the first text training vector and the second media training vector, and determine an image-text contrast loss value based on the first candidate similarity, the second candidate similarity, and the third candidate similarity.

[0194] In some embodiments, the server combines the visual loss value, the language loss value, the knowledge retrieval loss value, and the image-text contrast loss value to obtain a total loss value, and adjusts parameters of the pre-training feature extraction model, the pre-training knowledge retrieval model, and the pre-training knowledge enrichment model according to the total loss value to obtain an enriched vector extraction model until the pre-training feature extraction model, the pre-training knowledge retrieval model, and the pre-training knowledge enrichment model converge.

[0195] In the above embodiment, the parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model are adjusted in accordance with the visual loss value, the language loss value, the knowledge retrieval loss value, and the image-text contrast loss value. In this way, during the parameter adjustment process, content information between the first media sample data and the first text sample can be learned, entity information related to the first media sample data can be learned, and similar content between the first media sample data and the first text sample can be learned. This improves the quality of the enriched vector extraction model obtained by training, and further improves the quality of the recommendation model including the enriched vector extraction model. Target media data to be recommended to a target subject can be determined based on the recommendation model, and the media recommendation effect can be improved. In this way, the target subject can obtain and view media data that interests them without having to perform manual searches, thereby reducing resource consumption caused by manual searches and improving resource utilization.

[0196] In some embodiments, determining an image-text contrast loss value based on the second media training vector, the second text training vector, the first media training vector, and the first text training vector includes determining a first similarity based on the first media training vector and the second text training vector; determining a second similarity based on the first text training vector and the second media training vector; and determining an image-text contrast loss value based on the first similarity, the second similarity, the first similarity label, and the second similarity label.

[0197] Here, the first similarity label may be the similarity between the first media training vector and the first text training vector, and the second similarity label may be the similarity between the first text training vector and the first media training vector.

[0198] Illustratively, the first similarity label is:

number

number

number

number

number

number

number

[0199] In some embodiments, the number of second media sample data included in the sample set is multiple, and correspondingly, the number of second sample texts included in the sample set is multiple, and further, the number of second media training vectors is multiple, and the number of second text training vectors is multiple, and for a first media training vector, the server determines a first similarity between the first media training vector and each of the multiple second text training vectors, and for a first text training vector, the server determines a second similarity between the first text training vector and each of the multiple second media training vectors.

[0200] The server determines a first target similarity based on each first similarity between the first media training vector and the plurality of second text training vectors, and determines a second target similarity based on each second similarity between the first text training vector and the plurality of second media training vectors.

[0201] For example, it is shown in equation (9).

[0202]

number

number

number

[0203] For example, it is shown in equation (10).

[0204]

number

number

number

[0205] The server calculates a loss value between the first target similarity and the first similarity label using a cross-entropy loss function, calculates a loss value between the second target similarity and the second similarity label using a cross-entropy loss function, and determines an image-text contrast loss value based on the loss value between the first target similarity and the first similarity label and the loss value between the second target similarity and the second similarity label. The image-text contrast loss value is added to the model training process, so that the similarity between the media data and the expression vector of the description text for which an extracted correspondence exists is relatively large, and the similarity between the media data and the expression vector of the description text for which no correspondence exists is relatively small.

[0206] For example, it is shown in equation (11).

[0207]

number

number

number

number

number

number

[0208] In the above embodiment, an image contrast loss value is determined based on the first similarity, the second similarity, the first similarity label, and the second similarity label, and the image contrast loss value is added to the process of adjusting the parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enhancement model, and then training is performed to obtain an enhanced vector extraction model, and the quality of the enhanced vector extraction model is improved.

[0209] In some embodiments, as shown in FIG. 13, the training process of the enhanced vector extraction model includes: Dividing the initial sample image into N sample images, performing mask processing on the N sample images using a visual mask model to obtain first sample media data, and processing the first sample media data using a pre-training image feature extraction model to obtain a first media training vector; The initial sample is subjected to word segmentation, and N t We obtain N words and use the text mask model to t masking the words and adding a start mark before the masked sample text to obtain a first sample text; processing the first sample text through a pre-training text feature extraction model to obtain a first text training vector; inputting the knowledge graph and the first media training vector into a pre-training knowledge retrieval model, and determining a training sub-graph and an entity training vector of the training sub-graph through the pre-training knowledge retrieval model, wherein the pre-training knowledge retrieval model includes a pre-training retrieval sub-model, a pre-training sub-graph construction network, and a pre-training graph neural network; Combining the first media training vector, the first text training vector, and the entity training vector, and during the combining process, adding separators between the first media training vector and the first text training vector, and between the first text training vector and the entity training vector, to obtain a combined training vector; performing feature fusion on the training combined vectors by a pre-training knowledge-enhanced model to obtain knowledge-enhanced training vectors, where the pre-training knowledge-enhanced model includes a pre-training regularization layer, a pre-training encoder, and a pre-training feedforward layer, and the knowledge-enhanced training vectors include a media-enhanced training vector, a text-enhanced training vector, and an entity-enhanced training vector; The method includes steps of determining a visual loss value and a linguistic loss value based on the knowledge-enhanced training vector and the sample label, determining a knowledge retrieval loss value based on the knowledge-enhanced training vector and the training subgraph, and adjusting parameters of the feature extraction model, the knowledge retrieval model, and the knowledge-enhanced model based on the visual loss value, the linguistic loss value, and the knowledge retrieval loss value to obtain an enhanced vector extraction model.

[0210] In some embodiments, as shown in FIG. 14, the processing method of the recommendation model includes: Step 1401: extracting a first media training vector and a first text training vector from first sample media data and a first sample text of the first sample media data based on the feature extraction model; Step 1402: performing a knowledge retrieval process on the first media training vector and the knowledge graph based on the knowledge retrieval model to obtain a training subgraph of the first sample media data, and determine an entity training vector of the training subgraph; Step 1403: performing a feature fusion process on the first media training vector, the first text training vector, and the entity training vector based on the knowledge-enhanced model to obtain a knowledge-enhanced training vector, where the knowledge-enhanced training vector includes a media-enhanced training vector, a text-enhanced training vector, and an entity-enhanced training vector; Step 1404: obtain a hidden sub-image enhanced vector from the media-enhanced training vector, and determine a visual loss value based on the hidden sub-image enhanced vector and the hidden sub-image label; perform a classification process on the text-enhanced training vector to obtain a hidden word prediction probability; and determine a language loss value based on the hidden word prediction probability and the hidden word label; Step 1405: obtaining entity-positive sample pairs in the training sub-graph, and determining a first score for the entity-positive sample pairs based on the entity-enriched training vectors; obtaining entity-negative sample pairs in the training sub-graph, and determining a second score for the entity-negative sample pairs based on the entity-enriched training vectors, where the entity-negative sample pairs include two training entities that do not have an entity relationship in the training sub-graph; and determining a knowledge retrieval loss value based on the first score and the second score; In step 1406, feature extraction is performed on the second sample media data and the second sample text to obtain second media training vectors and second text training vectors, a first similarity is determined based on the first media training vectors and the second text training vectors, a second similarity is determined based on the first text training vectors and the second media training vectors, and an image-text contrast loss value is determined based on the first similarity, the second similarity, the first similarity label, and the second similarity label; Step 1407: Adjusting parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model according to the visual loss value, the language loss value, the knowledge retrieval loss value, and the image-text contrast loss value to obtain an enrichment vector extraction model; Step 1408 includes determining a recommendation model based on the enrichment vector extraction model and the classification model, in which the recommendation model is used to extract a knowledge enrichment vector based on the media data, the description text, and the knowledge graph, determine an interest type based on the knowledge enrichment vector, obtain target media data based on the interest type, and recommend the target media data to the target subject.

[0211] In some embodiments, the quality of the enhanced vector extraction model is detected and the enhanced vector extraction model is compared with other models in the related art, and the comparison results are shown in Table 1.

[0212] [Table 1-1] [Table 1-2]

[0213] Here, KAT (Knowledge Augmented Transformer, a knowledge transformation model), REVIVE (a type of visual question answering model), ALBEF (Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, a momentum distillation-based visual language representation model), BLIP (a type of visual language multimodal model), REVEAL (Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory, a multi-source multimodal visual language pre-training model), and VL-BERT are universal visual language models, while UNITER (UNiversal Image-TExt Representation Learning, a multimodal pre-training model), OSCAR (Object-Semantics Aligned Pre-training for Vision-Language Tasks, a type of multimodal pre-training model), and SimVLM are simple visual language pre-training models under weak supervision.

[0214] Wiki data is Wikidata, #image 12M is 12 million images, #image 129M is 129 million images, and the remaining content related to #image is similar, CC12M is 12 million image-text pairs, and WIT (Wikipedia-based Image Text Dataset-GitHub) is a Wiki-based image-text set.

[0215] In conjunction with downstream knowledge-based tasks, on the OK-VQA (Outside Knowledge-Visual Question Answering) dataset, the enhanced vector extraction model achieved higher accuracy rates than KAT, REVIVE, ALBEFF, BLIP, and REVEAL, all of which had relatively high relative accuracy gains over the currently relatively advanced REVIVE and BLIP. Compared to REVEAL, the enhanced vector extraction model can achieve better performance with fewer knowledge graph resources. On the AOK-VQA dataset, the enhanced vector extraction model also achieved higher accuracy rates than ALBEF, BLIP, and REVEAL.

[0216] In conjunction with downstream tasks of universal visual language, when training on the VQA-v2 (Visual Question Answering-v2) dataset and with a basic amount of data, the enhanced vector extraction model also achieves improved accuracy compared to VL-BERT, UNITER, OSCAR, and ALBEF. When training on the VQA-v2 dataset and with a large amount of data, the enhanced vector extraction model also has good competitiveness.

[0217] On the Stanford Natural Language Inference-Visual Entailment (SNLI-VE) dataset, along with downstream tasks for universal visual language, the enhanced vector extraction model achieved similar accuracy improvements compared to VL-BERT, UNITER, OSCAR, and ALBEF.

[0218] In some examples, the ability of the enhanced vector extraction model to search for entities is examined, and the enhanced vector extraction model is compared with a traditional multimodal entity search model, and the comparison results are shown in Table 2.

[0219] [Table 2]

[0220] Here, ViT+BERT (Vision Transformer+ Bidirectional Encoder Representation from Transformers) is a vision transformer + language representation model, ResNet is a residual network, and CLIP is training a transferable vision model using text as a supervised signal. As can be seen, in the scores of each model under six indicators, the enhanced vector extraction model outperforms the traditional multimodal entity search model in five indicators.

[0221] In the processing method of the recommendation model, a first media training vector and a first text training vector are extracted by a feature extraction model, an entity training vector is searched in a knowledge graph based on the first media training vector, and a feature fusion process is performed on the first media training vector, the first text training vector, and the entity training vector to obtain a knowledge-enhanced training vector, that is, an entity related to the first media sample data is searched in the knowledge graph, and the entity training vector, the first media training vector, and the first text amount of the related entity are fused to realize multi-modal data exchange, and the expression of the first media sample data, the first text sample, and the related entity is enhanced, and the quality of the knowledge-enhanced training vector is improved. is improved, and the parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model are adjusted in conjunction with the visual loss value, the language loss value, and the knowledge retrieval loss value, so that in the process of parameter adjustment, content information of the first media sample data and the first text sample can be learned, and entity information related to the first media sample data can be learned, thereby improving the quality of the enriched vector extraction model obtained by training, and further improving the quality of the recommendation model including the enriched vector extraction model, so that target media data to be recommended to the target subject can be determined based on the recommendation model, and the media recommendation effect can be improved. In this way, the target subject can obtain and view media data that interests them without needing to perform manual searches, thereby reducing resource consumption caused by manual searches and improving resource utilization.

[0222] As should be understood, although each step in the flowcharts relating to each of the above-described embodiments is shown in order according to the direction of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated otherwise in this specification, the execution of these steps is not limited to a strict order, and these steps may be executed in other orders. Furthermore, at least some of the steps in the flowcharts relating to each of the above-described embodiments may include multiple steps or multiple stages, and these steps or stages may not necessarily be completed at the same time but may be executed at different times. The order in which these steps or stages are executed is also not necessarily sequential, and they may be executed sequentially or alternately with other steps or at least some of the steps or stages in other steps.

[0223] Based on the same inventive concept, the embodiments of the present application further provide a media data recommendation device used to realize the media data recommendation method related to the above. The implementation means for solving the problem provided in the device are similar to the implementation means described in the above method, so that the specific limitations of one or more embodiments of the media data recommendation device provided below can refer to the limitations of the media data recommendation method in the above specification.

[0224] In some embodiments, as shown in FIG. 15 , a media data recommendation device is provided, which includes a vector extraction module 1501, a first knowledge retrieval module 1502, a first fusion module 1503, and a recommendation module 1504, wherein: The vector extraction module 1501 is used to extract media representation vectors and text representation vectors from media data and description text of the media data; the first knowledge retrieval module 1502 is used for performing knowledge retrieval in the knowledge graph based on the media expression vector, obtaining an entity subgraph of the media data, and determining an entity expression vector of the entity subgraph; The first fusion module 1503 is used for performing feature fusion processing on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge-enhanced vector; The recommendation module 1504 is used to obtain target media data based on the knowledge-enhanced vector and recommend the target media data to the target subject.

[0225] In some embodiments, the vector extraction module 1501 includes a media representation vector extraction unit and a text representation vector extraction unit; The media representation vector extraction unit is used to perform feature extraction on the media data according to the image feature extraction model to obtain a media representation vector; The text representation vector extraction unit is used to extract text representation vectors from the description text of the media data according to the text feature extraction model.

[0226] In some embodiments, the media representation vector extraction unit is further used for: when the media data is a video, extracting features from a plurality of image frames in the video using the image feature extraction model to obtain a media representation vector; and when the media data is an image, performing feature extraction on a plurality of image blocks of the image using the image feature extraction model to obtain a media representation vector.

[0227] In some embodiments, the first knowledge retrieval module 1502: a target entity determination unit used for searching for a plurality of target entities related to the media data in the knowledge graph according to the media representation vector; an entity subgraph determination unit used for determining an entity subgraph of the media data according to a plurality of target entities and the knowledge graph; an entity representation vector determining unit used for extracting features of the entities in the entity subgraph to obtain an entity representation vector.

[0228] In some embodiments, the media representation vector includes at least two image sub-representation vectors, and the target entity determination unit is further used for obtaining a plurality of initial entity vectors in the knowledge graph, searching for candidate entities in the knowledge graph based on the plurality of initial entity vectors and the at least two image sub-representation vectors, and selecting a plurality of target entities related to the media data from the candidate entities.

[0229] In some embodiments, the target entity determination unit further includes a candidate entity search subunit used for: determining a relevance set of the at least two image sub-representation vectors based on the plurality of initial entity vectors and each of the at least two image sub-representation vectors, where the relevance set includes respective relevances between the image sub-representation vectors and the initial entity vector; and selecting a candidate entity related to the at least two image sub-representation vectors in the knowledge graph based on the relevance set.

[0230] In some embodiments, the entity subgraph determination unit is further used for: determining adjacent nodes of each target entity in the knowledge graph; determining extended entities based on the multiple target entities and the adjacent nodes of each target entity, and determining relationships between the extended entities in the knowledge graph; and determining an entity subgraph of the media data based on the extended entities and the relationships between the extended entities. In some embodiments, the target entity determination unit is further used for searching for multiple target entities related to the media data in the knowledge graph based on the media representation vectors through a search submodel of the knowledge retrieval model; the entity subgraph determination unit is further used for determining an entity subgraph corresponding to the media data based on the multiple target entities and the knowledge graph through a subgraph construction network of the knowledge retrieval model; and the entity representation vector determination unit is further used for extracting features of the multiple entities in the entity subgraph to obtain entity representation vectors through a graph neural network of the knowledge retrieval model.

[0231] In some embodiments, the first fusion module 1503 is further used to combine the media representation vector, the text representation vector, and the entity representation vector, and in the process of combining, add a separator between the media representation vector and the text representation vector, and add a separator between the text representation vector and the entity representation vector, to obtain a combined vector; and perform feature fusion processing on the combined vector using a knowledge enrichment model to obtain a knowledge enrichment vector.

[0232] In some embodiments, the recommendation module 1504 is further used to perform a classification process on the knowledge-enhanced vector to obtain an interest type of the target subject, obtain target media data based on the interest type, and recommend the target media data to the target subject.

[0233] The whole or part of each module in the media data recommendation device can be realized by software, hardware, or a combination thereof. Each module can be integrated into or independent from a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, which facilitates the processor to call and execute the operations corresponding to each module.

[0234] Based on the same inventive concept, the embodiments of the present application further provide a media data recommendation device used to realize the media data recommendation method related to the above. The implementation means for solving the problem provided in the device are similar to the implementation means described in the above method, so that the specific limitations of one or more embodiments of the media data recommendation device provided below can refer to the limitations of the media data recommendation method in the above specification.

[0235] In some embodiments, as shown in FIG. 16 , a recommendation model processing device is provided, which includes a training vector extraction module 1601, a second knowledge retrieval module 1602, a second fusion module 1603, a first loss value determination module 1604, a second loss value determination module 1605, a parameter adjustment module 1606, and a recommendation model determination module 1607, wherein: The training vector extraction module 1601 is used to extract a first media training vector and a first text training vector from the first sample media data and the first sample text of the first sample media data according to the feature extraction model; The second knowledge retrieval module 1602 is used to perform a knowledge retrieval process on the first media training vector and the knowledge graph according to the knowledge retrieval model, obtain a training subgraph of the first sample media data, and determine an entity training vector of the training subgraph; The second fusion module 1603 is used for performing feature fusion processing on the first media training vector, the first text training vector, and the entity training vector according to the knowledge enrichment model to obtain a knowledge enrichment training vector; The first loss value determination module 1604 is used to determine a visual loss value and a linguistic loss value based on the knowledge-enhanced training vector and the sample label; The second loss value determination module 1605 is used to determine a knowledge retrieval loss value according to the knowledge enrichment training vector and the training subgraph; The parameter adjustment module 1606 is used to adjust parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model according to the visual loss value, the language loss value, and the knowledge retrieval loss value, to obtain an enriched vector extraction model; The recommendation model determination module 1607 is used to determine a recommendation model based on the reinforcement vector extraction model and the classification model. The recommendation model extracts a knowledge reinforcement vector based on the media data, the description text, and the knowledge graph, and determines an interest type based on the knowledge reinforcement vector, thereby obtaining target media data based on the interest type and recommending the target media data to the target subject.

[0236] In some embodiments, the knowledge-enhanced training vectors include media-enhanced training vectors and text-enhanced training vectors, and the sample labels include hidden sub-image labels and hidden word labels. The first loss value determination module 1604 is used to obtain hidden sub-image enhanced vectors from the media-enhanced training vectors and determine a visual loss value based on the hidden sub-image enhanced vectors and the hidden sub-image labels; and to perform a classification process on the text-enhanced training vectors to obtain hidden word prediction probabilities and determine a language loss value based on the hidden word prediction probabilities and the hidden word labels.

[0237] In some embodiments, the second loss value determination module 1605 is further used to obtain entity-positive sample pairs in the training subgraph and determine a first score for the entity-positive sample pairs based on the entity-enriched training vectors; obtain entity-negative sample pairs in the training subgraph and determine a second score for the entity-negative sample pairs based on the entity-enriched training vectors, where the entity-negative sample pairs include two training entities that do not have an entity relationship in the training subgraph; and determine a knowledge search loss value based on the first score and the second score.

[0238] In some embodiments, the first sample media data and the first sample text belong to a sample set, and the sample set further includes second sample media data and second sample text, and the processing device for the recommendation model further includes a third loss value determination module used for: performing feature extraction on the second sample media data and the second sample text to obtain second media training vectors and second text training vectors; and determining an image-text contrast loss value based on the second media training vectors, the second text training vectors, the first media training vectors, and the first text training vectors.

[0239] Correspondingly, the parameter adjustment module 1606 is used to adjust the parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enhancement model based on the visual loss value, the language loss value, the knowledge retrieval loss value, and the image-text contrast loss value, to obtain an enhanced vector extraction model.

[0240] In some embodiments, the third loss value determination module includes an image-text contrast loss value determination unit used to determine a first similarity based on the first media training vector and the second text training vector, determine a second similarity based on the first text training vector and the second media training vector, and determine an image-text contrast loss value based on the first similarity, the second similarity, the first similarity label, and the second similarity label.

[0241] The whole or part of each module in the processing device of the recommendation model can be realized by software, hardware, or a combination thereof. Each module may be integrated into or independent from the processor in the computer device in the form of hardware, or may be stored in the memory of the computer device in the form of software, which facilitates the processor to call and execute the operations corresponding to each module.

[0242] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be shown in FIG. 17. The computer device includes a processor, a memory, an input / output interface (abbreviated as Input / Output, I / O), and a communication interface. Here, the processor, memory, and I / O interface are connected by a system bus, and the communication interface is connected to the system bus by the I / O interface. Here, the processor of the computer device is used to provide calculation and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. An operating system, a computer program, and a database are stored in the non-volatile storage medium. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used for recommendation models, target media data, and sample sets. The I / O interface of the computer device is used for information exchange between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals via a network connection. When the computer program is executed by a processor, it realizes a method for recommending media data or a method for processing a recommendation model.

[0243] As will be understood by those skilled in the art, the structure shown in FIG. 17 is merely a block diagram of some structures related to the present invention and does not constitute a limitation on the computer device on which the present invention is applied; a specific computer device may include more or fewer components than those shown, or may combine some components, or have a different component arrangement.

[0244] In one embodiment, a computer device is provided, including a memory and a processor, a computer program is stored in the memory, and the above-mentioned media data recommendation method or recommendation model processing method is realized when the processor executes the computer program. In one embodiment, a computer-readable storage medium is provided, a computer program is stored in the computer device, and the above-mentioned media data recommendation method or recommendation model processing method is realized when the processor executes the computer program.

[0245] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, realizes the above-mentioned media data recommendation method or recommendation model processing method.It should be noted that the user information (including, but not limited to, user device information and user personal information, etc.) and data (including, but not limited to, data used for analysis, data used for storage, data used for display, etc.) involved in this application are all information and data authorized by the user or fully authorized by each party, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0246] As will be understood by those skilled in the art, all or part of the processes in the methods of the above embodiments can be achieved by issuing instructions to associated hardware via a computer program. The computer program may be stored in a non-volatile computer-readable storage medium, and when executed, the computer program may implement the processes of the above method embodiments. Any references to memory, database, or other media used in the embodiments provided herein may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM), external cache memory, etc. By way of illustration and not limitation, RAM may be in multiple forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). Databases involved in the embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases.The processors involved in each embodiment provided herein may be, but are not limited to, general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, and data processing logic devices based on quantum computing.

[0247] The technical features of the above embodiments can be combined in any manner, and for the sake of brevity, not all possible combinations of the technical features in the above embodiments are described, but combinations of these technical features should be considered to be within the scope of the present specification unless there is a contradiction.

[0248] The above-described examples only show some embodiments of the present application, and the descriptions are relatively specific and detailed, but should not be understood as limiting the scope of the patent of the present application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application should be determined based on the scope of the attached claims. [Explanation of symbols]

[0249] 102 terminals 104 Server 129 images 1501 Vector Extraction Module 1502 First Knowledge Search Module 1503 First Fusion Module 1504 Recommendation Module 1601 Training Vector Extraction Module 1602 Second Knowledge Retrieval Module 1603 Second Fusion Module 1604 First loss value determination module 1605 Second loss value determination module 1606 Parameter Adjustment Module 1607 Recommendation Model Decision Module

Claims

1. 1. A method for recommending media data, executed by a server, the method comprising: extracting media representation vectors and text representation vectors from media data and description text of the media data; performing a knowledge search in a knowledge graph based on the media representation vector to obtain an entity subgraph of the media data, and determining an entity representation vector of the entity subgraph; performing a feature fusion process on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge-enhanced vector; obtaining target media data based on the knowledge-enhanced vector, and recommending the target media data to a target subject.

2. The step of extracting a media representation vector and a text representation vector from the media data and the description text of the media data includes: performing feature extraction on the media data using an image feature extraction model to obtain a media representation vector; and extracting text representation vectors from the descriptive text of the media data by a text feature extraction model.

3. The step of extracting features from media data using an image feature extraction model to obtain a media representation vector includes: When the media data is a video, extracting features of a plurality of image frames in the video using an image feature extraction model to obtain a media representation vector; The method of claim 2 , further comprising: when the media data is an image, performing feature extraction on a plurality of image blocks of the image using the image feature extraction model to obtain a media representation vector.

4. The step of performing knowledge search in a knowledge graph based on the media representation vector to obtain an entity subgraph of the media data and determining an entity representation vector of the entity subgraph includes: searching for a plurality of target entities related to the media data in a knowledge graph based on the media representation vector; determining an entity subgraph of the media data based on the plurality of target entities and the knowledge graph; The method of any one of claims 1 to 3, further comprising the step of: extracting features of a plurality of entities in the entity subgraph to obtain entity representation vectors.

5. The media representation vector includes at least two image sub-representation vectors, and the step of searching for a plurality of target entities related to the media data in a knowledge graph based on the media representation vectors includes: obtaining initial entity vectors of a plurality of entities in a knowledge graph; searching for candidate entities in the knowledge graph based on the plurality of initial entity vectors and the at least two image sub-representation vectors; and selecting, from the candidate entities, a plurality of target entities related to the media data.

6. The step of searching for candidate entities in the knowledge graph based on the plurality of initial entity vectors and the at least two image sub-representation vectors includes: determining a set of associations of the at least two image sub-representation vectors based on a plurality of the initial entity vectors and the at least two image sub-representation vectors, the set of associations including respective associations between the image sub-representation vectors and the initial entity vectors; and selecting candidate entities in the knowledge graph that are associated with the at least two image sub-representation vectors based on the set of associations.

7. determining an entity subgraph of the media data based on the plurality of target entities and the knowledge graph, determining neighboring nodes of each one of the target entities in the knowledge graph; determining an extended entity based on the plurality of target entities and the neighboring nodes of each one of the target entities; determining relationships between the extended entities in the knowledge graph; and determining an entity subgraph of the media data based on the extended entities and relationships between the extended entities.

8. The step of searching for a plurality of target entities related to the media data in a knowledge graph based on the media representation vectors includes: searching for a plurality of target entities related to the media data in a knowledge graph based on the media representation vectors by a search sub-model of a knowledge search model; The step of determining an entity subgraph corresponding to the media data based on the target entity and the knowledge graph includes: determining an entity subgraph corresponding to the media data based on the plurality of target entities and the knowledge graph through a subgraph construction network of the knowledge retrieval model; and extracting features of a plurality of entities in the entity subgraph to obtain an entity representation vector includes: The method of claim 4 , further comprising: extracting features of a plurality of entities in the entity subgraph by a graph neural network of the knowledge retrieval model to obtain entity representation vectors.

9. performing a feature fusion process on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge-enhanced vector; Combining the media representation vector, the text representation vector, and the entity representation vector, and in the process of combining, adding a separator between the media representation vector and the text representation vector, and adding a separator between the text representation vector and the entity representation vector, to obtain a combined vector; and performing a feature fusion process on the combined vector through a knowledge-enhanced model to obtain a knowledge-enhanced vector.

10. The step of obtaining target media data based on the knowledge-enhanced vector and recommending the target media data to the target subject includes: performing a classification process on the knowledge-enhanced vector to obtain interest types of the target subject; The method according to any one of claims 1 to 9, further comprising: obtaining target media data based on the interest type; and recommending the target media data to the target subject.

11. 1. A method for processing a recommendation model, the method comprising: extracting first media training vectors and first text training vectors from first sample media data and first sample text of the first sample media data based on the feature extraction model; performing a knowledge retrieval process on the first media training vector and the knowledge graph based on a knowledge retrieval model to obtain a training subgraph of the first sample media data, and determine an entity training vector of the training subgraph; performing a feature fusion process on the first media training vector, the first text training vector, and the entity training vector based on a knowledge-enhanced model to obtain a knowledge-enhanced training vector; determining visual and linguistic loss values ​​based on the knowledge-enhanced training vectors and sample labels; determining a knowledge retrieval loss value based on the knowledge-enhanced training vector and the training subgraph; adjusting parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model according to the visual loss value, the language loss value, and the knowledge retrieval loss value to obtain an enriched vector extraction model; and determining a recommendation model based on the enrichment vector extraction model and the classification model, wherein the recommendation model is used to extract a knowledge enrichment vector based on media data, description text, and knowledge graph, determine an interest type based on the knowledge enrichment vector, obtain target media data based on the interest type, and recommend the target media data to a target subject.

12. the knowledge-enhanced training vectors include media-enhanced training vectors and text-enhanced training vectors, and the sample labels include hidden sub-image labels and hidden word labels; The step of determining visual loss values ​​and linguistic loss values ​​based on the knowledge-enhanced training vectors and sample labels includes: obtaining a hidden sub-image enhancement vector in the media-enhanced training vector, and determining a visual loss value based on the hidden sub-image enhancement vector and the hidden sub-image label; performing a classification process on the text-enhanced training vectors to obtain hidden word prediction probabilities; and determining a language loss value based on the hidden word prediction probabilities and the hidden word labels.

13. The knowledge-enhanced training vectors further include entity-enhanced training vectors, and the step of determining a knowledge retrieval loss value based on the knowledge-enhanced training vectors and the training subgraphs includes: obtaining entity-positive sample pairs in the training subgraph, and determining a first score for the entity-positive sample pairs based on the entity-enriched training vectors; obtaining entity-negative sample pairs in the training sub-graph and determining a second score for the entity-negative sample pairs based on the entity-enriched training vectors, the entity-negative sample pairs including two training entities that do not have an entity relationship in the training sub-graph; and determining a knowledge search loss value based on the first score and the second score.

14. The first sample media data and the first sample text belong to a sample set, and the sample set further includes second sample media data and second sample text, and the method further comprises: performing feature extraction on the second sample media data and the second sample text to obtain second media training vectors and second text training vectors; determining an image-text contrast loss value based on the second media training vector, the second text training vector, the first media training vector, and the first text training vector; adjusting parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model based on the visual loss value, the language loss value, and the knowledge retrieval loss value to obtain an enriched vector extraction model; The method according to any one of claims 11 to 13, comprising: adjusting parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model according to the visual loss value, the language loss value, the knowledge retrieval loss value, and the image-text contrast loss value to obtain an enriched vector extraction model.

15. determining an image-text contrast loss value based on the second media training vector, the second text training vector, the first media training vector, and the first text training vector; determining a first similarity measure based on the first media training vector and the second text training vector; determining a second similarity measure based on the first text training vector and the second media training vector; and determining an image-text contrast loss value based on the first similarity measure, the second similarity measure, the first similarity label, and the second similarity label.

16. 1. A media data recommendation device, comprising: a vector extraction module used for extracting media representation vectors and text representation vectors from media data and description text of the media data; a first knowledge retrieval module, which is used for performing knowledge retrieval in a knowledge graph based on the media representation vector, obtaining an entity subgraph of the media data, and determining an entity representation vector of the entity subgraph; a first fusion module for performing feature fusion processing on the media representation vector, the text representation vector, and the entity representation vector to obtain a knowledge-enriched vector; a recommendation module adapted to obtain target media data based on the knowledge-enhanced vector and recommend the target media data to a target subject.

17. A processing device for a recommendation model, said device comprising: a training vector extraction module used to extract a first media training vector and a first text training vector from the first sample media data and the corresponding first sample text according to the feature extraction model; a second knowledge retrieval module, which is used to perform a knowledge retrieval process on the first media training vector and the knowledge graph according to a knowledge retrieval model, to obtain a training subgraph of the first sample media data, and to determine an entity training vector of the training subgraph; a second fusion module for performing a feature fusion process on the first media training vector, the first text training vector, and the entity training vector according to a knowledge-enhanced model to obtain a knowledge-enhanced training vector; a first loss value determination module, used for determining a visual loss value and a linguistic loss value according to the knowledge-enhanced training vector and the sample label; a second loss value determination module, used for determining a knowledge retrieval loss value according to the knowledge-enriched training vector and the training subgraph; a parameter adjustment module for adjusting parameters of the feature extraction model, the knowledge retrieval model, and the knowledge enrichment model according to the visual loss value, the language loss value, and the knowledge retrieval loss value to obtain an enriched vector extraction model; and a recommendation model determination module used to determine a recommendation model based on the reinforcement vector extraction model and the classification model, wherein the recommendation model is used to extract a knowledge reinforcement vector based on media data, a description text, and a knowledge graph, determine an interest type based on the knowledge reinforcement vector, obtain target media data based on the interest type, and recommend the target media data to a target subject.

18. A computing device comprising a memory and one or more processors, wherein computer readable instructions are stored in the memory and cause the one or more processors to perform the steps of the method of any one of claims 1 to 15, when the computer readable instructions are executed by the processors.

19. One or more non-volatile readable storage media having computer readable instructions stored thereon that, when executed by the processor, cause the one or more processors to perform the steps of the method of any one of claims 1 to 15.

20. A computer readable instruction product comprising computer readable instructions which, when executed by a processor, implement the steps of the method of any one of claims 1 to 15.

Citation Information

Patent Citations

  • Information processor, information processing method and program

    JP2019160064A

  • Providing a response in a session

    US20200327327A1

  • System and method for product recommendation based on multimodal fashion knowledge graph

    US20220207587A1