Method, device, electronic device and storage medium for recommending media data

CN122548035APending Publication Date: 2026-08-11HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]有鉴于此,本申请提供了一种媒体数据的推荐方法、装置、电子设备及存储介质,以解决媒体数据的推荐效果不佳的问题

Benefits of technology

[0011]本申请实施例提供的媒体数据的推荐方法,通过媒体查询信息和历史行为数据获取推荐媒体数据集,基于历史行为数据有利于准确理解媒体查询信息背后的深层语义需求,使得推荐媒体数据在匹配媒体查询信息的同时也符合用户行为偏好,提高了媒体数据的推荐效果;而且,基于媒体查询信息包含的关键词,在媒体知识图谱中进行检索,操作简便,提高了媒体数据的推荐效率,且媒体知识图谱用于表征不同实体节点之间的关联关系,且实体节点对应有属性信息,基于媒体知识图谱进行检索,提高了推荐媒体数据的准确性和全面性;而且,基于媒体查询信息的语义复杂度所指示的多个指标阈值,对第二候选媒体数据集进行筛选,得到推荐媒体数据集,一方面,以多个指标阈值为约束进一步对第二候选媒体数据集进行精细筛选,使得推荐媒体数据集能够满足多个维度(指标阈值)要求,提高了媒体数据的推荐准确度,另一方面,由语义复杂度灵活指示多个指标阈值,提高了针对媒体数据的推荐过程的灵活性,不同媒体查询信息对应有不同的语义复杂度,即不同媒体查询信息对应有不同的指标阈值,提高了媒体数据的推荐多样性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548035A_ABST
    Figure CN122548035A_ABST
Patent Text Reader

Abstract

The application discloses a media data recommendation method and device, electronic equipment and storage medium, and relates to the technical field of computer processing. The method comprises the following steps: acquiring media query information corresponding to a target user identifier and historical behavior data corresponding to the target user identifier; searching a media knowledge graph based on at least one keyword contained in the media query information to obtain a first candidate media data set; filtering the first candidate media data set based on the media query information and the historical behavior data to obtain a second candidate media data set; and filtering the second candidate media data set based on a plurality of index thresholds indicated by the semantic complexity of the media query information to obtain a recommended media data set corresponding to the media query information. The recommended media data set is obtained through the media query information and the historical behavior data, and the recommendation effect of the media data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer processing technology, and more specifically to recommended methods, apparatus, electronic devices, and storage media for media data. Background Technology

[0002] Music search and recommendation is a feature of online music service platforms.

[0003] In related technologies, the music query text entered by the user is used as the index basis. Keywords are extracted from the music query text, and the keywords are matched with song information (such as song name) to obtain the recommended media dataset corresponding to the music query text.

[0004] However, because the above methods rely solely on literal keyword matching for retrieval, they struggle to accurately understand the deeper semantic needs behind user queries. This necessitates the server frequently expanding the retrieval scope and performing multiple rounds of re-ranking calculations during the recommendation process. This not only significantly increases the server's concurrent computational pressure and data storage redundancy but also results in high recommendation response latency and limited content recall, making it difficult to balance semantic relevance with personalized needs. Summary of the Invention

[0005] In view of this, this application provides a method, apparatus, electronic device, and storage medium for recommending media data to solve the problem of poor recommendation performance of media data.

[0006] Firstly, this application provides a method for recommending media data, the method comprising: Obtain media query information corresponding to the target user identifier, as well as historical behavior data corresponding to the target user identifier; Based on at least one keyword contained in the media query information, the media knowledge graph is retrieved to obtain a first candidate media dataset; wherein, the media knowledge graph is used to represent the relationship between different entity nodes, and the entity nodes correspond to attribute information; Based on the media query information and the historical behavior data, the first candidate media dataset is filtered to obtain the second candidate media dataset; Based on multiple indicator thresholds indicated by the semantic complexity of the media query information, the second candidate media dataset is filtered to obtain the recommended media dataset corresponding to the media query information.

[0007] Secondly, this application provides a media data recommendation device, the device comprising: The information acquisition module is used to acquire media query information corresponding to the target user identifier, as well as historical behavior data corresponding to the target user identifier; The graph retrieval module is used to retrieve the media knowledge graph based on at least one keyword contained in the media query information to obtain a first candidate media dataset; wherein, the media knowledge graph is used to represent the association relationship between different entity nodes, and the entity nodes correspond to attribute information; The first filtering module is used to filter the first candidate media dataset based on the media query information and the historical behavior data to obtain a second candidate media dataset. The second filtering module is used to filter the second candidate media dataset based on multiple indicator thresholds indicated by the semantic complexity of the media query information, so as to obtain the recommended media dataset corresponding to the media query information.

[0008] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the recommended method for media data of the first aspect or any corresponding embodiment described above.

[0009] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to perform the recommended method for media data of the first aspect or any corresponding embodiment described above.

[0010] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute a recommended method for media data in the first aspect or any corresponding embodiment described above.

[0011] The media data recommendation method provided in this application obtains a recommended media dataset through media query information and historical behavior data. Based on historical behavior data, it is beneficial to accurately understand the deep semantic needs behind the media query information, ensuring that the recommended media data matches both the media query information and user behavioral preferences, thus improving the recommendation effect. Furthermore, based on the keywords contained in the media query information, retrieval is performed in the media knowledge graph, which is simple to operate and improves the recommendation efficiency. The media knowledge graph is used to represent the relationships between different entity nodes, and each entity node corresponds to attribute information; retrieval based on the media knowledge graph improves the accuracy and comprehensiveness of the recommended media data. Moreover, based on multiple indicator thresholds indicated by the semantic complexity of the media query information, a second candidate media dataset is filtered to obtain the recommended media dataset. On the one hand, the second candidate media dataset is further refined using multiple indicator thresholds as constraints, ensuring that the recommended media dataset meets multiple dimensional (indicator threshold) requirements, thus improving the recommendation accuracy. On the other hand, the semantic complexity flexibly indicates multiple indicator thresholds, improving the flexibility of the media data recommendation process. Different media query information corresponds to different semantic complexities, i.e., different media query information corresponds to different indicator thresholds, thus improving the diversity of media data recommendations. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0013] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this application; Figure 2 This is a flowchart illustrating a media data recommendation method according to an embodiment of this application; Figure 3 An illustrative diagram illustrating the method for obtaining multimodal feature vectors is provided. Figure 4 An illustrative diagram of a media recommendation model is shown below; Figure 5 This is a structural block diagram of a media data recommendation device according to an embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] It should be noted that the information (including but not limited to user input information, such as information entered by the user into input boxes), data (including but not limited to data used for analysis, stored data, and displayed data, such as context code, all code of the current project, the service pressure corresponding to operations performed on all code of the current project, and the code development status of the current project), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant laws, regulations, and standards. For example, the context code, operations performed on all code of the current project, the corresponding service pressure, and the code development status involved in this application were all obtained with full authorization.

[0016] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0017] As one optional application scenario in the embodiments of this application, such as Figure 1 As shown, the system may include at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.

[0018] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.

[0019] For example, the terminal device includes an application. This application can be one that requires downloading and installation, or it can be an application that is available instantly. For example, the application can be any application that provides a specific media data recommendation function. For example, the media data includes, but is not limited to, at least one of the following: audio information, text information, and image information. For example, the media data is a song, which includes audio information, text information (lyrics, etc.), and image information (song cover and / or music video, etc.); another example is a video, which includes audio information (background music, actors' dialogue, etc.), text information (video description, subtitles, etc.), and image information (video frames, video cover, etc.).

[0020] For example, the server is the backend server of the application. For instance, after obtaining media recommendation information from the terminal device, the server obtains a recommended media dataset based on that information and then sends the media dataset to the terminal device. Of course, if the terminal device has sufficient computing resources, it can also calculate the recommended media dataset based on media query information; this application does not limit this approach.

[0021] In related technologies, user-input music query text is used as the index basis. Keywords are extracted from the music query text and matched with song information (such as song titles) to obtain a recommended media dataset corresponding to the music query text. However, in the aforementioned related technologies, song recommendation based solely on keywords can only determine the degree of matching between songs and keywords, resulting in poor song recommendation performance.

[0022] According to an embodiment of this application, a recommended method embodiment for media data is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0023] This embodiment provides a method for recommending media data, which can be used in the aforementioned terminal devices and / or servers (hereinafter collectively referred to as electronic devices). Figure 2 This is a flowchart of a media data recommendation method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the media query information corresponding to the target user identifier, and the historical behavior data corresponding to the target user identifier.

[0024] A user identifier is used to identify a unique user. A target user identifier refers to the user identifier corresponding to a target user. For example, a target user identifier includes, but is not limited to, at least one of the following used by the target user in the above application: account, name, identity identifier (such as mobile phone number, ID card number), etc.

[0025] Media query information refers to the input information corresponding to a media query operation. A media query operation is an operation triggered by a target user. For example, a target user triggers a media query operation by performing a trigger action on a query control.

[0026] In this embodiment, upon detecting a media query operation, the electronic device acquires media query information corresponding to the target user identifier, as well as historical behavior data corresponding to the target user identifier. For example, the electronic device acquires media query information from the information input area of ​​the query page. Optionally, in the information input area, the target user inputs media query information through text input, voice input, image input, or other methods. The information input area can be displayed in any style and at any location on the query page; the display position and style of the information input area can be flexibly set and adjusted according to actual conditions.

[0027] Historical behavior data is used to indicate historical behaviors triggered by a target user identifier. For example, historical behaviors include, but are not limited to, at least one of the following: liking, saving, forwarding, looping playback, and early exiting playback. For example, historical behavior data is stored data, and the electronic device retrieves the historical behavior data corresponding to the target user identifier from the stored information based on the target user identifier.

[0028] Step S202: Based on at least one keyword contained in the media query information, the media knowledge graph is searched to obtain the first candidate media dataset.

[0029] A media knowledge graph is used to represent the relationships between different entity nodes, and each entity node corresponds to attribute information. In one possible implementation, attribute information is stored in the media knowledge graph in the form of nodes; that is, the media knowledge graph includes entity nodes, attribute nodes, edges connecting entity nodes, and edges connecting entity nodes and attribute nodes. In another possible implementation, attribute information is stored in the media knowledge graph in a non-node form; that is, the media knowledge graph includes entity nodes and edges connecting entity nodes, and each entity node corresponds to attribute information. In one embodiment of this disclosure, the media knowledge graph is constructed based on massive music metadata and user behavior logs. Its entity types include at least singers, songs, albums, music styles, emotional tags, and scene tags; its relationship types include at least singing, inclusion, similarity, genre, and applicable scenarios; and its entity attribute information includes at least release time, beats per minute, language, and audio fingerprint. This knowledge graph supports batch construction in the offline indexing stage and incremental updates in the online retrieval stage. For example, when a new song is added to the database, the system automatically extracts its metadata and links it to existing entity nodes to maintain the timeliness and integrity of the graph.

[0030] In this embodiment, after obtaining the aforementioned media query information, the electronic device retrieves the media knowledge graph based on at least one keyword contained in the media query information to obtain a first candidate media dataset. The first candidate media dataset includes at least one first candidate media data. For example, at least one keyword includes an entity obtained from the media query information.

[0031] For example, the electronic device extracts keywords from the media query information to obtain at least one keyword; further, in the media knowledge graph, based on the similarity between keywords and entities, entity retrieval is performed on each keyword to obtain a first dataset; and, in the media knowledge graph, based on the similarity between keywords and attribute information, attribute retrieval is performed on each keyword to obtain a second dataset; then, the union of the first dataset and the second dataset is taken to obtain a first candidate media dataset.

[0032] Step S203: Based on media query information and historical behavior data, the first candidate media dataset is filtered to obtain the second candidate media dataset.

[0033] In this embodiment of the application, after obtaining the first candidate media dataset, the electronic device filters the first candidate media dataset based on media query information and historical behavior data to obtain a second candidate media dataset. The second candidate media dataset includes at least one second candidate media data set.

[0034] For example, since historical behavioral data is used when filtering the first candidate media dataset, this "filtering of the first candidate media dataset" can also be called "personalized filtering of the first candidate media dataset".

[0035] Step S204: Based on multiple indicator thresholds indicated by the semantic complexity of the media query information, the second candidate media dataset is filtered to obtain the recommended media dataset corresponding to the media query information.

[0036] Semantic complexity refers to the inherent complexity of media query information at the semantic level, used to measure the richness, abstractness, or cumbersome organizational structure of the media query information. In one embodiment of this disclosure, the electronic device first performs word segmentation and named entity recognition on the media query information, counts the number of semantic entities, adjectives, and scene modifiers contained in the media query text, and calculates a semantic complexity score based on preset weight rules, then maps this score to three complexity levels: low, medium, and high. In another embodiment, the electronic device inputs the media query information into a pre-trained semantic understanding model (e.g., a BERT model fine-tuned with music domain corpus), and the model directly outputs the probability distribution of the media query information belonging to each complexity level, taking the highest probability as the semantic complexity level of the media query information. Different complexity levels correspond to different matching index thresholds, personalization index thresholds, and diversity index thresholds, and the above mapping relationship can be dynamically corrected based on historical recommendation feedback data.

[0037] In this embodiment, after obtaining the second candidate media dataset, the electronic device filters the second candidate media dataset based on multiple indicator thresholds indicated by the semantic complexity of the media query information to obtain a recommended media dataset corresponding to the media query information. The recommended media dataset includes at least one recommended media data point.

[0038] The media data recommendation method provided in this embodiment obtains a recommended media dataset through media query information and historical behavior data. Based on historical behavior data, it is beneficial to accurately understand the deep semantic needs behind the media query information, ensuring that the recommended media data matches both the media query information and user behavioral preferences, thus improving the recommendation effect. Furthermore, based on the keywords contained in the media query information, retrieval is performed in the media knowledge graph, which is simple to operate and improves the recommendation efficiency. The media knowledge graph is used to represent the relationships between different entity nodes, and each entity node corresponds to attribute information; retrieval based on the media knowledge graph improves the accuracy and comprehensiveness of the recommended media data. Moreover, based on multiple indicator thresholds indicated by the semantic complexity of the media query information, the second candidate media dataset is filtered to obtain the recommended media dataset. On the one hand, the second candidate media dataset is further refined using multiple indicator thresholds as constraints, ensuring that the recommended media dataset meets multiple dimensional (indicator threshold) requirements, thus improving the recommendation accuracy. On the other hand, the semantic complexity flexibly indicates multiple indicator thresholds, improving the flexibility of the media data recommendation process. Different media query information corresponds to different semantic complexities, i.e., different media query information corresponds to different indicator thresholds, thus improving the diversity of media data recommendations.

[0039] In an exemplary embodiment, step S204 includes: Step S2041: Based on semantic complexity, determine the threshold for matching index, personalization index, and diversity index.

[0040] In this embodiment, after obtaining the aforementioned media query information, the electronic device determines a matching degree threshold, a personalization threshold, and a diversity threshold based on the semantic complexity of the media query information. The matching degree threshold measures the relevance between the recommended media data and the media query information; the personalization threshold measures the relevance between the media data and historical behavior data; and the diversity threshold measures the similarity between different media data sets.

[0041] For example, an electronic device determines a matching index threshold based on a first mapping relationship and semantic complexity; determines a personalization index threshold based on a second mapping relationship and semantic complexity; and determines a diversity index threshold based on a third mapping relationship and semantic complexity. The first mapping relationship characterizes the mapping relationship between semantic complexity levels and matching index thresholds; the second mapping relationship characterizes the mapping relationship between semantic complexity levels and personalization indexes; and the third mapping relationship characterizes the mapping relationship between semantic complexity levels and diversity index thresholds.

[0042] For example, the semantic complexity levels include three levels: "low", "medium", and "high". For example, the first mapping relationship, the second mapping relationship, and the third mapping relationship are pre-set mapping relationships.

[0043] For example, if the semantic complexity level is "low," it indicates that the media query information is precise. In this case, it can be understood that the query has a precise query purpose, and priority should be given to satisfying the query purpose indicated by the media query text. The high matching index threshold is determined based on the first mapping relationship, the low personalization index threshold is determined based on the second mapping relationship, and the low diversity index threshold is determined based on the third mapping relationship. If the semantic complexity level is "medium," it indicates that the media query information is relatively precise. In this case, it can be understood that the query has a relatively precise query purpose, and the query purpose indicated by the media query text and the user preferences indicated by the target user identifier should be balanced. The medium matching index threshold is determined based on the first mapping relationship, the medium personalization index threshold is determined based on the second mapping relationship, and the low diversity index threshold is determined based on the third mapping relationship. If the semantic complexity level is "high," it indicates that the media query information is ambiguous. In this case, it can be understood that the query does not have a precise query purpose, and it is necessary to recommend diverse media data to the target user. The low matching index threshold is determined based on the first mapping relationship, the low personalization index threshold is determined based on the second mapping relationship, and the high diversity index threshold is determined based on the third mapping relationship.

[0044] It's important to clarify that the "low," "medium," and "high" thresholds mentioned above refer only to the threshold itself and are unrelated to other thresholds. For example, due to the characteristics of media data recommendations, the low matching degree threshold is higher than the high personalization threshold and also higher than the high diversity threshold.

[0045] It should also be noted that the above description of the mapping relationship is only exemplary and explanatory. In practical applications, the first, second, and third mapping relationships can be flexibly set and adjusted according to the actual situation.

[0046] Step S2042: Based on the matching degree index threshold, the personalization index threshold, and the diversity index threshold, the second candidate media dataset is filtered to obtain the recommended media dataset.

[0047] In this embodiment, after obtaining the second candidate media data and the threshold values ​​for each indicator, the electronic device filters the second candidate media dataset based on the matching degree threshold, the personalization threshold, and the diversity threshold to obtain a recommended media dataset. The recommended media dataset includes at least one recommended media data point.

[0048] For example, to further improve the screening effect for the second candidate media data, in addition to the matching degree index threshold, the personalization index threshold, and the diversity index threshold, the above-mentioned multiple index thresholds also include an environmental index threshold. Specifically, step S2042 includes: Step S2042a: Obtain environmental information based on the environment where the target user is located.

[0049] In this embodiment of the application, after obtaining the aforementioned media query information, the electronic device obtains environmental information based on the environment in which the target user is located.

[0050] For example, environmental information includes, but is not limited to, at least one of the following: temperature, humidity, light intensity, walking speed, driving speed, driving environment, etc.

[0051] Step S2042b: Obtain the environmental indicator threshold indicated by the environmental information.

[0052] In this embodiment, after acquiring the aforementioned environmental information, the electronic device acquires an environmental indicator threshold indicated by the environmental information. The environmental indicator threshold is used to measure the degree of matching between the media data and the environmental information.

[0053] For example, an electronic device determines environmental indicator thresholds based on a fourth mapping relationship and environmental information. The fourth mapping relationship characterizes the mapping relationship between environmental information and environmental indicator thresholds. For example, the fourth mapping relationship is a pre-set mapping relationship.

[0054] For example, if the environmental information includes high driving speed and congested driving environment, then a safety indicator threshold is introduced based on the fourth mapping relationship. For example, the safety indicator threshold includes: a tempo intensity threshold (BPM, beats per minute) <120, low emotional intensity, and a volume dynamic range <40 dB.

[0055] It should be noted that the above introduction to environmental indicator thresholds and the fourth mapping relationship is only exemplary and explanatory. In practical applications, the fourth mapping relationship can be flexibly set and adjusted according to the actual situation.

[0056] Step S2042c: Based on the matching degree index threshold, the personalization index threshold, the diversity index threshold, and the environmental index threshold, the second candidate media dataset is filtered to obtain the recommended media dataset.

[0057] In this embodiment of the application, after obtaining the second candidate media data and the threshold values ​​of each indicator, the electronic device filters the second candidate media dataset based on the matching degree indicator threshold, the personalization indicator threshold, the diversity indicator threshold, and the environmental indicator threshold to obtain the recommended media dataset.

[0058] This embodiment provides a media data recommendation method. By determining the matching degree threshold, personalization threshold, and diversity threshold through semantic complexity, a recommended media dataset is obtained based on these thresholds. The second candidate media dataset is then finely screened from three dimensions: matching degree, personalization, and diversity. On one hand, a reasonable adjustment of the matching degree threshold can improve the relevance between the recommended media data and the media query information, facilitating the recommendation of media data that matches the query information. On the other hand, a reasonable adjustment of the personalization threshold can improve the relevance between the media data and user behavior data, facilitating the recommendation of media data that aligns with user behavior preferences. Furthermore, a reasonable adjustment of the diversity threshold can reduce the similarity between different recommended media data within the same recommended media dataset, facilitating the recommendation of a richer and more diverse range of media data.

[0059] In addition, by obtaining environmental indicator thresholds based on the user's environment, and by reasonably adjusting these thresholds, the matching degree between media data and environmental information can be improved. This is beneficial for recommending media data that matches the environmental information, and further improves the recommendation effect of media data.

[0060] In an exemplary embodiment, step S203 includes: Step S2031: After feature encoding of media query information and historical behavior data respectively, the query feature vector is obtained by feature concatenation.

[0061] In this embodiment of the application, after obtaining the above-mentioned media query information and the above-mentioned historical behavior data, the electronic device performs feature encoding on the media query information and the historical behavior data respectively, and obtains the query feature vector by feature concatenation.

[0062] For example, the electronic device encodes the music query information to obtain a semantic feature vector; it encodes the historical behavior data to obtain a historical feature vector; and further, it concatenates the semantic feature vector and the historical behavior data to obtain a query feature vector.

[0063] Step S2032: Obtain the multimodal feature vector of the first candidate media data.

[0064] In this embodiment of the application, before acquiring the second candidate media dataset, the electronic device acquires the multimodal feature vector of the first candidate media data. The first candidate media dataset includes at least one first candidate media data.

[0065] In this embodiment, the multimodal feature vector includes media text information based on the first candidate media data and a text feature vector obtained from at least one response query text. The response query text refers to the media query text that, in a media data recommendation scenario, indicates an action taken by a user identifier regarding the first candidate media data; the media data recommendation scenario is triggered by the media query text; the user identifier includes the aforementioned target user identifier, as well as other user identifiers besides the target user identifier. For example, the action includes, but is not limited to, at least one of the following: liking, playing, or favoriteing.

[0066] For example, the multimodal feature vector also includes audio feature vectors and visual feature vectors. For example, for the first candidate media data, after acquiring the first candidate media data, the electronic device acquires the audio information, text information, image information, and at least one response query text of the first candidate media data; further, it performs feature encoding on the audio information to obtain an audio feature vector; and performs feature encoding on the text information and at least one response query text to obtain a text feature vector; and performs feature encoding on the image information to obtain a visual feature vector; then, it performs cross-modal attention fusion on the audio feature vector, text feature vector, and visual feature vector to obtain a multimodal feature vector.

[0067] For example, audio feature vectors include rhythm intensity components, timbre characteristics components, pitch variation components, and acoustic structure components; text feature vectors include lyrics feature components, emotion feature components, theme feature components, keyword feature components, and sentence style feature components; visual feature vectors include color style components, composition feature components, and visual element (such as cover image) components. Among these, emotion feature components and theme feature components are obtained by fusing the response query text with the text information.

[0068] For example, the electronic device obtains the multimodal feature vector of the first candidate media data through residual quantization variational autoencoder, semantic similarity selection, and semantic compression. Taking a song as an example, the first candidate media data... Figure 3 As shown, the song includes audio information, text information (lyrics), and image information (song cover and / or music video, etc.). The electronic device uses a residual quantization variational autoencoder 31 to decompose the audio information, text information (lyrics), response query text, and image information; then, semantic similarity is calculated on the decomposed feature components, and the K feature components with the highest relevance are selected as effective feature components; then, semantic compression is performed on the K feature components to obtain audio feature vectors, text feature vectors, and visual feature vectors; finally, multimodal feature vectors are obtained through cross-modal attention fusion. Here, K is a positive integer, and K can be adjusted according to actual conditions; this embodiment does not limit this adjustment.

[0069] Step S2033: Based on the similarity between the query feature vector and the multimodal feature vectors of each first candidate media data, the first candidate media dataset is filtered to obtain the second candidate media dataset.

[0070] In this embodiment of the application, after obtaining the query feature vector and the multimodal feature vector, the electronic device filters the first candidate media dataset based on the similarity between the query feature vector and the multimodal feature vector of each first candidate media data, and obtains the second candidate media dataset.

[0071] Specifically, step S2033 includes: Step S2033a: Obtain the historical interaction time features corresponding to the target user identifier from the historical behavior data corresponding to the target user identifier.

[0072] In this embodiment of the application, after obtaining the aforementioned historical behavior data, the electronic device obtains the historical interaction time features corresponding to the target user identifier from the historical behavior data corresponding to the target user identifier.

[0073] Step S2033b: Obtain time decay weights based on historical interaction time characteristics.

[0074] In this embodiment of the application, after obtaining the aforementioned historical interaction time characteristics, the electronic device obtains a time decay weight based on the historical interaction time characteristics.

[0075] For example, the electronic device obtains the last interaction time corresponding to the target user identifier based on historical interaction time characteristics; further, it obtains a time decay weight based on a preset time decay coefficient, the current time, and the last interaction time. The time decay weight is used to characterize the degree of attention paid to the personalized starting point defined by the historical feature vector, which is obtained by feature encoding of historical behavioral data.

[0076] For example, the time decay coefficient is a model parameter of the media recommendation model described below. For example, the personalization starting point refers to the initial preferences of the target user identifier for media data as defined by the historical feature vector.

[0077] Step S2033c: Based on the time decay weight, the personalized starting point, and the similarity between the query feature vector and the multimodal feature vectors of each first candidate media data, the first candidate media dataset is filtered to obtain the second candidate media dataset.

[0078] In this embodiment of the application, after obtaining the first candidate media data, the electronic device filters the first candidate media dataset based on the time decay weight, the personalization starting point, and the similarity between the query feature vector and the multimodal feature vector of each first candidate media data to obtain the second candidate media dataset.

[0079] For example, the electronic device uses a hippocampal memory retrieval function to determine the similarity between the query feature vector and the multimodal feature vectors of each first candidate media data. For example, the hippocampal memory retrieval function is the model function of the media recommendation model described below.

[0080] The media data recommendation method provided in this embodiment filters the first candidate media dataset by the similarity between the query feature vector and the multimodal feature vector. The multimodal feature vector includes media text information based on the first candidate media data and at least one text feature vector obtained from the response query text. The response query text refers to the media query text obtained from the user's operation behavior on the first candidate media data in the media data recommendation scenario. That is, when obtaining the multimodal feature vector, the historical recommendation effect of the first candidate media data obtained from the user's perspective is integrated, which improves the accuracy of the multimodal feature vector, which is conducive to improving the filtering effect of the first candidate media dataset, and thus improving the recommendation effect of media data.

[0081] In addition, by using time decay weights to filter the first candidate media datasets, the time decay weights can balance the personalization and diversity of the recommendation process, which is conducive to recommending personalized and diverse media data and improving the recommendation effect of media data.

[0082] In an exemplary embodiment, step S202 includes: Step S2021: Obtain at least one keyword contained in the media query information.

[0083] In this embodiment of the application, after obtaining the above-mentioned media query information, the electronic device obtains at least one keyword contained in the media query information.

[0084] Step S2022: Expand at least one keyword based on the synonyms of each keyword to obtain a set of search terms.

[0085] In this embodiment of the application, after obtaining at least one keyword, the keyword is expanded based on synonyms of each keyword to obtain a search term set. The search term set includes at least one search term.

[0086] Step S2023: In the media knowledge graph, each search term is searched separately to obtain the first candidate media dataset.

[0087] In this embodiment of the application, after obtaining the above-mentioned set of search terms, the electronic device searches for each search term in the media knowledge graph to obtain the first candidate media dataset.

[0088] For example, similar to the introduction of keyword retrieval above, the electronic device performs entity retrieval for each search term in the media knowledge graph based on the similarity between the search term and the entity, and obtains the third dataset; and, in the media knowledge graph, performs attribute retrieval for each search term based on the similarity between the search term and attribute information, and obtains the fourth dataset; then, the union of the third dataset and the fourth dataset is taken to obtain the first candidate media dataset.

[0089] The media data recommendation method provided in this embodiment obtains search terms by expanding keywords with synonyms, and then performs searches in the media knowledge graph based on the search terms, which improves the accuracy and comprehensiveness of the search and helps to improve the overall recommendation effect of media data.

[0090] Optionally, the media data recommendation method described above is implemented through a media recommendation model. In this embodiment, media query information and historical behavior data are the input information for the media recommendation model. Figure 4 As shown, the media recommendation model 40 includes a retrieval layer 41, a personalized filtering layer 42, and a diversity filtering layer 43.

[0091] The retrieval layer 41 retrieves the media knowledge graph based on at least one keyword contained in the media query information to obtain the first candidate media dataset. The personalized filtering layer 42 filters the first candidate media dataset based on media query information and historical behavior data to obtain the second candidate media dataset.

[0092] The diversity filtering layer 43 filters the second candidate media dataset based on multiple indicator thresholds indicated by the semantic complexity of media query information, and obtains the recommended media dataset.

[0093] In addition, the specific processing methods of the media recommendation model based on media query information and historical behavior data are as shown above, and will not be repeated here.

[0094] For example, the expression P of the personalized filtering layer personalized Heavy ; in, Here, A represents the adjacency matrix and P represents the probability transition matrix, which are the time decay weights mentioned above. As the starting point for the above personalization, These are the model weight coefficients. For the above hippocampal memory retrieval function, H query This refers to the feature vector for the above query.

[0095] For example, time decay weight The calculation formula is: ; in, Based on the base time decay weight, The above time decay coefficient is t1, the above current time is t2, and the above last interaction time is t2.

[0096] For example, the media recommendation model is a contrastive learning model. For example, after acquiring the aforementioned recommended media dataset, the electronic device obtains positive training samples and negative training samples based on the target user account's actions related to the recommended media data. The media recommendation model is then incrementally trained based on these positive and negative training samples. The positive training samples include recommended media data for which actions were detected, and the negative training samples include recommended media data for which actions were not detected.

[0097] The media data recommendation method provided in this embodiment uses a media recommendation model to input media query information and historical behavior data, and outputs a recommended media dataset. It is easy to operate and improves the efficiency of media data recommendation.

[0098] As one or more specific application embodiments of this application, the optimal implementation scheme or the scheme that the inventors most want to embody is described in combination with the specific application scenario.

[0099] This embodiment also provides a media data recommendation device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0100] This embodiment provides a media data recommendation device, such as... Figure 5 As shown, it includes: The information acquisition module 501 is used to acquire media query information corresponding to the target user identifier, as well as historical behavior data corresponding to the target user identifier.

[0101] The graph retrieval module 502 is used to retrieve the media knowledge graph based on at least one keyword contained in the media query information to obtain a first candidate media dataset; wherein, the media knowledge graph is used to represent the relationship between different entity nodes, and the entity nodes have corresponding attribute information.

[0102] The first filtering module 503 is used to filter the first candidate media dataset based on media query information and historical behavior data to obtain the second candidate media dataset.

[0103] The second filtering module 504 is used to filter the second candidate media dataset based on multiple indicator thresholds indicated by the semantic complexity of the media query information, so as to obtain the recommended media dataset corresponding to the media query information.

[0104] In some alternative implementations, the second screening module 504 includes: The indicator acquisition unit is used to determine the thresholds for matching degree indicator, personalization indicator, and diversity indicator based on semantic complexity. Among them, the matching degree indicator is used to measure the relevance between recommended media data and media query information, the personalization indicator is used to measure the relevance between media data and historical behavior data, and the diversity indicator is used to measure the similarity between various media data. The second filtering unit is used to filter the second candidate media dataset based on the matching degree index threshold, the personalization index threshold, and the diversity index threshold to obtain the recommended media dataset; wherein the recommended media dataset includes at least one recommended media data.

[0105] In some alternative implementations, the second filtering unit is used for: Obtain environmental information based on the environment in which the target user is located; The environmental indicator thresholds indicated by the environmental information are obtained; whereby the environmental indicator thresholds are used to measure the degree of matching between media data and environmental information. Based on the matching degree threshold, personalization threshold, diversity threshold, and environmental threshold, the second candidate media dataset is filtered to obtain the recommended media dataset.

[0106] In some alternative implementations, the first screening module 503 includes: The feature acquisition unit is used to encode the media query information and historical behavior data separately, and then obtain the query feature vector by concatenating the features. The vector acquisition unit is used to acquire the multimodal feature vector of the first candidate media data. The first candidate media data includes at least one first candidate media data. The multimodal feature vector includes media text information based on the first candidate media data and at least one text feature vector obtained from the response query text. The response query text refers to the media query text that has been obtained from the user's identification of the operation behavior for the first candidate media data in the media data recommendation scenario. The first filtering unit is used to filter the first candidate media dataset based on the similarity between the query feature vector and the multimodal feature vector of each first candidate media data, so as to obtain the second candidate media dataset.

[0107] In some alternative implementations, the first screening unit is used for: Obtain the historical interaction time characteristics corresponding to the target user identifier from the historical behavior data corresponding to the target user identifier; Based on historical interaction time characteristics, obtain time decay weights; Based on time decay weights, personalized starting points, and the similarity between the query feature vector and the multimodal feature vectors of each first candidate media data, the first candidate media dataset is filtered to obtain the second candidate media dataset.

[0108] In some alternative implementations, the map retrieval module 502 includes: Key extraction unit, used to obtain at least one keyword contained in media query information; A key expansion unit is used to expand at least one keyword based on the synonyms of each keyword to obtain a set of search terms; wherein the set of search terms includes at least one search term; The graph retrieval unit is used to retrieve each search term in the media knowledge graph to obtain the first candidate media dataset.

[0109] In some optional implementations, media query information and historical behavior data serve as input information for the media recommendation model; wherein, the media recommendation model includes a retrieval layer, a personalized filtering layer, and a diversity filtering layer; The retrieval layer retrieves the media knowledge graph based on at least one keyword contained in the media query information to obtain the first candidate media dataset; The personalized filtering layer filters the first candidate media dataset based on media query information and historical behavior data to obtain the second candidate media dataset; The diversity filtering layer filters the second candidate media dataset based on multiple indicator thresholds indicated by the semantic complexity of the media query information, thus obtaining the recommended media dataset.

[0110] The media data recommendation apparatus provided in this application embodiment can execute the media data recommendation method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0111] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0112] The following is a detailed reference. Figure 6 This diagram illustrates a suitable structural schematic for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from memory 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the electronic device. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0113] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0114] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a memory 608, or installed from a ROM 602. When the computer program is executed by the processor 601, it performs the functions defined in the recommended method for media data of embodiments of this application.

[0115] Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0116] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the recommended method for media data shown in the above embodiments is implemented.

[0117] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0118] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A method for recommending media data, characterized in that, The method includes: Obtain media query information corresponding to the target user identifier, as well as historical behavior data corresponding to the target user identifier; Based on at least one keyword contained in the media query information, the media knowledge graph is retrieved to obtain a first candidate media dataset; wherein, the media knowledge graph is used to represent the relationship between different entity nodes, and the entity nodes correspond to attribute information; Based on the media query information and the historical behavior data, the first candidate media dataset is filtered to obtain the second candidate media dataset; Based on multiple indicator thresholds indicated by the semantic complexity of the media query information, the second candidate media dataset is filtered to obtain the recommended media dataset corresponding to the media query information.

2. The method of claim 1, wherein, The second candidate media dataset is filtered based on multiple threshold indicators indicated by the semantic complexity of the media query information to obtain a recommended media dataset corresponding to the media query information, including: Based on the semantic complexity, a matching degree threshold, a personalization threshold, and a diversity threshold are determined; wherein, the matching degree threshold is used to measure the relevance between the recommended media data and the media query information, the personalization threshold is used to measure the relevance between the media data and the historical behavior data, and the diversity threshold is used to measure the similarity between the various media data. Based on the matching degree threshold, the personalization threshold, and the diversity threshold, the second candidate media dataset is filtered to obtain the recommended media dataset; wherein, the recommended media dataset includes at least one recommended media data.

3. The method of claim 2, wherein, The second candidate media dataset is filtered based on the matching degree threshold, the personalization threshold, and the diversity threshold to obtain the recommended media dataset, including: Environmental information is obtained based on the environment in which the target user identifier is located; Obtain the environmental indicator threshold indicated by the environmental information; wherein the environmental indicator threshold is used to measure the degree of matching between the media data and the environmental information; Based on the matching degree threshold, the personalization threshold, the diversity threshold, and the environmental threshold, the second candidate media dataset is filtered to obtain the recommended media dataset.

4. The method of claim 1, wherein, The process of filtering the first candidate media dataset based on the media query information and the historical behavior data to obtain a second candidate media dataset includes: After encoding the media query information and the historical behavior data separately, the query feature vector is obtained by concatenating the features. A multimodal feature vector of first candidate media data is obtained, wherein the first candidate media dataset includes at least one first candidate media data; the multimodal feature vector includes media text information based on the first candidate media data and a text feature vector obtained from at least one response query text; wherein the response query text refers to the media query text that has been obtained from user identification of operation behavior for the first candidate media data in the media data recommendation scenario; Based on the similarity between the query feature vector and the multimodal feature vectors of each of the first candidate media data, the first candidate media dataset is filtered to obtain the second candidate media dataset.

5. The method of claim 4, wherein, The second candidate media dataset is obtained by filtering the first candidate media dataset based on the similarity between the query feature vector and the multimodal feature vectors of each of the first candidate media data, including: Obtain the historical interaction time characteristics corresponding to the target user identifier from the historical behavior data corresponding to the target user identifier; Based on the historical interaction time characteristics, obtain the time decay weight; Based on the time decay weight, the personalized starting point, and the similarity between the query feature vector and the multimodal feature vectors of each of the first candidate media data, the first candidate media dataset is filtered to obtain the second candidate media dataset.

6. The method of claim 1, wherein, The first candidate media dataset is obtained by retrieving the media knowledge graph based on at least one keyword contained in the media query information, including: Obtain at least one keyword contained in the media query information; The at least one keyword is expanded based on the synonyms of each keyword to obtain a search term set; wherein the search term set includes at least one search term; In the media knowledge graph, each of the search terms is searched to obtain the first candidate media dataset.

7. The method according to any one of claims 1 to 6, characterized in that, The media query information and the historical behavior data are the input information for the media recommendation model; wherein, the media recommendation model includes a retrieval layer, a personalized filtering layer, and a diversity filtering layer; The retrieval layer retrieves the media knowledge graph based on at least one keyword contained in the media query information to obtain the first candidate media dataset; The personalized filtering layer filters the first candidate media dataset based on the media query information and the historical behavior data to obtain the second candidate media dataset; The diversity filtering layer filters the second candidate media dataset based on multiple indicator thresholds indicated by the semantic complexity of the media query information to obtain the recommended media dataset.

8. A device for recommending media data, characterized by The device includes: The information acquisition module is used to acquire media query information corresponding to the target user identifier, as well as historical behavior data corresponding to the target user identifier; The graph retrieval module is used to retrieve the media knowledge graph based on at least one keyword contained in the media query information to obtain a first candidate media dataset; wherein, the media knowledge graph is used to represent the association relationship between different entity nodes, and the entity nodes correspond to attribute information; The first filtering module is used to filter the first candidate media dataset based on the media query information and the historical behavior data to obtain a second candidate media dataset. The second filtering module is used to filter the second candidate media dataset based on multiple indicator thresholds indicated by the semantic complexity of the media query information, so as to obtain the recommended media dataset corresponding to the media query information.

9. An electronic device, comprising: include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the recommended method for media data according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the recommended method for the media data as described in any one of claims 1 to 7.