Short video recommendation method and device, equipment and medium

By building a user-short video interactive graph and using a multi-modal graphic diffusion model, users' preference weight for short videos is determined, and the problem of low accuracy of personalized recommendations for short videos in the existing technology is solved, and a more efficient personalized recommendation effect is achieved.

CN120030228APending Publication Date: 2025-05-23PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510039128.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to improve the accuracy of personalized recommendations for short videos, and it is impossible to effectively model interaction with user short videos with multimodal context.

Method used

By obtaining user portrait features and short video data of different modalities, a user-short video interaction diagram is constructed, and inputting it into a multimodal graph diffusion model based on the modal perceptual signal injection mechanism, the final embedding vector of the user and the short video is calculated, the user's preference weight for short videos is determined, and finally personalized recommendation is made.

Benefits of technology

It effectively improves the accuracy and recommendation speed of personalized short video recommendations and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030228A_ABST
    Figure CN120030228A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video recommendation, in particular to a short video recommendation method, device and equipment and a medium, which are applied to a medical scene, and can recommend a short video according to user portrait features and short video data in different modes by obtaining the user portrait features and the short video data in different modes. Respectively constructing user-short video interaction diagrams by utilizing a preset construction method, inputting the user-short video interaction diagrams of each mode into a multi-mode graph diffusion model based on a mode perception signal injection mechanism, and calculating to obtain a final embedding vector of a user and a final embedding vector of a short video; and according to the final embedding vector of the user and the final embedding vector of the short video, determining the preference weight of the user for the short video, and further selecting a preset number of short videos with the top preference weight according to the high-low order of the preference weight of the user for the short video so as to obtain a short video recommendation result. Therefore, the accuracy and recommendation speed in personalized recommendation of the short video are effectively improved, and the user experience is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video recommendation, and in particular to a short video recommendation method, device, equipment and medium. Background Art

[0002] On social networking platforms, a huge amount of social data is generated every day, which contains extremely valuable information. For example, short videos can be used to show real clinical cases, such as emergency treatment for heart disease and surgical operations, combined with the background of medical technology, so that users can understand the application and effect of medical technology. However, on the one hand, these social data lack semantic content descriptions, and on the other hand, social data in a cross-platform context is often not limited to one form, involving multiple modalities such as text, images, and videos, which leads to huge challenges for users in the process of browsing, searching, and managing resources.

[0003] Existing technologies usually use self-supervised learning techniques for video recommendation. However, these methods often rely on simple random enhancements or intuitive cross-view information, which may introduce irrelevant noise and fail to accurately combine multimodal context with user short video interaction modeling, resulting in ineffective and poorly accurate personalized short video recommendations. Therefore, how to improve the accuracy of personalized short video recommendations to users is a technical problem that needs to be solved urgently. Summary of the invention

[0004] Based on this, it is necessary to address the above technical problems. The embodiments of the present invention provide a short video recommendation method, device, equipment and medium to solve the technical problem in the prior art that the accuracy of personalized short video recommendations to users cannot be improved.

[0005] A first aspect of an embodiment of the present application provides a short video recommendation method, the short video recommendation method comprising: Acquire user portrait features and short video data of different modalities, wherein the user portrait features include user operation features and user preference features, and the different modalities include text modality, visual modality, and audio modality; According to the user portrait features and the short video data of different modalities, a user-short video interaction graph is constructed respectively using a preset construction method; The user-short video interaction graph of each modality is input into the multimodal graph diffusion model based on the modality perception signal injection mechanism, and the final embedding vector of the user and the final embedding vector of the short video are calculated; Determining the user's preference weight for the short video according to the final embedding vector of the user and the final embedding vector of the short video; According to the ranking of the user's preference weights for short videos, a predetermined number of short videos with high preference weights are selected to obtain short video recommendation results.

[0006] A second aspect of an embodiment of the present application provides a short video recommendation device, the short video recommendation device comprising: An acquisition module, used to acquire user portrait features and short video data of different modalities, wherein the user portrait features include user operation features and user preference features, and the different modalities include text modality, visual modality and audio modality; A construction module, used to construct user-short video interaction graphs respectively according to the user portrait features and the short video data of different modalities using a preset construction method; A calculation module is used to input the user-short video interaction graph of each modality into a multimodal graph diffusion model based on a modality perception signal injection mechanism, and calculate the final embedding vector of the user and the final embedding vector of the short video; A determination module, used to determine the user's preference weight for the short video according to the final embedding vector of the user and the final embedding vector of the short video; The recommendation module is used to select a predetermined number of short videos with high preference weights according to the user's preference weights for short videos, so as to obtain short video recommendation results.

[0007] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the short video recommendation method as described in the first aspect is implemented.

[0008] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the short video recommendation method as described in the first aspect is implemented.

[0009] In summary, the present invention provides a short video recommendation method, device, equipment and medium, which obtains user portrait features and short video data of different modalities, wherein the user portrait features include user operation features and user preference features, and the different modalities include text modality, visual modality and audio modality. According to the user portrait features and the short video data of the different modalities, a preset construction method is used to respectively construct a user-short video interaction graph, and the user-short video interaction graph of each modality is input into a multimodal graph diffusion model based on a modal perception signal injection mechanism, and the final embedding vector of the user and the final embedding vector of the short video are calculated. According to the final embedding vector of the user and the final embedding vector of the short video, the user's preference weight for the short video is determined, and then a predetermined number of short videos with a high preference weight are selected according to the user's preference weight for the short video to obtain a short video recommendation result, thereby effectively improving the accuracy and recommendation speed in personalized short video recommendation and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technical businesses in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0011] Figure 1 This is an application environment diagram of a short video recommendation method provided by an embodiment of the present invention; Figure 2 It is a flowchart of a short video recommendation method provided by an embodiment of the present invention; Figure 3 is a structural schematic diagram of a short video recommendation device provided by an embodiment of the present invention; Figure 4 It is a structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0012] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technical personnel in this field without creative work are within the scope of protection of the present invention.

[0013] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0014] It should also be understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0015] As used in the present specification and the appended claims, the term “if” may be interpreted as “when” or “uponce” or “in response to determining”, depending on the context. Similarly, the phrases “if it is determined” or “if matched to [described condition or event]” may be interpreted as meaning “upon determination” or “in response to determination” or “uponce matched to [described condition or event]” or “in response to matching to [described condition or event]”, depending on the context.

[0016] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0017] References to "one embodiment" or "some embodiments" etc. described in the present specification mean that one or more embodiments of the present invention include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0018] It should be understood that the order of execution of the steps in the following embodiments does not imply a precedence of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0019] In order to illustrate the technical solution of the present invention, specific embodiments are provided below for illustration.

[0020] See also Figure 1 , is an application environment diagram of a short video recommendation method provided by an embodiment of the present invention. A short video recommendation method provided by an embodiment of the present invention can be applied in Figure 1 In the application environment, the client communicates with the server. The client includes but is not limited to PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs) and other computer devices. The client here is replaced by a business system. The server can be implemented with an independent server or a server cluster composed of multiple servers. The short video recommendation method is used through the server. The application of the short video recommendation method in medical scenarios is crucial. It helps to improve patients' acceptance and understanding of medical knowledge, enhance their self-care awareness, and ensure the safety and reliability of the medical process. It can also push relevant academic lecture short videos, clinical operation demonstration videos, etc. according to the professional fields and learning history of medical staff.

[0021] See also Figure 2 , is a flow chart of a short video recommendation method provided by an embodiment of the present invention, such as Figure 2As shown, the short video recommendation method can be implemented through the following steps.

[0022] S201: Acquire user portrait features and short video data of different modalities, wherein the user portrait features include user operation features and user preference features, and the different modalities include text modality, visual modality and audio modality.

[0023] In step S201, short video data and user behavior data of different modes are obtained from the short video platform by using a crawler program, an API interface or a third-party tool, and then the user behavior data is analyzed to obtain user portrait features. Taking the type of medical short videos watched by users as an example, if users often browse videos about cancer treatment, it can be inferred that users may be more concerned about cancer-related topics. Among them, short video data of different modes specifically refers to short video data that contains text mode, visual mode and audio mode at the same time, and user portrait features include one or more of the following features, such as: user operation features, user preference features, age, gender, occupation, region, whether it is a member, activity, video first and second level classification preferences, preferred time period, preferred day of the week or winter and summer vacations, video title, time of occurrence, scene or a combination of the above features. Among them, user operation characteristics can be related behaviors such as search behavior, playback channel and viewing time; the time of occurrence can be weekends, weekdays, winter vacation, summer vacation, or school or work time; user preference characteristics are the interests and preferences of the target user when watching videos, and the user preference characteristics can be inferred from the text read by the target user, the pictures watched, and the audio listened to, etc., and this application is not limited to this. For example, if the medical knowledge sharing platform brings together a large number of various medical science short videos produced by professional doctors and health experts to help users obtain accurate and practical health knowledge, users are accustomed to logging into the platform between 8 and 10 o'clock in the evening, and are more inclined to find short videos under the "sports rehabilitation" section through the classification directory. After watching the video, they will actively participate in the interactive question and answer session of the platform, ask questions or share their own experiences on the content such as rehabilitation training methods after sports injuries, or the short videos about "Chinese medicine therapy" that users like and comment the most are, such as acupuncture, massage principle and efficacy explanation videos, etc., and they will also actively find relevant offline Chinese medicine therapy institution recommendation information on the platform, and obviously have a strong personal preference for the field of Chinese medicine therapy.

[0024] It should be noted that the short video data can be a lecture video of a medical expert or a related popular science video, or it can be PPT data of a medical expert's related courses, etc. It can be set according to actual conditions. This application does not limit the specific content and specific form of expression of the short video data.

[0025] In one embodiment of the invention, obtaining user portrait features includes: A convolutional neural network is used to extract features from the covers of short videos clicked by each user, and a feature vector sequence is obtained according to the click time sequence; The feature vector sequence is converted into a time series graph, and a graph-based recurrent neural network is used for modeling to obtain user portrait features.

[0026] Specifically, a convolutional neural network is used to extract features from the cover (user behavior data) of each short video clicked by each user to obtain a feature vector, and a feature vector sequence is obtained according to the click time sequence, and the feature vector sequence is converted into a time sequence diagram, wherein the time sequence diagram is obtained by connecting the short video corresponding to each feature vector in the feature vector sequence with the short video corresponding to the previous feature vector of the sequence, and at the same time, finding the short video most similar to it in all previous short videos to connect, that is, for each short video, it is connected to the previous short video of the sequence, and at the same time, a short video most similar to it is found in all previous short videos and connected to it, thereby obtaining a time sequence diagram, and for this time sequence diagram, a graph-based recurrent neural network modeling is used to obtain user portrait features. Through the above steps, from the perspective of cover image features, the order in which users browse short videos and the similar associations between short videos are taken into account, and the behavioral information naturally generated by users on the platform is utilized to the greatest extent, so that the constructed user portrait features can be more in line with the user's real interests and browsing habits, so as to accurately recommend short videos that meet his interests and habits to the user in the future.

[0027] In the embodiment of the present application, by obtaining user portrait features and short video data of different modalities, it is possible to better understand the needs of users, so as to subsequently provide personalized short video recommendations and increase user viewing interest and stickiness.

[0028] S202: Based on the user portrait features and the short video data of different modalities, a user-short video interaction graph is constructed respectively using a preset construction method.

[0029] In step S202, after obtaining the user portrait features and short video data of different modes, a node is created for each user and a node is created for each short video according to the user portrait features and the short video data of different modes. For example: user nodes: U1, U2, ...; short video nodes: V1, V2, .... Then define the interactive relationship between the user and the short video, and then use machine learning or deep learning models (such as neural networks in the figure) to construct a more complex user-short video interaction graph under different modes according to the interactive relationship between the user and the short video. In this application, since the short video data contains multiple modes and the number of modes contained in each short video data is different, different short video data can be interactively processed through the user-short video interaction graph, fully explore the correlation between different short video data, fill in the missing information between each other, and eliminate the semantic gap between different modes, so as to facilitate subsequent data processing and better complete the subsequent short video personalized recommendation process.

[0030] In one embodiment of the invention, based on the user portrait features and the short video data of different modalities, a user-short video interaction graph is constructed respectively using a preset construction method, including: Extracting interactive information from the user portrait features and extracting multimodal information from the short video data of different modalities, and calculating the edge relationships and sampling neighbor nodes of the short videos through the extracted interactive information and multimodal information to obtain a short video graph under each relationship; For the short video graph under each relationship, the modality-similar semantic short video graphs and co-occurrence collaborative short video graphs under different modalities are fused to obtain user-short video interaction graphs under different modalities.

[0031] Specifically, the present invention uses the user interaction information reflected in the user portrait features and the multimodal information of the short video data as input to construct a user-short video interaction graph containing modal similarity semantic information and co-occurrence collaborative information under different modes. That is, the interaction information and multimodal information are extracted through data processing, and then the edge relationship of the short video is calculated through the extracted interaction information and multimodal information, the edge relationship between the user and the short video is calculated, and the neighbor node of each short video is obtained according to the calculated edge relationship, and then the interaction edge between the user and the short video is integrated according to the edge relationship calculation and the neighbor node sampling to form a short video graph for each relationship, and each short video graph only contains a specific type of interaction. For example, when constructing a short video graph with the disease theme as the relationship, all videos about diabetes are used as nodes, and the edge relationship between them is determined by the similarity and sampling calculated above, so that a short video graph with the diabetes theme as the relationship is obtained. Finally, based on modal features (such as text, vision, and audio), the similarity between short videos is calculated and a modal similarity graph is constructed. Based on the user's common viewing history or interaction, a co-occurrence graph of short videos is constructed. Then, the modal similarity semantic short video graphs and co-occurrence collaborative short video graphs under different modalities are fused to obtain user-short video interaction graphs under different modalities. The user-short video interaction graph is realized by the following formula: Among them, u represents the user node, i represents the short video node, and the edge It means that there is a connection interaction relationship between the user short video pair, that is, the user has played the short video, and the side It means that there is no connection interaction relationship between the user's short video pair, that is, the user has not played the short video. In this way, information of different modes can be integrated together to realize the construction of user-short video interaction graph, and information of different modes can be integrated together to more comprehensively reflect the complex interaction relationship between users and short videos in different modes, and provide richer information for applications such as short video recommendation.

[0032] In this embodiment, by constructing user-short video interaction graphs based on user portrait features and short video data of different modalities using a preset construction method, the recommendation performance of the short video platform and user satisfaction can be significantly improved, and the accuracy of recommendations can be greatly improved.

[0033] S203: Input the user-short video interaction graph of each modality into the multimodal graph diffusion model based on the modal perception signal injection mechanism, and calculate the final embedding vector of the user and the final embedding vector of the short video.

[0034] In step S203, the modality-aware signal injection mechanism is a mechanism that can identify information from different modalities (such as text, vision, and audio), and can help the model better understand and integrate information from user-short video interaction graphs of different modalities. By performing feature normalization on each modality data in the user-short video interaction graph, the standardized user-short video interaction graph data of each modality is converted into a tensor form suitable for model input. Taking a common deep learning framework as an example, the data of each modality is organized into a three-dimensional or higher-dimensional tensor, and then a special modality feature extractor is used to extract features from the tensor of each modality. For text modality, an encoder with a Transformer architecture can be used to extract semantic features of the text. In the Transformer encoder, a multi-head attention mechanism is used to capture the semantic associations of different positions in the text. For example, in the text part of the user-short video interaction graph, the semantic association between the keywords in the user comments and the keywords in the short video title and description can be captured; for the visual modality, the feature extraction layer of the convolutional neural network (such as ResNet, etc.) is used to extract image features. These feature extraction layers can capture visual features such as the shape, color, and texture of objects in the video. Taking medical short videos as an example, features such as the shape of surgical instruments and the color of human organs can be extracted; for the audio modality, an audio feature extraction network (such as an audio classification network based on a convolutional neural network) is used to extract audio features such as timbre, rhythm, and emotion. For example, in psychological counseling short videos, features such as the counselor's voice timbre and tone of voice can be extracted.

[0035] Specifically, by designing a modality-aware signal injection (MSI) mechanism, the multimodal graph diffusion model is guided to generate a multi-user-video graph with corresponding modalities. The multimodal graph diffusion model is usually implemented based on graph neural networks and diffusion processes. During the diffusion process, the model simulates the propagation of information in the user-short video interaction graph. For example, the state of the node is updated between the user node and the short video node through a message passing mechanism. In each diffusion iteration, the embedding vectors of the user and short video nodes are updated according to the received information. After multiple diffusion iterations, the embedding vectors of the user node and the short video node will gradually stabilize. At this time, the embedding vector of the user node is the final embedding vector of the user, and the embedding vector of the short video node is the final embedding vector of the short video. These final embedding vectors contain the complex interaction relationship between the user and the short video and information of different modalities, which can be used for subsequent personalized recommendations of short videos.

[0036] In one embodiment of the invention, the user-short video interaction graph of each modality is input into a multimodal graph diffusion model based on a modality perception signal injection mechanism, and the final embedding vector of the user and the final embedding vector of the short video are calculated, including: Interactively process the user-short video interaction graphs of each modality to obtain node embedding vectors of users and short videos of each modality; Noise filtering is performed on the node embedding vectors of the users and short videos of each modality to obtain filtered node embedding vectors of the users and short videos of each modality; The node embedding vectors of the users and short videos of each modality after filtering are weighted summed to obtain the final embedding vector of the user and the final embedding vector of the short video.

[0037] Specifically, in a multimodal graph diffusion model based on a modal perception signal injection mechanism, the user-short video interaction graph of each modality is interactively processed to obtain node embedding vectors of users and short videos of each modality. These node embedding vectors contain the user's preferences in each modality and the characteristics of short videos in each modality. Various noise detection algorithms, such as statistical-based methods, machine learning-based methods, etc., are used to identify the noise in the node embedding vectors of users and short videos of each modality. Then, based on the results of noise detection, the node embedding vectors of users and short videos of each modality are noise filtered, and smaller embedding vectors (noise) are removed or similar vectors are merged through clustering and other methods to obtain the filtered node embedding vectors of users and short videos of each modality. Finally, the weighted summation of the filtered node embedding vectors of users and short videos of each modality is performed to obtain the final embedding vector of the user and the final embedding vector of the short video. In the above way, through the fusion and noise filtering of multimodal data, more accurate and comprehensive user and short video embedding vectors can be obtained, so that user interests and short video features can be captured more comprehensively, so as to subsequently recommend videos that are more in line with their interests and needs to users, thereby improving the accuracy of recommendations.

[0038] In one embodiment of the invention, the multimodal graph diffusion model is trained in the following manner: Obtain training sample data set; Inputting the training sample data set into a preset multimodal graph diffusion model for training to obtain a final embedding vector of the user and a final embedding vector of the short video; Based on a preset contrast loss function, calculating a contrast loss value between a final embedding vector of the user and a final embedding vector of the short video; According to the contrast loss value, the weights of the preset multimodal graphic diffusion model are iteratively updated and optimized through the forward propagation algorithm and the back propagation algorithm until the preset training stop condition is reached, thereby completing the training.

[0039] Specifically, the training sample data set is obtained through various channels, such as public data sets, web crawlers, professional data annotation service providers, etc. The training sample data set should contain interaction information between users and short videos, which contains user portrait features, as well as multiple modal features of users and short videos (such as text, vision, audio, etc.). The preset multimodal graph diffusion model includes multiple modules, such as a modal perception signal injection module, a graph convolutional network module, etc. The user portrait features and multimodal features of short videos in the training sample data set are converted into an embedding form suitable for model input, and then according to the importance of different modes and the specific needs of users, weights are assigned to the embedding vectors of each mode, and weighted summation is performed to obtain the final embedding vector of the user and the final embedding vector of the short video, and then the InfoNCE loss function is used to calculate the contrast loss value between the final embedding vector of the user and the final embedding vector of the short video, wherein the purpose of the InfoNCE loss function is to make the similarity of the positive sample pair (the short video that the user prefers) higher than the similarity of the negative sample pair (the short video that the user does not prefer), and then according to the contrast loss value, the weights of the preset multimodal graph diffusion model are iteratively updated and optimized through the forward propagation algorithm and the back propagation algorithm until the preset training stop condition is reached to complete the training. This application uses the contrast loss function and the iterative update optimization algorithm to more effectively train the multimodal graph diffusion model, thereby improving the model training efficiency and model performance, and through the training of the multimodal graph diffusion model, more accurate and comprehensive user and short video embedding vectors can be obtained, which can better reflect the user's preferences and the characteristics of the short video, and provide valuable reference information for the recommendation system.

[0040] In this embodiment, by introducing a modal perception signal injection mechanism, the multimodal graph diffusion model can more accurately capture user-short video interaction information under different modalities, so that the final embedding vector of the generated user and the final embedding vector of the short video can provide the recommendation system with more accurate user portraits and content features, so as to more accurately recommend suitable short video content according to the user's interests and needs in the future, thereby improving the accuracy and diversity of recommendations.

[0041] S204: Determine the user's preference weight for the short video according to the final embedding vector of the user and the final embedding vector of the short video.

[0042] In step S204, the preference weight is used as a quantitative indicator to measure the user's preference for short videos, that is, the larger the preference weight, the higher the corresponding short video preference; for example, the preference weight of short video A is greater than that of short video B, which can determine that short video A is more preferred by the user. This application uses the entropy weight method to construct a standard decision matrix for the user's final embedding vector and the short video's final embedding vector according to the user's final embedding vector and the short video's final embedding vector. Based on the standard decision matrix, the contribution of the user's final embedding vector and the short video's final embedding vector is obtained, and then the information entropy corresponding to the user's final embedding vector and the short video's final embedding vector is calculated according to the contribution of the user's final embedding vector and the short video's final embedding vector. According to the discrete degree of the user's final embedding vector and the short video's final embedding vector reflected by each information entropy, the user's preference weight for the short video is determined. It can be seen that this application determines the user's preference weight for short videos so that the short videos can be personalized recommended according to the user's preference weight for short videos in the future, thereby improving the overall recommendation effect and improving user satisfaction and retention rate.

[0043] It should be noted that the preset algorithm can not only be the entropy method, but also the hierarchical analysis method, the principal component analysis method, the CRITIC weight method, etc., which can be selected according to the application scenario and needs, and this application does not make any limitation on this.

[0044] In one embodiment of the invention, before determining the user's preference weight for the short video according to the final embedding vector of the user and the final embedding vector of the short video, the method includes: Calculating the similarity between the final embedding vector of the user and the final embedding vector of the short video; Determining whether the similarity is greater than a preset similarity threshold; If the similarity is greater than a preset similarity threshold, the step of determining the user's preference weight for the short video based on the user's final embedding vector and the short video's final embedding vector is performed.

[0045] Specifically, by using similarity calculation methods such as cosine similarity and Euclidean distance, the similarity between the final embedding vector of the user and the final embedding vector of the short video is calculated. By setting a preset similarity threshold, it is used to determine whether the similarity between the user and the short video is high enough to indicate the user's preference for the short video, and then determine whether the similarity is greater than the preset similarity threshold. If the similarity is greater than the preset similarity threshold, the user's preference weight for the short video is determined based on the final embedding vector of the user and the final embedding vector of the short video. If the similarity is not greater than the preset similarity threshold, the steps of the short video recommendation method are re-executed. For example, if the preset similarity threshold is 0.5, the calculated cosine similarity is 0.8, and 0.8 is greater than 0.5, the user's preference weight for the short video is determined. It can be seen that by calculating the similarity between the user and the short video, and determining the preference weight according to the similarity, personalized recommendation services can be implemented in the future to improve the accuracy of the recommendation and user satisfaction.

[0046] It should be noted that the preset similarity threshold can be set according to actual conditions, and this application does not impose any limitation on this.

[0047] S205: Selecting a predetermined number of short videos with higher preference weights according to the user's preference weights for short videos, to obtain a short video recommendation result.

[0048] In one embodiment of the invention, obtaining a short video recommendation result includes: According to the user's preference weight for short videos, each short video in the short video set to be recommended is sorted from high to low to obtain a sorted recommended video set; A predetermined number of short videos with high preference weights are selected from the sorted recommended video set for recommendation to obtain a short video recommendation result.

[0049] Specifically, according to the user's preference weight for short videos, each short video in the short video set to be recommended is sorted from high to low to obtain a sorted recommended video set, and then a predetermined number of short videos with high preference weights are selected from the sorted recommended video set for recommendation, that is, a predetermined number of short videos with the highest priority are recommended first, thereby obtaining a short video recommendation result. For example, suppose we have a short video set to be recommended, which contains the following 5 short videos and their corresponding preference weights: short video A: weight 0.85; short video B: weight 0.70; short video C: weight 0.90; short video D: weight 0.60; short video E: weight 0.80. Now, we set the number of recommendations to 3. After sorting in order of preference weight from high to low, the short video set becomes: short video C (weight 0.90); short video A (weight 0.85); short video E (weight 0.80); short video B (weight 0.70); Short video D (weight 0.60), select the top 3 short videos (C, A, E) from the sorted set as the recommendation results. By using the preference weights of users for short videos for sorting and recommendation, we can achieve personalized recommendation services, meet the unique needs of different users, and thus improve the accuracy of short video personalized recommendations.

[0050] Exemplarily, if the user is an elderly person who is concerned about health preservation and disease prevention, and has browsed a large number of short videos on the prevention and treatment methods of chronic diseases such as hypertension and diabetes on the platform, and often likes and comments on these videos. By analyzing the user portrait characteristics and different modalities of medical and health short video data, the platform constructs a user portrait that includes the types of diseases he is concerned about, the preferred content forms (such as animation demonstrations, expert explanations, etc.), and interaction habits. It is determined that the user is particularly interested in short videos on hypertension prevention and has a high weight. After sorting according to the preference weights, the top 10 short videos on hypertension prevention are selected as the recommendation results, and these short videos are displayed to the user in a list form, and information such as the viewing duration, number of likes, and number of comments of each video is marked, so that he can quickly understand the content and popularity of the videos.

[0051] It should be noted that the predetermined quantity can be specifically set according to the actual situation, and this application does not make any limitations on this.

[0052] In this embodiment, after obtaining the preference weights of users for short videos, then select a predetermined number of short videos with higher preference weights according to the ranking of the preference weights of users for short videos, so as to quickly and accurately obtain the short video recommendation results. By sorting the videos to be recommended according to the priority, the videos that users prefer more can be placed in the front, which is convenient for users to watch according to their preferences, thereby improving the accuracy of the recommendation and user satisfaction, and bringing a better experience to users.

[0053] In summary, the present invention provides a short video recommendation method, device, equipment and medium, which obtains user portrait features and short video data of different modalities, wherein the user portrait features include user operation features and user preference features, and the different modalities include text modality, visual modality and audio modality. According to the user portrait features and the short video data of the different modalities, a preset construction method is used to respectively construct a user-short video interaction graph, and the user-short video interaction graph of each modality is input into a multimodal graph diffusion model based on a modal perception signal injection mechanism, and the final embedding vector of the user and the final embedding vector of the short video are calculated. According to the final embedding vector of the user and the final embedding vector of the short video, the user's preference weight for the short video is determined, and then a predetermined number of short videos with a high preference weight are selected according to the user's preference weight for the short video to obtain a short video recommendation result, thereby effectively improving the accuracy and recommendation speed in personalized short video recommendation and enhancing the user experience.

[0054] See also Figure 3 , Figure 3 is a schematic diagram of the structure of the short video recommendation device provided by an embodiment of the present invention. The short video recommendation device corresponds to the short video recommendation method in the above embodiment. Figure 2 as well as Figure 2 For the convenience of explanation, only the parts related to this embodiment are shown. Figure 3 The short video recommendation device 30 includes: an acquisition module 31, a construction module 32, a calculation module 33, a determination module 34, and a recommendation module 35.

[0055] An acquisition module 31 is used to acquire user portrait features and short video data of different modes, wherein the user portrait features include user operation features and user preference features, and the different modes include text mode, visual mode and audio mode; A construction module 32 is used to construct user-short video interaction graphs respectively according to the user portrait features and the short video data of different modalities using a preset construction method; A calculation module 33 is used to input the user-short video interaction graph of each modality into the multimodal graph diffusion model based on the modality perception signal injection mechanism, and calculate the final embedding vector of the user and the final embedding vector of the short video; A determination module 34, configured to determine the user's preference weight for the short video according to the final embedding vector of the user and the final embedding vector of the short video; The recommendation module 35 is used to select a predetermined number of short videos with higher preference weights according to the user's preference weights for short videos, so as to obtain short video recommendation results.

[0056] Optionally, the acquisition module 31 is specifically used for: A convolutional neural network is used to extract features from the covers of short videos clicked by each user, and a feature vector sequence is obtained according to the click time sequence; The feature vector sequence is converted into a time series graph, and a graph-based recurrent neural network is used for modeling to obtain user portrait features.

[0057] Optionally, the building module 32 is specifically used for: Extracting interactive information from the user portrait features and extracting multimodal information from the short video data of different modalities, and calculating the edge relationships and sampling neighbor nodes of the short videos through the extracted interactive information and multimodal information to obtain a short video graph under each relationship; For the short video graph under each relationship, the modality-similar semantic short video graphs and co-occurrence collaborative short video graphs under different modalities are fused to obtain user-short video interaction graphs under different modalities.

[0058] Optionally, the calculation module 33 is specifically used for: Interactively process the user-short video interaction graphs of each modality to obtain node embedding vectors of users and short videos of each modality; Noise filtering is performed on the node embedding vectors of the users and short videos of each modality to obtain filtered node embedding vectors of the users and short videos of each modality; The node embedding vectors of the users and short videos of each modality after filtering are weighted summed to obtain the final embedding vector of the user and the final embedding vector of the short video.

[0059] Optionally, the calculation module 33 is further used for: Obtain training sample data set; Inputting the training sample data set into a preset multimodal graph diffusion model for training to obtain a final embedding vector of the user and a final embedding vector of the short video; Based on a preset contrast loss function, calculating a contrast loss value between a final embedding vector of the user and a final embedding vector of the short video; According to the contrast loss value, the weights of the preset multimodal graphic diffusion model are iteratively updated and optimized through the forward propagation algorithm and the back propagation algorithm until the preset training stop condition is reached, thereby completing the training.

[0060] Optionally, the determination module 34 is previously specifically used for: Calculating the similarity between the final embedding vector of the user and the final embedding vector of the short video; Determining whether the similarity is greater than a preset similarity threshold; If the similarity is greater than a preset similarity threshold, the step of determining the user's preference weight for the short video based on the user's final embedding vector and the short video's final embedding vector is performed.

[0061] Optionally, the recommendation module 35 is specifically used for: According to the user's preference weight for short videos, each short video in the short video set to be recommended is sorted from high to low to obtain a sorted recommended video set; A predetermined number of short videos with high preference weights are selected from the sorted recommended video set for recommendation to obtain a short video recommendation result.

[0062] It should be noted that the information interaction, execution process and other contents between the above-mentioned units are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0063] Figure 4 Schematic diagram of the structure of a computer device provided by an embodiment of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps of any of the above-mentioned short video recommendation method embodiments are implemented.

[0064] The computer device may include, but is not limited to, a processor and a memory. It can be understood by those skilled in the art that Figure 4 These are merely examples of computer devices and do not constitute limitations on the computer devices. The computer devices may include more or fewer components than those shown in the figure, or a combination of certain components, or different components. For example, they may also include a network interface, a display screen, and an input system.

[0065] In one embodiment, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor in a computer device, the computer device can perform the steps of any embodiment of a short video recommendation method disclosed in the present invention, which will not be repeated here. The computer-readable storage medium can be non-volatile or volatile.

[0066] The processor may be a CPU, or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0067] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory may be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be a hard disk of a computer device, and in other embodiments, it may also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Further, the memory may also include both an internal storage unit of the computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program, etc. The memory may also be used to temporarily store data that has been output or is to be output.

[0068] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0069] The technical business in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0070] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A short video recommendation method, characterized in that: include: Acquire user portrait features and short video data of different modalities, wherein the user portrait features include user operation features and user preference features, and the different modalities include text modality, visual modality, and audio modality; According to the user portrait features and the short video data of different modalities, a user-short video interaction graph is constructed respectively using a preset construction method; The user-short video interaction graph of each modality is input into the multimodal graph diffusion model based on the modality perception signal injection mechanism, and the final embedding vector of the user and the final embedding vector of the short video are calculated; Determining the user's preference weight for the short video according to the final embedding vector of the user and the final embedding vector of the short video; According to the ranking of the user's preference weights for short videos, a predetermined number of short videos with high preference weights are selected to obtain short video recommendation results.

2. The short video recommendation method according to claim 1, characterized in that: The method of constructing user-short video interaction graphs respectively according to the user portrait features and the short video data of different modes using a preset construction method includes: Extracting interactive information from the user portrait features and extracting multimodal information from the short video data of different modalities, and calculating the edge relationships and sampling neighbor nodes of the short videos through the extracted interactive information and multimodal information to obtain a short video graph under each relationship; For the short video graph under each relationship, the modality-similar semantic short video graphs and co-occurrence collaborative short video graphs under different modalities are fused to obtain user-short video interaction graphs under different modalities.

3. The short video recommendation method according to claim 1, characterized in that: The method of inputting the user-short video interaction graph of each modality into the multimodal graph diffusion model based on the modality perception signal injection mechanism, and calculating the final embedding vector of the user and the final embedding vector of the short video, includes: Interactively process the user-short video interaction graphs of each modality to obtain node embedding vectors of users and short videos of each modality; Noise filtering is performed on the node embedding vectors of the users and short videos of each modality to obtain filtered node embedding vectors of the users and short videos of each modality; The node embedding vectors of the users and short videos of each modality after filtering are weighted summed to obtain the final embedding vector of the user and the final embedding vector of the short video.

4. The short video recommendation method according to claim 1, characterized in that: Before determining the user's preference weight for the short video according to the final embedding vector of the user and the final embedding vector of the short video, the method includes: Calculating the similarity between the final embedding vector of the user and the final embedding vector of the short video; Determining whether the similarity is greater than a preset similarity threshold; If the similarity is greater than a preset similarity threshold, the step of determining the user's preference weight for the short video based on the user's final embedding vector and the short video's final embedding vector is performed.

5. The short video recommendation method according to claim 1, characterized in that: The multimodal graph diffusion model is trained in the following way: Obtain training sample data set; Inputting the training sample data set into a preset multimodal graph diffusion model for training to obtain a final embedding vector of the user and a final embedding vector of the short video; Based on a preset contrast loss function, calculating a contrast loss value between a final embedding vector of the user and a final embedding vector of the short video; According to the contrast loss value, the weights of the preset multimodal graphic diffusion model are iteratively updated and optimized through the forward propagation algorithm and the back propagation algorithm until the preset training stop condition is reached, and the training is completed.

6. The short video recommendation method according to claim 1, characterized in that: The obtaining of user portrait features includes: A convolutional neural network is used to extract features from the covers of short videos clicked by each user, and a feature vector sequence is obtained according to the click time sequence; The feature vector sequence is converted into a time series graph, and a graph-based recurrent neural network is used for modeling to obtain user portrait features.

7. The short video recommendation method according to claim 1, characterized in that: The selecting a predetermined number of short videos with higher preference weights according to the user's preference weights for the short videos to obtain a short video recommendation result includes: According to the user's preference weight for short videos, each short video in the short video set to be recommended is sorted from high to low to obtain a sorted recommended video set; A predetermined number of short videos with high preference weights are selected from the sorted recommended video set for recommendation to obtain a short video recommendation result.

8. A short video recommendation device, characterized in that: include: An acquisition module, used to acquire user portrait features and short video data of different modalities, wherein the user portrait features include user operation features and user preference features, and the different modalities include text modality, visual modality and audio modality; A construction module, used to construct user-short video interaction graphs respectively according to the user portrait features and the short video data of different modalities using a preset construction method; A calculation module is used to input the user-short video interaction graph of each modality into a multimodal graph diffusion model based on a modality perception signal injection mechanism, and calculate the final embedding vector of the user and the final embedding vector of the short video; A determination module, used to determine the user's preference weight for the short video according to the final embedding vector of the user and the final embedding vector of the short video; The recommendation module is used to select a predetermined number of short videos with high preference weights according to the user's preference weights for short videos, so as to obtain short video recommendation results.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the short video recommendation method as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the short video recommendation method as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Quantum evolution guidance-based short video recommendation method suitable for aging

    CN120429467A

  • An aging-friendly short video recommendation method based on quantum evolution guidance

    CN120429467B

  • Short video recommendation method and system based on graph contrast learning

    CN121210713A