Content Recommendation Method and System Based on Semantic Recognition

Through semantic recognition technology and multi-layer graph convolution networks, the user interest map is dynamically constructed and combined with the Transformer model and multimodal information fusion, the shortcomings of the existing recommendation methods in data sparseness and cold start problems are solved, and efficient, personalized and diversified content recommendations are achieved.

CN119089398BActive Publication Date: 2025-06-27HANGZHOU SCENE INNOVATION CENTER CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411580772.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-06-27
Estimated Expiration
2044-11-07

AI Technical Summary

Technical Problem

The existing recommendation methods based on graph collaborative filtering and knowledge graph attention network are insufficiently solved in data sparseness and cold start problems, and low training efficiency, resulting in slow response speed of recommendation systems, prone to homogeneity, lack of diversity and personalization, and poor user experience.

Method used

Through semantic recognition technology, analyzing user needs, dynamically building user interest maps, using multi-layer graph convolution network and Transformer model to perform node and edge representation learning, combining multimodal information fusion and adaptive learning algorithms to ensure the diversity, personalization and real-timeness of recommended content.

Benefits of technology

It improves the overall effect of the recommendation system, enhances the relevance and accuracy of the recommended content, solves the problems of sparse data and cold start, and improves the response speed and user experience of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119089398B_ABST
    Figure CN119089398B_ABST
Patent Text Reader

Abstract

The present invention discloses a content recommendation method and system based on semantic recognition. By using natural language processing technology to analyze the user's text, voice or video input, user semantic representation data is generated; combined with the user's historical behavior data, a user interest graph is dynamically constructed, and node and edge representation learning is performed through a multi-layer graph convolutional network to generate time-series user interest graph representation data; by using the self-attention mechanism and multi-head attention mechanism of the Transformer model, combined with a contrastive learning module to optimize the matching degree between the user interest representation and the recommended content, personalized recommendation content data is generated; multi-modal information of the user is collected and fused with the personalized recommendation content data, and optimized multi-modal recommendation content data is generated through a multi-modal fusion algorithm; user feedback feature data is collected, and the generation strategy of the recommended content is dynamically adjusted and optimized through an adaptive learning algorithm. The present invention realizes efficient, accurate and personalized recommendation of the recommendation system, and improves the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information recommendation, and particularly to a content recommendation method and system based on semantic recognition. Background Art

[0002] In modern information society, users are faced with a vast amount of information. How to effectively recommend content that users are interested in to them has become an important research topic. Semantic recognition technology can understand users' real needs by parsing users' text, speech, and video inputs, thereby improving the relevance and accuracy of recommended content.

[0003] Current content recommendation technologies mainly rely on traditional methods such as collaborative filtering and content filtering. However, these methods have some inherent problems. Although the collaborative filtering method can capture the relationship between users and content, it performs poorly in dealing with data sparsity and cold start problems. The content filtering method depends on the explicit features of content and is difficult to capture users' implicit interests. In addition, existing technologies also have obvious deficiencies in dealing with multi-modal information fusion and real-time feedback optimization.

[0004] The prior art (Chinese invention patent, publication number: CN116340648A, title: A Recommendation Method of Knowledge Graph Attention Network Based on Graph Collaborative Filtering) mainly adopts a recommendation method of a knowledge graph attention network based on graph collaborative filtering. This method includes constructing a user-item bipartite graph and using a collaborative filtering layer and a knowledge graph attention embedding layer for recommendation. However, this method has the following defects:

[0005] Existing recommendation methods based on graph collaborative filtering and knowledge graph attention network do not solve the data sparsity and cold start problems sufficiently; when the user or item data is scarce, the accuracy of the recommendation system drops significantly;

[0006] Existing graph convolutional networks have low training efficiency and slow convergence speed when dealing with large-scale graph data, resulting in a slow response speed of the recommendation system;

[0007] When existing recommendation systems generate recommended content, it is easy to show homogenization, lack of diversity and personalization, and the user experience is poor;

[0008] Combining the embeddings of the collaborative filtering layer and the knowledge graph attention layer, although it increases the amount of information, it also increases the complexity of the model and the difficulty of optimization. Summary of the Invention

[0009] In view of the many problems existing in the above-mentioned prior art, the present invention provides a content recommendation method and system based on semantic recognition. The present invention analyzes user needs through semantic recognition technology, dynamically constructs a user interest graph, and uses a multi-layer graph convolutional network and a Transformer model for node and edge representation learning and recommended content generation. By introducing multi-modal information fusion and an adaptive learning algorithm, the diversity, personalization, and real-time nature of the recommended content are ensured, thereby improving the overall effect of the recommendation system.

[0010] A content recommendation method based on semantic recognition includes the following steps:

[0011] Parse the needs input by the user through text, voice, or video through natural language processing technology to generate user semantic representation data;

[0012] Based on the user semantic representation data and the user's historical behavior data, dynamically construct a user interest graph, and perform node and edge representation learning through a multi-layer graph convolutional network to generate time-series user interest graph representation data;

[0013] Utilize the time-series user interest graph representation data, through the self-attention mechanism and the multi-head attention mechanism of the Transformer model, and combine with a contrast learning module to optimize the matching degree between the user interest representation and the recommended content, and generate personalized recommended content data;

[0014] Collect the multi-modal information of the user, fuse it with the personalized recommended content data, use a multi-modal fusion algorithm to generate multi-modal recommendation data, and generate optimized multi-modal recommended content data through an optimization algorithm;

[0015] Realtime display the optimized multi-modal recommended content data to the user, collect user feedback feature data, and dynamically adjust and optimize the generation strategy of the recommended content through an adaptive learning algorithm to ensure the continuous accuracy and personalization of the recommended content.

[0016] Preferably, the parsing of the needs input by the user through text, voice, or video through natural language processing technology includes:

[0017] Perform speech recognition on the speech data and convert it into text data;

[0018] Perform speech and image analysis on the video data, extract speech and image features, and generate video semantic data;

[0019] Perform word segmentation, part-of-speech tagging, and semantic parsing on the text data to generate user semantic representation data.

[0020] Preferably, the dynamic construction of the user interest graph includes:

[0021] Extract the user's points of interest and generate point-of-interest data; combine the user's historical behavior data to construct a user interest graph and generate user interest graph data;

[0022] Apply a multi-layer graph convolutional network to the user interest graph for node and edge representation learning to generate user interest graph representation data.

[0023] Preferably, the node and edge representation learning through the multi-layer graph convolutional network includes:

[0024] Initialize the feature vectors of each node in the user interest graph;

[0025] Apply multi-layer graph convolutional operations to aggregate the information of neighboring nodes layer by layer and update the feature vectors of each node. The convolutional operation of each layer is expressed as:

[0026]

[0027] Among them, represents the feature vector of node at the th layer; is the set of neighboring nodes of node ; is the normalization coefficient; is the weight matrix of the th layer; is the activation function; u represents the neighboring node of node v, that is, the node directly connected to node v; represents the feature vector of node u at the (k - 1)th layer, the feature obtained through the calculation of the previous layer of the graph convolutional network;

[0028] Introduce an attention mechanism to dynamically adjust the weights of the edges in the user interest graph and generate adjusted user interest graph representation data. The edge weight calculation formula is:

[0029]

[0030] Among them, is the weight of edge ; and respectively represent the feature vectors of nodes and in the graph; is the attention weight vector; represents the transpose operation, that is, represents the transpose of vector ; is the learnable parameter matrix for linearly transforming the node feature vectors; ‖ represents the concatenation operation, concatenating the node feature vectors after being transformed by the weight matrix ; is an activation function used to enhance the non - linear characteristics of the model;

[0031] Utilize a temporal convolutional network to capture the time - series information of user interests and generate time - series feature data;

[0032] Fuse the time - series feature data with the adjusted user interest graph representation data to generate time - series user interest graph representation data.

[0033] Preferably, the self - attention mechanism and multi - head attention mechanism of the Transformer model include:

[0034] Convert the time - series user interest graph representation data into the input format of the Transformer model;

[0035] Apply the self - attention mechanism to calculate the importance of each node, and weight the encoded user interest data. The calculation formula of the self - attention mechanism is:

[0036]

[0037] Among them, represents the query vector; represents the key vector; represents the value vector; is the vector dimension; represents the transpose operation; Softmax represents the normalization function, which converts the input into a probability distribution to ensure that the sum of the output weights is 1;

[0038] Utilize the multi - head attention mechanism to capture information at different levels and dimensions in the user interest graph. The calculation formula of the multi - head attention mechanism is:

[0039]

[0040] where each attention head is calculated as:

[0041]

[0042] Among them, is the weight matrix of the query vector; is the weight matrix of the key vector; is the weight matrix of the value vector; is the output weight matrix to generate comprehensive user interest data; Concat represents the concatenation operation, which combines multiple vectors or matrices by dimension; Attention represents the self - attention mechanism, which is used to calculate the similarity between the query vector and the key vector to generate a weighted value vector.

[0043] Preferably, the optimization of the matching degree between the user interest representation and the recommended content by the combined contrast learning module includes:

[0044] Construct positive and negative sample pairs to generate positive and negative sample data;

[0045] Calculate the distance between the positive and negative sample data to generate sample distance data;

[0046] Calculate the loss value according to the loss function of contrast learning, and the loss function is defined as:

[0047]

[0048] where, is the distance between samples and ; is a binary label; is the margin distance for generating loss data;

[0049] Utilize the optimization result of the contrast learning module to adjust the comprehensive user interest data and optimize the matching degree between it and the recommended content;

[0050] Generate personalized recommended content data according to the optimized comprehensive user interest data.

[0051] Preferably, the collection of the user's multimodal information and its fusion with the personalized recommended content data includes:

[0052] Collect the user's geographical location, device usage, and social network data as multimodal information;

[0053] Preprocess the multimodal information to generate standardized multimodal feature data;

[0054] Utilize a multimodal fusion algorithm to fuse the standardized multimodal feature data and the personalized recommended content data to generate multimodal recommendation data;

[0055] Utilize an optimization algorithm to optimize the multimodal recommendation data to ensure the relevance and personalization of the recommended content, and generate optimized multimodal recommended content data, where the optimization algorithm includes: gradient descent method, genetic algorithm, and simulated annealing algorithm.

[0056] Preferably, the real-time display of the optimized multimodal recommended content data to the user and the collection of user feedback feature data includes:

[0057] Real-time display the optimized multimodal recommended content data;

[0058] Collect the user's interaction data with the recommended content, including click-through rate, dwell time, number of likes, and comment content;

[0059] Preprocess the user interaction data, extract key metrics and features, and generate user feedback feature data.

[0060] Preferably, the method of dynamically adjusting and optimizing the generation strategy of recommended content through the adaptive learning algorithm includes:

[0061] According to the user feedback feature data, use the adaptive learning algorithm to dynamically adjust and optimize the generation strategy of recommended content to generate optimized recommended content data;

[0062] Use the optimized recommended content data to update the parameters and strategies of the recommended content generation strategy, including the adjustment of model weights, learning rates, and loss functions, to generate updated multi-modal recommended content data;

[0063] Apply the updated multi-modal recommended content data to the real-time recommendation generation module, dynamically adjust the recommendation strategy, and generate optimized real-time recommendation data for presenting and providing recommendation services to users.

[0064] And a system for implementing the content recommendation method based on semantic recognition. Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:

[0065] The present invention analyzes user needs through semantic recognition technology and realizes more accurate user semantic representation data;

[0066] The present invention constructs a dynamic user interest graph through a multi-layer graph convolutional network and realizes efficient learning of user interest representation;

[0067] The present invention combines a contrast learning module to optimize the matching degree between user interest representation and recommended content, generates personalized recommended content data, and improves the relevance of recommended content;

[0068] The present invention collects and fuses user multi-modal information, uses a multi-modal fusion algorithm to generate optimized multi-modal recommended content data, and ensures the diversity and accuracy of recommended content;

[0069] The present invention dynamically adjusts and optimizes the generation strategy of recommended content through an adaptive learning algorithm, realizes continuous optimization of the recommendation model, and ensures the timeliness and personalization of recommended content. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is a flow chart of the method of the present invention;

[0071] Figure 2 It is a flow chart of user input parsing in the present invention;

[0072] Figure 3 It is a flow chart of user interest graph construction in the present invention;

[0073] Figure 4 This is the optimization flowchart of the Transformer model in the present invention;

[0074] Figure 5 This is the multi-modal information fusion flowchart in the present invention;

[0075] Figure 6 This is the real-time feedback and adaptive learning flowchart in the present invention;

[0076] Figure 7 This is the structural block diagram of the system of the present invention. Detailed implementation manners

[0077] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0078] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0079] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0080] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but is not limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0081] Some block diagrams and / or flowcharts are shown in the accompanying drawings. It should be understood that some blocks or combinations of blocks in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when executed by the processor, these instructions can create a device for implementing the functions / operations illustrated in these block diagrams and / or flowcharts. The technology of the present disclosure can be implemented in the form of hardware and / or software (including firmware, microcode, etc.). Additionally, the technology of the present disclosure can take the form of a computer program product on a computer-readable storage medium storing instructions, which can be used by or in conjunction with an instruction execution system.

[0082] As Figure 1 shown, a content recommendation method based on semantic recognition includes the following steps:

[0083] As Figure 2 shown, parse the requirements input by the user through text, voice, or video through natural language processing technology to generate user semantic representation data;

[0084] Preferably, the parsing of the requirements input by the user through text, voice, or video through natural language processing technology includes:

[0085] Perform speech recognition on the speech data and convert it into text data;

[0086] Perform speech and image analysis on the video data, extract speech and image features, and generate video semantic data;

[0087] Perform word segmentation, part-of-speech tagging, and semantic parsing on the text data to generate user semantic representation data.

[0088] In the present invention, parsing the requirements input by the user through text, voice, or video is a key step, and its principle and effect are described as follows. This method uses natural language processing (NLP) technology to process different forms of data (text, voice, video) input by the user to generate user semantic representation data. This involves the integration and optimization of multiple technologies, aiming to extract semantic information from multi-modal data and then achieve accurate content recommendation.

[0089] For speech data, the system first performs speech recognition to convert the speech content into processable text data. This step utilizes speech recognition technology, which converts speech signals into text representations through acoustic models and language models. For example, it converts the requirements dictated by the user into text data that can be understood by a computer. The processing of video data is more complex. The system needs to analyze both the speech content and image features in the video simultaneously. The speech content analysis uses speech recognition technology to extract the speech information in the video, while the image feature analysis extracts the visual information in the video through computer vision technologies such as object recognition and action detection. These analysis results are combined to generate the semantic data of the video, that is, the semantic representation of the video content.

[0090] For text data, the system performs word segmentation, part-of-speech tagging, and semantic parsing. Word segmentation divides the text into meaningful word units; part-of-speech tagging determines the part of speech of each word; and semantic parsing analyzes the semantic relationships between words to generate semantic representation data of the text. This process relies on natural language processing technologies, including word embedding models, part-of-speech taggers, and semantic analyzers, to effectively capture the semantic information of the text.

[0091] The user semantic representation data generated through the above steps has the following effects and advantages: It can process various input forms of users (text, speech, video), thus achieving unified understanding and processing of multimodal data; it improves the understanding accuracy and precision of the system for user input requirements, avoiding information loss or misunderstanding under a single data form; based on accurate semantic representation data, the system can more precisely construct user interest models and behavior patterns to achieve personalized content recommendations; due to its ability to process multiple data forms, the system shows good flexibility and adaptability in the process of real-time recommendation and feedback adjustment. For example, when a user inputs a question through speech, the system first performs speech recognition to convert the speech into text. Subsequently, through word segmentation and semantic parsing, the system can understand and precisely capture the meaning of the question raised by the user. Finally, by combining the user's historical behavior data and other context information, the system generates personalized recommendation content suitable for the user's needs, thereby enhancing the user experience and recommendation accuracy.

[0092] As Figure 3 shown, based on the user semantic representation data and the user's historical behavior data, a user interest graph is dynamically constructed, and node and edge representation learning is performed through a multi-layer graph convolutional network to generate time-series user interest graph representation data;

[0093] First, the system dynamically constructs a user interest graph by combining the user's semantic representation data and historical behavior data. The user semantic representation data comes from the parsing of the user's text, voice, or video input, reflecting the user's current interests and needs. The user historical behavior data records the user's interaction history on the platform, including behaviors such as clicks, views, likes, and comments. These two types of data together form the basis of the user interest graph. By integrating these data, the system can form a dynamic and comprehensive user interest graph.

[0094] Next, through a multi-layer graph convolutional network (GCN), the system performs representation learning on the nodes and edges in the user interest graph. The graph convolutional network is a deep learning model that can directly perform convolutional operations on graph-structured data and is suitable for processing non-Euclidean space data such as the user interest graph. Each layer of the GCN aggregates the information of the nodes and their neighbor nodes, gradually updating the feature representations of the nodes. Specifically, the graph convolution operation captures the relationships between nodes and complex patterns in the graph structure by performing a weighted sum of the feature vectors of each node and its neighbor nodes and then applying an activation function. For example, the interest nodes of user A and its neighbor nodes (such as other users with similar interests or relevant content) are subjected to convolutional operations through multiple layers of GCN, gradually extracting and fusing features, and finally generating a high-dimensional node representation.

[0095] To capture the time-series changes in user interests, the system introduces time-series user interest graph representation data. By adding a time dimension to the GCN, the system can track the changing trajectory of user interests over time. For example, a user may be particularly interested in a certain topic during a certain period and show more attention to other topics during another period. Through time-series analysis, the system can identify and predict the changing trends of user interests, thereby achieving more accurate and timely content recommendations.

[0096] The effect of this process is that by dynamically constructing and learning the user interest graph, the system can capture the multi-dimensional interests and behavior patterns of users and update the user's interest status in a timely manner. Combining the powerful representation learning ability of the GCN and time-series analysis, the system can not only understand the user's current interests but also predict the possible future changes in the user's interests, thereby providing personalized and forward-looking content recommendations. For example, when a user frequently browses a certain type of content during a certain period, the system can identify this change and recommend relevant content in a timely manner, improving the user experience and the accuracy of recommendations.

[0097] In summary, by dynamically constructing a user interest graph based on user semantic representation data and historical behavior data, and using a multi-layer graph convolutional network for node and edge representation learning, the system can generate time-series user interest graph representation data, thereby achieving a deep understanding and accurate prediction of user interests, and ultimately improving the personalization and effectiveness of content recommendation.

[0098] Preferably, the dynamic construction of the user interest graph includes:

[0099] Extracting the user's interest points to generate interest point data; combining the user's historical behavior data to construct a user interest graph and generate user interest graph data;

[0100] Applying a multi-layer graph convolutional network to the user interest graph for node and edge representation learning to generate user interest graph representation data.

[0101] First of all, extracting the user's interest points is the first step in dynamically constructing a user interest graph. Interest points refer to the attention and preferences shown by users for certain topics or content within a specific time period. For example, when a user frequently browses articles, videos, or pictures related to travel, "travel" can be identified as an interest point of this user. The methods for extracting interest points include natural language processing (NLP) analysis of the user's text, speech, or video input, and through techniques such as keyword extraction and topic modeling, identifying the topics that the user is currently interested in and generating interest point data. These interest point data constitute the basic nodes of the user interest graph.

[0102] Next, the system combines the user's historical behavior data to construct a user interest graph. Historical behavior data includes all interaction records of the user on the platform, such as clicks, views, likes, comments, etc. These data can reflect the user's long-term interest preferences and behavior patterns. By combining the interest point data with the historical behavior data, the system can construct a comprehensive user interest graph and generate user interest graph data. In this graph, nodes represent the user's interest points, and edges represent the relationships between these interest points, such as the degree of association between multiple interest points of the same user, or the connection between similar interest points of different users.

[0103] To further enhance the expressive power of the user interest graph, the system applies a multi-layer graph convolutional network (GCN) to the user interest graph for node and edge representation learning. The graph convolutional network is a deep learning model capable of processing graph-structured data. By aggregating and learning the information of nodes and their neighbor nodes in the graph, it realizes an efficient representation of node features. In the present invention, the GCN aggregates the information of nodes and their neighbor nodes in the user interest graph layer by layer through multi-layer convolutional operations, continuously updating the feature representation of the nodes. For example, in the first-layer convolutional operation, the feature vector of a node is weighted and summed with the feature vectors of its direct neighbors, and then a new node feature is generated through an activation function; in the second-layer convolutional operation, the new node features and the information of their neighbor nodes are aggregated again, and so on until the multi-layer convolutional operations are completed. This process not only captures the local relationships between nodes but also learns the deep patterns in the graph structure from a global perspective, finally generating the user interest graph representation data.

[0104] Through this process of dynamically constructing the user interest graph, the system can achieve a comprehensive, dynamic, and accurate modeling of user interests. The user interest graph not only reflects the user's current interest points but also combines the user's historical behavior data, providing insights into the changes and development trends of user interests. At the same time, through the representation learning of the multi-layer graph convolutional network, the system can extract high-dimensional features from the complex graph structure, realizing a deep understanding and representation of user interests. The effect of this process is that through efficient user interest modeling, the system can more accurately predict the user's content needs, thereby providing more personalized and accurate content recommendations.

[0105] For example, if user A has frequently browsed technology-related articles in the past period and recently started to pay attention to content in the field of artificial intelligence, the system uses NLP technology to identify "technology" and "artificial intelligence" as user A's interest points and constructs a user interest graph containing these interest points in combination with their historical behavior data. Then, through the GCN for representation learning of this graph, the system can capture the upward trend of user A's interest in "artificial intelligence" and, in subsequent recommendations, preferentially recommend more high-quality content related to artificial intelligence to user A, thereby improving the user experience and the effectiveness of recommendations.

[0106] Preferably, the representation learning of nodes and edges through the multi-layer graph convolutional network includes:

[0107] Initializing the feature vectors of each node in the user interest graph;

[0108] Applying multi-layer graph convolutional operations to aggregate the information of neighbor nodes layer by layer and update the feature vectors of each node. The convolutional operation of each layer is expressed as:

[0109]

[0110] Among them, represents the feature vector of node at the layer; is the set of neighbor nodes of node ; is the normalization coefficient; is the weight matrix of the layer; is the activation function; u represents the neighbor node of node v, that is, the node directly connected to node v; represents the feature vector of node u at the (k - 1)th layer, which is the feature obtained through the calculation of the previous layer of the graph convolutional network;

[0111] Introduce the attention mechanism to dynamically adjust the weights of the edges in the user interest graph, and generate the adjusted user interest graph representation data. The weight calculation formula of the edge is:

[0112]

[0113] Among them, is the weight of the edge ; and respectively represent the feature vectors of nodes and in the graph; is the attention weight vector; represents the transpose operation, that is, represents the transpose of the vector ; is the learnable parameter matrix used for linearly transforming the node feature vectors; ‖ represents the concatenation operation, which concatenates the node feature vectors transformed by the weight matrix ; is the activation function used to enhance the non - linear characteristics of the model;

[0114] Use the temporal convolutional network to capture the time - series information of the user interest and generate the time - series feature data;

[0115] Fuse the time - series feature data with the adjusted user interest graph representation data to generate the time - series user interest graph representation data.

[0116] In the present invention, representing nodes and edges through a multi - layer graph convolutional network (GCN) is a key step in achieving accurate content recommendation. This process includes initializing the feature vectors of each node in the user interest graph, and then applying multi - layer graph convolutional operations to aggregate the information of neighbor nodes layer by layer and update the feature vectors of each node. In each layer, the feature vector of the node The calculation formula is as shown above. In this way, GCN can effectively aggregate the information of each node and its neighbor nodes, thereby capturing the complex relationships in the graph structure.

[0117] Based on GCN, an attention mechanism is introduced to further dynamically adjust the weights of the edges in the user interest graph, generating the adjusted user interest graph representation data (using the edge weight calculation formula). The attention mechanism enables the model to dynamically adjust the edge weights according to the importance of the nodes, thus more accurately reflecting the relationships between the nodes.

[0118] In addition, a Temporal Convolutional Network (TCN) is used to capture the time series information of user interests, generating time series feature data. TCN effectively captures the dynamic patterns of user interests changing over time through convolutional operations in the time dimension. This process ensures that the system can not only understand the user's current interests but also predict the future trend of the user's interest changes.

[0119] Finally, the time series feature data is fused with the adjusted user interest graph representation data to generate the time series user interest graph representation data. In this way, the system can integrate the time dynamics and structural relationships of user interests, thereby providing more personalized and accurate content recommendations.

[0120] First of all, the combination of GCN and the attention mechanism enables the system to extract high-dimensional features from the complex graph structure and accurately capture the multi-dimensional relationships of user interests. Secondly, the introduction of TCN enables the system to dynamically track and predict the changes in user interests. Overall, by fusing time series features with graph structure features, the system can achieve a comprehensive, dynamic, and in-depth modeling of user interests, thereby providing personalized and accurate content recommendations.

[0121] For example, if user A has shown a high degree of interest in technology and artificial intelligence-related content in the past few months and has recently started to pay attention to the field of quantum computing, the system will first parse the user's text, speech, and video inputs through NLP technology to identify the current interest points of user A. Then, by combining the historical behavior data of user A, a user interest graph containing interest points such as "technology", "artificial intelligence", and "quantum computing" is constructed. With the help of GCN and the attention mechanism, the system can dynamically adjust the weights of the nodes and edges in the graph, accurately reflecting the interest structure and intensity of user A. Finally, by capturing the time dynamics of user interests through TCN, the system can predict the content that user A may be more interested in in the future, thereby preferentially displaying high-quality content related to quantum computing in the recommendation to improve the user experience and recommendation accuracy.

[0122] Such as Figure 4As shown, time series user interest graphs are used to represent data. Through the self-attention mechanism and multi-head attention mechanism of the Transformer model, combined with a contrastive learning module, the matching degree between user interest representation and recommended content is optimized to generate personalized recommendation content data;

[0123] The system uses time series user interest graphs to represent data and processes it through the self-attention mechanism of the Transformer model. The core of the self-attention mechanism is that it can calculate the importance of each element (in this context, the nodes in the user interest graph) to other elements and perform weighted summation. This mechanism allows the model to simultaneously focus on the entire time series data when processing the current data, thereby capturing the dynamic changes in user interests. For example, when calculating the user's current interest in a certain piece of content, the self-attention mechanism can refer to the relevant interest changes of the user in the past period of time and comprehensively evaluate the intensity and relevance of the current interest.

[0124] The multi-head attention mechanism further enhances this ability. By dividing the attention mechanism into multiple heads, each head independently learns different relationships and features, and then combines this information. The model can more comprehensively capture the complex patterns in the user interest graph. Each attention head independently processes a part of the information, and finally the outputs of all heads are concatenated together to form the final representation. This way enables the model to process different levels and different dimensions of information in parallel, thereby improving the accuracy and diversity of recommendations.

[0125] To optimize the matching degree between user interest representation and recommended content, the system combines a contrastive learning module. Contrastive learning learns better representations by constructing positive and negative sample pairs. Positive sample pairs refer to the relationship between the content that the user is actually interested in and its interest representation, while negative sample pairs are the relationship between the content that the user is not interested in and its interest representation. By calculating the distance between positive and negative sample pairs, the system can optimize the representation, making the distance between positive sample pairs closer and the distance between negative sample pairs farther, thereby improving the accuracy of recommendations. Specifically, the loss function is used to measure the performance of the model in distinguishing positive and negative sample pairs. By minimizing the loss function, the model can continuously optimize its own parameters to achieve a better matching effect.

[0126] The effect of this process is that through the self-attention mechanism and multi-head attention mechanism of the Transformer model, the system can extract accurate interest representations from the time-series user interest graph, and combine contrastive learning to optimize the matching degree of recommended content, and finally generate personalized recommendation content data. For example, when user B has frequently browsed health-related articles in the past period of time and recently started to pay attention to mental health content, the system captures this interest change trend through the self-attention mechanism and comprehensively processes the user's multi-dimensional interest data through the multi-head attention mechanism. Combining with the contrastive learning module, the system can optimize the interest representation of user B, making the recommended mental health content more in line with the user's needs and preferences, thereby improving the accuracy of recommendations and user satisfaction.

[0127] In this way, the system can not only dynamically track the changes in user interests, but also significantly improve the accuracy and relevance of recommended content, and finally provide personalized high-quality recommendation services.

[0128] Preferably, the self-attention mechanism and multi-head attention mechanism through the Transformer model include:

[0129] Convert the time-series user interest graph representation data into the input format of the Transformer model;

[0130] Apply the self-attention mechanism to calculate the importance of each node and weight the encoded user interest data. The calculation formula of the self-attention mechanism is:

[0131]

[0132] Among them, represents the query vector; represents the key vector; represents the value vector; is the vector dimension; represents the transpose operation; Softmax represents the normalization function, which converts the input into a probability distribution to ensure that the sum of the output weights is 1;

[0133] Use the multi-head attention mechanism to capture information at different levels and dimensions in the user interest graph. The calculation formula of the multi-head attention mechanism is:

[0134]

[0135] Among them, each attention head is calculated as:

[0136]

[0137] Among them, is the weight matrix for the query vector; is the weight matrix for the key vector; is the weight matrix for the value vector; is the output weight matrix, generating comprehensive user interest data; Concat represents the concatenation operation, combining multiple vectors or matrices by dimension; Attention represents the self-attention mechanism, which is used to calculate the similarity between the query vector and the key vector to generate a weighted value vector.

[0138] The time series user interest graph representation data is converted into the input format of the Transformer model. The input of the Transformer model usually includes a query vector (Q), a key vector (K), and a value vector (V), which represent the nodes and their relationships in the user interest graph. The self-attention mechanism weights the encoded user interest data by calculating the importance of each node. The self-attention mechanism can refer to the information of other nodes in the entire graph when processing the information of the current node, thereby capturing global relationships and patterns.

[0139] The multi-head attention mechanism further enhances this ability. By dividing the attention mechanism into multiple heads, each head independently learns different relationships and features, and then combines this information. The model can more comprehensively capture the complex patterns in the user interest graph. In this way, the multi-head attention mechanism can process information at different levels and dimensions in parallel, thereby generating comprehensive user interest data.

[0140] Through the self-attention mechanism and the multi-head attention mechanism of the Transformer model, the system can extract accurate interest representations from the time series user interest graph. This process allows the model to comprehensively consider the information of all nodes in the entire time series when calculating the representation of the current node, thereby capturing the dynamic changes of user interests. For example, when a user shows different interests in different categories of content in the past few months, the model can identify the trends of these interest changes through the self-attention mechanism and comprehensively process the user's multi-dimensional interest data through the multi-head attention mechanism.

[0141] Combined with the contrastive learning module, the system can further optimize the matching degree between the user interest representation and the recommended content. Contrastive learning learns better representations by constructing positive and negative sample pairs. The positive sample pair refers to the relationship between the content that the user is actually interested in and its interest representation, while the negative sample pair is the relationship between the content that the user is not interested in and its interest representation. By calculating the distance between the positive and negative sample pairs, the system can optimize the representation, making the distance between the positive sample pairs closer and the distance between the negative sample pairs farther, thereby improving the accuracy of recommendations.

[0142] The effect of this method is remarkable. Through the self-attention mechanism and the multi-head attention mechanism, the system can extract accurate interest representations from the time-series user interest graph, and combine with the contrastive learning module to optimize the matching degree between the user interest representation and the recommended content, thereby generating personalized recommendation content data. For example, when a user frequently browses technology and health-related content in the past period, the system can identify the multi-dimensional interests of the user through the above mechanisms and preferentially display high-quality content in these categories in the recommendation, thus improving the user experience and the effectiveness of the recommendation.

[0143] In summary, through the self-attention mechanism and the multi-head attention mechanism of the Transformer model, combined with the contrastive learning module, the system can achieve a deep understanding of user interests and accurate recommendations, and finally generate personalized recommendation content.

[0144] Preferably, the optimization of the matching degree between the user interest representation and the recommended content by combining the contrastive learning module includes:

[0145] Construct positive and negative sample pairs to generate positive and negative sample data;

[0146] Calculate the distance between the positive and negative sample data to generate sample distance data;

[0147] Calculate the loss value according to the loss function of contrastive learning. The loss function is defined as:

[0148]

[0149] where, is the distance between sample and ; is the binary label; is the margin distance, used to generate loss data;

[0150] Utilize the optimization result of the contrastive learning module to adjust the comprehensive user interest data and optimize its matching degree with the recommended content;

[0151] Generate personalized recommendation content data according to the optimized comprehensive user interest data.

[0152] The system generates positive and negative sample data by constructing positive and negative sample pairs. A positive sample pair refers to the relationship between the content that the user is actually interested in and their interest representation, while a negative sample pair is the relationship between the content that the user is not interested in and their interest representation. In this way, the system can learn from the user's actual feedback to distinguish between the content that the user is interested in and not interested in. For example, user A clicks on and browses an article about health, and this article forms a positive sample pair with user A's interest representation; while user A shows no interest in another article about sports, and this article forms a negative sample pair with user A's interest representation.

[0153] The system calculates the distance between the positive and negative sample data to generate sample distance data. The distance between samples reflects the matching degree between the user's interest representation and the recommended content. For positive sample pairs, the closer the distance, the higher the matching degree; for negative sample pairs, the farther the distance, the lower the matching degree. This step is achieved by calculating the distance matrix to ensure that the system can quantify the similarity between each pair of samples.

[0154] The system calculates the loss value according to the loss function of contrastive learning. This loss function optimizes the model parameters by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs. Specifically, for positive sample pairs, if the distance is less than the boundary distance , the loss value is zero; if the distance is greater than the boundary distance, the loss value increases. For negative sample pairs, the greater the distance, the smaller the loss value. This design enables the model to effectively learn to distinguish between positive and negative sample pairs, thereby improving the accuracy of recommendations.

[0155] The system uses the optimization result of the contrastive learning module to adjust the user interest data. By minimizing the loss function, the system can continuously optimize the user's interest representation to make it more accurately reflect the user's actual interests. The adjusted user interest data is further used to generate personalized recommended content, thereby enhancing the effect of the recommendation system. For example, after being optimized by contrastive learning, the system can more accurately identify user A's interest in health-related content and recommend more relevant high-quality content.

[0156] By combining the optimization of the user interest representation and the matching degree between the recommended content in the contrastive learning module, and by constructing positive and negative sample pairs, calculating sample distances, optimizing the loss function, and adjusting the user interest data, the system can achieve a deep understanding of the user's interests and accurate recommendations. This method not only improves the relevance and personalization of the recommended content but also continuously enhances the performance of the recommendation system through continuous learning and optimization. In this way, users can obtain personalized recommended content that better meets their interests and needs, improving the user experience and satisfaction.

[0157] Collect the user's multimodal information, fuse it with the personalized recommendation content data, generate multimodal recommendation data using the multimodal fusion algorithm, and generate the optimized multimodal recommendation content data through the optimization algorithm;

[0158] Preferably, the collecting the user's multimodal information and fusing it with the personalized recommendation content data includes:

[0159] Collect the user's geographical location, device usage, and social network data as multimodal information;

[0160] Preprocess the multimodal information to generate standardized multimodal feature data;

[0161] Use the multimodal fusion algorithm to fuse the standardized multimodal feature data and the personalized recommendation content data to generate multimodal recommendation data;

[0162] Use the optimization algorithm to optimize the multimodal recommendation data to ensure the relevance and personalization of the recommended content, and generate the optimized multimodal recommendation content data, where the optimization algorithm includes: gradient descent method, genetic algorithm, and simulated annealing algorithm.

[0163] The system will collect the user's multimodal information, which includes but is not limited to the user's geographical location, device usage, and social network data. Geographical location data can reflect the user's current activity area and location preferences. For example, the user may be more inclined to receive content recommendations related to their current location. Device usage data involves the type of device, operating system, and usage habits used by the user, and this information can help the system understand the user's technical preferences and usage environment. Social network data includes the user's activities, interest groups, and interaction records on social platforms, and this information can reveal the user's social behavior and interest preferences.

[0164] Preprocess the collected multimodal information to generate standardized multimodal feature data. The preprocessing steps include data cleaning, normalization, and feature extraction, etc. The purpose is to convert the multimodal information into a unified feature representation form for subsequent fusion and processing. For example, geographical location data can be represented by coordinate normalization, device usage can be represented by device type encoding, and social network data can be represented by means of interest tag extraction and user relationship graph construction, etc.

[0165] Using a multimodal fusion algorithm, the standardized multimodal feature data is fused with personalized recommendation content data to generate multimodal recommendation data. The core of the multimodal fusion algorithm lies in its ability to effectively integrate feature information from different sources, thereby capturing the comprehensive interests and needs of users. Common multimodal fusion methods include feature-level fusion, decision-level fusion, and hybrid-level fusion, etc. For example, feature-level fusion forms a comprehensive feature vector by concatenating feature vectors of different modalities, representing the overall interests of users.

[0166] The multimodal recommendation data is optimized through an optimization algorithm to ensure the relevance and personalization of the recommended content. The optimization algorithm can include gradient descent method, genetic algorithm, and simulated annealing algorithm, etc. The goal of these algorithms is to find the optimal recommendation parameters and strategies through iterative optimization, thereby generating optimized multimodal recommendation content data. For example, the gradient descent method continuously adjusts the parameters of the recommendation model by minimizing the recommendation error function to make the recommendation results closer to the actual needs of users; the genetic algorithm searches for the globally optimal recommendation strategy by simulating natural selection and genetic variation; the simulated annealing algorithm finds the optimal solution in a high-dimensional search space by simulating the physical annealing process.

[0167] By collecting and fusing the multimodal information of users, the system can comprehensively understand the interests and needs of users, thereby providing more accurate and personalized content recommendations. For example, when a user frequently participates in travel-related discussions on a social network, and at the same time their geographical location data indicates that the user is planning a trip, and the device usage shows that the user often uses a mobile device to browse travel information, the system can integrate this information, generate personalized recommendation content related to travel through the multimodal fusion algorithm, and ensure the relevance and personalization of the recommended content through the optimization algorithm, thereby enhancing the user experience and satisfaction.

[0168] In this way, the system can achieve a comprehensive understanding and dynamic adjustment of users' interests, thereby providing personalized and high-precision content recommendation services to meet the diverse needs of users in different situations.

[0169] The optimized multimodal recommendation content data is displayed to users in real time, user feedback feature data is collected, and the generation strategy of the recommended content is dynamically adjusted and optimized through an adaptive learning algorithm to ensure the continuous accuracy and personalization of the recommended content.

[0170] Such as Figure 5As shown, the system will display the optimized multi-modal recommendation content data to users in real time. This step involves the design of the user interface and the presentation method of the recommended content. The optimized multi-modal recommendation content data is the result of multi-modal information fusion and optimization algorithms, ensuring the high relevance and personalization of the recommended content. For example, when a user opens a content recommendation platform, the system will generate and display in real time the content that best matches the user's current interests based on the user's historical behavior, current environment, and multi-modal information. The advantage of real-time display is that it can instantly respond to changes in user needs, improving user satisfaction and engagement.

[0171] While displaying the recommended content, the system will collect the user's feedback feature data. These data include the user's click-through rate, dwell time, number of likes, comment content, etc. Through these data, we can understand the user's actual reaction and satisfaction with the recommended content. For example, if a user frequently clicks on and browses a certain type of recommended content, the system will consider that this type of content matches the user's interests and record these behavior data as feedback features. The collected feedback feature data provides valuable reference information for subsequent recommendation optimization.

[0172] Through the adaptive learning algorithm, the system dynamically adjusts and optimizes the generation strategy of the recommended content. The core of the adaptive learning algorithm is that it can adjust the model parameters and recommendation strategies in real time according to the user feedback feature data to continuously improve the recommendation effect. Specifically, the adaptive learning algorithm will analyze the user's feedback feature data to identify which recommendation strategies are effective and which need to be adjusted. For example, the system may find that certain content is more popular among users during a specific period, and thus increase the weight of this type of content in future recommendations. Through continuous iteration and optimization, the adaptive learning algorithm enables the recommendation system to dynamically adapt to changes in user interests and generate more accurate and personalized recommended content.

[0173] Through this process, the system can achieve the following effects: First, by displaying the optimized multi-modal recommendation content data in real time, the system can instantly meet the user's personalized needs, improving user engagement and satisfaction; Second, by collecting and analyzing the user feedback feature data, the system can deeply understand the user's interests and behavior patterns, providing data support for subsequent recommendation optimization; Finally, through the dynamic adjustment and optimization of the adaptive learning algorithm, the system can continuously improve the accuracy and personalization of the recommended content, ensuring the continuous optimization of the recommendation effect.

[0174] For example, when user C opens the content recommendation platform, the system generates and displays a series of content related to user C's interests in real time based on user C's historical behavior and current environment. During the browsing process, user C shows strong interest in some recommended content and conducts multiple interactions (such as liking and commenting). The system collects these feedback feature data, analyzes user C's behavior pattern, and adjusts the recommendation strategy through an adaptive learning algorithm. In the next recommendation, the system will, according to the optimized strategy, give priority to recommending content that better suits user C's interests, thereby further enhancing the user experience and satisfaction.

[0175] In summary, by displaying optimized multimodal recommendation content data in real time, collecting user feedback feature data, and dynamically adjusting and optimizing the generation strategy of recommended content through an adaptive learning algorithm, the system can achieve continuous, accurate, and personalized content recommendation, effectively enhancing user satisfaction and loyalty.

[0176] Preferably, the real-time display of optimized multimodal recommendation content data to users and the collection of user feedback feature data include:

[0177] Real-time display of optimized multimodal recommendation content data;

[0178] Collecting user interaction data for recommended content, including click-through rate, dwell time, number of likes, and comment content;

[0179] Preprocessing the user interaction data, extracting key metrics and features, and generating user feedback feature data.

[0180] The system will display optimized multimodal recommendation content data to users in real time. The multimodal recommendation content data is generated by fusing the user's multimodal information (such as geographical location, device usage, social network data, etc.) and personalized recommendation content data and then processed by an optimization algorithm. These data reflect the user's current interests and needs, ensuring the relevance and personalization of the recommended content. For example, when a user opens the content recommendation platform, the system will generate and display the content that best suits the user's interests in real time based on the user's historical behavior and current environment, such as the latest news, interesting articles, or videos.

[0181] During the process of users browsing and interacting with the recommended content, the system will collect the users' interaction data. The interaction data includes click-through rate, dwell time, number of likes, comment content, etc. These data can reflect the users' actual reactions and satisfaction with the recommended content. For example, when a user clicks on an article, the system records the click behavior; the time the user stays on the article page reflects the attractiveness of the content; likes and comments indicate the users' recognition and participation in the content. These interaction data provide an important basis for the system to understand the users' interest preferences and behavior patterns.

[0182] Next, the system preprocesses the collected user interaction data, extracts key metrics and features, and generates user feedback feature data. The preprocessing steps include data cleaning, normalization, and feature extraction, etc. The purpose is to convert the original interaction data into a standardized and structured feature representation. For example, the click-through rate can be obtained by calculating the ratio of the number of times a recommended content is clicked to the number of times it is displayed; the dwell time can be represented by the number of seconds the user stays on the page; and the number of likes and comments can be directly used as indicators of user engagement. Through these preprocessing steps, the system can extract the feature data that is most valuable for recommendation optimization.

[0183] By presenting the optimized multi-modal recommendation content data in real time, collecting user feedback feature data, and performing preprocessing, the system can dynamically capture and accurately model user interests. The effect of this process is that by presenting the most relevant and personalized content, the system can immediately meet the user's needs, improve user engagement and satisfaction; by collecting and analyzing the user's interaction data, the system can continuously learn and optimize the recommendation strategy to ensure the continuous accuracy and personalization of the recommended content.

[0184] For example, after user D opens the content recommendation platform, the system will, based on user D's historical behavior and multi-modal information, present a series of personalized recommendation content in real time. User D clicks on and browses an article about artificial intelligence and stays on the page for a long time, showing strong interest. At the same time, user D likes and comments on this article. The system will collect these interaction data and extract key metrics such as click-through rate, dwell time, number of likes, and comment content through preprocessing. By analyzing these user feedback feature data, the system identifies user D's high interest in artificial intelligence and preferentially presents more relevant content in future recommendations, thereby improving the accuracy of the recommendation and user satisfaction.

[0185] In summary, by presenting the optimized multi-modal recommendation content data in real time, collecting user feedback feature data, and performing preprocessing, the system can dynamically track and accurately model user interests, thereby providing continuous accurate and personalized content recommendation services. This process not only improves the effectiveness of the recommendation system but also significantly enhances the user experience and satisfaction.

[0186] As Figure 6 shown, preferably, the generation strategy for dynamically adjusting and optimizing the recommended content through the adaptive learning algorithm includes:

[0187] According to the user feedback feature data, use the adaptive learning algorithm to dynamically adjust and optimize the generation strategy of the recommended content to generate optimized recommended content data;

[0188] Using the optimized recommended content data, update the parameters and strategies of the recommended content generation strategy, including the adjustment of model weights, learning rates, and loss functions, to generate updated multi-modal recommended content data;

[0189] Apply the updated multi-modal recommended content data to the real-time recommendation generation module, dynamically adjust the recommendation strategy, and generate optimized real-time recommendation data for presenting and providing recommendation services to users.

[0190] Based on the user feedback feature data, the system dynamically adjusts and optimizes the recommended content generation strategy using an adaptive learning algorithm. User feedback feature data includes the click-through rate, dwell time, number of likes, comment content, etc. of the user on the recommended content. Through these data, the actual response and satisfaction of the user with the recommended content can be understood. The core of the adaptive learning algorithm lies in its ability to adjust the model parameters and recommendation strategy in real time according to this feedback data. For example, if the system detects that the user shows a higher interest in a certain type of content, it will increase the weight of this type of content in the recommendation, generating optimized recommended content data. This step enables the system to dynamically respond to changes in user interests by analyzing the real-time feedback of the user, thereby improving the relevance and personalization of the recommendation.

[0191] Using the optimized recommended content data, the system updates the parameters and strategies of the recommended content generation strategy. Specifically, this includes the adjustment of model weights, learning rates, and loss functions. Model weights determine the influence degree of each feature in the recommendation model; the learning rate controls the pace of model parameter updates; the loss function is used to measure the accuracy of the recommendation results. Through the optimized recommended content data, the system can identify and adjust these parameters to make the recommendation model better adapt to the changes in user interests and needs. For example, the system may increase the weights of certain features to make them occupy a more important position in the recommendation results; or adjust the learning rate to make the model updates more stable and efficient.

[0192] Apply the updated multi-modal recommended content data to the real-time recommendation generation module, dynamically adjust the recommendation strategy, and generate optimized real-time recommendation data. The role of the real-time recommendation generation module is to generate and display the content that best matches the user's current interests according to the latest recommendation strategy. By applying the updated multi-modal recommended content data to this module, the system can ensure the continuous accuracy and personalization of the recommended content. For example, every interaction of the user on the platform will be recorded and analyzed in real time, and the system immediately adjusts the recommendation results according to the latest feedback data and optimization strategy, presenting the most relevant and personalized content to the user.

[0193] The effect of this process is remarkable. Through the adaptive learning algorithm, the system can continuously learn and dynamically adjust the user's interests and behaviors. First, based on the user feedback feature data, the system can immediately identify and respond to the changes in the user's interests, improving the relevance and satisfaction of the recommended content. Second, by optimizing the parameters and strategies of the recommended content data update generation strategy, the system can continuously improve the performance and accuracy of the recommendation model. Finally, by applying the updated recommended content data to the real-time recommendation generation module, the system can ensure that each recommendation can immediately reflect the user's latest interests and needs, thus providing continuous accurate and personalized recommendation services.

[0194] For example, when user E browses a series of articles on technological innovation on the content recommendation platform and likes and comments on several of them, the system will collect this feedback feature data, analyze the changes in user E's interests through the adaptive learning algorithm, and dynamically adjust the recommendation strategy. The optimized recommended content data will reflect user E's high interest in technological innovation. The system will update the parameters and strategies of the recommendation model and increase the weight of content related to technological innovation. When user E visits the platform next time, the system will, through the real-time recommendation generation module, give priority to recommending more high-quality content related to technological innovation to improve the user experience and satisfaction.

[0195] Generally speaking, by dynamically adjusting and optimizing the generation strategy of the recommended content through the adaptive learning algorithm, the system can continuously learn and dynamically adjust the user's interests, ensure the continuous accuracy and personalization of the recommended content, and effectively improve the user's satisfaction and loyalty.

[0196] As Figure 7 shown, a system for implementing the content recommendation method based on semantic recognition includes:

[0197] A natural language processing module configured to parse the user's requirements input through text, voice, or video and generate user semantic representation data. The natural language processing module is configured to parse the user's requirements input through text, voice, or video and generate user semantic representation data. This module uses natural language processing technology to parse and process the user's input data. For example, it performs word segmentation, part-of-speech tagging, and semantic parsing on the user's input text data to generate structured user semantic representation data. For voice input, this module uses speech recognition technology to convert the voice into text and further parse its semantic content. For video input, it extracts voice and image features through voice and image analysis to generate video semantic data. In this way, the natural language processing module can accurately capture the user's intention and improve the recommendation system's ability to understand and respond to the user's needs.

[0198] The user interest graph construction module dynamically constructs a user interest graph based on user semantic representation data and user historical behavior data, and performs representation learning on nodes and edges through a multi-layer graph convolutional network to generate time-series user interest graph representation data. The user interest graph construction module dynamically constructs a user interest graph based on user semantic representation data and user historical behavior data, and performs representation learning on nodes and edges through a multi-layer graph convolutional network to generate time-series user interest graph representation data. This module first extracts the user's interest point data and constructs a user interest graph in combination with the user's historical behavior data. On the user interest graph, a multi-layer graph convolutional network (GCN) is applied to perform representation learning on nodes and edges. Through multi-layer convolutional operations, the information of neighbor nodes is aggregated layer by layer to update the feature vector of each node. At the same time, an attention mechanism is introduced to dynamically adjust the weights of the edges in the user interest graph to generate the adjusted user interest graph representation data. A temporal convolutional network (TCN) is used to capture the time-series information of user interests, and the time-series feature data is fused with the adjusted user interest graph representation data to generate time-series user interest graph representation data. This process not only improves the accuracy of user interest representation but also effectively addresses the problems of data sparsity and cold start.

[0199] The Transformer model module uses the time-series user interest graph representation data to optimize the matching degree between the user interest representation and the recommended content through the self-attention mechanism and the multi-head attention mechanism, in combination with the contrastive learning module, to generate personalized recommended content data. The Transformer model module uses the time-series user interest graph representation data to optimize the matching degree between the user interest representation and the recommended content through the self-attention mechanism and the multi-head attention mechanism, in combination with the contrastive learning module, to generate personalized recommended content data. This module first converts the time-series user interest graph representation data into the input format of the Transformer model. Through the self-attention mechanism, the importance of each node is calculated, and the encoded user interest data is weighted. The multi-head attention mechanism further captures the information at different levels and dimensions in the user interest graph to generate comprehensive user interest data. During the training process, a contrastive learning module is introduced. By constructing positive and negative sample pairs and calculating their distances, the matching degree between the user interest representation and the recommended content is optimized to ensure a high similarity match between the recommended content and the user's preferences, generating personalized recommended content data. This process not only improves the relevance and personalization of the recommended content but also solves the problems of homogenization and high model optimization difficulty in the recommendation system.

[0200] The multimodal information collection and fusion module collects the user's multimodal information, fuses it with the personalized recommendation content data, generates multimodal recommendation data using multimodal fusion algorithms, and generates optimized multimodal recommendation content data through optimization algorithms. The multimodal information collection and fusion module collects the user's multimodal information, fuses it with the personalized recommendation content data, generates multimodal recommendation data using multimodal fusion algorithms, and generates optimized multimodal recommendation content data through optimization algorithms. This module first collects multimodal information such as the user's geographical location, device usage, and social network data, and preprocesses this data to generate standardized multimodal feature data. Next, using multimodal fusion algorithms, the standardized multimodal feature data is fused with the personalized recommendation content data to generate multimodal recommendation data. Through optimization algorithms (such as gradient descent, genetic algorithms, and simulated annealing algorithms), the multimodal recommendation data is optimized to ensure the relevance and personalization of the recommended content, and finally optimized multimodal recommendation content data is generated. This process can effectively improve the diversity and accuracy of the recommended content and enhance the user experience.

[0201] The real-time display and feedback collection module real-time displays the optimized multimodal recommendation content data to the user and collects user feedback feature data. The real-time display and feedback collection module is responsible for real-time displaying the optimized multimodal recommendation content data to the user and collecting user feedback feature data. This module first presents the optimized multimodal recommendation content data to the user in real time and collects user feedback feature data through the user's interaction data with the recommended content (such as click-through rate, dwell time, number of likes, comment content, etc.). The user interaction data is preprocessed to extract key metrics and features, generating user feedback feature data. This process can help the system timely understand the user's satisfaction and preferences for the recommended content, providing basic data support for subsequent adaptive learning and recommendation strategy optimization.

[0202] The adaptive learning module, based on the user feedback feature data, dynamically adjusts and optimizes the generation strategy of the recommended content through adaptive learning algorithms to ensure the continuous accuracy and personalization of the recommended content. The adaptive learning module, based on the user feedback feature data, dynamically adjusts and optimizes the generation strategy of the recommended content through adaptive learning algorithms to ensure the continuous accuracy and personalization of the recommended content. This module first uses the user feedback feature data to dynamically adjust the parameters and strategies of the recommended content generation strategy, including the adjustment of model weights, learning rates, and loss functions, to generate optimized recommended content data. Next, the optimized recommended content data is applied to the real-time recommendation generation module to dynamically adjust the recommendation strategy and generate optimized real-time recommendation data for presenting and providing recommendation services to the user. Through this process, the adaptive learning module can continuously optimize the recommendation model and strategy, ensuring the real-time, accurate, and personalized nature of the recommended content and improving the user's satisfaction and usage experience.

[0203] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0204] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0205] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0206] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0207] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0208] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0209] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0210] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0211] The above are only examples of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A content recommendation method based on semantic recognition, characterized in that: The following steps are involved: Analyze the needs of users through text, voice or video input through natural language processing technology to generate user semantic representation data; Based on the user semantic representation data and the user historical behavior data recording the user's interaction history on the platform, a user interest graph is dynamically constructed, and a multi-layer graph convolutional network is applied to the user interest graph to perform node and edge representation learning to generate user interest graph representation data. The node and edge representation learning by the multi-layer graph convolutional network includes: Initialize the feature vector of each node in the user interest graph; Apply multi-layer graph convolution operations to aggregate the information of neighbor nodes layer by layer and update the feature vector of each node; Introduce the attention mechanism to dynamically adjust the weights of the edges in the user interest graph and generate the adjusted user interest graph representation data; Use the time convolutional network to capture the time series information of user interests and generate time series feature data; The time series feature data is merged with the adjusted user interest graph representation data to generate the time series user interest graph representation data; The user interest graph of time series is used to represent data. Through the self-attention mechanism and multi-head attention mechanism of the Transformer model, combined with the contrastive learning module, the matching degree between the user interest representation and the recommended content is optimized to generate personalized recommended content data. Collecting multimodal information of users and fusing it with personalized recommendation content data, generating multimodal recommendation data using a multimodal fusion algorithm, and generating optimized multimodal recommendation content data using an optimization algorithm, wherein the user's geographic location, device usage, and social network data are used as multimodal information; Display optimized multimodal recommended content data to users in real time, collect user feedback feature data, and dynamically adjust and optimize the generation strategy of recommended content through adaptive learning algorithms to ensure the continuous accuracy and personalization of recommended content. The generation strategy of dynamically adjusting and optimizing recommended content through adaptive learning algorithms includes: Based on user feedback feature data, the adaptive learning algorithm is used to dynamically adjust and optimize the generation strategy of recommended content to generate optimized recommended content data; Using the optimized recommended content data, updating the parameters and strategies of the recommended content generation strategy, including the adjustment of model weights, learning rates, and loss functions, to generate updated multimodal recommended content data; Apply the updated multimodal recommendation content data to the real-time recommendation generation module, dynamically adjust the recommendation strategy, and generate optimized real-time recommendation data for displaying and providing recommendation services to users.

2. The method according to claim 1, characterized in that The requirements of analyzing the user's input through text, voice or video by natural language processing technology include: Perform speech recognition on the speech data and convert it into text data; Perform speech and image analysis on video data, extract speech and image features, and generate video semantic data; Perform word segmentation, part-of-speech tagging and semantic analysis on text data to generate user semantic representation data.

3. The method according to claim 1, characterized in that The dynamically constructing user interest graph includes: Extract the user's points of interest and generate point of interest data; combine the user's historical behavior data to build a user interest graph and generate user interest graph data.

4. The method according to claim 1, characterized in that: The updating of the feature vector of each node and the convolution operation of each layer are expressed as: in, Representation Node In the The feature vector of the layer; For Node The set of neighbor nodes of is the normalization coefficient; For the The weight matrix of the layer; is the activation function; u represents the neighbor node of node v, that is, the node directly connected to node v; u represents the neighbor node of node v, that is, the node directly connected to node v; represents the feature vector of node u in the k-1th layer, the feature calculated by the previous layer of the graph convolutional network; In the dynamic adjustment of the edge weights in the user interest graph, the edge weight calculation formula is: in, For edge The weight of and Represents the nodes in the graph and The eigenvector of is the attention weight vector; represents the transpose operation, that is Representation vector The transpose of is a learnable parameter matrix used to linearly transform node feature vectors; ‖ represents a concatenation operation, which will be passed through the weight matrix The transformed node feature vectors are concatenated; It is an activation function used to enhance the nonlinear characteristics of the model.

5. The method according to claim 1, characterized in that The self-attention mechanism and multi-head attention mechanism through the Transformer model include: Convert the time series user interest graph representation data into the input format of the Transformer model; The self-attention mechanism is applied to calculate the importance of each node and weight the encoded user interest data. The calculation formula of the self-attention mechanism is: in, represents the query vector; represents the key vector; represents a value vector; is the vector dimension; represents the transposition operation; Softmax represents the normalization function, which converts the input into a probability distribution to ensure that the sum of the output weights is 1; The multi-head attention mechanism is used to capture information of different levels and dimensions in the user interest graph. The calculation formula of the multi-head attention mechanism is: Among them, each attention head The calculation method is: in, is the weight matrix of the query vector; is the weight matrix of the key vector; is the weight matrix of the value vector; To output the weight matrix, comprehensive user interest data is generated; Concat represents the concatenation operation, which merges multiple vectors or matrices by dimension; Attention represents the self-attention mechanism, which is used to calculate the similarity between the query vector and the key vector and generate a weighted value vector.

6. The method according to claim 5, characterized in that The optimization of the matching degree between the user interest representation and the recommended content by combining the contrastive learning module includes: Construct positive and negative sample pairs to generate positive and negative sample data; Calculate the distance between positive and negative sample data to generate sample distance data; The loss value is calculated according to the loss function of contrastive learning, which is defined as: in, For sample and The distance between is a binary label; is the boundary distance, used to generate loss data; Using the optimization results of the comparative learning module, adjust the comprehensive user interest data to optimize the match between it and the recommended content; Generate personalized recommendation content data based on the optimized comprehensive user interest data.

7. The method according to claim 1, characterized in that The collecting of multimodal information of users and fusing it with personalized recommendation content data includes: Preprocess the multimodal information to generate standardized multimodal feature data; Use a multimodal fusion algorithm to fuse standardized multimodal feature data and personalized recommendation content data to generate multimodal recommendation data; The multimodal recommendation data is optimized by using an optimization algorithm to ensure the relevance and personalization of the recommended content, and to generate optimized multimodal recommendation content data, wherein the optimization algorithm includes: a gradient descent method, a genetic algorithm, and a simulated annealing algorithm.

8. The method according to claim 1, characterized in that The real-time display of optimized multi-modal recommended content data to the user and the collection of user feedback feature data include: Real-time display of optimized multi-modal recommended content data; Collect user interaction data on recommended content, including click-through rate, dwell time, number of likes, and comment content; Preprocess user interaction data, extract key indicators and features, and generate user feedback feature data.

9. A system for implementing the content recommendation method based on semantic recognition according to any one of claims 1 to 8, characterized in that: include: A natural language processing module configured to parse user requirements inputted via text, voice or video and generate user semantic representation data; The user interest graph construction module dynamically constructs the user interest graph based on the user semantic representation data and the user historical behavior data, and learns the representation of nodes and edges through a multi-layer graph convolutional network to generate time series user interest graph representation data; The Transformer model module uses the time series user interest graph to represent data, and optimizes the matching degree between user interest representation and recommended content through self-attention mechanism and multi-head attention mechanism combined with the contrastive learning module to generate personalized recommended content data; The multimodal information collection and fusion module collects the multimodal information of users and fuses it with the personalized recommendation content data, generates multimodal recommendation data using the multimodal fusion algorithm, and generates optimized multimodal recommendation content data through the optimization algorithm; Real-time display and feedback collection module, which displays optimized multi-modal recommended content data to users in real time and collects user feedback feature data; The adaptive learning module dynamically adjusts and optimizes the generation strategy of recommended content based on user feedback feature data through an adaptive learning algorithm to ensure the continued accuracy and personalization of recommended content.

Citation Information

Patent Citations

  • Knowledge graph attention network recommendation method based on graph collaborative filtering

    CN116340648A

  • Text data processing method and device, computer equipment and storage medium

    CN114138985A

  • Online interest group recommendation method based on self-attention and contrast learning

    CN117171447A