Data theme recommendation method fusing Ochiai index and network representation learning

By integrating Ochiai index and network representation learning methods, semantic similarity networks and multimodal fusion networks are constructed, combined with user interest modeling, the problem of difficulty in mining data semantic information in the existing technology is solved, and the high accuracy and personalized topic recommendation effect is achieved.

CN119989249APending Publication Date: 2025-05-13NANJING UNIV +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411811925.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology is difficult to deeply explore the semantic information and potential theme structure behind the data, resulting in insufficiently accurate recommendation results.

Method used

The data topic recommendation method that integrates Ochiai index and network representation learning is adopted, and the semantic similarity is calculated through the Ochiai index, a semantic similarity network is constructed, and a multimodal fusion network representation learning and user interest modeling is used to personalize the topic recommendation.

Benefits of technology

It can calculate similarity from a deeper semantic level, accurately discover the correlation between data, improve the accuracy and quality of recommendations, meet users' needs for accurate topic recommendations, and adjust the recommendation results in real time through a dynamic recommendation mechanism to meet user interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989249A_ABST
    Figure CN119989249A_ABST
Patent Text Reader

Abstract

The invention discloses a data theme recommendation method fusing an Ochiai index and network representation learning, and relates to the technical field of recommendation systems.The method comprises the steps of S1, data collection, S2, data preprocessing, S3, semantic similarity network construction based on the Ochiai index, S4, multi-mode fusion network representation learning and S5, personalized theme recommendation. According to the method, semantic association between data objects can be deeply mined through semantic similarity network construction of Ochiai indexes, semantic weight adjustment factors are introduced into an improved Ochiai index formula, similarity calculation better fits semantic importance of data, deep-level semantic feature vectors are extracted in combination with a topic model, and the semantic similarity between the data objects is obtained. Therefore, the constructed semantic similarity network can accurately reflect the topic association relationship between the data, meanwhile, the attention mechanism is utilized to automatically learn the feature weight of each modal, the limitation of single modal information is avoided, and the demand of a user for accurate topic recommendation can be better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of recommendation systems, and in particular relates to a data topic recommendation method integrating Ochiai index and network representation learning. Background Art

[0002] With the rapid development of information technology, the amount of data is growing explosively. How to mine valuable information from massive data and accurately recommend it to users has become a key issue that needs to be solved. Among the many fields of data processing and recommendation technology, topic recommendation, as an important application method, aims to help users quickly discover topics of interest to them. It has broad application prospects in many fields such as academic research, news information, e-commerce, and social media.

[0003] Traditional content-based recommendation methods mainly measure similarity based on the surface features of text and other data, such as simple keyword matching, which makes it difficult to deeply explore the semantic information and potential topic structure behind the data. To address the above problems, the following solutions are proposed. Summary of the invention

[0004] The purpose of the present invention is to provide a data topic recommendation method that integrates the Ochiai index and network representation learning. By combining the Ochiai index with network representation learning, similarity can be calculated from a deeper semantic level, and the relationship between data can be discovered more accurately, solving the problem that the existing technology is difficult to deeply mine the semantic information and potential topic structure behind the data.

[0005] To solve the above technical problems, the present invention is achieved through the following technical solutions:

[0006] The present invention is a data topic recommendation method integrating Ochiai index and network representation learning, and the recommendation method comprises the following steps:

[0007] Step S1, data collection;

[0008] Step S2, data preprocessing;

[0009] Step S3, constructing a semantic similarity network based on the Ochiai index;

[0010] Step S4, multimodal fusion network representation learning;

[0011] Step S5: Personalized topic recommendation.

[0012] Preferably, the step S1, data collection, is used to collect data related to the topic from structured databases, semi-structured data and unstructured data.

[0013] Preferably, the step S2, data preprocessing, is used to perform cleaning operations on different types of data to remove noise, outliers and irrelevant information.

[0014] Preferably, the step S3, constructing a semantic similarity network based on the Ochiai index, comprises the following steps:

[0015] Step S31, feature vector construction: extracting deep semantic features based on the data conversion into feature vectors;

[0016] Step S32, Ochiai index calculation and network construction: Calculate the semantic similarity between data objects and introduce the semantic weight adjustment factor λ, the formula is as follows:

[0017]

[0018] In the formula, f ik and f jk is the value of the kth semantic feature of data objects i and j, λ k is the weight determined by the feature importance, and the similarity S is calculated based on ij ,Set the threshold θ to construct a semantic similarity network, and establish edges between the data objects whose similarity is greater than θ, and the data objects are used as nodes.

[0019] Preferably, the step S4, multimodal fusion network representation learning, specifically comprises the following steps:

[0020] Step S41: Each node in the network contains not only the semantic features calculated previously, but also modal information. The importance weights of different modal features to node representation are automatically learned through the attention mechanism. Suppose the multimodal feature vector of node i is:

[0021] M i =(m i1 ,m i2 ,…,m ip ), where p is the number of modes;

[0022] Step S42, network representation learning: Use a random walk strategy to perform walk sampling on the constructed multimodal fusion network to obtain a node sequence, and input the node sequence into the attention-based multimodal fusion network representation learning model for training. The model structure includes multiple fully connected layers and attention layers. The model parameters are optimized through the back propagation algorithm to obtain a low-dimensional vector representation V of the node. i =(v i1 ,v i2 ,…v id ), where d is the dimension of the low-dimensional vector.

[0023] Preferably, the step S5, personalized topic recommendation, includes the following steps:

[0024] Step S51, user interest modeling: collect user historical behavior data, analyze and mine the user behavior data, and construct a user interest vector I u =(i u1 ,i u2 ,…,i ud ), obtain the user's interest weight distribution on different topics as the interest vector;

[0025] Step S52, recommendation score calculation and ranking: calculate the recommendation score between the user and each topic.

[0026] Preferably, in step S52, the recommendation score calculation and ranking, the recommendation score is calculated by a method based on vector similarity combined with a user interest dynamic attenuation factor γ, and the formula is as follows:

[0027]

[0028] Where V t is the node vector corresponding to topic t, Δt is the time interval from the user's last behavior related to the topic to the current one, the topics are sorted according to the recommendation scores, and the top-ranked topics are selected as the recommendation results to generate a personalized recommendation list for the user.

[0029] The present invention has the following beneficial effects:

[0030] 1. The present invention constructs a semantic similarity network through the Ochiai index, which can deeply explore the semantic association between data objects. The improved Ochiai index formula introduces a semantic weight adjustment factor, so that the similarity calculation is more in line with the semantic importance of the data. The deep semantic feature vector is extracted in combination with the topic model, so that the constructed semantic similarity network can accurately reflect the topic association relationship between the data. At the same time, the attention mechanism is used to automatically learn the weights of each modal feature, so that the node representation integrates multiple key information, avoiding the limitations of single modal information, and capturing the inherent structure and semantics of the data in an all-round way, providing a rich and accurate information basis for subsequent recommendations, greatly improving the accuracy and quality of recommendations, and better meeting the user's needs for accurate topic recommendations.

[0031] 2. The present invention conducts in-depth analysis and mining based on the user's rich historical behavior data in the user interest modeling step, extracts the user's interest weight distribution on different topics and constructs an interest vector, which truly reflects the user's interest preferences. In addition, the recommendation score calculation adopts a method combining the user's interest dynamic attenuation factor, which can keenly capture the dynamic changes of user interests over time. As the user behavior is updated, the interest vector and the recommendation score can be adjusted in real time to ensure that the recommendation list is always highly consistent with the user's current interests, providing highly personalized recommendation services. This dynamic recommendation mechanism effectively avoids the problem of delayed or irrelevant recommendation results due to changes in user interests in traditional recommendation methods, significantly improves user experience, and enhances the interactivity and stickiness between users and the recommendation system.

[0032] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0034] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] See also Figure 1 As shown, the present invention is a data topic recommendation method integrating Ochiai index and network representation learning, and the recommendation method comprises the following steps:

[0037] Step S1, data collection;

[0038] Step S2, data preprocessing;

[0039] Step S3, constructing a semantic similarity network based on the Ochiai index;

[0040] Step S4, multimodal fusion network representation learning;

[0041] Step S5: Personalized topic recommendation.

[0042] Step S1, data collection is used to collect data related to the topic from structured databases, semi-structured data and unstructured data.

[0043] Step S2, data preprocessing is used to perform cleaning operations on different types of data to remove noise, outliers and irrelevant information.

[0044] Step S3, constructing a semantic similarity network based on the Ochiai index, includes the following steps:

[0045] Step S31, feature vector construction: extracting deep semantic features based on the data conversion into feature vectors;

[0046] Step S32, Ochiai index calculation and network construction: Calculate the semantic similarity between data objects and introduce the semantic weight adjustment factor λ, the formula is as follows:

[0047]

[0048] In the formula, f ik and f jk is the value of the kth semantic feature of data objects i and j, λ k is the weight determined by the feature importance, and the similarity S is calculated based on ij ,Set the threshold θ to construct a semantic similarity network, and establish edges between the data objects whose similarity is greater than θ, and the data objects are used as nodes.

[0049] Step S4, multimodal fusion network representation learning specifically includes the following steps:

[0050] Step S41: Each node in the network contains not only the semantic features calculated previously, but also modal information. The importance weights of different modal features to node representation are automatically learned through the attention mechanism. Suppose the multimodal feature vector of node i is:

[0051] M i =(m i1 ,m i2 ,…,m ip ), where p is the number of modes;

[0052] Step S42, network representation learning: Use a random walk strategy to perform walk sampling on the constructed multimodal fusion network to obtain a node sequence, and input the node sequence into the attention-based multimodal fusion network representation learning model for training. The model structure includes multiple fully connected layers and attention layers. The model parameters are optimized through the back propagation algorithm to obtain a low-dimensional vector representation V of the node. i =(v i1 ,v i2,…v id ), where d is the dimension of the low-dimensional vector.

[0053] Step S5, personalized topic recommendation includes the following steps:

[0054] Step S51, user interest modeling: collect user historical behavior data, analyze and mine the user behavior data, and construct a user interest vector I u =(i u1 ,i u2 ,…,i ud ), obtain the user's interest weight distribution on different topics as the interest vector;

[0055] Step S52, recommendation score calculation and ranking: calculate the recommendation score between the user and each topic.

[0056] Step S52, in which the recommendation score is calculated and sorted, the recommendation score is calculated by a method based on vector similarity combined with a user interest dynamic attenuation factor γ, and the formula is as follows:

[0057]

[0058] Where V t is the node vector corresponding to topic t, Δt is the time interval from the user's last behavior related to the topic to the current one, the topics are sorted according to the recommendation scores, and the top-ranked topics are selected as the recommendation results to generate a personalized recommendation list for the user.

[0059] A specific application of this embodiment is:

[0060] Step S1, data collection: collecting data related to the topic from structured databases (such as relational databases), semi-structured data (such as XML files), and unstructured data (such as social media texts, web page content);

[0061] Step S2, data preprocessing: perform cleaning operations on different types of data to remove noise, outliers, and irrelevant information. For example, for text data, use regular expressions to remove special symbols and HTML tags; for numerical data, process missing values ​​(such as mean filling, median filling, or model-based filling), and then convert all types of data into feature vector form. For text, a word vector model (such as Word2Vec) can be used for pre-training to obtain a vector representation of the text. Numerical data can be directly used as features or standardized.

[0062] Step S3: constructing a semantic similarity network based on the Ochiai index:

[0063] Step S31, feature vector construction: based on the data conversion into feature vectors, further extract deep semantic features, for example, perform topic model (such as LDA) analysis on the text vector to obtain the probability distribution vector of each data object under different topics as the semantic feature vector;

[0064] Step S32, Ochiai index calculation and network construction: The semantic similarity between data objects is calculated using the improved Ochiai index formula, and a semantic weight adjustment factor λ is introduced. The formula is as follows:

[0065]

[0066] In the formula, f ik and f jk is the value of the kth semantic feature of data objects i and j, λ k It is the weight determined by the importance of the feature (which can be calculated by methods such as information gain), and the similarity S ij , set a threshold θ (which can be adaptively determined through data distribution analysis, such as a density-based threshold determination method) to build a semantic similarity network, and establish edges between data objects with similarity greater than θ, with the data objects serving as nodes;

[0067] Step S4: Multimodal fusion network representation learning:

[0068] Step S41: For each node in the network, its features include not only the semantic features calculated previously, but also other modal information (such as the timestamp and source channel of the data). The importance weights of different modal features to the node representation are automatically learned through the attention mechanism. Let the multimodal feature vector of node i be:

[0069] M i =(m i1 ,m i2 ,…,m ip ), where p is the number of modes;

[0070] Step S42, network representation learning: Use the random walk strategy to perform walk sampling on the constructed multimodal fusion network to obtain a node sequence, and then input the node sequence into the attention-based multimodal fusion network representation learning model for training. The model structure includes multiple fully connected layers and attention layers. The model parameters are optimized through the back propagation algorithm to obtain the low-dimensional vector representation V of the node. i =(v i1 ,v i2 ,…v id ), where d is the dimension of the low-dimensional vector;

[0071] Step S5: Personalized topic recommendation:

[0072] Step S51, user interest modeling: collect user historical behavior data (such as browsing history, search keywords, likes and comments, etc.), analyze and mine the user behavior data, and construct the user interest vector I u =(i u1 ,i u2 ,…,i ud ), for example, extracting topics from the text data browsed by the user, and obtaining the user's interest weight distribution on different topics as an interest vector;

[0073] Step S52, calculation and ranking of recommendation scores: Calculate the recommendation scores between the user and each topic, using a method based on vector similarity (such as cosine similarity) combined with a user interest dynamic attenuation factor γ, the formula is as follows:

[0074]

[0075] Where V t is the node vector corresponding to topic t, Δt is the time interval from the user's last behavior related to the topic to the current one, the topics are sorted according to the recommendation scores, and the top-ranked topics are selected as the recommendation results to generate a personalized recommendation list for the user.

[0076] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0077] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A data topic recommendation method integrating Ochiai index and network representation learning, characterized in that: The recommended method comprises the following steps: Step S1, data collection; Step S2, data preprocessing; Step S3, constructing a semantic similarity network based on the Ochiai index; Step S4, multimodal fusion network representation learning; Step S5: Personalized topic recommendation.

2. According to claim 1, a data topic recommendation method integrating Ochiai index and network representation learning is characterized in that: The step S1, data collection, is used to collect data related to the subject from structured databases, semi-structured data and unstructured data.

3. According to claim 1, a data topic recommendation method integrating Ochiai index and network representation learning is characterized in that: The step S2, data preprocessing, is used to perform cleaning operations on different types of data to remove noise, outliers and irrelevant information.

4. According to claim 1, a data topic recommendation method integrating Ochiai index and network representation learning is characterized in that: The step S3, constructing a semantic similarity network based on the Ochiai index, comprises the following steps: Step S31, feature vector construction: extracting deep semantic features based on the data conversion into feature vectors; Step S32, Ochiai index calculation and network construction: Calculate the semantic similarity between data objects and introduce the semantic weight adjustment factor λ, the formula is as follows: In the formula, f ik and f jk is the value of the kth semantic feature of data objects i and j, λ k is the weight determined by the feature importance, and the similarity S is calculated based on ij ,Set the threshold θ to construct a semantic similarity network, and establish edges between the data objects whose similarity is greater than θ, and the data objects are used as nodes.

5. According to claim 1, a data topic recommendation method integrating Ochiai index and network representation learning is characterized in that: The step S4, multimodal fusion network representation learning, specifically includes the following steps: Step S41: Each node in the network contains not only the semantic features calculated previously, but also modal information. The importance weights of different modal features to node representation are automatically learned through the attention mechanism. Suppose the multimodal feature vector of node i is: M i =(m i1 ,m i2 ,…,m ip ), where p is the number of modes; Step S42, network representation learning: Use a random walk strategy to perform walk sampling on the constructed multimodal fusion network to obtain a node sequence, and input the node sequence into the attention-based multimodal fusion network representation learning model for training. The model structure includes multiple fully connected layers and attention layers. The model parameters are optimized through the back propagation algorithm to obtain a low-dimensional vector representation V of the node. i =(v i1 ,v i2 ,…v id ), where d is the dimension of the low-dimensional vector.

6. The data topic recommendation method integrating Ochiai index and network representation learning according to claim 1, characterized in that: The step S5, personalized topic recommendation, includes the following steps: Step S51, user interest modeling: collect user historical behavior data, analyze and mine the user behavior data, and construct a user interest vector I u =(i u1 ,i u2 ,…,i ud ), obtain the user's interest weight distribution on different topics as the interest vector; Step S52, recommendation score calculation and ranking: calculate the recommendation score between the user and each topic.

7. A data topic recommendation method integrating Ochiai index and network representation learning according to claim 6, characterized in that: In step S52, the recommendation score calculation and ranking, the recommendation score is calculated by a method based on vector similarity combined with a user interest dynamic attenuation factor γ, and the formula is as follows: Where V t is the node vector corresponding to topic t, Δt is the time interval from the user's last behavior related to the topic to the current one, the topics are sorted according to the recommendation scores, and the top-ranked topics are selected as the recommendation results to generate a personalized recommendation list for the user.