Literature information retrieval and analysis system and method based on AI intelligence
Through the AI-based literature information retrieval and analysis system, multimodal feature extraction and dynamic knowledge graph construction, the problem of inefficient existing literature search methods is solved, personalized recommendation and efficient information screening are realized, and the accuracy and intelligence of literature search are improved.
Patent Information
- Application Number
- CN202510651990.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing literature search methods are inefficient, difficult to accurately reflect user needs, and lack intelligence and personalized recommendations, so valuable information cannot be effectively screened out.
The literature information retrieval and analysis system based on AI intelligence is adopted, including characterization vector output module, preliminary search result output module, update search result module, dynamic knowledge graph construction module, enhanced literature characterization output module and knowledge graph feedback update module. Through multimodal feature extraction, semantic enhancement search, dynamic knowledge graph construction and personalized recommendation, the knowledge graph is optimized with user interaction data.
It improves the accuracy and intelligence of literature search, can dynamically adapt to user needs, provide personalized recommendations, reduce the time for screening low-quality literature, quickly focus on valuable content, and improve information screening efficiency and accuracy.
Smart Images

Figure CN120492636A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of document information retrieval technology, and specifically to a document information retrieval and analysis system and method based on AI intelligence. Background Art
[0002] With technological advancements and the widespread use of the internet, the volume of global scientific research data is growing exponentially, and the number of documents in various databases has skyrocketed. Traditional keyword-based search methods are unable to cope with this massive amount of data and fail to fully and accurately reflect user needs. This leads to inefficient information retrieval, difficulty in filtering out valuable information, and users are easily overwhelmed by irrelevant or low-quality content.
[0003] Current mainstream technologies include: Keyword matching retrieval: Based on Boolean logic and inverted indexes (such as Elasticsearch), but unable to capture semantic associations; Citation network analysis: This uses algorithms such as PageRank to mine the influence of literature, but ignores content semantics and interdisciplinary connections; Machine learning model: uses LDA topic model or Word2Vec word vector, but is limited by shallow feature representation capabilities.
[0004] Therefore, existing literature retrieval methods are not only inefficient but also prone to missing important information. They require manual intervention when determining retrieval strategies and adjusting results, which may lead to human errors. They also lack intelligence and cannot make personalized recommendations based on user retrieval history and behavior patterns. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a document information retrieval and analysis system and method based on AI intelligence, which comprehensively improves the accuracy, intelligence and dynamic adaptability of document retrieval and can deeply meet user needs.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: An AI-based literature information retrieval and analysis system, including: The representation vector output module is used to extract the text features of the document to be retrieved, build a multimodal document representation model, and output the document representation vector; The preliminary search result output module is used to perform semantic enhanced search, expand user queries, and output the preliminary search result set and the corresponding similarity score set; An update retrieval result module is used to optimize the preliminary retrieval result set based on the similarity score to obtain an updated retrieval result set; Dynamic knowledge graph construction module, used to build a dynamic knowledge graph based on the updated search result set; Enhanced document representation output module, which is used to perform cross-disciplinary association analysis based on dynamic knowledge graphs and output enhanced document representations by combining document representation vectors; The recommendation list output module is used to obtain user portrait data and combine it with enhanced document representation and dynamic knowledge graph to output a personalized recommendation list; The knowledge graph feedback update module is used to collect user interaction flow data and output the updated knowledge graph in combination with the personalized recommendation list.
[0007] Preferably, in the representation vector output module, the process of outputting the document representation vector is: Use the BERT model to extract [CLS] token vectors for document titles and abstracts , expressed as: ; Where, To input the BERT model into the document title and abstract text content, The original output vector after extracting the [CLS] tag vector of the document title and abstract using the BERT model is denoted as the BERT output vector; Output vector for BERT Perform whitening processing, the steps include: Step 1: Calculate the mean: ; Where i is the sample number, i=1,2,3,...,N, N is the number of samples, for The i-th element in For the calculated mean of a vector; Step 2: Centralized processing: ; Where, For The vector after centralization; Step 3: Eigenvalue decomposition covariance matrix: ; Where, is the covariance matrix, is the eigenvector matrix, is the diagonal eigenvalue matrix, Represents the transpose of a vector; Step 4: Dimensionality reduction: ; Where, is the text feature vector after dimensionality reduction; Extract visual features and use the CLIP model to analyze graphs in the literature and formula Encode and obtain the visual feature vector : ; Integrate metadata features and extract metadata features M, including author index H, journal impact factor I and citation network weight C; ; Output document representation vector by weighted fusion of text features, visual features and metadata features : ; in, is the ELU activation function, 、 、 is a trainable weight matrix, which is learned through the training dataset to minimize the loss function of the retrieval task.
[0008] Preferably, in the preliminary search result output module, the process of outputting the preliminary search result set and the corresponding similarity score set is: Based on the term correlation matrix S and combined with the TF-IDF value, the cosine similarity of BERT embedding is calculated, which is expressed as: ; in, is an element in the term association matrix S, representing the mth query term and the nth query word The correlation between For query words and The cosine similarity between the embedding vectors obtained after processing by the BERT model, For query words The term frequency-inverse document frequency value in the document set D, For query words The term frequency-inverse document frequency value in the document set D; Based on the term relevance matrix S, the top five terms with the highest relevance are selected to expand the original query; Introduce an abnormal term filtering mechanism and calculate the frequency deviation of the query word q : ; when > , remove the query term q from the expanded query. is the standard deviation; For each query word q, find the document d with the largest dot product, and use the ColBERT architecture to calculate the similarity score between the query statement Q and the document set D : ; In the formula, Q is the query statement, which is the search content input by the user, and D is the document collection, that is, the document library to be searched. and Both are embedding functions used to convert query words q and document d into vector representations. Represents the transpose of a vector; According to the output similarity score, the corresponding documents are used as preliminary search results; Output the preliminary search result set and the corresponding similarity score set.
[0009] Preferably, in the update retrieval result module, the process of obtaining the updated retrieval result set is: Constructing a credibility assessment matrix: ; Where, is the credibility index of the document, is the number of citations of the document, is the author ranking based on the author H-index level, is the journal's impact factor, For the setting The weight factor, For the setting The weight factor, For the setting The weight factor, e is a natural constant; Dynamically modify the similarity score based on the document credibility index: ; Where, is the modified similarity score between statement Q and document set D, which serves as the retrieval score. is the activation function; when When the value is less than 0.6 and the article is from a non-core journal, it will be removed from the results; Get the updated search result set and the corresponding search score.
[0010] Preferably, in the dynamic knowledge graph construction module, the process of constructing the dynamic knowledge graph is: Establishing entity nodes, including subject entity nodes and object entity nodes; Based on the updated retrieval result set, the improved OpenNRE model is used to combine BiLSTM and contextual attention features to calculate the entity relationship probability. : ; In the formula, s is the subject entity corresponding to the subject entity node, o is the object entity corresponding to the object entity node, r is the relationship type between s and o, c is the context attention feature, is the weight matrix associated with relation r; If the entity relationship probability If the probability is greater than the set threshold, a graph edge is established between s and o, and an initial weight is assigned to the graph edge. If the entity relationship probability If it is not greater than the set probability threshold, no graph edge will be established between s and o; Apply temporal decay to the initial weights of graph edges: ; Where, is the weight of the graph edge at time t, is the initial weight of the graph edge, that is, the time is The weight of is the attenuation coefficient, the value is 0.05, and e is a natural constant; output is a dynamic knowledge graph.
[0011] Preferably, the process of acquiring the contextual attention feature c is: Get BiLSTM hidden state sequence , a=1,2,3,...,z, z is the number of elements in the hidden state sequence, is the hidden state vector of the ath element in the hidden state sequence; Calculate the attention weight of the ath element in the hidden state sequence : ; Where, is the query vector, is the dimension of the vector; Generate contextual attention features: .
[0012] Preferably, in the enhanced document representation output module, the process of outputting the enhanced document representation is: Input the dynamic knowledge graph and extract the subject label distribution matrix: ; Where, is an element in the subject label distribution matrix, used to measure the subject and disciplines The relationship between For the subject and disciplines The co-occurrence frequency of For the subject The information entropy of , g and h are both subject numbers; Computing interdisciplinary : ; Where, is the PageRank value of the entity in document d, which is used as the weight when calculating the interdisciplinary degree; Generate enhanced document representation: ; Where, To enhance the literature representation, For the original literature representation, is the interdisciplinary projection matrix, is an adjustable coefficient.
[0013] Preferably, in the recommendation list output module, the process of outputting the personalized recommendation list is: Construct a user feature vector U, including search behavior, reading history, citation pattern, collaboration network, field preference, and freshness requirement: ; Where, is the weight factor of the search behavior set, is the weight factor of the set reading record, is the weight factor of the set citation mode, is the weight factor of the set cooperation network, is the weight factor of the set field preference, is the weight factor of the set freshness requirement; The recommendation score is obtained by combining the similarity between the user feature vector and the document representation vector, as well as the similarity between similar user sets: ; Where, is the recommendation score of user u for document d, is the feature vector of user u, is the enhanced document representation of document d, is the user feature vector Enhanced document representation with document d The similarity between is a set of similar users, v is one of the similar users, is the feature vector of similar user v Enhanced document representation with document d The similarity between Similar user groups The cardinality, that is, the number of users in the set; According to the calculated Sort all the literature; Select the top x documents with the highest scores as a personalized recommendation list; Output a personalized recommendation list.
[0014] Preferably, in the knowledge graph feedback update module, the process of outputting the updated knowledge graph is: Capture user interaction behaviors in real time and construct a user interaction behavior set F: The user interaction behavior set F includes click, dwell_time, download, and citation, which can be expressed as: ; Calculating behavior weights : ; Where, is the indicator function; Analyze and update entity influence weights : ; Where, is the old value of the entity's influence weight; When the node weight change rate is higher than 15%, the local graph topology is rebuilt. After the local graph topology reconstruction is completed, it is integrated into the global knowledge graph and the updated knowledge graph is output.
[0015] An AI-based literature information retrieval and analysis method, based on the above system, includes the following steps: S1. Extract the text features of the document to be retrieved, build a multimodal document representation model, and output the document representation vector; S2. Perform semantic enhanced retrieval, expand the user query, and output a preliminary retrieval result set and a corresponding similarity score set; S3. Based on the similarity score, the preliminary search result set is optimized to obtain an updated search result set; S4. Build a dynamic knowledge graph based on the updated search result set; S5. Perform interdisciplinary association analysis based on dynamic knowledge graphs, combine document representation vectors, and output enhanced document representations. S6. Obtain user profile data and combine it with enhanced document representation and dynamic knowledge graph to output a personalized recommendation list; S7. Collect user interaction flow data, combine it with the personalized recommendation list, and output the updated knowledge graph.
[0016] The present invention has the following beneficial effects: This invention provides an AI-based document information retrieval and analysis system and method. The representation vector output module extracts document text features, and the multimodal document representation model comprehensively characterizes the document. The preliminary search result output module performs semantically enhanced retrieval and query expansion, improving retrieval accuracy and enabling users to access more relevant documents. The knowledge graph feedback update module updates the knowledge graph based on user interaction stream data, enabling the system to continuously learn and optimize, adapting to new documents and changing user needs, and maintaining the accuracy and effectiveness of retrieval and analysis.
[0017] This invention optimizes preliminary results by updating the search results module, eliminating low-quality or irrelevant documents, reducing user screening time and effort and quickly focusing on valuable content. The dynamic knowledge graph construction module constructs a graph based on the search results, displaying entities and relationships within the documents. Interdisciplinary association analysis explores interdisciplinary knowledge connections and provides new insights. The recommendation list output module combines enhanced document representation, knowledge graphs, and user profiles to provide personalized recommendations, meeting the needs and preferences of different users and discovering potentially useful documents. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a schematic diagram of system module connections of the present invention; Figure 2 Schematic diagram of the method steps of the present invention; Figure 3 It is a logic block diagram of the system operation of the present invention. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0020] Example 1: Figure 1 As shown, an AI-based literature information retrieval and analysis system includes: The representation vector output module is used to extract the text features of the document to be retrieved, build a multimodal document representation model, and output the document representation vector; The preliminary search result output module is used to perform semantic enhanced search, expand user queries, and output the preliminary search result set and the corresponding similarity score set; An update retrieval result module is used to optimize the preliminary retrieval result set based on the similarity score to obtain an updated retrieval result set; Dynamic knowledge graph construction module, used to build a dynamic knowledge graph based on the updated search result set; Enhanced document representation output module, which is used to perform cross-disciplinary association analysis based on dynamic knowledge graphs and output enhanced document representations by combining document representation vectors; The recommendation list output module is used to obtain user portrait data and combine it with enhanced document representation and dynamic knowledge graph to output a personalized recommendation list; The knowledge graph feedback update module is used to collect user interaction flow data and output the updated knowledge graph in combination with the personalized recommendation list.
[0021] In the representation vector output module, the process of outputting the document representation vector is as follows: Use the BERT model to extract [CLS] token vectors for document titles and abstracts , expressed as: ; Where, To input the BERT model into the document title and abstract text content, The original output vector after extracting the [CLS] tag vector of the document title and abstract using the BERT model is denoted as the BERT output vector; Output vector for BERT Perform whitening processing, the steps include: Step 1: Calculate the mean: ; Where i is the sample number, i=1,2,3,...,N, N is the number of samples, for The i-th element in For the calculated mean of a vector; Step 2: Centralized processing: ; Where, For The vector after centralization; Step 3: Eigenvalue decomposition covariance matrix: ; Where, is the covariance matrix, is the eigenvector matrix, is the diagonal eigenvalue matrix, Represents the transpose of a vector; Step 4: Dimensionality reduction: ; Where, is the text feature vector after dimensionality reduction; Extract visual features and use the CLIP model to analyze graphs in the literature and formula Encode and obtain the visual feature vector : ; Integrate metadata features and extract metadata features M, including author index H, journal impact factor I and citation network weight C; ; Output document representation vector by weighted fusion of text features, visual features and metadata features : ; in, is the ELU activation function, 、 、 is a trainable weight matrix, which is learned through the training dataset to minimize the loss function of the retrieval task.
[0022] Combining text, visual, and metadata features, this multimodal approach comprehensively characterizes documents. Text feature extraction utilizes the BERT model to capture semantic information, visual features are encoded using the CLIP model to create graphical formulas, and metadata features encompass key information such as author and journal. The resulting document representation vector is more complete and accurate.
[0023] The BERT output vector is whitened (calculating the mean, centering, eigenvalue decomposition and dimensionality reduction) to reduce the data dimension, remove redundancy and correlation, improve computational efficiency, highlight key text features, and enhance the accuracy of subsequent retrieval analysis. It should be noted that a cross-modal alignment layer (such as CCA or cross-modal attention) needs to be introduced to normalize and project the features into a unified semantic space.
[0024] When weighted fusion of different features, the trainable weight matrix can be used to learn the training data set, minimize the loss function according to the retrieval task, flexibly adjust the weight of each feature, adapt to different retrieval analysis requirements, and improve system performance.
[0025] All parameters in this embodiment are dimensionless, and the weight factors in this embodiment are obtained based on data stored in the database.
[0026] In the preliminary search result output module, the process of outputting the preliminary search result set and the corresponding similarity score set is as follows: Based on the term correlation matrix S and combined with the TF-IDF value, the cosine similarity of BERT embedding is calculated, which is expressed as: ; in, is an element in the term association matrix S, representing the mth query term and the nth query word The correlation between For query words and The cosine similarity between the embedding vectors obtained after processing by the BERT model, For query words The term frequency-inverse document frequency value in the document set D, For query words The term frequency-inverse document frequency value in the document set D; Based on the term relevance matrix S, the top five terms with the highest relevance are selected to expand the original query; Introduce an abnormal term filtering mechanism and calculate the frequency deviation of the query word q : ; when > , remove the query term q from the expanded query. is the standard deviation; For each query word q, find the document d with the largest dot product, and use the ColBERT architecture to calculate the similarity score between the query statement Q and the document set D : ; In the formula, Q is the query statement, which is the search content input by the user, and D is the document collection, that is, the document library to be searched. and Both are embedding functions used to convert query words q and document d into vector representations. Represents the transpose of a vector; According to the output similarity score, the corresponding documents are used as preliminary search results; Output the preliminary search result set and the corresponding similarity score set.
[0027] Based on the term relevance matrix, combined with the cosine similarity and TF-IDF value of BERT embedding, it can more accurately measure the relevance between query terms, select the top 5 related terms to expand the original query, supplement and refine the user's query content, improve the comprehensiveness and accuracy of the query, and help users find documents that better meet their needs.
[0028] An abnormal term filtering mechanism is introduced to remove abnormal terms by calculating the frequency deviation of query terms, avoiding irrelevant or low-frequency abnormal terms from interfering with retrieval results, purifying query content and making retrieval results more relevant and reliable.
[0029] The ColBERT architecture is used to calculate the similarity score between the query statement and the document collection. This takes into account the vector relationship between the query terms and the words in the documents, and can measure the degree of match between the two at a fine-grained level. Compared with simple matching methods, it can more accurately find documents that are highly relevant to the query, thereby improving the accuracy and quality of retrieval.
[0030] In the update retrieval result module, the process of obtaining the updated retrieval result set is as follows: Constructing a credibility assessment matrix: ; Where, is the credibility index of the document, is the number of citations of the document, is the author ranking based on the author H-index level, is the journal's impact factor, For the setting The weight factor, For the setting The weight factor, For the setting The weight factor, e is a natural constant; Dynamically modify the similarity score based on the document credibility index: ; Where, is the modified similarity score between statement Q and document set D, which serves as the retrieval score. is the activation function; when When the value is less than 0.6 and the article is from a non-core journal, it will be removed from the results; Get the updated search result set and the corresponding search score.
[0031] By constructing a credibility assessment matrix and comprehensively considering multi-dimensional factors such as the number of citations, author rankings, and journal impact factors of the document, we can more comprehensively and scientifically evaluate the credibility of each document and provide a reliable basis for subsequent processing.
[0032] Dynamically modifying search scores based on document credibility indicators, incorporating the quality of the document itself into the search scoring system. Compared with scoring based solely on similarity, this can more reasonably reflect the match between the document and user needs, improving the quality of search results.
[0033] By setting clear screening conditions, that is, removing documents from the results when their credibility index is low and they come from non-core journals, we can effectively filter out possible low-quality or unreliable documents, making the final updated search result set more accurate and valuable, saving users time and energy in screening documents.
[0034] In the dynamic knowledge graph construction module, the process of building a dynamic knowledge graph is as follows: Establishing entity nodes, including subject entity nodes and object entity nodes; Based on the updated retrieval result set, the improved OpenNRE model is used to combine BiLSTM and contextual attention features to calculate the entity relationship probability. : ; In the formula, s is the subject entity corresponding to the subject entity node, o is the object entity corresponding to the object entity node, r is the relationship type between s and o, c is the context attention feature, is the weight matrix associated with relation r; If the entity relationship probability If the probability is greater than the set threshold, a graph edge is established between s and o, and an initial weight is assigned to the graph edge. If the entity relationship probability If it is not greater than the set probability threshold, no graph edge will be established between s and o; Apply temporal decay to the initial weights of graph edges: ; Where, is the weight of the graph edge at time t, is the initial weight of the graph edge, that is, the time is The weight of is the attenuation coefficient, the value is 0.05, and e is a natural constant; output is a dynamic knowledge graph.
[0035] The process of obtaining the contextual attention feature c is as follows: Get BiLSTM hidden state sequence , a=1,2,3,...,z, z is the number of elements in the hidden state sequence, is the hidden state vector of the ath element in the hidden state sequence; Calculate the attention weight of the ath element in the hidden state sequence : ; Where, is the query vector, is the dimension of the vector; Generate contextual attention features: .
[0036] With the help of the improved OpenNRE model, the BiLSTM and contextual attention features are integrated to calculate the entity relationship probability, which can deeply mine the semantic information in the text, more accurately judge the relationship type between the subject entity and the object entity in the knowledge graph, and improve the construction quality of the knowledge graph.
[0037] Applying time decay to graph edge weights automatically adjusts edge weights over time to reflect the timeliness of knowledge. This effectively reduces the impact of outdated information, allowing the knowledge graph to better reflect the dynamic changes in knowledge and ensure that the content of the knowledge graph is consistent with current reality.
[0038] By obtaining the BiLSTM hidden state sequence and calculating the attention weights to generate contextual attention features, the system fully considers the contextual information of the text and enhances the understanding of the text's semantics. This helps to more accurately capture the semantic relationships between entities and build a more structured and semantically rich knowledge graph.
[0039] In the enhanced document representation output module, the process of outputting enhanced document representation is as follows: Input the dynamic knowledge graph and extract the subject label distribution matrix: ; Where, is an element in the subject label distribution matrix, used to measure the subject and disciplines The relationship between For the subject and disciplines The co-occurrence frequency of For the subject The information entropy of , g and h are both subject numbers; Computing interdisciplinary : ; Where, is the PageRank value of the entity in document d, which is used as the weight when calculating the interdisciplinary degree; Generate enhanced document representation: ; Where, To enhance the literature representation, For the original literature representation, is the interdisciplinary projection matrix, is an adjustable coefficient.
[0040] Interdisciplinary association analysis based on dynamic knowledge graphs can explore potential connections between different disciplines and discover the integration points of interdisciplinary knowledge by extracting the discipline label distribution matrix and calculating the discipline intersection. This will help scientific researchers conduct interdisciplinary research and expand their research horizons.
[0041] Integrating interdisciplinary information into document representation and combining it with the original document representation vector to generate enhanced document representation makes the document representation more comprehensive and rich. It not only includes the content characteristics of the document itself, but also reflects its association information with other disciplines, providing more valuable information for subsequent document retrieval, recommendation and other tasks.
[0042] The setting of the adjustable coefficient enables the module to flexibly adjust the influence of interdisciplinary information on enhanced literature representation according to different application scenarios and needs, thereby enhancing the adaptability and versatility of the module.
[0043] In the recommendation list output module, the process of outputting a personalized recommendation list is as follows: Construct a user feature vector U, including search behavior, reading history, citation pattern, collaboration network, field preference, and freshness requirement: ; Where, is the weight factor of the search behavior set, is the weight factor of the set reading record, is the weight factor of the set citation mode, is the weight factor of the set cooperation network, is the weight factor of the set field preference, is the weight factor of the set freshness requirement; Among them, search behavior is the number of search keywords used by the user in a period of time, reading history is the total number of documents read by the user, citation pattern is the number of documents cited by the user in the content created, collaboration network is the number of different fields involved in the collaboration, field preference is the maximum number of searches in each field by the user, and freshness demand is the proportion of new documents read by the user; The recommendation score is obtained by combining the similarity between the user feature vector and the document representation vector, as well as the similarity between similar user sets: ; Where, is the recommendation score of user u for document d, is the feature vector of user u, is the enhanced document representation of document d, is the user feature vector Enhanced document representation with document d The similarity between is a set of similar users, v is one of the similar users, is the feature vector of similar user v Enhanced document representation with document d The similarity between Similar user groups The cardinality, that is, the number of users in the set; According to the calculated Sort all the literature; Select the top x documents with the highest scores as a personalized recommendation list; Output a personalized recommendation list.
[0044] Constructing a user feature vector with multiple dimensions including search behavior, reading records, etc., and setting a weight factor for each dimension, it can comprehensively and accurately characterize the user's behavior patterns, interest preferences and demand characteristics, laying a solid foundation for personalized recommendations.
[0045] The recommendation score is calculated by combining the similarity between the user feature vector and the document enhanced representation vector, as well as the similarity between the set of similar users and the document. This not only takes into account the matching degree between the user and the document, but also draws on the preferences of similar users to make the recommendation results more in line with the user's actual needs. It should be noted that when U sparseness is less than the threshold, the document features are directly matched with the query similarity, that is, content-based recommendation.
[0046] In the knowledge graph feedback update module, the process of outputting the updated knowledge graph is as follows: Capture user interaction behaviors in real time and construct a user interaction behavior set F: The user interaction behavior set F includes click, dwell_time, download, and citation, which can be expressed as: ; Calculating behavior weights : ; Where, is the indicator function; Analyze and update entity influence weights : ; Where, is the old value of the entity's influence weight; When the node weight change rate is higher than 15%, the local graph topology is rebuilt. After the local graph topology reconstruction is completed, it is integrated into the global knowledge graph and the updated knowledge graph is output.
[0047] Capture user interactions in personalized recommendation lists in real time, calculate weights based on clicks, dwell time, and other behaviors, and use this to update entity influence weights. This approach closely integrates actual user feedback, making the knowledge graph more tailored to user needs and interests, and improving its practicality for users.
[0048] When the rate of change of node weights is high, the local graph topology is rebuilt and integrated into the global knowledge graph to promptly reflect the dynamic changes in knowledge and user interests. The knowledge graph structure and content are continuously optimized to maintain timeliness and accuracy, adapting to the ever-evolving information environment.
[0049] Adjusting entity influence weights based on user interaction behavior helps more accurately reflect the relationships and importance between knowledge. This makes the knowledge relationships in the knowledge graph more realistic, providing a more reliable knowledge foundation for subsequent tasks such as literature retrieval and recommendation.
[0050] Combined with the above description, the system operation logic block diagram of the present invention is as follows Figure 3 shown.
[0051] Compared with traditional retrieval and analysis systems, the advantages of this application are summarized in Table 1.
[0052] Table 1 Summary of system advantages
[0053] Example 2: Based on Example 1, after the updated knowledge graph is output based on the personalized recommendation list and user interaction flow data is collected, the updated knowledge graph is fed back to the enhanced document representation output module, and cross-disciplinary association analysis is performed again based on the updated knowledge graph. The enhanced document representation is output in combination with the output document representation vector, and the recommendation list output module and the knowledge graph feedback update module are continued to be executed until the knowledge graph feedback update module reaches the maximum number of iterations, and the knowledge graph updated after the last iteration is output and stored.
[0054] Example 3: Figure 2 As shown, a literature information retrieval and analysis method based on AI intelligence, based on the system of Example 1, includes the following steps: S1. Extract the text features of the document to be retrieved, build a multimodal document representation model, and output the document representation vector; S2. Perform semantic enhanced retrieval, expand the user query, and output a preliminary retrieval result set and a corresponding similarity score set; S3. Based on the similarity score, the preliminary search result set is optimized to obtain an updated search result set; S4. Build a dynamic knowledge graph based on the updated search result set; S5. Perform interdisciplinary association analysis based on dynamic knowledge graphs, combine document representation vectors, and output enhanced document representations. S6. Obtain user profile data and combine it with enhanced document representation and dynamic knowledge graph to output a personalized recommendation list; S7. Collect user interaction flow data, combine it with the personalized recommendation list, and output the updated knowledge graph.
Claims
1. A literature information retrieval and analysis system based on AI intelligence, characterized by: include: The representation vector output module is used to extract the text features of the document to be retrieved, build a multimodal document representation model, and output the document representation vector; The preliminary search result output module is used to perform semantic enhanced search, expand user queries, and output the preliminary search result set and the corresponding similarity score set; An update retrieval result module is used to optimize the preliminary retrieval result set based on the similarity score to obtain an updated retrieval result set; Dynamic knowledge graph construction module, used to build a dynamic knowledge graph based on the updated search result set; Enhanced document representation output module, which is used to perform cross-disciplinary association analysis based on dynamic knowledge graphs and output enhanced document representations by combining document representation vectors; The recommendation list output module is used to obtain user portrait data and combine it with enhanced document representation and dynamic knowledge graph to output a personalized recommendation list; The knowledge graph feedback update module is used to collect user interaction flow data and output the updated knowledge graph in combination with the personalized recommendation list.
2. The AI-based document information retrieval and analysis system according to claim 1, characterized in that: In the representation vector output module, the process of outputting the document representation vector is as follows: Use the BERT model to extract [CLS] token vectors for document titles and abstracts , expressed as: ; Where, To input the BERT model into the document title and abstract text content, The original output vector after extracting the [CLS] tag vector of the document title and abstract using the BERT model is denoted as the BERT output vector; Output vector for BERT Perform whitening processing, the steps include: Step 1: Calculate the mean: ; Where i is the sample number, i=1,2,3,...,N, N is the number of samples, for The i-th element in For the calculated mean of a vector; Step 2: Centralized processing: ; Where, For The vector after centralization; Step 3: Eigenvalue decomposition covariance matrix: ; Where, is the covariance matrix, is the eigenvector matrix, is the diagonal eigenvalue matrix, Represents the transpose of a vector; Step 4: Dimensionality reduction: ; Where, is the text feature vector after dimensionality reduction; Extract visual features and use the CLIP model to analyze graphs in the literature and formula Encode and obtain the visual feature vector : ; Integrate metadata features and extract metadata features M, including author index H, journal impact factor I and citation network weight C; ; Output document representation vector by weighted fusion of text features, visual features and metadata features : ; in, is the ELU activation function, 、 、 is a trainable weight matrix, which is learned through the training dataset to minimize the loss function of the retrieval task.
3. The AI-based document information retrieval and analysis system according to claim 1, characterized in that: In the preliminary search result output module, the process of outputting the preliminary search result set and the corresponding similarity score set is as follows: Based on the term correlation matrix S and combined with the TF-IDF value, the cosine similarity of BERT embedding is calculated, which is expressed as: ; in, is an element in the term association matrix S, representing the mth query term and the nth query word The correlation between For query words and The cosine similarity between the embedding vectors obtained after processing by the BERT model, For query words The term frequency-inverse document frequency value in the document set D, For query words The term frequency-inverse document frequency value in the document set D; Based on the term relevance matrix S, the top five terms with the highest relevance are selected to expand the original query; Introduce an abnormal term filtering mechanism and calculate the frequency deviation of the query word q : ; when > , remove the query term q from the expanded query. is the standard deviation; For each query word q, find the document d with the largest dot product, and use the ColBERT architecture to calculate the similarity score between the query statement Q and the document set D : ; In the formula, Q is the query statement, which is the search content input by the user, and D is the document collection, that is, the document library to be searched. and Both are embedding functions used to convert query words q and document d into vector representations. Represents the transpose of a vector; According to the output similarity score, the corresponding documents are used as preliminary search results; Output the preliminary search result set and the corresponding similarity score set.
4. The AI-based document information retrieval and analysis system according to claim 3, characterized in that: In the update retrieval result module, the process of obtaining the updated retrieval result set is as follows: Constructing a credibility assessment matrix: ; Where, is the credibility index of the document, is the number of citations of the document, is the author ranking based on the author H-index level, is the journal's impact factor, For the setting The weight factor, For the setting The weight factor, For the setting The weight factor, e is a natural constant; Dynamically modify the similarity score based on the document credibility index: ; Where, is the modified similarity score between statement Q and document set D, which serves as the retrieval score. is the activation function; when When the value is less than 0.6 and the article is from a non-core journal, it will be removed from the results; Get the updated search result set and the corresponding search score.
5. The AI-based document information retrieval and analysis system according to claim 1, characterized in that: In the dynamic knowledge graph construction module, the process of constructing the dynamic knowledge graph is as follows: Establishing entity nodes, including subject entity nodes and object entity nodes; Based on the updated retrieval result set, the improved OpenNRE model is used to combine BiLSTM and contextual attention features to calculate the entity relationship probability. : ; In the formula, s is the subject entity corresponding to the subject entity node, o is the object entity corresponding to the object entity node, r is the relationship type between s and o, c is the context attention feature, is the weight matrix associated with relation r; If the entity relationship probability If the probability is greater than the set threshold, a graph edge is established between s and o, and an initial weight is assigned to the graph edge. If the entity relationship probability If it is not greater than the set probability threshold, no graph edge will be established between s and o; Apply temporal decay to the initial weights of graph edges: ; Where, is the weight of the graph edge at time t, is the initial weight of the graph edge, that is, the time is The weight of is the attenuation coefficient, the value is 0.05, and e is a natural constant; output is a dynamic knowledge graph.
6. The AI-based document information retrieval and analysis system according to claim 5, characterized in that: The acquisition process of the contextual attention feature c is: Get BiLSTM hidden state sequence , a=1,2,3,...,z, z is the number of elements in the hidden state sequence, is the hidden state vector of the ath element in the hidden state sequence; Calculate the attention weight of the ath element in the hidden state sequence : ; Where, is the query vector, is the dimension of the vector; Generate contextual attention features: 。 7. The AI-based document information retrieval and analysis system according to claim 1, characterized in that: In the enhanced document representation output module, the process of outputting the enhanced document representation is as follows: Input the dynamic knowledge graph and extract the subject label distribution matrix: ; Where, is an element in the subject label distribution matrix, used to measure the subject and disciplines The relationship between For the subject and disciplines The co-occurrence frequency of For the subject The information entropy of , g and h are both subject numbers; Computing interdisciplinary : ; Where, is the PageRank value of the entity in document d, which is used as the weight when calculating the interdisciplinary degree; Generate enhanced document representation: ; Where, To enhance the literature representation, For the original literature representation, is the interdisciplinary projection matrix, is an adjustable coefficient.
8. The AI-based document information retrieval and analysis system according to claim 1, characterized in that: In the recommendation list output module, the process of outputting the personalized recommendation list is as follows: Construct a user feature vector U, including search behavior, reading history, citation pattern, collaboration network, field preference, and freshness requirement: ; Where, is the weight factor of the search behavior set, is the weight factor of the set reading record, is the weight factor of the set citation mode, is the weight factor of the set cooperation network, is the weight factor of the set field preference, is the weight factor of the set freshness requirement; The recommendation score is obtained by combining the similarity between the user feature vector and the document representation vector, as well as the similarity between similar user sets: ; Where, is the recommendation score of user u for document d, is the feature vector of user u, is the enhanced document representation of document d, is the user feature vector Enhanced document representation with document d The similarity between is a set of similar users, v is one of the similar users, is the feature vector of similar user v Enhanced document representation with document d The similarity between Similar user groups The cardinality, that is, the number of users in the set; According to the calculated Sort all the literature; Select the top x documents with the highest scores as a personalized recommendation list; Output a personalized recommendation list.
9. The AI-based document information retrieval and analysis system according to claim 1, characterized in that: In the knowledge graph feedback update module, the process of outputting the updated knowledge graph is as follows: Capture user interaction behaviors in real time and construct a user interaction behavior set F: The user interaction behavior set F includes click, dwell_time, download, and citation, which can be expressed as: ; Calculating behavior weights : ; Where, is the indicator function; Analyze and update entity influence weights : ; Where, is the old value of the entity's influence weight; When the node weight change rate is higher than 15%, the local graph topology is rebuilt. After the local graph topology reconstruction is completed, it is integrated into the global knowledge graph and the updated knowledge graph is output.
10. A method for literature information retrieval and analysis based on AI intelligence, based on the system according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1. Extract the text features of the document to be retrieved, build a multimodal document representation model, and output the document representation vector; S2. Perform semantic enhanced retrieval, expand the user query, and output a preliminary retrieval result set and a corresponding similarity score set; S3. Based on the similarity score, the preliminary search result set is optimized to obtain an updated search result set; S4. Build a dynamic knowledge graph based on the updated search result set; S5. Perform interdisciplinary association analysis based on dynamic knowledge graphs, combine document representation vectors, and output enhanced document representations. S6. Obtain user profile data and combine it with enhanced document representation and dynamic knowledge graph to output a personalized recommendation list; S7. Collect user interaction flow data, combine it with the personalized recommendation list, and output the updated knowledge graph.
Citation Information
Patent Citations
Semantic similarity vector re-sparse coding indexing and retrieval method
CN114860868A
Vector database-based question and answer reasoning method and device
CN118839005A
Construction method and device of knowledge base question-answering system, equipment and storage medium
CN119293164A
Application method for realizing travel-related enterprise assistant based on large model
CN119961527A
Easy moving sample collection and analysis device
KR1020220016725A
Cited By
Historical literature version knowledge ontology dynamic collaborative construction method and system
CN120893547A
Retrieval and tracking method and device for scientific research information
CN121233697A
Intelligent retrieval recommendation method and system for cross-modal data
CN121681934A
Knowledge service platform thematic set construction method
CN122021842A