A heuristic knowledge navigation recommendation method integrating user search intention
By integrating the heuristic knowledge navigation recommendation method of user search intentions, combining the user intention understanding model and literature correlation search model, the problem that existing recommendation methods cannot accurately understand user intentions is solved, efficient and personalized knowledge recommendations are achieved, and scientific research efficiency is improved.
Patent Information
- Application Number
- CN202410998133.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-24
AI Technical Summary
The existing recommendation methods cannot accurately understand the actual intentions of users, resulting in most recommendation results staying at the level of articles or journals, which makes users still need to manually screen and reduce scientific research efficiency.
It provides a heuristic knowledge navigation recommendation method that integrates user search intentions. It can obtain real-time search text through interaction, analyzes the input text to obtain real-time search types, pre-constructs the user's intention understanding model and literature association search model, combines user intentions to conduct accurate literature search, perform fine-grained knowledge extraction, and filter out the associated knowledge element set that matches the user's intention.
It enables users to directly obtain the information they really need without manually screening a large amount of literature data, improves scientific research efficiency, and ensures the accuracy and personalization of recommended results.
Smart Images

Figure CN118939787B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, specifically to the field of knowledge recommendation technology, and in particular to a heuristic knowledge navigation recommendation method integrating user search intent. Background Art
[0002] With the rapid development of information technology and the increasing complexity of the scientific research environment, researchers are facing unprecedented challenges in efficiently and accurately obtaining the required information from massive knowledge resources. Although traditional search recommendation algorithms have made significant progress in text matching and relevance ranking, they generally have the problem of focusing on matching rather than intent, that is, they rely too much on the text content input by users, while ignoring the real purpose and deep-seated needs behind the user's query. Although this recommendation method guided by literal similarity can meet basic search needs to a certain extent, it is difficult to accurately capture users' personalized preferences and potential needs, resulting in recommendation results that often deviate from user expectations. Summary of the invention
[0003] This application provides a heuristic knowledge navigation recommendation method that integrates user search intentions, aiming to solve the technical problem that existing recommendation methods are unable to accurately understand the user's actual intentions, resulting in most recommendation results remaining at the article or journal level, requiring users to still manually screen, reducing scientific research efficiency.
[0004] In view of the above problems, the present application provides a heuristic knowledge navigation recommendation method that integrates user search intent.
[0005] The present application provides a heuristic knowledge navigation recommendation method that integrates user search intent, and the method includes: interactively obtaining input text of a real-time search user, parsing the input text to obtain a real-time search type; pre-building a user intent understanding model; performing search intent recognition by inputting the real-time search type into the user intent understanding model, and outputting a user intent recognition result; pre-building a document association retrieval model; performing document association retrieval by synchronizing the input text and the user intent recognition result to the document association retrieval model to obtain a target retrieval result; performing fine-grained knowledge extraction on the target retrieval result to obtain a key knowledge element set; filtering and matching from the key knowledge element set according to the user intent recognition result to obtain a related knowledge element set; and feeding back the related knowledge element set to the real-time search user.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0007] The above-mentioned heuristic knowledge navigation recommendation method integrating user search intention receives the user's input text in real time and preliminarily determines the user's search type based on the text content. Subsequently, the user's input is analyzed using a pre-built user intention understanding model to identify the user's real purpose and intention of this search. After that, the user's input text and the identified intention are input into a pre-built document association search model. This document association search model not only analyzes the relevance of the text content, but also combines the user's intention to conduct a more accurate document search, thereby obtaining a set of target search results that highly match the user's query intention. Then, these target search results are analyzed to extract the key knowledge elements therein to form a key knowledge element set. Finally, according to the identified user intention, the knowledge elements most relevant to the user's intention are screened from the key knowledge element set to form a set of associated knowledge elements, and this set is fed back to the user. In this way, users can directly obtain the information they really need without having to manually screen and process a large amount of document data, thereby improving scientific research efficiency.
[0008] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0010] Figure 1 A schematic diagram of a flow chart of a heuristic knowledge navigation recommendation method integrating user search intention in one embodiment;
[0011] Figure 2 The present invention is a flowchart of obtaining real-time search types in a heuristic knowledge navigation recommendation method integrating user search intention in one embodiment. DETAILED DESCRIPTION
[0012] The embodiment of the present application provides a heuristic knowledge navigation recommendation method that integrates user search intent, thereby solving the technical problem that existing recommendation methods are unable to accurately understand the user's actual intent, resulting in most recommendation results remaining at the article or journal level, requiring users to still manually screen, thereby reducing scientific research efficiency.
[0013] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0014] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules that are not explicitly listed or inherent to these processes, methods, products or devices.
[0015] Examples, such as Figure 1 As shown, the present application provides a heuristic knowledge navigation recommendation method integrating user search intention, the method comprising:
[0016] The user's input text is interactively obtained for real-time retrieval, and the input text is parsed to obtain the real-time retrieval type.
[0017] In an embodiment of the present application, the system terminal first interacts with the user terminal in real time to obtain the text information input by the real-time retrieval user. Subsequently, the input text obtained is analyzed, and the type of retrieval the user wants to perform is understood and parsed through the constructed retrieval type recognition model. For example, keyword-based search, name-based search, etc. Through this parsing process, the system terminal can more accurately understand the user's retrieval needs and provide more targeted input for the subsequent user intent understanding model. For example, if the real-time retrieval type is a name, the user intent understanding model can focus more on learning the intention pattern behind the name query from the training data, so as to select the corresponding database and retrieval field when searching, narrow the search scope, and improve the retrieval efficiency.
[0018] Further, if Figure 2 As shown, the present application provides interactive acquisition of real-time retrieval of user input text, parsing the input text to obtain real-time retrieval type, and further includes:
[0019] Interactively obtain a preset search type set, and collect search text samples with the preset search type set as a constraint to obtain a search text sample set; perform entity annotation on multiple search text samples in the search text sample set to obtain multiple search type samples;
[0020] Preferably, the system terminal pre-sets a search type set according to business needs. This preset search type set clearly lists the search types that need to be identified and supported, including names, time, keywords, specific sentences, etc. Subsequently, the system terminal uses the preset search type set as a constraint to collect search text samples from the search log in a targeted manner to obtain a search text sample set. These samples can widely cover the preset search types to ensure that the search type recognition model trained subsequently can comprehensively and accurately identify various situations. Afterwards, the collected search text sample set is entity labeled, that is, the search text sample set is type identified using the preset search type set, and each search text sample is labeled with a corresponding search type label to obtain multiple search type samples. Through entity labeling, the search type information contained in each sample can be more clearly understood, providing strong support for subsequent analysis and modeling.
[0021] A retrieval type recognition model is constructed based on a conditional random field; the multiple retrieval text samples and the multiple retrieval type samples are used as training data to optimize the retrieval type recognition model until the retrieval type recognition accuracy of the retrieval type recognition model meets the preset requirements; the retrieval type recognition model is used to parse the input text to obtain the real-time retrieval type.
[0022] Preferably, conditional random field (CRF) is a powerful sequence labeling model suitable for handling label dependency problems in sequence data, for example, identifying different entity types in text. The system terminal uses conditional random field to initialize a retrieval type recognition model, including setting key components such as feature template, state transition matrix, feature weight and output layer configuration of the retrieval type recognition model. Subsequently, the system terminal divides multiple retrieval text samples and multiple retrieval type samples to obtain training set, validation set and test set. After that, each sample in the training set is input into the retrieval type recognition model to calculate the feature function scores corresponding to all possible label sequences. These scores are based on the features of the current position and the labels of the previous position. The state transition matrix and feature function scores are then used to calculate the path scores of all possible label sequences. The path score is the sum of all label transition probabilities and feature scores in the sequence. After obtaining the path score. By normalizing all path scores, the probability distribution of each label sequence is obtained. Then, the cross entropy loss function is used to calculate the loss between the probability distribution predicted by the model and the retrieval type sample. The gradient of the loss function with respect to the model parameters, such as state transition probability and feature weight, is calculated by the back propagation algorithm. Repeat the forward propagation and back propagation process until the performance on the validation set is no longer significantly improved. To prevent overfitting, the performance of the model can be monitored on the validation set. If the performance of the model on the validation set begins to decline, stop training in advance. After the training is completed, the system terminal uses the test set to evaluate the accuracy of the model. When the accuracy of the model meets the preset requirements, the current retrieval type recognition model is output. Conversely, according to the performance of the retrieval type recognition model on the validation set and the test set, adjust the learning rate, regularization term, batch size and other hyperparameters, or adjust the retrieval type recognition model structure, such as adding hidden layers, changing activation functions, etc. After obtaining the retrieval type recognition model, the system terminal inputs the input text into the retrieval type recognition model. The retrieval type recognition model will use its learned knowledge to automatically identify the real-time retrieval types in the text, such as names, time, keywords, etc. This process is automated and can respond quickly, providing users with instant retrieval type recognition services.
[0023] Pre-built user intent understanding models.
[0024] In one embodiment, when the amount of historical data of the real-time retrieval user reaches the amount that can analyze the user search rules, that is, when the amount of historical data meets the preset data volume threshold, the system terminal obtains the real-time retrieval user's historical data from the system log, and based on these historical data, pre-builds a user intention understanding model based on RNN, which is intended to automatically parse and analyze the deep meaning of the user input to identify the user's true intention. When the amount of historical data of the real-time retrieval user does not meet the preset data volume threshold, the system terminal performs search rule analysis on the real-time retrieval user's historical data, that is, counts the number of times each keyword appears, and identifies multiple search topics that the user is most concerned about. Subsequently, the system terminal retrieves users with these search topics from the user preference database based on the identified multiple search topics to form a similar user set. Afterwards, the system terminal inputs the real-time retrieval user and the similar user set into the user similarity model, and the user similarity model can match a most similar user based on the learned knowledge. Then, the system terminal migrates the historical data of this most similar user to build a user intention understanding model for the real-time retrieval user.
[0025] For the user similarity model, the system terminal first obtains a sample similarity data set of multiple pairs of sample similar users. This similarity data set includes multiple groups of similar data, each group of similar data includes the first sample historical data, the second sample historical data and the sample similarity, and the first sample historical data and the second sample historical data correspond to each pair of sample similar users. Subsequently, the system terminal uses the aforementioned similar forward propagation and back propagation to train a user similarity model for implementing the historical data migration application of similar users.
[0026] Furthermore, the present application provides a pre-built user intent understanding model, which also includes:
[0027] The interactive system log obtains the click log data of the real-time retrieval user; constructs query-document pairs, and converts the click log data into multiple groups of sample query-document pairs with reference to the query-document pairs; configures intent labels for the multiple groups of sample query-document pairs to obtain multiple sample intent labels; adds the multiple sample intent labels to the multiple groups of sample query-document pairs to obtain an instruction fine-tuning dataset, wherein the instruction fine-tuning dataset includes multiple groups of sample query-document-intent label pairs.
[0028] Preferably, the system terminal interacts with the system log and collects the user's real-time click log data from it. This data covers the user's query information, click behavior, historical browsing history, and dwell time, etc., as the basic source of the data set. Subsequently, the original click log is cleaned to remove abnormal data and noise, including robot clicks, malicious clicks, etc., to ensure the quality and reliability of the data, thereby extracting high-quality click samples. After that, the cleaned click log is converted into multiple sets of sample query-document pairs, each of which consists of the user's query term (query) and the clicked document ID (doc id ) to form a preliminary two-tuple data structure (query, doc id Then, based on the query-document pair, we configure the intent label according to historical experience. These intent labels reflect the user’s intention behind clicking on the document, such as obtaining research background, finding technical methods, etc., providing rich semantic information for the dataset. Finally, we append the configured intent label to the query-document pair to form a triple (query, document ID, and intent label) containing the query term, document ID, and intent label. id , intent), which is the final instruction fine-tuning dataset. This instruction fine-tuning dataset not only records the user's query behavior and click results, but also reveals the user's retrieval intent, providing strong support for training and optimizing the user intent understanding model.
[0029] The user intent understanding model is constructed based on RNN, and the instruction fine-tuning data set is used as training data to optimize the user intent understanding model until the performance of the user intent understanding model meets preset requirements.
[0030] Preferably, the system terminal uses a recurrent neural network (RNN) as the infrastructure to build a user intent understanding model. RNN is particularly suitable for such tasks because it can process sequence data and capture the temporal dependencies therein. In this model, the input is the real-time retrieval type of the real-time retrieval user, and the user intent understanding model analyzes these texts through the learned knowledge to predict the real intention of the real-time retrieval user. The system terminal trains and optimizes the constructed user intent understanding model by using the instruction fine-tuning data set as training data, combined with the LoRA low-rank projection matrix and loss function. When the training and optimization process reaches the maximum number of iterations, the system terminal uses the instruction fine-tuning data that is not used for training to evaluate the current user intent understanding model and calculate the prediction accuracy of the current user intent understanding model. If the prediction accuracy meets the preset requirements, that is, greater than or equal to the preset accuracy, the system terminal outputs the current user intent understanding model. Otherwise, the training parameters are continuously adjusted, the model structure is optimized, or more training data is added to gradually improve the prediction ability and robustness of the model.
[0031] Furthermore, the present application provides a method of constructing the user intent understanding model based on RNN, and optimizing the user intent understanding model using the instruction fine-tuning data set as training data until the performance of the user intent understanding model meets preset requirements, and further includes:
[0032] Pre-introducing a LoRA low-rank projection matrix into each of the multiple network layers of the user intent understanding model; presetting a loss function threshold and a maximum number of iterations; optimizing the multiple LoRA low-rank projection matrices of the multiple network layers with the loss function threshold and / or the maximum number of iterations as constraints to complete the optimization of the user intent understanding model.
[0033] Preferably, in the process of training and optimizing the user intent understanding model, the system terminal uses LoRA (Low-Rank Adaptation) technology for fine-tuning. The core strategy of LoRA is to pre-introduce a set of low-rank projection matrices in each layer of the user intent understanding model. These low-rank matrices are smaller in scale than the original user intent understanding model parameters, but can efficiently capture and adjust the behavior of the user intent understanding model to adapt to new data. By optimizing only these low-rank matrices instead of all parameters of the entire user intent understanding model, LoRA achieves fast and effective model adaptation while keeping most of the knowledge of the user intent understanding model unchanged. In order to implement this process, the system terminal pre-sets the loss function threshold and the maximum number of iterations as constraints of the optimization process. The loss function threshold is used to ensure that the performance of the optimized model meets the required standards, while the maximum number of iterations limits the computing resource consumption and time cost of the optimization process. In the optimization stage, by adjusting the LoRA low-rank projection matrix in multiple network layers, iterative optimization is performed with the goal of minimizing the loss function. This process is limited by the preset loss function threshold and / or the maximum number of iterations to ensure that the user intent understanding model achieves the expected optimization effect at a reasonable time and resource consumption.
[0034] Further, the present application provides a method of optimizing multiple LoRA low-rank projection matrices of the multiple network layers with the loss function threshold and / or the maximum number of iterations as constraints to complete the optimization of the user intent understanding model, and also includes:
[0035] The i-th network layer is obtained based on the extraction of the multiple network layers, wherein the i-th LoRA low-rank projection matrix of the i-th network layer includes the projection matrix A i and B i , the i-th network layer is any one of the multiple network layers; the i-th network layer is modified by adding a correction term to perform a forward propagation process, and the multiple network layers are modified by analogy to perform a forward propagation process; a predefined loss function is as follows: in, is the sum of the cross entropy loss values, are the multiple LoRA low-rank projection matrices, θ is a fixed parameter of the user intent understanding model, is the instruction fine-tuning dataset, l is the task-related cross entropy loss; with the loss function threshold and / or the maximum number of iterations as constraints, multiple cycle iterations are performed on the user intent understanding model, and after each cycle iteration, the loss function is used to obtain multiple gradients of the multiple LoRA low-rank projection matrices, and the multiple LoRA low-rank projection matrices are updated based on the multiple gradients; and so on, until the calculation result of the loss function falls into the loss function threshold and / or the cycle iteration frequency meets the maximum number of iterations, and the user intent understanding model is output.
[0036] Optionally, during the optimization process of the user intent understanding model, the system terminal obtains multiple network layers of the user intent understanding model and extracts these network layers sequentially. For the i-th network layer of the user intent understanding model, the system terminal uses LoRA to define two projection matrices A i and B i , the dimensions are (d, r) and (r, d), where d is the model hidden layer dimension, r is the projection dimension, and r<<d. Subsequently, the system terminal uses the instruction fine-tuning dataset to train the user intent understanding model. During forward propagation, LoRA adds a correction term based on the projection matrix to the original calculation result of the i-th network layer. The original forward calculation of the i-th network layer can be expressed as: h i =f i (x i ), where x i is the input of the i-th network layer, f i is the forward calculation function of the ith network layer. In LoRA, the modified forward calculation formula is: ‘ i =f i (x i )+A i B i f i (x i )=f i (x i )+Δ i ; Among them, Δ i =A i B i f i (x i ) represents the correction term introduced by LoRA. This correction term is the original i-th network layer output f i (x i), a low-rank perturbation is added. During the training phase, LoRA fixes the original parameters of the user intent understanding model and only optimizes the projection matrix A of each layer. i and B i , that is, a restricted parameter search is performed on the user intent understanding model space to find a low-rank perturbation direction that is suitable for the new task. The optimization goal of LoRA is to minimize the loss function of the modified user intent understanding model on the new task: in, is the sum of the cross entropy loss values, are multiple LoRA low-rank projection matrices, that is, all the projection matrices introduced, θ is a fixed parameter of the user intent understanding model, is the training data set for the new task, i.e., the instruction fine-tuning data set, and l is the task-related cross entropy loss. During the optimization process, the system terminal only updates the LoRA low-rank projection matrix While keeping the fixed parameter θ unchanged. This makes the training overhead of LoRA much smaller than the traditional full parameter fine-tuning. At the same time, since the rank r of the projection matrix is much smaller than the dimension d of the user intent understanding model, the amount of additional parameters introduced by LoRA is also much smaller than the user intent understanding model. Therefore, LoRA achieves efficient model adaptation by introducing low-rank correction terms at each layer of the user intent understanding model. It finds an optimal perturbation direction that adapts to new tasks by optimizing a small number of projection matrices while keeping the structure of the user intent understanding model unchanged. While greatly reducing the amount of training parameters, LoRA can still achieve an adaptation effect comparable to full parameter fine-tuning. By analogy, the system terminal repeats the above training and optimization process until the calculation result of the defined loss function falls into the pre-set loss function threshold and / or the number of cycle iterations reaches the set maximum number of iterations. At this time, the system terminal will output the current user intent understanding model as the final model.
[0037] The real-time retrieval type is input into the user intention understanding model to perform retrieval intention recognition, and a user intention recognition result is output.
[0038] In one embodiment, after constructing the user intention understanding model, the system terminal inputs the real-time search type of the real-time search user into the user intention understanding model. The user intention understanding model can automatically analyze the real-time search type of the real-time search user for the user's search intention. For example, if the real-time search user inputs a certain subject word, the user intention understanding model infers that the real-time search user's intention may be to query papers, well-known scholars, research results or other relevant information related to the subject word; if the input is time, the user intention understanding model may believe that the real-time search user intends to query meetings, major technological breakthroughs, historical events, etc. that occurred at that time node. Based on the above analysis, the user intention understanding model will output a clear user intention recognition result. This user intention recognition result directly reflects the type of information or target that the real-time search user hopes to obtain through the search, thereby improving the accuracy and efficiency of the search.
[0039] Pre-built literature association retrieval model.
[0040] In one embodiment, the system terminal encapsulates multiple algorithms to form multiple search branches, and trains each search branch in combination with corresponding training data, gradually optimizing the performance of each search branch. Subsequently, the system terminal connects these search branches in parallel to construct a document association search model, providing accurate search services for real-time search users.
[0041] Furthermore, the present application provides a pre-built document association retrieval model, including:
[0042] Based on the keyword matching algorithm, a precise keyword retrieval branch is generated; based on the cognate expansion algorithm, an extended synonym discovery branch is generated; based on the subject classification matching algorithm, a topic-oriented retrieval branch is generated; based on the citation analysis algorithm, a citation association discovery branch is generated; based on the collaborative filtering algorithm, a personalized collaborative recommendation branch is generated; based on the knowledge graph retrieval algorithm, a structured knowledge navigation branch is generated.
[0043] Preferably, the system terminal pre-builds a training database, which contains keyword retrieval training data, synonym discovery training data, topic retrieval training data, citation association discovery training data, collaborative recommendation training data and knowledge navigation training data. After building the training database, the system terminal extracts the keyword retrieval training data from it, introduces natural language processing (NLP) technology, and builds a BERT pre-trained language model as a keyword matching algorithm to perform semantic analysis on keywords. Subsequently, the BERT pre-trained language model is trained using the keyword retrieval training data, and the importance of keywords is dynamically adjusted to improve the accuracy of the retrieval results. Afterwards, the system terminal encapsulates the trained BERT pre-trained language model, that is, connects it to the knowledge base, and builds a precise keyword retrieval branch for keyword matching of the search terms entered by the user with the documents in the knowledge base, and returns documents containing these keywords.
[0044] The system terminal extracts the synonym discovery training data, trains and builds a deep learning model as a cognate expansion algorithm, analyzes the context of the user's query, dynamically generates synonyms and cognate words related to the user's query context, and improves the accuracy and relevance of synonym expansion. Subsequently, the same method is used for encapsulation to build an extended synonym discovery branch, which is used to expand the user's search terms into a set of synonyms or cognate words, and then matches them with the knowledge base to improve the recall rate of the search.
[0045] The system terminal extracts the subject retrieval training data, uses the machine learning algorithm to perform subject classification training on the documents in the knowledge base, generates a subject classification model, and then designs and trains the subject matching model by using the subject retrieval training data, and connects the subject classification model and the subject matching model in parallel to form a subject classification matching algorithm. Subsequently, the same method is used for encapsulation to construct a subject-oriented retrieval branch, which is used to pre-classify the documents in the knowledge base. When the user searches, the most likely subject category is predicted according to his intention, and then a more accurate match is performed under the subject to narrow the search scope and improve the search efficiency and accuracy.
[0046] The system terminal extracts the citation association discovery training data from it, and extracts the citation relationship between documents from the knowledge base to construct a citation network. Subsequently, the citation association discovery training data is used to design and train a citation analysis algorithm based on a neural network, which can recommend references and citing documents based on the citation network. After that, the same method is used for encapsulation to construct a citation association discovery branch for retrieval using the citation relationship between documents. When a user searches for a document, the references of the document and other documents that cite the document are recommended. This method can help users quickly find a group of documents that are highly relevant to a certain topic.
[0047] The system terminal extracts the collaborative recommendation training data from it, and designs and trains the collaborative filtering algorithm based on users and items in combination with the deep learning algorithm. The collaborative filtering algorithm can recommend documents of interest to these users by finding other users with similar interests to the current user, or it can recommend other documents similar to the document currently browsed by the user by calculating the similarity between the documents, and supports the use of user feedback and real-time behavior data to dynamically adjust the recommendation results and improve the degree of personalization. Subsequently, the same method is used for encapsulation to construct a personalized collaborative recommendation branch, which is used to make personalized recommendations for the current user by utilizing the search history and click behavior of other users.
[0048] The system terminal extracts the knowledge navigation training data from it, and then uses technologies such as knowledge extraction and knowledge fusion to build a domain knowledge graph from the knowledge base. Subsequently, using the knowledge navigation training data, the neural network designs and trains a query algorithm for the knowledge graph, which can handle the user's complex query requirements and return structured query results. Subsequently, the same method is used for encapsulation to construct a structured knowledge navigation branch, which is used to structure the knowledge base and build a domain knowledge graph. The nodes in the graph can be authors, institutions, papers, keywords, etc., and the edges represent the relationship between them. The user's search terms are located on the knowledge graph, and then the search results are expanded through the link relationship of the graph. Knowledge graph retrieval can provide richer and more accurate structured search results.
[0049] The precise keyword search branch, the extended synonym discovery branch, the subject-oriented search branch, the citation association discovery branch, the personalized collaborative recommendation branch and the structured knowledge navigation branch are connected in parallel to generate a multi-dimensional document search layer; a document aggregation layer is pre-built, and the multi-dimensional document search layer and the document aggregation layer are cascaded to obtain the document association retrieval model.
[0050] Preferably, after constructing the precise keyword search branch, the extended synonym discovery branch, the theme-oriented search branch, the citation association discovery branch, the personalized collaborative recommendation branch and the structured knowledge navigation branch, the system terminal connects these branches in parallel to form a multi-dimensional document search layer. After connecting these search branches in parallel, the system terminal builds a document aggregation layer. The function of this document aggregation layer is to aggregate, remove and sort the results from each search branch to present a unified, orderly and high-quality document list to the real-time search user. Through the document aggregation layer, the interference of duplicate information can be avoided and the accuracy and availability of the search results can be improved. Subsequently, the system terminal connects the multi-dimensional document search layer and the document aggregation layer in parallel to form a complete document association search model. This model can make full use of the advantages of various search algorithms to provide comprehensive, in-depth and personalized document search services for real-time search users. Real-time search users can select suitable search branches for search according to their needs and preferences, and obtain the final results through the document aggregation layer. Such a design not only improves the efficiency and accuracy of the search, but also enhances the real-time search user experience and satisfaction.
[0051] The input text and the user intention recognition result are synchronized to the document association retrieval model to perform document association retrieval and obtain target retrieval results.
[0052] In one embodiment, the system terminal synchronizes the input text and the user intention recognition result to the constructed document association retrieval model, and the document association retrieval model performs document association retrieval on the input text and the user intention recognition result according to multiple internal branches to obtain multiple branch retrieval results. Subsequently, these branch retrieval results are transmitted to the internal document aggregation layer to generate and output the target retrieval results.
[0053] Fine-grained knowledge extraction is performed on the target retrieval results to obtain a set of key knowledge elements.
[0054] In one embodiment, after obtaining the target search results, in order to further mine the information in these results, the system terminal performs a fine-grained knowledge extraction process. This process aims to extract named entities, keywords, abstracts, paragraphs and other parts from the retrieved documents in a more in-depth manner to form a set of key knowledge elements. Through fine-grained knowledge extraction, the originally large and complex document data can be converted into a series of refined, orderly, easy-to-understand and easy-to-apply knowledge elements. These knowledge elements can not only help real-time retrieval users grasp the essence of the documents more quickly, but also provide strong support for subsequent research, learning and innovation.
[0055] Furthermore, the present application provides a method for performing fine-grained knowledge extraction on the target search results to obtain a set of key knowledge elements, including:
[0056] Perform text preprocessing on multiple retrieved papers in the target search results to obtain multiple formatted papers; perform named entity recognition on the multiple formatted papers to obtain multiple groups of paper entity information; perform keyword extraction on the multiple formatted papers to obtain multiple groups of subject keywords-research focus keywords; extract multiple paper abstracts from the multiple formatted papers, and perform sentence step recognition on the multiple paper abstracts to obtain multiple groups of abstract sentence categories; extract multiple key paragraphs from the multiple formatted papers, and perform sentence step recognition on the multiple key paragraphs to obtain multiple groups of key paragraph categories; extract multiple citation information of the multiple formatted papers, and perform sentiment analysis on the multiple citation information to obtain multiple sentiment tendencies of cited documents; based on the knowledge graph, associate and store the multiple groups of paper entity information, multiple groups of subject keywords-research focus keywords, multiple groups of abstract sentence categories, multiple groups of key paragraph categories and multiple sentiment tendencies of cited documents to obtain the key knowledge element set.
[0057] Preferably, the system terminal performs text preprocessing on multiple retrieved papers in the target search results, aiming to eliminate noise in the text, such as HTML tags, irrelevant symbols, etc., unify the text format, such as converting it into plain text or a document in a specific format, and perform basic processing such as word segmentation and part-of-speech tagging, so that subsequent steps can process text data more accurately. After preprocessing, these retrieved papers are converted into formatted papers with a unified format and clear structure.
[0058] On the basis of formatting the paper, Named Entity Recognition (NER) technology is used to extract important entity information such as names of people, organization names, and professional terms from the paper. This entity information is the cornerstone of the paper's knowledge framework and is of great significance for understanding the content of the paper and mining potential relationships. Through NER technology, the system terminal obtains multiple sets of paper entity information, providing rich entity nodes for the subsequent knowledge graph construction.
[0059] Perform keyword extraction on formatted papers. Different from simple word frequency statistics, the system terminal uses a keyword recognition algorithm, combined with context information, to accurately identify keywords in paragraphs and distinguish between subject keywords and research focus keywords. These keywords not only reflect the main research content of the paper, but also reveal the focus and depth of the research. At the same time, by classifying and sorting keywords, multiple groups of subject keyword-research focus keyword pairs are obtained, providing users with more refined knowledge tags.
[0060] In order to understand the structure and content of the paper, sentence recognition was performed on the abstracts and key paragraphs of multiple formatted papers. The system terminal uses sentence classification technology to identify different categories of sentences in the abstract, such as research question sentences, method sentences, and conclusion sentences. The classification of these sentence categories helps users quickly locate the core part of the paper and understand the background, purpose, method, and results of the research. At the same time, similar sentence recognition was performed on key paragraphs, and multiple groups of key paragraph categories were obtained, which further refined the understanding of the content of the paper.
[0061] In order to evaluate the academic influence and credibility of the paper, the system terminal extracted the citation information of multiple formatted papers and performed sentiment analysis. Then, using the sentiment recognition algorithm, it judged the sentiment tendency of the cited literature, such as positive citation, neutral citation or negative citation, based on the expression and context of the citation content. This step not only helps users understand the acceptance of the paper in the academic community, but also provides important sentiment dimension information for subsequent citation relationship analysis and knowledge graph construction.
[0062] Finally, based on the multiple sets of paper entity information, multiple sets of subject keywords-research focus keywords, multiple sets of abstract sentence categories, multiple sets of key paragraph categories and multiple reference literature sentiment tendencies extracted in the above steps, the system terminal uses knowledge graph technology for associated storage. The knowledge graph represents the relationship between entities in the form of a graph, and can intuitively display complex information such as citation relationships, subject associations, and sentiment tendencies between papers. By constructing such a knowledge graph, the system terminal obtains a set of key knowledge elements, providing users with a comprehensive, in-depth, and intuitive knowledge service experience. Users can quickly locate the required knowledge elements in the knowledge graph according to their needs and interests, thereby improving the efficiency of knowledge acquisition and scientific research.
[0063] According to the user intention recognition result, the key knowledge element set is screened and matched to obtain the associated knowledge element set.
[0064] In one embodiment, after obtaining the user intent recognition result, the system terminal uses this result as a screening condition to screen and match from the key knowledge element set that has been constructed, obtain knowledge elements that are highly relevant to the user intent, and form a set of related knowledge elements. The knowledge elements in this set not only directly respond to the user's needs, but also provide more relevant background information, extended reading, and in-depth understanding through the association relationship of the knowledge graph, thereby improving the efficiency and quality of scientific research work.
[0065] Further, the present application provides a method for obtaining a set of associated knowledge elements by screening and matching the key knowledge element set according to the user intention recognition result, including:
[0066] According to the user intention recognition result, elements are extracted from the key knowledge element set to obtain the associated knowledge element set; a display form is preset, and the associated knowledge element set is visualized based on the preset display form to obtain a visualization interface; and the visualization interface is fed back to the real-time retrieval user.
[0067] Optionally, the system terminal defines a set of extraction rules based on the user intent to filter out knowledge elements related to the user intent from the key knowledge element set. These rules include keyword matching, entity type filtering, sentiment tendency filtering, etc. For example, if the user intent is to understand the method of a certain experiment, the extraction rules may focus on matching paragraphs or sentences containing keywords such as experimental steps, experimental materials, and experimental conditions. Subsequently, according to the defined extraction rules, traversal and search are performed in the key knowledge element set. By applying the extraction rules, the system terminal identifies knowledge elements that are highly relevant to the user intent, such as specific entities, keywords, sentences, or paragraphs. After all relevant knowledge elements are extracted, these elements are sorted to construct a set of related knowledge elements with a clear structure and logical coherence. This set of related knowledge elements contains multiple levels of information, such as topic overview, detailed steps, case analysis, cited literature, etc., to meet the different needs of users. Afterwards, in order to improve the user experience, the system terminal presets one or more display forms, which are intended to present knowledge in the most intuitive and easy-to-understand way. Based on these preset display forms, the system terminal visualizes the associated knowledge element set, including displaying the data in the form of charts, graphs, timelines, tree diagrams, etc., making the originally boring or complex knowledge easy to understand. Then, the associated knowledge element set after visualization will be encapsulated into a visualization display interface, which can clearly display the knowledge information required by the user. Finally, the system terminal feeds back this designed visualization display interface to the user who is conducting real-time retrieval, so that the user can immediately obtain the required knowledge content, thereby improving retrieval efficiency and satisfaction.
[0068] The associated knowledge element set is fed back to the real-time search user.
[0069] In one embodiment, after obtaining the associated knowledge element set, the system terminal presents the selected associated knowledge element set in a visual form through the user terminal display interface. This presentation method is not only intuitive and easy to understand, but also helps users quickly grasp the key points, providing users with an efficient, accurate and convenient way to acquire knowledge.
[0070] In summary, the embodiments of the present application have at least the following technical effects:
[0071] The embodiment of the present application receives and analyzes the text input by the user to clarify its search type. Subsequently, the pre-built user intent understanding model is used. The user intent understanding model is based on the user behavior data in the system log, and is optimized through the RNN architecture and the LoRA low-rank projection matrix to accurately identify the user's search intent. After that, the document association retrieval model is used to search in combination with the text input by the user and the intention recognition result. The document association retrieval model integrates a variety of retrieval algorithms, such as keyword matching, cognate word expansion, subject classification, etc., to form a multi-dimensional retrieval layer to ensure the comprehensiveness and accuracy of the retrieval results. After obtaining the target retrieval results, fine-grained knowledge extraction is performed, including text preprocessing, named entity recognition, keyword and abstract extraction, etc., to construct a set of key knowledge elements. Then, according to the user intent recognition results, a set of associated knowledge elements that highly match the user's intent is selected from the key knowledge element set, and visualized, and presented to the user in an intuitive and easy-to-understand form, completing the whole process of knowledge navigation recommendation. These technical effects jointly solve the technical problem that existing recommendation methods are unable to accurately understand the actual intentions of users, resulting in most recommendation results remaining at the article or journal level, requiring users to still manually screen, reducing scientific research efficiency. It achieves the effect of improving the accuracy of retrieval intent recognition and knowledge recommendation based on user intention understanding models and fine-grained knowledge extraction.
[0072] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned specific embodiments of this specification are described. The processes depicted in the accompanying drawings do not necessarily require the specific order and continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0073] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0074] This specification and drawings are merely exemplary illustrations of the present application and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, a person skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalents, the present application intends to include these modifications and variations.
Claims
1. A heuristic knowledge navigation recommendation method integrating user search intention, characterized in that: The method comprises: Interactively obtain input text from a real-time search user, and parse the input text to obtain a real-time search type; Pre-built user intent understanding model; Retrieval intent is identified by inputting the real-time retrieval type into the user intent understanding model, and outputting a user intent identification result; Pre-build literature association retrieval model; By synchronizing the input text and the user intention recognition result to the document association retrieval model, a document association retrieval is performed to obtain a target retrieval result; Performing fine-grained knowledge extraction on the target search results to obtain a set of key knowledge elements; According to the user intention recognition result, the key knowledge element set is selected and matched to obtain a related knowledge element set; Feeding back the associated knowledge element set to the real-time search user; Pre-building a user intention understanding model, the method further includes: The interactive system log obtains the click log data of the real-time retrieval user; Constructing query-document pairs, and converting the click log data into a plurality of groups of sample query-document pairs with reference to the query-document pairs; Configuring intent labels for the multiple groups of sample query-document pairs to obtain multiple sample intent labels; Adding the multiple sample intent labels to the multiple groups of sample query-document pairs to obtain an instruction fine-tuning dataset, wherein the instruction fine-tuning dataset includes multiple groups of sample query-document-intent label pairs; Building the user intention understanding model based on RNN, and optimizing the user intention understanding model using the instruction fine-tuning data set as training data until the performance of the user intention understanding model meets preset requirements; Pre-constructing a document association retrieval model, the method further comprises: Parallel precise keyword search branch, extended synonym discovery branch, subject-oriented search branch, citation association discovery branch, personalized collaborative recommendation branch and structured knowledge navigation branch to generate a multi-dimensional literature search layer; Pre-constructing a document aggregation layer, and cascading the multi-dimensional document retrieval layer and the document aggregation layer to obtain the document association retrieval model; Performing fine-grained knowledge extraction on the target search result to obtain a key knowledge element set, the method further includes: Performing text preprocessing on multiple searched papers in the target search results to obtain multiple formatted papers; Performing named entity recognition on the multiple formatted papers to obtain multiple groups of paper entity information; Perform keyword extraction on the plurality of formatted papers to obtain a plurality of groups of subject keywords-research focus keywords; Extracting and obtaining a plurality of paper abstracts from the plurality of formatted papers, and performing sentence step recognition on the plurality of paper abstracts to obtain a plurality of groups of abstract sentence categories; Extracting a plurality of key paragraphs from the plurality of formatted papers, and performing sentence step recognition on the plurality of key paragraphs to obtain a plurality of groups of key paragraph categories; Extracting and obtaining multiple citation information of the multiple formatted papers, and performing sentiment analysis on the multiple citation information to obtain sentiment tendencies of multiple citation documents; The key knowledge element set is obtained by associating and storing the multiple groups of paper entity information, the multiple groups of subject keywords-research focus keywords, the multiple groups of abstract sentence categories, the multiple groups of key paragraph categories and the sentiment tendencies of multiple cited documents based on the knowledge graph.
2. The heuristic knowledge navigation recommendation method integrating user search intention as claimed in claim 1, characterized in that: Interactively obtaining input text of a real-time search user, parsing the input text to obtain a real-time search type, the method further comprising: Interactively obtain a preset search type set, and collect search text samples with the preset search type set as a constraint to obtain a search text sample set; Performing entity annotation on a plurality of search text samples in the search text sample set to obtain a plurality of search type samples; Construct a retrieval type recognition model based on conditional random fields; Optimizing the retrieval type recognition model by using the multiple retrieval text samples and the multiple retrieval type samples as training data until the retrieval type recognition accuracy of the retrieval type recognition model meets preset requirements; The input text is parsed using the retrieval type recognition model to obtain the real-time retrieval type.
3. The heuristic knowledge navigation recommendation method integrating user search intention as claimed in claim 1, characterized in that: The user intention understanding model is constructed based on the RNN, and the instruction fine-tuning data set is used as training data to optimize the user intention understanding model until the performance of the user intention understanding model meets the preset requirements, and the method further includes: Pre-introducing a LoRA low-rank projection matrix in each of the multiple network layers of the user intent understanding model; Preset loss function threshold and maximum number of iterations; The multiple LoRA low-rank projection matrices of the multiple network layers are optimized with the loss function threshold and / or the maximum number of iterations as constraints to complete the optimization of the user intent understanding model.
4. The heuristic knowledge navigation recommendation method integrating user search intention as described in claim 3 is characterized in that: The method further comprises optimizing the multiple LoRA low-rank projection matrices of the multiple network layers with the loss function threshold and / or the maximum number of iterations as constraints to complete the optimization of the user intent understanding model. The i-th network layer is obtained based on the extraction of the multiple network layers, wherein the i-th LoRA low-rank projection matrix of the i-th network layer includes the projection matrix and , the i-th network layer is any one of the multiple network layers; Modify the forward propagation process of the i-th network layer by adding a correction term, and modify the forward propagation process of the multiple network layers by analogy; A loss function is predefined, which is as follows: ; in, is the sum of the cross entropy loss values, are the multiple LoRA low-rank projection matrices, are fixed parameters of the user intent understanding model, Fine-tune the dataset for the instructions, is the task-related cross entropy loss; Taking the loss function threshold and / or the maximum number of iterations as constraints, performing multiple cycle iterations on the user intent understanding model, and using the loss function for each cycle iteration, multiple gradients of the multiple LoRA low-rank projection matrices, and updating the multiple LoRA low-rank projection matrices based on the multiple gradients; And so on, until the calculation result of the loss function falls within the loss function threshold and / or the periodic iteration frequency meets the maximum number of iterations, the user intent understanding model is output.
5. The heuristic knowledge navigation recommendation method integrating user search intention as claimed in claim 1, characterized in that: Pre-constructing a document association retrieval model, the method further comprises: Generate precise keyword search branches based on keyword matching algorithm; Based on the same root word expansion algorithm, an extended synonym discovery branch is generated; Construct and generate topic-oriented retrieval branches based on topic classification matching algorithm; Based on the citation analysis algorithm, a citation association discovery branch is constructed; Generate personalized collaborative recommendation branches based on collaborative filtering algorithms; Based on the knowledge graph retrieval algorithm, structured knowledge navigation branches are generated.
6. The heuristic knowledge navigation recommendation method integrating user search intention as claimed in claim 1, characterized in that: According to the user intention recognition result, the key knowledge element set is screened and matched to obtain a related knowledge element set, and the method further includes: Extracting elements from the key knowledge element set according to the user intention recognition result to obtain the associated knowledge element set; Preset a display format, and perform visualization processing on the associated knowledge element set based on the preset display format to obtain a visualization display interface; The visual display interface is fed back to the real-time search user.
Citation Information
Patent Citations
Optimization method and device for knowledge graph question-answering system
CN118093842A
Power knowledge acquisition method and system combined with plug-in professional knowledge base
CN118260381A