Electric power scientific research information retrieval and recommendation expansion method and system

By constructing a knowledge graph and user interest vectors in the power field, and combining parallel retrieval and personalized recommendation, the problem of difficulty in grasping users' personalized needs in power scientific research information retrieval and recommendation is solved, and more accurate retrieval results and recommendation effects are achieved.

CN120804287APending Publication Date: 2025-10-17STATE GRID JIANGSU ELECTRIC POWER CO LTD NANTONG POWER SUPPLY BRANCH +2
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510905892.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing methods for retrieving and recommending power research information are insufficient to fully understand users' query intent, fail to fully explore the connections between information, and lack a precise grasp of users' personalized needs, resulting in inaccurate search results and low matching degree between recommended content and user needs.

Method used

A knowledge graph in the power field is constructed based on natural language processing. Extended queries are generated through query expansion and parallel retrieval, and interest vectors are constructed based on users' historical retrieval records for result screening, including entity concept extraction, semantic relationship recognition, multi-dimensional retrieval result fusion and personalized recommendation.

Benefits of technology

It achieves a more accurate understanding of user query intent, outputs more comprehensive and accurate search results, and improves the accuracy of power scientific research information retrieval and the matching degree of personalized recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804287A_ABST
    Figure CN120804287A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power scientific research information retrieval and recommendation expansion method and system, and relates to the field related to intelligent information retrieval recommendation, and the method comprises the steps: building an electric power field knowledge graph based on natural language processing; receiving a user query, performing query expansion according to the user query and the power field knowledge graph, and generating an expansion query; the expansion query is submitted to N databases for parallel retrieval, a multi-dimensional retrieval result is output, and N is an integer larger than or equal to 3; constructing an interest vector based on the historical retrieval record of the user; and screening the multi-dimensional retrieval results according to the interest vectors to generate a personalized recommendation list. The technical problem that the accuracy of the retrieval result is poor due to the fact that the personalized demand of the user is difficult to accurately grasp in the existing power scientific research information retrieval and recommendation is solved, and the technical effect that the query intention of the user is more accurately understood, so that the more comprehensive and accurate retrieval result is output is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent information retrieval and recommendation, and particularly relates to an electric power scientific research information retrieval and recommendation expansion method and system. BACKGROUND

[0002] In the field of electric power scientific research, efficiently and accurately obtaining relevant information and making personalized recommendations are crucial for promoting scientific research progress and improving scientific research efficiency. At present, the main method to solve the problem of electric power scientific research information retrieval and recommendation is based on traditional keyword retrieval and combined with certain rules for simple recommendation. However, the current method is difficult to fully understand the user query intention, cannot fully mine the association between information, and lacks precise grasp of user personalized needs, resulting in less comprehensive and accurate retrieval results, and the matching degree of recommended content with user actual needs is low, which cannot effectively meet the diversified information needs of electric power researchers.

[0003] At present, in the related technology, the electric power scientific research information retrieval and recommendation cannot accurately grasp the user's personalized needs, resulting in the technical problem of poor accuracy of retrieval results. SUMMARY

[0004] The present application provides an electric power scientific research information retrieval and recommendation expansion method and system, which adopts technical means such as constructing an electric power field knowledge graph based on natural language processing, query expansion, parallel retrieval, and constructing an interest vector based on user historical retrieval records for result screening, solves the technical problem that the existing electric power scientific research information retrieval and recommendation cannot accurately grasp the user's personalized needs, resulting in poor accuracy of retrieval results, and achieves the technical effect of more accurately understanding the user's query intention, thereby outputting more comprehensive and accurate retrieval results.

[0005] The present application provides an electric power scientific research information retrieval and recommendation expansion method, which includes: based on natural language processing, constructing an electric power field knowledge graph; receiving a user query, expanding the query according to the user query and the electric power field knowledge graph, and generating an expanded query; submitting the expanded query to N databases for parallel retrieval, and outputting multi-dimensional retrieval results, wherein N is an integer greater than or equal to 3; based on user historical retrieval records, constructing an interest vector; and screening the multi-dimensional retrieval results according to the interest vector, and generating a personalized recommendation list.

[0006] In a possible implementation manner, based on natural language processing, an electric power field knowledge graph is constructed, and the following processing is performed: through natural language processing, the title, abstract and keywords of electric power field literature are analyzed, entity concepts are extracted as nodes to obtain a node set; using dependency syntax analysis, the semantic relationship between the entity concepts is identified as edges to obtain an edge set; and according to the node set and the edge set, the electric power field knowledge graph is constructed.

[0007] In a possible implementation, the user query is received, query expansion is performed according to the user query and the power domain knowledge graph, an expanded query is generated, and the following processing is performed: the user query is received, subgraph matching is performed in the power domain knowledge graph according to the user query, and k-hop neighborhood nodes connected to all keywords in the user query are found out, where k is an integer greater than or equal to 1 and less than or equal to 3; the comprehensive semantic similarity between the k-hop neighborhood nodes and the user query is calculated, and an expanded word with a similarity value greater than a first threshold value is screened out; and the expanded query is generated based on the expanded word.

[0008] In a possible implementation, the comprehensive semantic similarity between the k-hop neighborhood nodes and the user query is calculated, and the following processing is performed: a path similarity component between the k-hop neighborhood nodes and the user query is calculated; a word vector similarity component between the k-hop neighborhood nodes and the user query is calculated; and the path similarity component and the word vector similarity component are weighted and summed to obtain the comprehensive semantic similarity.

[0009] In a possible implementation, the expanded query is submitted to N databases for parallel retrieval, and a multi-dimensional retrieval result is output, and the following processing is performed: the expanded query is simultaneously submitted to N databases for parallel retrieval, and N types of retrieval results are obtained; the relevance of the N types of retrieval results to the user query is calculated, the N types of retrieval results are sorted based on the relevance, and a sorting result is obtained; the sorting result is fused and sorted again by using the Borda counting method, and a comprehensive sorting list is obtained; and the multi-dimensional retrieval result is output based on the comprehensive sorting list.

[0010] In a possible implementation, the multi-dimensional retrieval result is filtered according to the interest vector, and a personalized recommendation list is generated, and the following processing is performed: a topic distribution of each result is extracted from the multi-dimensional retrieval result, and a plurality of topic distributions are obtained; cosine similarity between the interest vector and the plurality of topic distributions is calculated, and a plurality of cosine similarity values are obtained; a result corresponding to a cosine similarity value greater than a second threshold value in the plurality of cosine similarity values is filtered out from the multi-dimensional retrieval result, and an interest-related result is obtained; and the interest-related result is arranged in descending order of cosine similarity value, and a personalized recommendation list is generated.

[0011] In a possible implementation, after the personalized recommendation list is generated by screening the multi-dimensional search results according to the interest vector, the following processing is further performed: recording the click behavior of the user on the personalized recommendation list to obtain a click behavior data set; extracting the click times of each result from the click behavior data set, and extracting a corresponding high click rate result if the click times are greater than a third threshold; extracting a high click rate entity concept from the high click rate result; and enhancing the weights of nodes and edges related to the high click rate entity concept in the power field knowledge graph, and performing incremental optimization of search.

[0012] The application also provides an electric power scientific research information search and recommendation expansion system, which comprises: an electric power field knowledge graph construction module, configured to construct an electric power field knowledge graph based on natural language processing; a query expansion module, configured to receive a user query, perform query expansion according to the user query and the electric power field knowledge graph, and generate an expanded query; a parallel search module, configured to submit the expanded query to N databases for parallel search, and output multi-dimensional search results, wherein N is an integer greater than or equal to 3; an interest vector construction module, configured to construct an interest vector based on user historical search records; and a personalized recommendation module, configured to screen the multi-dimensional search results according to the interest vector, and generate a personalized recommendation list.

[0013] The electric power scientific research information search and recommendation expansion method and system provided in the application first constructs an electric power field knowledge graph based on natural language processing, then receives a user query, performs query expansion according to the user query and the electric power field knowledge graph, generates an expanded query, submits the expanded query to N databases for parallel search, and outputs multi-dimensional search results, wherein N is an integer greater than or equal to 3, constructs an interest vector based on user historical search records, and finally screens the multi-dimensional search results according to the interest vector, and generates a personalized recommendation list. The technical effect of more accurately understanding the user query intention and outputting more comprehensive and accurate search results is achieved. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings of the embodiments of the application will be briefly introduced below. In the present application, a flowchart is used to illustrate the operations performed by the system according to the embodiments of the application. It should be understood that the foregoing or the following operations are not necessarily performed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously according to needs. Meanwhile, other operations can be added to these processes, or a step or several steps can be removed from these processes.

[0015] Figure 1 A flowchart of the electric power scientific research information search and recommendation expansion method provided by the embodiments of the application is shown.

[0016] Figure 2 A structural schematic diagram of an electric power scientific research information retrieval and recommendation expansion system provided by an embodiment of the present application is shown.

[0017] The reference signs are explained as follows: an electric power field knowledge graph construction module 10, a query expansion module 20, a parallel retrieval module 30, an interest vector construction module 40, and a personalized recommendation module 50. DETAILED DESCRIPTION

[0018] The above description is only a summary of the technical solutions of the present application. In order to make the technical means of the present application more clearly understood, and to be implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described.

[0019] In order to make the purposes, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the accompanying drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those skilled in the art without making creative labor, belong to the scope of protection of the present application.

[0020] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict. The term "first\second" is only to distinguish similar objects, and does not represent the specific order of the objects. The terms "include" and "have" and any variations, are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application.

[0021] The embodiments of the present application provide an electric power scientific research information retrieval and recommendation expansion method, as shown in the following method. Figure 1 The method comprises the following steps:

[0022] In step S100, an electric power field knowledge graph is constructed based on natural language processing.

[0023] Specifically, data is collected from text resources such as documents, reports, operation manuals, etc. in the power field. Preprocessing of the text is performed, including text cleaning (removing noise data such as HTML tags, special symbols, etc.), word segmentation, part-of-speech tagging, etc. Key entities in the power field (such as device names, technical parameters, fault types, etc.) are identified using Named Entity Recognition (NER) technology (a natural language processing technique used to identify entities with specific meanings from text, such as names, places, device names, etc.), and linked to corresponding entities in the knowledge base. Through dependency syntax analysis and pattern matching, etc., the relationships between entities (such as the connection between devices, the cause-and-effect relationship between faults and causes, etc.) are extracted from the text. The extracted entities and relationships are integrated into the knowledge graph, and redundancies and conflicts are eliminated to build a complete knowledge system. The knowledge graph is a structured semantic knowledge base that describes entities and their relationships. In the power field, the knowledge graph contains entities such as devices, technical parameters, fault types, and their relationships.

[0024] For example, text is extracted from a power device operation manual, HTML tags and special symbols are removed, and word segmentation is performed. For example, "The rated voltage of the transformer is 220kV" is segmented into "transformer / 's / rated voltage / is / 220kV". Using the NER model, "transformer" is identified as a device entity, and "220kV" is identified as a technical parameter entity. Through dependency syntax analysis, the "rated voltage" relationship between "transformer" and "220kV" is extracted. "Transformer", "220kV", and "rated voltage" relationship are stored in the knowledge graph database to form part of the knowledge graph.

[0025] In one possible implementation, based on natural language processing, the power field knowledge graph is constructed, and step S100 further includes step S110. The title, abstract and keywords of the power field literature are parsed by natural language processing to extract entity concepts as nodes, and a node set is obtained. Specifically, the title, abstract and keywords of the literature are extracted from the literature database in the power field. The text is cleaned and segmented. Using Named Entity Recognition (NER) technology, the entity concepts in the power field are identified, for example, "transformer" is identified as a power device entity, "insulation aging" is identified as a technical term entity, and "a university research team" is identified as a research subject entity. The identified entity concepts are stored as nodes of the knowledge graph in the node set, each node containing entity type and entity name.

[0026] For example, from the title "Research on the Aging of High Voltage Transformer Insulation", the abstract "This paper studies the causes and preventive measures of high voltage transformer insulation aging", and the keywords "high voltage transformer, insulation aging, preventive measures" are extracted. The entities "high voltage transformer" (power equipment entity), "insulation aging" (technical term entity), and "preventive measures" (technical term entity) are identified. The node set is generated as shown in Table 1.

[0027] Table 1: Node set example

[0028] Node ID Node Type Node Name 1 Power equipment entity high-voltage transformer 2 Technical term entity Insulation aging 3 Technical term entity Preventive measures

[0029] In step S120, the semantic relationship between the entity concepts is identified as an edge using dependency syntax analysis, and an edge set is obtained. Specifically, the title, abstract, and keywords of the literature are subjected to dependency syntax analysis using a dependency syntax analysis tool (such as Stanford NLP). Dependency syntax analysis is a natural language processing technique used to analyze the dependency relationships between words in a sentence, helping to understand the structure and semantics of the sentence. According to the dependency relationship, the relationship between entity concepts is extracted, such as causal relationship, association relationship, etc. The extracted relationship is stored as an edge of the knowledge graph in the edge set, and each edge contains the relationship type and the connected entity nodes.

[0030] For example, from the sentence "This paper studies the causes and preventive measures of high voltage transformer insulation aging", the "research" relationship between "high voltage transformer" and "insulation aging" and the "association" relationship between "insulation aging" and "preventive measures" are extracted. The edge set is generated as shown in Table 2.

[0031] Table 2: Edge set example

[0032] Edge ID Relationship Type Starting node ID Target node ID 1 Research 1 2 2 association 2 3

[0033] In step S130, the power field knowledge graph is constructed according to the node set and the edge set. Specifically, the graph database (such as Neo4j) is initialized, and the mode of nodes and edges is created. The node set and the edge set are imported into the graph database. The imported data is integrated to eliminate redundancy and conflicts, ensuring the integrity and consistency of the knowledge graph.

[0034] In step S200, a user query is received, and a query expansion is generated according to the user query and the power field knowledge graph.

[0035] Specifically, a user input query such as "high-voltage transformer failure" is received. The query is parsed using NLP tools to extract the query intent and keywords, such as extracting the keywords "high-voltage transformer" and "failure". Using the entities and relationships in the power domain knowledge graph, relevant expansion keywords are found, such as "insulation aging", "short circuit", etc. Combining the original query and the expansion keywords, a new expanded query is generated, such as "high-voltage transformer failure insulation aging short circuit", which is submitted to the retrieval system.

[0036] In one possible implementation, a user query is received, and query expansion is performed based on the user query and the power domain knowledge graph to generate an expanded query. Step S200 further includes step S210, receiving a user query, and performing subgraph matching in the power domain knowledge graph based on the user query to find k-hop neighborhood nodes connected to all keywords in the user query, where k is an integer greater than or equal to 1 and less than or equal to 3. Specifically, a user input query is received, and NLP tools are used to parse the user input query to extract keywords. In the power domain knowledge graph, subgraph matching is performed starting from the keywords of the user query to find k-hop neighborhood nodes connected to the keywords. For example, in the power domain knowledge graph, starting from "high-voltage transformer" and "failure", subgraph matching is performed to find k-hop neighborhood nodes connected to these keywords. For example, set k = 2, and find all nodes within 2 hops of "high-voltage transformer" and "failure", such as "insulation aging", "short circuit", "overload", etc. Subgraph matching refers to finding subgraph structures connected to certain nodes in the knowledge graph, and is used to find k-hop neighborhood nodes connected to user query keywords. K-hop neighborhood nodes refer to all nodes within k steps (hops) of a certain node in the knowledge graph. For example, 1-hop neighborhood nodes refer to directly connected nodes, and 2-hop neighborhood nodes refer to nodes connected through an intermediate node.

[0037] Step S220, calculate the comprehensive semantic similarity between the k-hop neighborhood nodes and the user query, and select expansion words with a similarity value greater than a first threshold. Specifically, for each k-hop neighborhood node, use a semantic similarity algorithm (such as Word2Vec, BERT) to calculate its comprehensive semantic similarity with the user query. Set a threshold to select nodes with a similarity greater than the threshold as expansion words. For example, the similarity between node "insulation aging" and "high-voltage transformer failure" is 0.8, the similarity between "short circuit" and "high-voltage transformer failure" is 0.75, and the similarity between "overload" and "high-voltage transformer failure" is 0.6. Set the threshold to 0.7, and select "insulation aging" and "short circuit" as expansion words.

[0038] Step S230, based on the expansion word, generate an expansion query. Specifically, combine the original user query and the filtered expansion word to generate a new expansion query. For example, combine the original query "high-voltage transformer failure" with the expansion words "insulation aging" and "short circuit" to generate a new expansion query "high-voltage transformer failure insulation aging short circuit", which is submitted to the retrieval system.

[0039] In one possible implementation, the comprehensive semantic similarity between the k-hop neighborhood node and the user query is calculated, and step S220 further includes step S221 of calculating a path similarity component between the k-hop neighborhood node and the user query. Specifically, in the power domain knowledge graph, all paths from the user query keyword to the k-hop neighborhood node are extracted. Path similarity algorithms (such as path length-based similarity calculation) are used to evaluate the similarity of these paths. Wherein, the shorter the path, the higher the similarity. For example, assuming that the user query keyword is "high-voltage transformer" and the k-hop neighborhood node is "insulation aging". In the power domain knowledge graph, the path from "high-voltage transformer" to "insulation aging" is extracted, such as "high-voltage transformer → research → insulation aging", and the similarity of the path "high-voltage transformer → research → insulation aging" is calculated as 0.8.

[0040] Step S222, calculate the word vector similarity component between the k-hop neighborhood node and the user query. Specifically, use word embedding models (such as Word2Vec, BERT) to generate word vectors for user query keywords and k-hop neighborhood nodes. Use cosine similarity algorithms and other algorithms to calculate the similarity between word vectors. For example, use the BERT model to generate word vectors for "high-voltage transformer" and "insulation aging", and calculate the cosine similarity of the word vectors of "high-voltage transformer" and "insulation aging" as 0.75.

[0041] Step S223, weight and sum the path similarity component and the word vector similarity component to obtain the comprehensive semantic similarity. Specifically, set the weights of the path similarity and the word vector similarity. For example, the path similarity weight is 0.6 and the word vector similarity weight is 0.4. According to the preset weight, weight and sum the path similarity component and the word vector similarity component to obtain the comprehensive semantic similarity. For example: comprehensive semantic similarity = 0.6 x 0.8 (path similarity) + 0.4 x 0.75 (word vector similarity) = 0.78.

[0042] Step S300, submit the expansion query to N databases for parallel retrieval and output multi-dimensional retrieval results, where N is an integer greater than or equal to 3.

[0043] Specifically, the multi-database parallel retrieval is implemented using a distributed retrieval framework such as an Elasticsearch cluster. Data synchronization and updates between multiple databases are ensured to avoid data inconsistency issues. The extended query is distributed to N databases (such as power equipment database, technical literature database, fault case database, academic database, patent library, project library, etc.). Each database independently performs retrieval operations and returns its own results. The results are integrated using a data fusion algorithm (such as a similarity-based fusion algorithm) to generate a unified search result list.

[0044] In one possible implementation, the extended query is submitted to N databases for parallel retrieval, and the multi-dimensional search results are output. Step S300 further includes step S310 of simultaneously submitting the extended query to N databases for parallel retrieval to obtain N types of search results. Specifically, the extended query is simultaneously submitted to multiple databases (such as power equipment database, technical literature database, fault case database, etc.). Each database independently performs retrieval operations and returns its own search results. For example, the extended query "high-voltage transformer fault insulation aging short circuit" is simultaneously submitted to three databases, and each database returns its own search results as shown in Table 3.

[0045] Table 3: Search result example 1

[0046] Database type Search Results Power Equipment Database High voltage transformer models and parameters Technical Literature Database Insulation aging research literature Fault case database Short circuit fault case analysis

[0047] Step S320, the relevance of the N types of search results to the user query is calculated, and the N types of search results are sorted based on the relevance size to obtain a sorted result. Specifically, a relevance algorithm (such as TF-IDF, cosine similarity) is used to calculate the relevance of each type of search result to the user query. Each type of search result is sorted according to the relevance size. For example, the relevance of the power equipment database result is 0.85, the relevance of the technical literature database result is 0.75, and the relevance of the fault case database result is 0.65. According to the relevance size, the sorting is: 1 power equipment database: high-voltage transformer model and parameters, 2 technical literature database: insulation aging research literature, 3 fault case database: short circuit fault case analysis.

[0048] Step S330, the ranking results are fused and ranked again by Borda counting method to obtain a comprehensive ranking list. Specifically, the Borda counting method is a multi-criteria decision method that assigns scores to each result to combine multiple ranking results to obtain a final comprehensive ranking list. For example, assuming that the ranking result of each database has 3, the scores are assigned as follows: 1st: 3 points, 2nd: 2 points, 3rd: 1 point. Assuming that each database returns the search results shown in Table 4, according to the ranking result of each database, each search result is assigned a score as shown in Table 5. The scores of each result are added to obtain a comprehensive ranking list as shown in Table 6.

[0049] Table 4: Search result example 2

[0050] Database type No. 1 No. 2 No. 3 Academic Paper Library Research on Transformer Insulation Aging Circuit breaker fault diagnosis Smart Grid Dispatch Algorithms Patent Database Circuit breaker fault diagnosis Research on Transformer Insulation Aging Power Internet of Things Sensors Technical Report Library Smart Grid Dispatch Algorithms Research on Transformer Insulation Aging High voltage cable temperature monitoring

[0051] Table 5: Score assignment example

[0052]

[0053] Table 6: Comprehensive ranking list example

[0054]

[0055]

[0056] Step S340, based on the comprehensive ranking list, output multi-dimensional search results. Specifically, according to the comprehensive ranking list, set a relevance threshold, only search results with a total score exceeding the threshold will be output. From the comprehensive ranking list, filter out search results with a total score exceeding the threshold. The filtered search results are output in the order of the comprehensive ranking list to form the final multi-dimensional search results.

[0057] Step S400, based on the user's historical search records, construct an interest vector.

[0058] Specifically, record the user's historical search records, such as search keywords, clicked document IDs, dwell time, etc. Extract features from user behavior data, such as search keywords, clicked document types, etc. Use machine learning algorithms (such as TF-IDF, Word2Vec) to construct a user interest vector. For example, the user's search records include keywords such as "high-voltage transformer fault" and "insulation aging", and the user has clicked on the "high-voltage transformer model and parameters" document with a dwell time of 30 seconds. Extract keywords such as "high-voltage transformer", "fault", and "insulation aging" as features. Use the TF-IDF algorithm to calculate the weight of each keyword to generate a user interest vector, such as a weight of 0.8 for high-voltage transformer, a weight of 0.7 for fault, and a weight of 0.6 for insulation aging.

[0059] Step S500, filtering the multi-dimensional search results according to the interest vector to generate a personalized recommendation list.

[0060] Specifically, the similarity between the multi-dimensional search results and the user interest vector is calculated, the multi-dimensional search results are sorted according to the similarity, and a personalized recommendation list is generated.

[0061] In one possible implementation, the multi-dimensional search results are filtered according to the interest vector to generate a personalized recommendation list, and step S500 further includes step S510 of extracting a topic distribution of each result from the multi-dimensional search results to obtain a plurality of topic distributions. Specifically, each search result is preprocessed, including text cleaning, word segmentation, and stop word removal. Key features such as keywords, abstracts, and content fragments are extracted from each search result. A topic modeling algorithm such as LDA (Latent Dirichlet Allocation) is used to model each search result and extract its topic distribution. Each search result is represented as a mixture distribution of multiple topics, and each topic is composed of a set of keywords. A topic distribution example is shown in Table 7.

[0062] Table 7: Topic distribution example

[0063] Search Results Topic distribution Document A Topic 1 (high-voltage transformer fault diagnosis): 0.8, Topic 2 (smart grid): 0.2 Document B Topic 1 (insulation aging): 0.7, Topic 2 (high-voltage transformer performance): 0.3 Document C Topic 1 (Power IoT Sensor): 0.9, Topic 2 (Smart Grid): 0.1 Document D Topic 1 (high-voltage cable temperature monitoring): 0.8, Topic 2 (smart grid): 0.2

[0064] Step S520, calculate the cosine similarity between the interest vector and the plurality of topic distributions to obtain a plurality of cosine similarity values. Specifically, the user interest vector is represented as a vector, where each dimension corresponds to a topic or keyword, and the weight represents the user's interest in the topic or keyword. The topic distribution of each search result is represented as a vector, where each dimension corresponds to the probability distribution of a topic. The cosine similarity algorithm is used to calculate the similarity between the interest vector and the topic distribution of each search result.

[0065] Step S530, filtering the results corresponding to the cosine similarity values greater than the second threshold value from the multi-dimensional search results to obtain interest-related results. Specifically, a cosine similarity threshold is set, and only search results with a similarity greater than the threshold are filtered out. The results with a similarity greater than the threshold are filtered out from the multi-dimensional search results to obtain the interest-related results.

[0066] Step S540, according to the interest-related results, arrange in descending order of cosine similarity value to generate a personalized recommendation list. Specifically, the filtered search results are sorted in descending order of cosine similarity value to generate the final personalized recommendation list.

[0067] In one possible implementation, after filtering the multi-dimensional retrieval results according to the interest vector and generating a personalized recommendation list, the method further includes: recording the user's click behavior on the personalized recommendation list to obtain a click behavior data set; extracting the number of clicks on each result from the click behavior data set, and if the number of clicks is greater than a third threshold, extracting the corresponding high-click rate result; extracting a high-click rate entity concept from the high-click rate result; enhancing the weights of nodes and edges related to the high-click rate entity concept in the power field knowledge graph to perform incremental retrieval optimization.

[0068] Specifically, behavior monitoring code is embedded in the user interface to record, in real time, the user's click behavior on each result in the personalized recommendation list. This user's click behavior data is stored in a database, forming a click behavior dataset. For example, a click event listener is set for each result item in the personalized recommendation list. When a user clicks a result, information such as the result ID, click time, and user ID is recorded.

[0069] Analyze the click behavior dataset and count the number of clicks on each result. Set a click threshold and select results with clicks greater than the threshold as high-click-through-rate results. Preprocess the literature content of high-click-through-rate results, including text cleaning and word segmentation. Use named entity recognition (NER) technology to identify key entity concepts in the literature. Extract the identified entity concepts to form a set of high-click-through-rate entity concepts.

[0070] Based on the number of clicks on high-click-rate entity concepts, we calculate the incremental weights of related nodes and edges. We then add these incremental weights to the weights of the relevant nodes and edges in the power sector knowledge graph, updating the knowledge graph. We then use this updated knowledge graph for subsequent searches and recommendations, achieving incremental optimization and improving the relevance and accuracy of search results.

[0071] The embodiments of the present application adopt technical means such as constructing a knowledge graph in the power field based on natural language processing, query expansion, parallel retrieval, and constructing interest vectors based on user historical search records for result screening. This solves the technical problem of the existing power scientific research information retrieval and recommendation that is difficult to accurately grasp the user's personalized needs, resulting in poor accuracy of retrieval results, and achieves the technical effect of more accurately understanding the user's query intention, thereby outputting more comprehensive and accurate retrieval results.

[0072] In the above, refer to Figure 1 The power research information retrieval and recommendation expansion method according to the embodiment of the present invention is described in detail. Figure 2 The following describes an electric power scientific research information retrieval and recommendation expansion system according to an embodiment of the present invention.

[0073] The power scientific research information retrieval and recommendation expansion system according to the embodiment of the present application is used to solve the technical problem that the existing power scientific research information retrieval and recommendation cannot accurately grasp the user's individualized needs, resulting in poor accuracy of the retrieval results, and achieves the technical effect of more accurately understanding the user's query intention, thereby outputting more comprehensive and accurate retrieval results. The power scientific research information retrieval and recommendation expansion system comprises a power field knowledge graph construction module 10, a query expansion module 20, a parallel retrieval module 30, an interest vector construction module 40, and a personalized recommendation module 50.

[0074] The power field knowledge graph construction module 10 is configured to construct a power field knowledge graph based on natural language processing. The query expansion module 20 is configured to receive a user query, perform query expansion based on the user query and the power field knowledge graph, and generate an expanded query. The parallel retrieval module 30 is configured to submit the expanded query to N databases for parallel retrieval and output multi-dimensional retrieval results, where N is an integer greater than or equal to 3. The interest vector construction module 40 is configured to construct an interest vector based on user historical retrieval records. The personalized recommendation module 50 is configured to filter the multi-dimensional retrieval results based on the interest vector and generate a personalized recommendation list.

[0075] In the following, the specific configuration of the power field knowledge graph construction module 10 will be described in detail. As described above, the power field knowledge graph is constructed based on natural language processing. The power field knowledge graph construction module 10 can further comprise an entity concept extraction unit configured to parse the title, abstract and keywords of power field literature through natural language processing, extract entity concepts as nodes, and obtain a node set; a semantic relationship identification unit configured to identify the semantic relationship between the entity concepts as edges using dependency syntax analysis, and obtain an edge set; and a power field knowledge graph construction unit configured to construct a power field knowledge graph according to the node set and the edge set.

[0076] In the following, the specific configuration of the query expansion module 20 will be described in detail. As described above, the query expansion module 20 receives a user query, performs query expansion based on the user query and the power field knowledge graph, and generates an expanded query. The query expansion module 20 can further comprise a subgraph matching unit configured to receive a user query, perform subgraph matching in the power field knowledge graph based on the user query, and find out k-hop neighborhood nodes connected to all keywords in the user query, where k is an integer greater than or equal to 1 and less than or equal to 3; a comprehensive semantic similarity calculation unit configured to calculate the comprehensive semantic similarity between the k-hop neighborhood nodes and the user query, and filter to obtain expanded words with a similarity value greater than a first threshold; and an expanded query generation unit configured to generate an expanded query based on the expanded words.

[0077] The integrated semantic similarity calculation unit can further include: a path similarity component calculation subunit configured to calculate a path similarity component of the k-hop neighbor node and the user query; a word vector similarity component calculation subunit configured to calculate a word vector similarity component of the k-hop neighbor node and the user query; and a weighted summation subunit configured to perform weighted summation on the path similarity component and the word vector similarity component to obtain the integrated semantic similarity.

[0078] Next, the specific configuration of the parallel retrieval module 30 will be described in detail. As described above, the expanded query is submitted to N databases for parallel retrieval, and the multi-dimensional retrieval result is output. The parallel retrieval module 30 can further include: a parallel retrieval unit configured to simultaneously submit the expanded query to N databases for parallel retrieval to obtain N types of retrieval results; a relevance calculation unit configured to calculate the relevance of the N types of retrieval results and the user query, sort the N types of retrieval results based on the relevance size to obtain a sorted result; a fusion sorting unit configured to perform Borda counting fusion sorting on the sorted result again to obtain an integrated sorting list; and a multi-dimensional retrieval result output unit configured to output the multi-dimensional retrieval result based on the integrated sorting list.

[0079] Next, the specific configuration of the personalized recommendation module 50 will be described in detail. As described above, the multi-dimensional retrieval result is filtered according to the interest vector to generate a personalized recommendation list. The personalized recommendation module 50 can further include: a topic distribution extraction unit configured to extract the topic distribution of each result from the multi-dimensional retrieval result to obtain a plurality of topic distributions; a cosine similarity calculation unit configured to calculate the cosine similarity of the interest vector and the plurality of topic distributions to obtain a plurality of cosine similarity values; a filtering unit configured to filter out the result corresponding to the cosine similarity value greater than a second threshold value from the multi-dimensional retrieval result to obtain an interest-related result; and a personalized recommendation list generation unit configured to generate a personalized recommendation list according to the interest-related result in descending order of the cosine similarity value.

[0080] After the multi-dimensional retrieval result is filtered according to the interest vector to generate a personalized recommendation list, the system can further include: a retrieval incremental optimization module configured to record the click behavior of the user on the personalized recommendation list to obtain a click behavior dataset, extract the number of clicks of each result from the click behavior dataset, and if the number of clicks is greater than a third threshold value, extract a high-click-rate result corresponding thereto, extract a high-click-rate entity concept from the high-click-rate result, enhance the weight of the node and edge related to the high-click-rate entity concept in the power domain knowledge graph, and perform retrieval incremental optimization.

[0081] The power scientific research information retrieval and recommendation expansion system provided by the embodiments of the present application can execute the power scientific research information retrieval and recommendation expansion method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0082] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or server, and the various units and modules are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual differentiation, and do not limit the protection scope of the present application.

[0083] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application. In some cases, the actions or steps described in the present application can be executed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.

Claims

1. The electric power scientific research information retrieval and recommendation expansion method is characterized by: The method comprises: Build a knowledge graph in the power field based on natural language processing; Receive a user query, expand the query based on the user query and the power field knowledge graph, and generate an expanded query; Submitting the expanded query to N databases for parallel search, and outputting multi-dimensional search results, where N is an integer greater than or equal to 3; Build interest vectors based on user historical search records; The multi-dimensional search results are screened according to the interest vector to generate a personalized recommendation list.

2. The electric power scientific research information retrieval and recommendation expansion method according to claim 1, characterized in that: Based on natural language processing, we build a knowledge graph in the power sector, including: Through natural language processing, the titles, abstracts and keywords of literature in the power field are parsed, entity concepts are extracted as nodes, and a node set is obtained; Using dependency parsing to identify semantic relationships between entity concepts as edges, and obtaining an edge set; A knowledge graph in the electric power field is constructed based on the node set and the edge set.

3. The electric power scientific research information retrieval and recommendation expansion method according to claim 1, characterized in that: Receive a user query, perform query expansion based on the user query and the power field knowledge graph, and generate an expanded query, including: Receive a user query, perform subgraph matching in the power field knowledge graph based on the user query, and find k-hop neighboring nodes connected to all keywords in the user query, where k is an integer greater than or equal to 1 and less than or equal to 3; Calculating the comprehensive semantic similarity between the k-hop neighborhood node and the user query, and screening out expansion words whose similarity values ​​are greater than a first threshold; Based on the expanded word, an expanded query is generated.

4. The electric power scientific research information retrieval and recommendation expansion method according to claim 3, characterized in that: Calculating the comprehensive semantic similarity between the k-hop neighborhood nodes and the user query, including: Calculating and obtaining a path similarity component between the k-hop neighborhood node and the user query; Calculate the similarity component of the word vector between the k-hop neighborhood node and the user query; The path similarity component and the word vector similarity component are weighted and summed to obtain the comprehensive semantic similarity.

5. The electric power scientific research information retrieval and recommendation expansion method according to claim 1, characterized in that: Submit the expanded query to N databases for parallel search, and output multi-dimensional search results, including: Submitting the expanded query to N databases simultaneously for parallel retrieval to obtain N types of retrieval results; Calculating the relevance between the N types of search results and the user query, and sorting the N types of search results based on the relevance to obtain a sorting result; The sorting results are again subjected to Borda counting fusion sorting to obtain a comprehensive sorting list; Based on the comprehensive ranking list, multi-dimensional search results are output.

6. The electric power scientific research information retrieval and recommendation expansion method according to claim 1, characterized in that: Filtering the multi-dimensional search results according to the interest vector to generate a personalized recommendation list includes: Extracting the topic distribution of each result from the multi-dimensional retrieval results to obtain multiple topic distributions; Calculating the cosine similarity between the interest vector and the plurality of topic distributions to obtain a plurality of cosine similarity values; Filtering the results corresponding to the cosine similarity values ​​greater than a second threshold among the multiple cosine similarity values ​​from the multi-dimensional search results to obtain interest-related results; According to the interest-related results, the cosine similarity values ​​are arranged in descending order to generate a personalized recommendation list.

7. The electric power scientific research information retrieval and recommendation expansion method according to claim 1, characterized in that: After filtering the multi-dimensional search results according to the interest vector to generate a personalized recommendation list, the method further includes: Recording the user's click behavior on the personalized recommendation list to obtain a click behavior dataset; Extracting the number of clicks for each result from the click behavior dataset, and if the number of clicks is greater than a third threshold, extracting the corresponding high click-through rate result; Extracting high-click-rate entity concepts from the high-click-rate results; The weights of nodes and edges related to the high-click-rate entity concepts in the power field knowledge graph are enhanced to perform incremental retrieval optimization.

8. The power research information retrieval and recommendation expansion system is characterized by: The system is used to implement the electric power scientific research information retrieval and recommendation expansion method according to any one of claims 1 to 7, and the system includes: The power sector knowledge graph construction module is used to construct the power sector knowledge graph based on natural language processing; A query expansion module is used to receive a user query, expand the query based on the user query and the power field knowledge graph, and generate an expanded query; A parallel search module, configured to submit the expanded query to N databases for parallel search and output multi-dimensional search results, where N is an integer greater than or equal to 3; Interest vector construction module, used to construct interest vectors based on user historical retrieval records; The personalized recommendation module is used to filter the multi-dimensional search results according to the interest vector and generate a personalized recommendation list.

Citation Information

Patent Citations

  • Multilayer quotation recommendation method based on literature content mapping knowledge domain

    CN105653706A

  • Information retrieval method, server, medium and product

    CN115470396A

  • Power grid regulation and control knowledge intelligent retrieval method fusing knowledge graph and recommendation algorithm

    CN118210900A

  • Power data intelligent search method based on resource map

    CN119336831A

  • Method for expanding user queries

    EP2482202A1