A method and system for optimizing text retrieval of electric power scientific research information

By dividing the electric power scientific research information database and identifying user intentions, a double helix path is constructed for retrieval, which solves the problem of accurate matching in existing electric power scientific research information retrieval methods and achieves efficient and accurate information acquisition.

CN120508607BActive Publication Date: 2025-09-30STATE GRID JIANGSU ELECTRIC POWER CO LTD NANTONG POWER SUPPLY BRANCH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511000364.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-30
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing power scientific research information retrieval methods cannot accurately match user search intent, resulting in low retrieval efficiency and inaccurate results, and it is difficult to process multi-level and multi-dimensional scientific research content and results information.

Method used

The electric power scientific research information database is divided into a scientific research text information database and an achievement text information database. A spiral path of scientific research content and a spiral path of scientific research achievement are constructed. The user intention mapping model is used to identify the user intention vector and locate the entry node of the double helix path for spiral retrieval.

Benefits of technology

It improves retrieval accuracy and efficiency, can more accurately match user intentions, and optimize the information acquisition process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508607B_ABST
    Figure CN120508607B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing technology, and provides a method and system for optimizing text retrieval of electric power scientific research information. The method comprises: dividing the electric power scientific research information database, outputting a scientific research text information database and an achievement text information database, and constructing a scientific research content spiral path and a scientific research achievement spiral path; associating the two to obtain a scientific research information double helix path; obtaining a user search statement, inputting a user intention mapping model for identification, and obtaining a user intention vector; locating the spiral entry node of the double helix path according to the user intention vector, performing a spiral search based on this, and outputting the search results. The present application solves the technical problem that the traditional scientific research information retrieval method cannot accurately match the user search intention and the information path, resulting in low retrieval efficiency and inaccurate results, and achieves the technical effect of optimizing the retrieval process through the double helix path and improving the accuracy and efficiency of the retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method and system for optimizing text retrieval of electric power scientific research information. Background Art

[0002] As the scale of electric power research continues to grow, the accumulation of information is becoming increasingly massive, complex, and diverse. Researchers often spend considerable time sifting through this vast amount of information to identify valuable insights, and struggle to capture the deep connections between research content and research results. Existing methods for retrieval of electric power research information typically rely on traditional keyword matching or simple categorized searches. These methods fail to fully understand user search intent and struggle to process multi-level and multi-dimensional research content and results information. These traditional methods, when faced with large and complex research information repositories, suffer from low retrieval accuracy, poor efficiency, and a poor user experience. Therefore, effectively storing, managing, and rapidly retrieving this research information has become a crucial task for improving the efficiency and innovation capabilities of electric power research. Summary of the Invention

[0003] This application provides an optimization method and system for electric power scientific research information text retrieval, aiming to solve the technical problem that traditional scientific research information retrieval methods cannot accurately match user search intentions and information paths, resulting in low retrieval efficiency and inaccurate results.

[0004] The first aspect disclosed in the present application provides a method for optimizing text retrieval of electric power scientific research information, the method comprising: dividing an electric power scientific research information database, outputting a scientific research text information database and an achievement text information database, and constructing a scientific research content spiral path and a scientific research achievement spiral path with the scientific research text information database and the achievement text information database; associating the scientific research content spiral path and the scientific research achievement spiral path to obtain a scientific research information double spiral path; obtaining a user search statement, inputting the user search statement into a user intention mapping model for identification, and obtaining a user intention vector; locating a spiral entry node of the scientific research information double spiral path according to the user intention vector, performing a spiral search based on the spiral entry node, and outputting a search return result.

[0005] Another aspect disclosed in the present application provides an electric power scientific research information text retrieval optimization system, the system comprising: a spiral path construction unit: dividing the electric power scientific research information database, outputting a scientific research text information database and an achievement text information database, and constructing a scientific research content spiral path and a scientific research achievement spiral path with the scientific research text information database and the achievement text information database; a spiral path association unit: associating the scientific research content spiral path and the scientific research achievement spiral path to obtain a scientific research information double spiral path; an intention recognition unit: obtaining a user search statement, inputting the user search statement into a user intention mapping model for identification, and obtaining a user intention vector; a spiral retrieval unit: locating the spiral entry node of the scientific research information double spiral path according to the user intention vector, performing a spiral search based on the spiral entry node, and outputting a search return result.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0007] The above-mentioned method for optimizing text retrieval of electric power research information first divides the electric power research information database into a research text information database and a research results text information database. These two parts of data are then used to construct a research content path and a research results path, respectively, and these are linked to form a double helix path. Subsequently, when the user searches, the search statement entered is recognized by the intent mapping model and converted into an intent vector. Then, based on the identified intent vector, the entry node of the double helix path is located, and an in-depth search is conducted from this node, ultimately returning research information and results relevant to the user's needs. This method can significantly improve retrieval accuracy and efficiency, more accurately match user intent, and optimize the information acquisition process.

[0008] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 The figure is a flow chart of a method for optimizing text retrieval of electric power scientific research information in one embodiment.

[0011] Figure 2This is an architecture diagram of an electric power scientific research information text retrieval optimization system in one embodiment.

[0012] Explanation of reference numerals: spiral path construction unit 11 , spiral path association unit 12 , intention recognition unit 13 , spiral retrieval unit 14 . DETAILED DESCRIPTION

[0013] The embodiments of the present application provide a method and system for optimizing electric power scientific research information text retrieval to solve the technical problem that traditional scientific research information retrieval methods cannot accurately match user search intentions and information paths, resulting in low retrieval efficiency and inaccurate results.

[0014] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0015] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.

[0016] Example 1, as Figure 1 As shown, the present application provides a method for optimizing text retrieval of electric power scientific research information, the method comprising:

[0017] The electric power scientific research information database is divided, and a scientific research text information database and an achievement text information database are output, and a scientific research content spiral path and a scientific research achievement spiral path are constructed with the scientific research text information database and the achievement text information database.

[0018] In the embodiment of the present application, a semantic classification model based on a binary classification network is first used to divide the scientific research text data stored in the power research information database, and the divided scientific research text information and result text information are stored separately to form a scientific research text information database and a result text information database. The scientific research text information database contains various research papers, academic articles and technical reports, etc., mainly focusing on the theory, method, experimental data and other contents of power research; the result text information database includes the specific results of scientific research projects, such as technical inventions, scientific research reports, patents and experimental results, etc., focusing on showing the actual results of scientific research work. Subsequently, according to the contents of the scientific research text information database and the result text information database, a spiral path of scientific research content and a spiral path of scientific research results are constructed respectively. The scientific research content spiral path is to form an orderly structure of a knowledge system by organizing and associating the contents in the scientific research text; the scientific research result spiral path describes the evolution process of scientific research results from the beginning to the final output based on the information of the result text. Through this structured division and path construction, not only can the power research information be effectively managed and classified, but also a clear framework and support can be provided for subsequent intelligent retrieval and knowledge mining.

[0019] Furthermore, the present application provides a method for dividing the electric power research information database and outputting a research text information database and an achievement text information database, including:

[0020] The text stored in the electric power scientific research information database is semantically encoded; the semantic vector encoding is binary labeled to train a binary classification network to obtain a semantic classification model, wherein the binary classification labeling includes semantic vector labeling of scientific research text and semantic vector labeling of achievement text; the semantic classification model classifies the text in the electric power scientific research information database according to the semantic vector encoding, and outputs a scientific research text information database and an achievement text information database.

[0021] Preferably, first, all text data in the power scientific research information database (such as scientific research articles, technical reports, scientific research achievements, etc.) are processed. Each text is converted into a vector through semantic vector encoding technology. This vector can effectively represent the semantic features of the text, capture the key information and context relationships in the text. Commonly used encoding methods include word embedding (such as Word2Vec, GloVe) or encoding based on deep learning models (such as BERT), etc. Taking Word2Vec as an example, first perform word segmentation on each text to convert the text into a sequence of words. For Chinese texts, tools such as Jieba segmentation can be used for Chinese word segmentation. Then, remove high-frequency words without substantial meaning such as "de", "shi", "zai", etc. to reduce interference in model training. At the same time, clean the invalid symbols and non-text characters in the text to ensure that the training data is clean and consistent. Subsequently, set the parameters of the Word2Vec tool, such as the dimension of the word vector (usually choose 100 to 300 dimensions), the size of the context window (usually choose 5 to 10 words), and the minimum word frequency (to avoid interference from too many low-frequency words during the training process). Then, use the preprocessed text data as the corpus and input it into the Skip-gram model of the Word2Vec tool. Each word will appear in the context in the form of a window, and the model trains the word vector by learning the lexical relationships in the context. In this process, the model will be iteratively optimized through the Stochastic Gradient Descent (SGD) algorithm and backpropagation. The model will gradually adjust the word vector so that words with similar semantics have closer representations in the vector space. After training, each vocabulary will be mapped to a semantic vector with a fixed dimension (for example, 100 dimensions). These semantic vectors all contain the semantic features of the words and can reflect the relationships between words. For example, electricity and energy may be close in the vector space. After the semantic vector encoding of the text information stored in the power scientific research information database is completed, binary classification annotation training will be performed on a part of the semantic vector encodings. That is, the semantic vectors of scientific research texts are labeled as one class, and the semantic vectors of achievement texts are labeled as another class. These labeled text information will form a training data set to train a binary classification network model. In this process, a Support Vector Machine (SVM) will be used to construct the semantic classification model structure, including the input layer, output layer, support vectors, kernel function, etc. Then, randomly initialize and select a suitable kernel function (such as linear kernel, radial basis kernel, etc.) for the SVM, and initialize the parameters of the support vector machine. Subsequently, input the text vectors of the training data set into the model for training. Determine the support vectors by calculating the distance between each data point and the decision hyperplane, and construct the classification boundary with the maximum margin. During the training process, use the hinge loss function to calculate the error between the model output result and the true label, and gradually adjust the model parameters through the gradient descent algorithm or the SMO (Sequential Minimal Optimization) algorithm to minimize the value of the loss function. The training process will go through multiple rounds of iteration to continuously optimize the parameters of the model until the loss function converges.After training, the model performance is evaluated using a validation dataset. Classification accuracy, precision, recall, and other indicators are calculated to assess the model's performance in the tasks of classifying scientific research texts and research results texts. If the accuracy meets expectations, the semantic classification model obtained through current training is output; if it does not meet expectations, hyperparameters such as the C value and kernel function type are adjusted to further optimize model performance. Afterwards, the trained semantic classification model automatically determines the category of the text based on the received semantic vector encoding. For each text in the electric power research information database, the model analyzes the semantic features of its semantic vector encoding based on the learned knowledge, and classifies it as scientific research text information or research results text information. The information is then added to the corresponding database to form a scientific research text information database (containing all scientific research-related documents and papers) and a research results text information database (containing all scientific research results and technical reports), providing a data foundation for subsequent intelligent retrieval.

[0022] Furthermore, the present application provides a method for constructing a spiral path of scientific research content and a spiral path of scientific research results using the scientific research text information library and the results text information library, the method comprising:

[0023] The texts in the scientific research text information database and the achievement text information database are segmented into paragraph-level structures, and content-structured paragraphs and achievement-structured paragraphs are output; the scientific research text information database is clustered with each content-structured paragraph, and the output content clustering results are connected in series to obtain a scientific research content spiral path; the achievement text information database is clustered with each achievement-structured paragraph, and the output achievement clustering results are connected in series to obtain a scientific research achievement spiral path.

[0024] Preferably, first, according to the natural breakpoints of each text information in the scientific research text information library and the achievement text information library, such as the line break, period or other structural markers of the paragraph, paragraph-level structure segmentation is performed, thereby splitting each text information into multiple structured paragraphs, forming content-structured paragraphs and achievement-structured paragraphs, each representing different content or achievement parts, for example, a paragraph representing the research background, a paragraph representing the technical route, etc. in the scientific research text, and a paragraph representing the application background, a paragraph representing the technical solution, etc. in the achievement text. Subsequently, the number of clusters is determined according to actual business needs. For example, the number of clusters can be 5. For scientific research texts, each cluster represents the research background, technical route, experimental design, analysis results, and stage conclusions. For achievement texts, each cluster represents the application background, technical solution, implementation method, performance results, and achievement archiving. Afterwards, the semantic vector codes of multiple paragraphs in the content-structured paragraphs are randomly selected as the initial cluster centers, and the Euclidean distance between the semantic vector code of each paragraph and the cluster center is calculated, and the paragraph corresponding to the semantic vector code is assigned to the nearest cluster. After this assignment is complete, the average semantic vector encoding of all paragraphs in each cluster is calculated and used as the new cluster center. The process of assigning paragraphs and updating cluster centers is repeated until the clustering results converge, that is, the cluster centers no longer change significantly. After clustering is completed, each cluster represents a specific stage in the scientific research content. For example, one cluster may contain multiple paragraphs related to research background, while another cluster may contain multiple paragraphs related to technical approach. The cluster results are then concatenated in a preset order to form a scientific research content spiral path. For example, the order of the scientific research content path in the scientific research content spiral path may be: research background → technical approach → experimental design → data analysis → stage conclusion. The scientific research content spiral path may be: Layer 1: Research Background (multiple documents) → Layer 2: Technical Approach (multiple documents) → Layer 3: Experimental Design (multiple documents) → Layer 4: Analysis Results (multiple documents) → Layer 5: Stage Conclusion (multiple documents). Similarly, a similar clustering operation is performed on the structured paragraphs of the research results, and the clustering results are concatenated in a preset order to form a research results spiral path. The order of the research results paths in this research results spiral path can be: application background → technical solution → implementation method → ​​performance results → results archiving. The research results spiral path can be: Layer 1: Application Background (multiple results) → Layer 2: Technical Solution (multiple results) → Layer 3: Implementation Method (multiple results) → Layer 4: Performance Results (multiple results) → Layer 5: Results Archiving (multiple results). These spiral paths provide a clear framework for subsequent research information retrieval. Each path accurately represents the different stages of the research process and the detailed content of the research results, providing users with more structured and in-depth search results.

[0025] The scientific research content spiral path and the scientific research results spiral path are associated to obtain a scientific research information double helix path.

[0026] In one embodiment, after obtaining the scientific research content spiral path and the scientific research results spiral path, coupling node pairs are extracted from the scientific research content spiral path and the scientific research results spiral path. These nodes represent the key content in the scientific research process and the corresponding scientific research results. Semantic similarity calculation is then performed on each coupling node pair to evaluate the similarity between the scientific research content and the results. Valid coupling pairs whose similarity exceeds a preset threshold are screened out. These valid coupling pairs are considered to be parts with strong correlations between the scientific research content and the scientific research results. Subsequently, connection paths are established based on the screened valid coupling pairs to form an association between the two paths. These connection paths organically combine the scientific research content and the scientific research results, ensuring a close and orderly flow between the scientific research process and the results. Finally, through these valid connection paths, the scientific research content spiral path and the scientific research results spiral path are linked together to obtain a scientific research information double helix path. This scientific research information double helix path can more comprehensively reflect the overall picture of the scientific research project, showing both the various links of the scientific research process and the results produced by each link, so that the scientific research content and the scientific research results form a close and organic combination, facilitating subsequent analysis and retrieval.

[0027] Furthermore, the present application provides a method for associating the scientific research content spiral path and the scientific research results spiral path to obtain a scientific research information double helix path, the method comprising:

[0028] Acquire the coupling node pairs of the scientific research content spiral path and the scientific research results spiral path, and output an initial coupling pair list; perform semantic similarity calculation on each coupling node in the coupling node pair, and output a semantic similarity set; screen the valid coupling pairs with a value greater than a preset semantic similarity in the semantic similarity set, and establish connection paths for the valid coupling pairs; associate the scientific research content spiral path and the scientific research results spiral path according to the connection paths of the valid coupling pairs, and obtain a double helix path of scientific research information.

[0029] Preferably, first, extract all nodes in the scientific research content spiral path and the scientific research achievement spiral path, and these nodes are randomly combined in pairs to form multiple coupling node pairs. By storing these coupling node pairs, an initial coupling pair list is constructed. This initial coupling pair list lists all possible matching node pairs between scientific research content and scientific research achievements. Subsequently, semantic similarity is calculated for the semantic vector encoding of the paragraphs involved in each pair of coupling nodes using cosine similarity, and the similarity between each pair of nodes is evaluated. The calculated semantic similarity is then stored in the order in the initial coupling pair list to construct a semantic similarity set. This semantic similarity set contains similarity information of all coupling node pairs, and corresponds one to one with the content in the initial coupling pair list. Afterwards, from the semantic similarity set, filter out those coupling node pairs whose semantic similarity is greater than the preset semantic similarity. Only those semantically highly correlated node pairs will be considered as valid coupling pairs. By screening, valid coupling pairs are obtained. These valid coupling pairs have a strong logical or content association between scientific research content and scientific research achievements. Then, using the semantic information of valid coupling pairs, connection paths are established—that is, mapping relationships between valid coupling pairs. These paths demonstrate the close interaction between research content and research results and guide the logical flow of research progress. Finally, based on the connection paths of valid coupling pairs, the research content spiral path is linked to the research results spiral path. By organically connecting the content and results of each stage, a complete double helix path of research information is formed. This path not only illustrates the evolution of content in the research process but also reflects the continuous output of results. It ensures the full traceability and precise location of research work, helping users better understand and retrieve information and results from the research process.

[0030] Furthermore, the present application provides a method of setting the spiral path of the scientific research content to a right spiral or a left spiral, setting the spiral path of the scientific research results to a left spiral or a right spiral, and setting the connection path of the effective coupling pair to a coupling path through a visualization processing module to obtain the double spiral path of the scientific research information.

[0031] Optionally, in the visualization module, the research content spiral path is first configured as either a right-handed or left-handed spiral based on its specific content and structure. This spiral structure represents different levels and stages of the research process, distinguished by its right-handed or left-handed spiral direction. In visualization, the spiral direction of the research content path helps intuitively depict the gradual progression of research content from research background to stage conclusions. Similarly, the research results spiral path also requires a spiral direction setting. Based on the logic and development sequence of the research results, it can be configured as either a left-handed or right-handed spiral. This setting ensures that the research result generation process (e.g., from application background to research result archiving) is symmetrical or complementary to the research content path in the visualization. Subsequently, valid coupling pairs in the research content spiral path and the research results spiral path are connected to form coupling paths. These coupling paths represent the relationship between each stage in the research process and its corresponding research results. By connecting these valid coupling pairs in the visualization, a three-dimensional spiral structure is formed, making the relationship between research content and research results more clearly visible. The visualization module then merges the research content spiral and the research results spiral according to the set direction and connection path, creating a double helix of research information. This double helix intuitively demonstrates the intertwined relationship between the research process and results, allowing users to clearly track the entire process from the beginning of research to the completion of results in a visual interface and quickly understand the connection between content and results. This visualization of the double helix not only enhances the sense of hierarchy of information but also provides an intuitive tool to help researchers quickly locate key nodes in complex research data, improving the efficiency of information retrieval and analysis.

[0032] A user search sentence is obtained, and the user search sentence is input into a user intention mapping model for recognition to obtain a user intention vector.

[0033] In one embodiment, the user's search query is first processed to extract key information. This key information represents the core content of the user's search, which may include technical terms, field names, research directions, etc. The extracted search keywords are then input into a user intent mapping model to determine multiple keyword vectors, such as technical keyword vectors and functional keyword vectors. The model then predicts user intent based on these multiple keyword vectors. By comprehensively analyzing information from different dimensions, it outputs a final user intent vector. This vector represents the user's specific search requirements, helps understand the user's intent, and thus provides more accurate search results.

[0034] Furthermore, the present application provides a method for obtaining a user search statement, inputting the user search statement into a user intent mapping model for identification, and obtaining a user intent vector, the method comprising:

[0035] Keywords are extracted from the user search statement to obtain search keywords; the search keywords are input into the user intent mapping model for multi-channel convolution, and a multi-channel keyword vector is output, including a technical keyword vector, a functional keyword vector, and a stage keyword vector; user intent labels are predicted according to the multi-channel keyword vector to obtain a user intent vector.

[0036] Preferably, first, a domain dictionary pre-built based on commonly used scientific research words in the power field is activated, and the user search sentence is processed using this domain dictionary. During this process, a pointer is initialized (for example, an index i=0 is set) to point to the starting position of the sentence, and the pointer is gradually moved to traverse each character. If a word is successfully matched, the pointer is moved to the end of the word (that is, the recognized word is skipped), and then the next word is matched from the new pointer position. If no matching word is found, the pointer moves back one character and tries to match again. Through this process, search keywords can be extracted from the user search sentence. These search keywords represent the core content of the user query, usually including specific technical terms, field names, product names, research directions and other key information. Subsequently, the extracted search keywords are input into the user intent mapping model, which uses multi-channel convolution technology to process different dimensions of the search keywords separately to generate a multi-channel keyword vector containing a technical keyword vector, a functional keyword vector and a stage keyword vector. Afterwards, the multi-channel keyword vector output by the fully connected layer will be processed by the activation function to predict the user's intention and ultimately map out the user intention vector. This user intention vector contains the user's complete search requirements and indicates the user's focus on specific fields, functions, and research stages, thereby helping to more accurately match and retrieve relevant information, ensuring that the search results can accurately meet the user's needs.

[0037] The user intent mapping model is a multi-channel convolutional model used to extract technology, function, and phase-related keyword information from user search queries. The model consists of an input layer, a word embedding layer, a convolutional layer (multiple convolutional channels), a pooling layer, and a fully connected layer. First, the word embedding layer converts each word in the search query into a word vector. This vector is then fed into the convolutional layer, where each convolutional channel focuses on extracting features related to a specific dimension, such as technology, function, and phase. The convolutional layer extracts a feature map for each channel using a convolution kernel. The pooling layer then compresses these feature maps. The pooled features are then fed into a fully connected layer, which fuses the feature information from different channels to generate corresponding technology keyword vectors, function keyword vectors, and phase keyword vectors. The outputs of each channel are then combined and processed using an activation function (such as ReLU) to generate a complete user intent vector. This user intent vector represents the user's search requirement, encompassing information about technology, function, and phase. To train the model, the cross-entropy loss function is used to calculate the loss between the predicted intent vector and the true label. The gradient of the loss with respect to the weights of each layer is calculated through backpropagation. The model parameters are then adjusted using an optimizer (such as Adam) to minimize the value of the loss function, thereby ensuring the predictive ability of the user intent mapping model.

[0038] The spiral entry node of the scientific research information double spiral path is located according to the user intention vector, a spiral search is performed based on the spiral entry node, and the search return result is output.

[0039] In one embodiment, after obtaining the user intention vector, the user intention vector and the semantic vector encoding of each node in the double helix path of scientific research information are compared using a similarity comparison method to find the spiral entry node that best matches the user's needs. These spiral entry nodes are the starting point for user retrieval. The system will start to retrieve relevant information based on the content and characteristics of the node. In this process, it will start from the spiral entry node and perform a spiral search along the double helix path of scientific research information, gradually going deeper into more relevant nodes and information. This spiral search can cover multiple levels and angles related to user needs, making the retrieval results more comprehensive. After completing the spiral search, the search return results are output based on the information matching in the retrieval process. These results are the scientific research content and scientific research results that are most relevant to the user's needs. Its entire retrieval process can help users efficiently find the required scientific research materials and results and meet the user's query needs.

[0040] Furthermore, the present application provides a method for locating a spiral entry node of the scientific research information double helix path according to the user intention vector, the method comprising:

[0041] Based on the user intention vector, vector similarity calculation is performed with the vectors of each node in the scientific research information double helix path, and a vector similarity set is output; based on the vector similarity set, the top K nodes with vector similarities greater than a preset value are selected as spiral entry nodes.

[0042] Preferably, first, the user intention vector is compared with the semantic vector encoding of each node in the scientific research information double helix path. The semantic vector encoding of each node represents the specific content or results of the node in the scientific research process, and the user intention vector represents the user's search needs. By using cosine similarity to calculate the similarity between the user intention vector and each node vector, the degree of match between each node and the user needs can be evaluated. After calculating the similarity, a vector similarity set will be obtained. This set contains the similarity values ​​of all scientific research information double helix path nodes. Each value represents the correlation strength between the corresponding node and the user intention. The higher the similarity value, the closer the node content is to the user's search needs. Subsequently, according to the vector similarity set, those nodes with similarity values ​​greater than the preset vector similarity will be screened out, and from these nodes, the top K nodes with the highest similarity values ​​are selected as spiral entry nodes. These nodes best represent the user's search needs and are the ideal starting point for starting spiral search. Through this process, the most relevant starting node can be accurately selected from the scientific research information double helix path, ensuring the accuracy and efficiency of the search.

[0043] Furthermore, the present application provides that each node of the double helix path of scientific research information includes an index module, wherein the index module includes a sequential index facing the next node, a reverse index facing the previous node, and a cross index for jumping to another spiral path; at the spiral entry node, the index module performs parallel spiral retrieval of multiple entries.

[0044] Optionally, each node in the double-helix path of scientific research information contains an index module. This module provides navigation for each node and guides the search process. The index module includes three main types of indexes: sequential index, reverse index, and cross index. Sequential indexes point to the next node immediately following the current node in the path and are typically used for sequential searches along the path. Reverse indexes point to the previous node in the path, allowing the search process to retrace back to previous nodes. Cross indexes are used to jump to nodes in another spiral path. This index can span different spiral paths, connecting research content paths and research output paths, thereby enabling information flow between paths. After locating the spiral entry node, the index module assists in performing multi-directional parallel searches. The index module allows not only downward search to the immediately following node (sequential search) but also upward search to the previous node (reverse search). Cross indexes can also be used to jump to related nodes in another spiral path. This parallel search method enables the system to simultaneously search for relevant information in multiple dimensions and directions. Through these three indexes, guided by spiral entry nodes, users can flexibly switch between research content and research results based on their intent, quickly uncovering relevant information and results. This parallel search ensures comprehensive and in-depth information acquisition, thereby determining search results. The collaboration of index modules improves search efficiency, ensuring comprehensive coverage of research information from multiple perspectives, allowing users to quickly find the most relevant research content and results.

[0045] Furthermore, the present application provides a method for performing parallel spiral retrieval of multiple entries by the index module at the spiral entry node, the method comprising:

[0046] Obtain candidate search points of any spiral entry node, wherein the candidate search points include a next node, a previous node, and a cross node; perform similarity calculations on the user intention vector and the candidate search points, respectively; and determine the next search direction from the candidate search points according to the similarity calculation results by the index module.

[0047] Optionally, starting from the spiral entry node, the indexing module retrieves candidate search points for that node. These candidate search points include the next node, the previous node, and the intersection node. The next node is the node immediately following the current node in the path and is used to continue the search forward in the path; the previous node is the previous node in the path and is used to retrace back to the previous part of the path; and the intersection node is a node that jumps from the current path to another spiral path, allowing the search to span different scientific research content or research results paths, expanding the search scope. For each candidate search point, the cosine similarity is calculated to measure the degree of match between the user intent vector and the semantic vector encoding of the node. The higher the similarity value, the more relevant the node is to the user's search requirement. Subsequently, based on the similarity calculation results for all candidate search points, the indexing module selects the node with the highest similarity value as the next search direction. These next search directions and the spiral entry node together form the optimal search path, guiding subsequent search steps, thereby improving the accuracy and efficiency of the search and ensuring the most relevant scientific research information is found.

[0048] In summary, the embodiments of the present application have at least the following technical effects:

[0049] The embodiment of the present application first divides the electric power scientific research information database, outputs the scientific research text information database and the achievement text information database, and constructs the scientific research content spiral path and the scientific research achievement spiral path with the scientific research text information database and the achievement text information database; then, the scientific research content spiral path and the scientific research achievement spiral path are associated to obtain the scientific research information double helix path; after that, the user search statement is obtained, and the user search statement is input into the user intention mapping model for identification to obtain the user intention vector; finally, the spiral entry node of the scientific research information double helix path is located according to the user intention vector, a spiral search is performed based on the spiral entry node, and the search return result is output. These technical effects jointly solve the technical problem that the traditional scientific research information retrieval method cannot accurately match the user search intention and the information path, resulting in low retrieval efficiency and inaccurate results, and achieve the technical effect of optimizing the retrieval process through the double helix path to improve the accuracy and efficiency of the retrieval.

[0050] Example 2, based on the same inventive concept as the above-mentioned method for optimizing text retrieval of electric power research information, Figure 2As shown, the present application provides an electric power scientific research information text retrieval optimization system, and the system includes: a spiral path construction unit 11: dividing the electric power scientific research information database, outputting a scientific research text information database and an achievement text information database, and constructing a scientific research content spiral path and a scientific research achievement spiral path with the scientific research text information database and the achievement text information database; a spiral path association unit 12: associating the scientific research content spiral path and the scientific research achievement spiral path to obtain a scientific research information double spiral path; an intention recognition unit 13: obtaining a user search statement, inputting the user search statement into a user intention mapping model for identification, and obtaining a user intention vector; a spiral retrieval unit 14: locating the spiral entry node of the scientific research information double spiral path according to the user intention vector, performing a spiral search based on the spiral entry node, and outputting the retrieval return result.

[0051] Furthermore, the spiral path construction unit 11 is further configured to execute the following method:

[0052] The text stored in the electric power scientific research information database is semantically encoded; the semantic vector encoding is binary labeled to train a binary classification network to obtain a semantic classification model, wherein the binary classification labeling includes semantic vector labeling of scientific research text and semantic vector labeling of achievement text; the semantic classification model classifies the text in the electric power scientific research information database according to the semantic vector encoding, and outputs a scientific research text information database and an achievement text information database.

[0053] Furthermore, the spiral path construction unit 11 is further configured to execute the following method:

[0054] The texts in the scientific research text information database and the achievement text information database are segmented into paragraph-level structures, and content-structured paragraphs and achievement-structured paragraphs are output; the scientific research text information database is clustered with each content-structured paragraph, and the output content clustering results are connected in series to obtain a scientific research content spiral path; the achievement text information database is clustered with each achievement-structured paragraph, and the output achievement clustering results are connected in series to obtain a scientific research achievement spiral path.

[0055] Furthermore, the spiral path association unit 12 is further configured to perform the following method:

[0056] Acquire the coupling node pairs of the scientific research content spiral path and the scientific research results spiral path, and output an initial coupling pair list; perform semantic similarity calculation on each coupling node in the coupling node pair, and output a semantic similarity set; screen the valid coupling pairs with a value greater than a preset semantic similarity in the semantic similarity set, and establish connection paths for the valid coupling pairs; associate the scientific research content spiral path and the scientific research results spiral path according to the connection paths of the valid coupling pairs, and obtain a double helix path of scientific research information.

[0057] Furthermore, the spiral path association unit 12 is further configured to perform the following method:

[0058] Through the visualization processing module, the spiral path of the scientific research content is set to a right spiral or a left spiral, the spiral path of the scientific research results is set to a left spiral or a right spiral, and the connection path of the effective coupling pair is set to a coupling path to obtain the double spiral path of the scientific research information.

[0059] Furthermore, the intention recognition unit 13 is further configured to perform the following method:

[0060] Keywords are extracted from the user search statement to obtain search keywords; the search keywords are input into the user intent mapping model for multi-channel convolution, and a multi-channel keyword vector is output, including a technical keyword vector, a functional keyword vector, and a stage keyword vector; user intent labels are predicted according to the multi-channel keyword vector to obtain a user intent vector.

[0061] Furthermore, the spiral retrieval unit 14 is further configured to perform the following method:

[0062] Based on the user intention vector, vector similarity calculation is performed with the vectors of each node in the scientific research information double helix path, and a vector similarity set is output; based on the vector similarity set, the top K nodes with vector similarities greater than a preset vector similarity are selected as spiral entry nodes.

[0063] Furthermore, the spiral retrieval unit 14 is further configured to perform the following method:

[0064] Each node of the double-helix path of scientific research information includes an index module, wherein the index module includes a sequential index facing the next node, a reverse index facing the previous node, and a cross index for jumping to another spiral path; at the spiral entry node, the index module performs parallel spiral retrieval of multiple entries.

[0065] Furthermore, the spiral retrieval unit 14 is further configured to perform the following method:

[0066] Obtain candidate search points of any spiral entry node, wherein the candidate search points include a next node, a previous node, and a cross node; perform similarity calculations on the user intention vector and the candidate search points, respectively; and determine the next search direction from the candidate search points according to the similarity calculation results by the index module.

[0067] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0068] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0069] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.

Claims

1. A method for optimizing text retrieval of electric power scientific research information, characterized in that: The method comprises: The electric power scientific research information database is divided, and a scientific research text information database and an achievement text information database are outputted, and a scientific research content spiral path and a scientific research achievement spiral path are constructed with the scientific research text information database and the achievement text information database; Associating the scientific research content spiral path with the scientific research results spiral path to obtain a scientific research information double helix path; Obtaining a user search statement, inputting the user search statement into a user intent mapping model for identification, and obtaining a user intent vector; Locating the spiral entry node of the scientific research information double spiral path according to the user intention vector, performing a spiral search based on the spiral entry node, and outputting the search return result; Obtaining a user search statement, inputting the user search statement into a user intent mapping model for identification, and obtaining a user intent vector, the method comprising: Extract keywords from the user's search statement to obtain search keywords; Inputting the search keywords into the user intent mapping model for multi-channel convolution, and outputting a multi-channel keyword vector, including a technical keyword vector, a functional keyword vector, and a stage keyword vector; Predicting user intent labels based on the multi-channel keyword vectors to obtain a user intent vector; The method for locating the spiral entry node of the scientific research information double-helix path according to the user intention vector includes: Calculate vector similarity between the user intention vector and the vectors of each node in the scientific research information double helix path, and output a vector similarity set; Selecting the first K nodes with a vector similarity greater than a preset vector similarity as spiral entry nodes according to the vector similarity set; Each node of the scientific research information double spiral path includes an index module, wherein the index module includes a sequential index facing the next node, a reverse index facing the previous node, and a cross index for jumping to another spiral path; At the spiral entry node, the index module performs a parallel spiral search of multiple entries; At the spiral entry node, the index module performs a parallel spiral search of multiple entries, the method comprising: Obtain candidate retrieval points of any spiral entry node, wherein the candidate retrieval points include a next node, a previous node, and an intersection node; Similarity calculations are performed on the user intention vector and the candidate search points respectively, and the indexing module determines the next search direction from the candidate search points according to the similarity calculation results.

2. The method according to claim 1, wherein The electric power research information database is divided and output into a research text information database and an achievement text information database. The method includes: Performing semantic vector encoding on the text stored in the electric power scientific research information database; Performing binary classification annotation on the semantic vector encoding and training a binary classification network to obtain a semantic classification model, wherein the binary classification annotation includes semantic vector annotation of the scientific research text and semantic vector annotation of the achievement text; The semantic classification model classifies the text in the electric power scientific research information database according to the semantic vector encoding, and outputs a scientific research text information database and an achievement text information database.

3. The method according to claim 1, wherein The scientific research content spiral path and the scientific research achievement spiral path are constructed using the scientific research text information database and the achievement text information database, and the method includes: Performing paragraph-level structural segmentation on the texts in the scientific research text information database and the achievement text information database, and outputting content-structured paragraphs and achievement-structured paragraphs; Clustering the scientific research text information database based on each content structured paragraph, and connecting the output content clustering results in series to obtain a scientific research content spiral path; The achievement text information database is clustered with each achievement structured paragraph, and the output achievement clustering results are connected in series to obtain a spiral path of scientific research achievements.

4. The method according to claim 1, wherein The scientific research content spiral path and the scientific research results spiral path are associated to obtain a scientific research information double helix path, the method comprising: Obtaining coupling node pairs of the scientific research content spiral path and the scientific research achievement spiral path, and outputting an initial coupling pair list; Calculating semantic similarity for each coupling node in the coupling node pair and outputting a semantic similarity set; Screening valid coupling pairs with a semantic similarity greater than a preset semantic similarity in the semantic similarity set, and establishing connection paths for the valid coupling pairs; The scientific research content spiral path and the scientific research achievement spiral path are associated with each other according to the connection path of the effective coupling pair to obtain a scientific research information double helix path.

5. The method according to claim 4, wherein Through the visualization processing module, the spiral path of the scientific research content is set to a right spiral or a left spiral, the spiral path of the scientific research results is set to a left spiral or a right spiral, and the connection path of the effective coupling pair is set to a coupling path to obtain the double spiral path of the scientific research information.

6. A power scientific research information text retrieval optimization system, characterized by: The system is used to execute the electric power scientific research information text retrieval optimization method according to any one of claims 1 to 5, and the system includes: A spiral path construction unit: divides the electric power scientific research information database, outputs a scientific research text information database and an achievement text information database, and constructs a scientific research content spiral path and a scientific research achievement spiral path with the scientific research text information database and the achievement text information database; Spiral path association unit: associates the scientific research content spiral path with the scientific research results spiral path to obtain a scientific research information double spiral path; Intent recognition unit: obtains a user search sentence, inputs the user search sentence into the user intent mapping model for recognition, and obtains a user intent vector; Spiral retrieval unit: locates the spiral entry node of the double spiral path of scientific research information according to the user intention vector, performs spiral retrieval based on the spiral entry node, and outputs the retrieval return result.

Citation Information

Patent Citations

  • Multi-intention recognition method and system and readable storage medium

    CN118861200A

  • Heuristic knowledge navigation recommendation method fusing user retrieval intention

    CN118939787A