Full-text retrieval enhancement method, system and equipment for hydraulic power plant knowledge and medium
By constructing a triple structure coding model and an enhanced contrast learning algorithm, decomposing and expanding user query requests, the shortcomings of traditional full-text search technology in hydropower plant knowledge retrieval are solved, and higher quality and accurate search results are achieved.
Patent Information
- Application Number
- CN202510592715.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
When dealing with the complex knowledge system of hydropower plants, traditional full-text search technology has problems such as overlapping vocabulary, inaccurate search and limited recall scope, which is difficult to meet the high-quality and high-precision needs of modern hydropower plants for knowledge retrieval.
By constructing a triple structure coding model and introducing an enhanced contrast learning algorithm, the user query request is decomposed and expanded, and converted into the target semantic vector for similarity calculation, combining vector database and synonym expansion, two similarity calculations and sorting are performed to improve the accuracy of the search results.
It effectively solves the vocabulary gap and semantic gap in traditional search methods, expands the scope of recall, and ensures the accuracy of the search results and the matching of user needs.
Smart Images

Figure CN120104777A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of full-text retrieval of hydropower plant knowledge, and in particular to a full-text retrieval enhancement method, system, equipment and medium for hydropower plant knowledge. Background Art
[0002] In the production management and operation process of hydropower plants, a large amount of unstructured data such as professional equipment knowledge, operation and maintenance experience, production logs, laws and regulations, and expert experience has been accumulated. Effectively preserving and utilizing this knowledge is crucial to improving the production efficiency of hydropower plants, ensuring stable operation of equipment, and promoting technological innovation. At present, traditional full-text retrieval technology has exposed many shortcomings when processing such complex knowledge systems, and it is difficult to meet the needs of modern hydropower plants for high-quality and high-precision knowledge retrieval.
[0003] Traditional retrieval methods, such as the statistically based TF-IDF, BM25 algorithm and its derivative tools Elasticsearch, Lucene, etc., mainly rely on the frequency and distribution of vocabulary for matching. Their core flaw is that they rely too much on the vocabulary overlap between user queries and documents. This mechanism often leads to inaccurate retrieval and limited recall range when facing user queries that are colloquial, diverse, or have vocabulary differences. For example, when users query the knowledge base in an informal or colloquial way, traditional methods may not be able to effectively match documents that are semantically related but have inconsistent vocabulary, resulting in the omission of important information.
[0004] With the accelerated advancement of my country's energy internet construction, the knowledge system in the power sector is increasingly showing an open, flat, and blurred development trend, and the complexity and professionalism of hydropower plant knowledge are further enhanced. In this context, the limitations of traditional retrieval technology are becoming increasingly prominent and can no longer meet the needs of modern hydropower plant knowledge management and retrieval. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a full-text retrieval enhancement method for hydropower plant knowledge, comprising constructing a triple structure encoding model and a target semantic recognition model, wherein the target semantic recognition model is obtained by introducing an enhanced contrastive learning algorithm training into the triple structure encoding model; Decomposing the acquired query request into at least two query keywords, and performing expansion processing on the decomposed query keywords to obtain an expanded query request; The expanded query request is converted into a target semantic vector through a triple structure encoding model, and the first similarity calculation is performed between the target semantic vector and the semantic vector in the vector database to obtain the similarity between the target semantic vector and the semantic vector in the vector database. The relevant documents are recalled according to the similarity, and the recalled documents are preliminarily sorted from large to small according to the similarity; The target semantic recognition model is used to perform a second similarity calculation on the recalled documents and the expanded query requests, and the recalled documents are sorted secondarily in descending order according to the calculated similarities to obtain the target retrieval sorting result, in which the first similarity calculation and the second similarity calculation are performed in different ways.
[0006] As a preferred solution of the full-text retrieval enhancement method of the hydropower plant knowledge of the present invention, the construction of the triple structure coding model includes: Preprocess the acquired hydropower plant knowledge documents to obtain target document data; Three pre-trained models with shared weights are used as the core network to construct the initial triple structure encoding model; Sentences are extracted from the target document data to form training samples. The initial triple structure encoding model is trained and optimized by using contrastive learning method and combined with vector-based dense retrieval loss function to obtain the triple structure encoding model.
[0007] As a preferred solution of the full-text retrieval enhancement method of the hydropower plant knowledge of the present invention, wherein: the semantic vector is obtained by encoding the target document data through a triple structure encoding model; A vector database is a database that stores semantic vectors.
[0008] As a preferred solution of the full-text retrieval enhancement method of the hydropower plant knowledge of the present invention, the expansion processing includes: According to the synonym dictionary and edit distance calculation, synonyms with similar semantics to the query keywords are selected to generate an expanded query matrix.
[0009] As a preferred solution of the full-text retrieval enhancement method of hydropower plant knowledge of the present invention, extracting sentences from target document data to form training samples includes extracting multiple groups of sentences from the target document data to form training samples, wherein each group of sentences includes: sentences used as anchors, positive sample sentences and negative sample sentences.
[0010] As a preferred solution of the full-text retrieval enhancement method of the hydropower plant knowledge of the present invention, the formula for the first similarity calculation is: , Where: Similarity is the similarity score; q is the target semantic vector; d is the semantic vector in the vector database; Represents the modulus length of the target semantic vector q; Represents the modulus length of the semantic vector d in the vector database; is the value of the i-th dimension of the target semantic vector; is the square of the value of the i-th dimension of the target semantic vector; is the value of the i-th dimension of a semantic vector in the vector database; is the square of the value of the i-th dimension of a semantic vector in the vector database; n is the dimension of the semantic vector, .
[0011] As a preferred solution of the full-text retrieval enhancement method of hydropower plant knowledge of the present invention, the enhanced contrastive learning algorithm loss function is: , Where: is the loss function of query q; q is the target semantic vector; It is a semantic vector representation that is positively correlated with q; is the semantic vector representation of q negative correlation; τ is a temperature hyperparameter, and K is the total number of negatively correlated samples; is a positively correlated sample pair Similarity score of are all negatively correlated sample pairs The sum of similarity scores.
[0012] In a second aspect, the present invention provides a full-text retrieval enhancement system for hydropower plant knowledge, comprising: a construction module for constructing a triple structure encoding model and a target semantic recognition model, wherein the target semantic recognition model is obtained by introducing an enhanced contrastive learning algorithm training into the triple structure encoding model; A decomposition and expansion module, used to decompose the acquired query request into at least two query keywords, and perform expansion processing on the decomposed query keywords to obtain an expanded query request; The initial sorting module is used to convert the expanded query request into a target semantic vector through a triple structure encoding model, and perform the first similarity calculation between the target semantic vector and the semantic vector in the vector database to obtain the similarity between the target semantic vector and the semantic vector in the vector database, recall related documents according to the similarity, and preliminarily sort the recalled documents according to the similarity from large to small; The refined sorting module is used to perform a second similarity calculation on the recalled documents and the expanded query requests through the target semantic recognition model and to perform a second sorting on the recalled documents in descending order according to the calculated similarities to obtain the target retrieval sorting result, wherein the first similarity calculation and the second similarity calculation adopt different methods.
[0013] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when the computer program is executed by a processor.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: by constructing a triple structure coding model and introducing an enhanced contrastive learning algorithm, the vocabulary gap and semantic gap problems existing in traditional retrieval methods are effectively solved. Furthermore, the triple structure coding model is used to encode knowledge documents into semantic vectors and store them in a vector database, while synonym expansion and edit distance calculation are performed on user queries to expand the recall range and ensure that more potentially relevant documents can be obtained. Furthermore, an enhanced contrastive learning algorithm is used to perform secondary sorting on the recalled documents, and through a momentum update mechanism and a dictionary maintenance strategy, the model's ability to understand and distinguish document semantics is further improved, so that the final retrieval results more accurately match user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0017] Figure 1 Flow chart of the full-text retrieval enhancement method for hydropower plant knowledge.
[0018] Figure 2 Schematic diagram of the triple structure encoding model.
[0019] Figure 3 Schematic diagram of the target semantic model structure.
[0020] Figure 4 The figure shows the evaluation results of the full-text retrieval enhancement method and BM25 algorithm for hydropower plant knowledge on the test set. DETAILED DESCRIPTION
[0021] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0022] Example 1, reference Figure 1 to Figure 3 , which is the first embodiment of the present invention, and which provides a full-text retrieval enhancement method for hydropower plant knowledge, comprising: S1. Construct a triple structure encoding model and a target semantic recognition model, wherein the target semantic recognition model is trained by introducing an enhanced contrastive learning algorithm into the triple structure encoding model.
[0023] Furthermore, a triple structure encoding model and a target semantic recognition model are constructed, including constructing a triple structure encoding model, and constructing a target semantic recognition model based on the constructed triple structure encoding model.
[0024] Furthermore, the enhanced contrastive learning algorithm is an algorithm for contrastive learning based on multiple sample pairs and multiple negative examples, wherein the sample pairs are generated according to a dictionary consisting of multiple batches.
[0025] Furthermore, the construction of the triple structure coding model includes: Preprocess the acquired hydropower plant knowledge documents to obtain target document data, wherein the preprocessing includes data cleaning of the hydropower plant knowledge documents, and the data cleaning specifically includes removing stop words, punctuation marks, special symbols, and converting numbers to uppercase, etc.; Three pre-trained models with shared weights are used as the core network to build an initial triple structure encoding model. The pre-trained model uses the BERT pre-trained model, which is an open source deep learning model based on the Transformer architecture. Sentences are extracted from the target document data to form training samples. The initial triple structure encoding model is trained and optimized by using contrastive learning method and combined with vector-based dense retrieval loss function to obtain the triple structure encoding model.
[0026] Furthermore, extracting sentences from the target document data to form training samples includes extracting multiple groups of sentences from the target document data to form training samples, wherein each group of sentences includes: sentences used as anchors, positive sample sentences, and negative sample sentences.
[0027] It should be noted that the expression of the vector-based dense retrieval loss function is: , , , , , Where: The average pooling result of the vector representing the anchor sentence; Represents the average pooling result of the vector of the positive sample sentence; The average pooling result of the vector representing the negative sample sentence; Represents distance; edge parameter express and The distance should be at least and The distance is close, is a constant with a value of 1, which ensures that the distance difference between positive and negative samples is not too large, but also not too small to be effectively distinguished.
[0028] S2. Decompose the acquired query request into at least two query keywords, and perform expansion processing on the decomposed query keywords to obtain an expanded query request.
[0029] It can be understood that the obtained query request is decomposed into at least two query keywords, and each query keyword in the multiple query keywords is expanded to obtain an extended keyword corresponding to each query keyword, and the expanded query request is obtained based on the multiple query keywords and the extended keywords corresponding to each query keyword.
[0030] Further, the expansion processing for each query keyword includes: According to the synonym dictionary and edit distance calculation, synonyms with similar semantics to the query keywords are selected to generate an expanded query matrix, where synonyms refer to words that can be used interchangeably in a specific context and have the same or very similar meanings. Furthermore, in the expansion process, synonyms are used to increase the vocabulary related to the original query request, thereby improving the recall rate of the retrieval system.
[0031] Furthermore, the synonyms with semantic closeness to the query keyword are calculated based on the edit distance of each query keyword, and s target keywords with the closest edit distance to each query keyword are selected as the synonyms with semantic closeness to the query keyword.
[0032] It should be noted that the criteria for selecting synonyms are as follows: 1. Semantic similarity: The selected synonyms should be highly similar to the original query request in semantics to ensure that the expanded query request can more accurately reflect the user's intention.
[0033] 2. Contextual relevance: The selection of synonyms should consider the meaning of the query request in a specific context and ensure that the synonyms can maintain contextual consistency when expanding the query.
[0034] 3. Frequency and generality: Synonyms with high frequency of use and high generality are usually preferred to improve the coverage of the expanded query.
[0035] 4. Discrimination: The selected synonyms should be discriminative and able to effectively distinguish different query intents.
[0036] It should be further explained that, assuming that the user's query request Q is decomposed into a query request consisting of n query keywords q, that is, Q=q 1 +q 2 +…+q n , and q 1 ≠q 2 ≠…≠q n (Where: Q is the query request, q n is the nth query keyword), the edit distance of each query keyword is calculated according to the edit distance formula, and s target keywords closest to each query keyword are selected, s=8 (where s is the number of target keywords selected that are closest to the query keyword).
[0037] Furthermore, the edit distance formula is: , Where: is the edit distance between the first i characters of string a and the first j characters of string b; Indicates when When , that is, when one of the strings is an empty string, the edit distance is the length of the other string; It means to delete the i-th character of string a, then calculate the edit distance between the first i-1 characters of string a and the first j characters of string b, plus 1 operation; It means to insert at the jth character of string b, then calculate the edit distance between the first i characters of string a and the first j-1 characters of string b, plus 1 operation; Indicates that a replacement operation is performed at the i-th character of string a and the j-th character of string b (if the characters are different) or no operation is performed (if the characters are the same). If the characters are different, one operation is added; if the characters are the same, no operation is added. It should be noted that when ai = bj, 1 ( ai ≠ bj ) is 0, otherwise it is 1.
[0038] Furthermore, let the s edit distances calculated by the i-th query keyword qi be ki1, ki2, …, kis, so the expanded query matrix is: , Where: Q-extend is the extended query matrix, the row index is i, the column index is j, where i represents the i-th query keyword and j represents the j-th target keyword; represents the edit distance between the i-th query keyword and the s-th target keyword.
[0039] S3. The expanded query request is converted into a target semantic vector through a triple structure encoding model, and the first similarity calculation is performed between the target semantic vector and the semantic vector in the vector database to obtain the similarity between the target semantic vector and the semantic vector in the vector database. Relevant documents are recalled based on the similarity, and the recalled documents are preliminarily sorted in descending order of similarity.
[0040] Furthermore, the semantic vectors in the vector database are encoded according to each target document data, wherein related documents are recalled according to similarity, and the recalled documents are preliminarily sorted in descending order according to similarity, including: sorting the semantic vectors in descending order according to similarity, and determining the target document data corresponding to the top N semantic vectors as the above-mentioned related documents.
[0041] Furthermore, the semantic vector is obtained by encoding the target document data through the triple structure encoding model; The vector knowledge base is a database that stores semantic vectors.
[0042] It should be noted that the target document data is encoded through the triple structure encoding model to obtain a semantic vector, and the obtained semantic vector is stored in the vector knowledge base, which specifically includes data cleaning of the acquired hydropower plant knowledge document to obtain the target document data, dividing each target document data into blocks, the size of which is set to 512 characters by default, and retaining an overlap of 36 characters between each block, and then encoding it in units of blocks through the triple structure encoding model with a dimension of 768, i.e., the output dimension of the model, and further storing the semantic vector formed after encoding in the vector knowledge base, and establishing the corresponding document index and quick index.
[0043] Furthermore, the formula for the first similarity calculation is: , Similarity is the similarity score; q is the target semantic vector; d is the semantic vector in the vector database; Represents the modulus length of the target semantic vector q; Represents the modulus length of the semantic vector d in the vector database; is the value of the i-th dimension of the target semantic vector; is the square of the value of the i-th dimension of the target semantic vector; is the value of the i-th dimension of a semantic vector in the vector database; is the square of the value of the i-th dimension of a semantic vector in the vector database; n is the dimension of the semantic vector, .
[0044] It should be noted that the search results are sorted according to the calculation results of the similarity, and the search results with a target number (set according to actual needs) at the front are recalled.
[0045] S4. Perform a second similarity calculation on the recalled documents and the expanded query request through the target semantic recognition model and sort the recalled documents in descending order according to the calculated similarities to obtain the target retrieval sorting result, wherein the first similarity calculation and the second similarity calculation adopt different methods.
[0046] It should be understood that, in one embodiment, the above-mentioned target retrieval ranking result includes the ranking result of the recalled documents obtained after the above-mentioned secondary sorting; in another embodiment, the first M documents of the recalled documents obtained after the above-mentioned secondary sorting can be selected as target documents according to the ranking, and the ranking result of the M target documents can be used as the above-mentioned target retrieval ranking result.
[0047] It should be noted that the formula for the second similarity calculation is:
[0048]
[0049] Where: W is the linear layer of the target semantic recognition model; S is the score, ctxt is the semantic vector of the expanded query request, and cand i is the semantic vector of the recalled document; first is a function used to obtain the first vector ([CLS] token) output by the last layer of the model; T(ctxt, cand i ) represents the semantic vector ctxt of the expanded query request and the semantic vector cand of the recalled document i The output layer obtained after inputting into the target semantic recognition model; represents the first vector of the output layer of the target semantic recognition model; y represents the first vector of the output layer of the target semantic recognition model.
[0050] Furthermore, the loss function of the enhanced contrastive learning algorithm is: Contrastive loss formula 1 Where: is the loss function of query q; q is the target semantic vector; It is a semantic vector representation that is positively correlated with q; is the semantic vector representation of q negative correlation; τ is a temperature hyperparameter, and K is the total number of negatively correlated samples; is a positively correlated sample pair Similarity score of are all negatively correlated sample pairs The sum of similarity scores.
[0051] It should be noted that the design idea of the enhanced contrastive learning algorithm is to use a dictionary (a dictionary is a data structure) to maintain the data sample queue. The dictionary consists of multiple batches, which is much larger than the sample pairs of a batch. The model can perform contrastive learning with a large number of negative examples at a time. The sample update in the dictionary adopts momentum: , Where: θ k is a negative sample state value in the dictionary; θ q is the state value of a positive sample in the dictionary; m ∈ [ 0 ,1 ] , m is a momentum coefficient.
[0052] Preferably, adding momentum m enables the model to perform effective contrastive learning on a large number of negative examples while maintaining a certain stability.
[0053] In summary, the beneficial effect of the full-text retrieval enhancement method for hydropower plant knowledge of the present invention is that by constructing a triple structure coding model and introducing an enhanced contrastive learning algorithm, the vocabulary gap and semantic gap problems existing in traditional retrieval methods are effectively solved. Furthermore, the knowledge document is encoded into a semantic vector using the triple structure coding model and stored in a vector database, while synonym expansion and edit distance calculation are performed on user queries to expand the recall range and ensure that more potentially relevant documents can be obtained. Furthermore, an enhanced contrastive learning algorithm is used to perform secondary sorting of the recalled documents, and through the momentum update mechanism and dictionary maintenance strategy, the model's ability to understand and distinguish document semantics is further improved, so that the final retrieval results more accurately match user needs.
[0054] Example 2, reference Figure 4 , which is the second embodiment of the present invention, provides a full-text retrieval enhancement method for hydropower plant knowledge. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0055] We extracted about 240 documents from the existing power plant knowledge documents and cleaned the documents, including removing stop words, punctuation, special symbols, and converting numbers to uppercase. Then, we used the code to extract about 20,000 sentences and shuffled their order. Then, we manually filtered out about 2,000 sentences with similar meanings or low quality, leaving 18,000 sentences as training samples, of which 14,000 were used as training sets and 4,000 were used as test sets. The training samples are grouped into three sentences, one for anchoring, one for positive samples, and one for negative samples. Using the contrastive learning method, for the same training batch, we randomly selected a sentence as the anchor sentence, and used the code to generate positive samples by synonym replacement and back translation. Any other sentence in the batch was used as a negative sample. By training the triple structure encoding model to minimize the loss function, the similarity between the anchor sentence and the positive sample sentence is close, while the similarity between the anchor sentence and the negative sample is increased. Specifically, the training parameter epoch = 6, selection basis: the number of epochs is determined by experimental verification. The reason for setting 6 epochs is that in preliminary experiments, it was found that the model can achieve good performance with this number while avoiding overfitting. If the epoch is set too large, the model will overfit on the training data, resulting in reduced generalization ability; if the epoch is too small, the model may not fully learn the patterns in the training data. Batchsize = 16, selection basis: the choice of batchsize is limited by the existing computing resources. 16 is a medium-sized batchsize, and increasing it will cause memory overflow in the experiment. Reducing the batchsize will increase the training time. Warmup = 0.1, selection basis: warmup is a ratio set based on experience, which is used to gradually increase the learning rate before the learning rate rises to a predetermined value. 0.1 means that the learning rate is gradually increased in the first 10% of the training steps. Warmup can help the model learn stably in the early stage of training and avoid model instability caused by too large an initial learning rate. LR (learning rate) = 2e-5. The learning rate is set to 2e-5 because it is found in experiments that this learning rate can make the model converge stably during training. A learning rate that is too large will make it difficult for the model to converge or even diverge; a learning rate that is too small will lead to a slow training process. Optimizer: ADAM, selection basis: ADAM is a commonly used optimizer that combines the advantages of momentum and adaptive learning rate (Adagrad / RMSprop). ADAM is usually selected because it can converge quickly and has excellent performance in experiments.
[0056] Furthermore, in the 4,000 test sets, synonym replacement was used in the code, and the synonym library used the extended version of the Harbin Institute of Technology Synonym Cilin, which was built by introducing the thesaurus package manager: from pyltp import Segmentor. (The Segmentor class in the pyltp library is imported to facilitate the use of the LTP (Language Technology Platform) word segmentation function in Python programs), and back translation was used to generate sentence pairs. After shuffling the order, 1,800 pairs were randomly selected as positive pairs, and the remaining pairs were manually screened to construct 1,800 negative pairs. Directly using the pre-trained model, the test set accuracy was 30%; training 4 iterations, the test set accuracy was 61%; training set 5 iterations, the test set accuracy was 68%; training set, 6 iterations, the test set accuracy was 71%.
[0057] Furthermore, the target semantic model is obtained by introducing an enhanced contrastive learning algorithm into the triple structure encoding model. The specific training parameters are epoch=6, batchsize=16, warmup= 0.1, LR=3e-6. The actual training adjusts the learning rate to obtain the performance of the model at the expense of time. Momentum parameter: 0.99. The momentum parameter is adjusted through experiments. The value of 0.99 can accelerate the movement of gradient descent in the relevant direction and reduce oscillation. Temperature coefficient: 0.07. The temperature coefficient is used to control the sensitivity of the model to the probability distribution. The temperature coefficient of this training is 0.07. The purpose is to make the output probability distribution sharper and the model can converge faster. Optimizer: ADAM. The training is independent of the retrieval process and is performed independently.
[0058] Furthermore, a comparative study was conducted on the BM25 algorithm (a traditional text retrieval algorithm) based on statistical TF-IDF and this method in a test environment. Documents were divided according to the self-built test set, imported into ElasticSearch, and the text embedding vector was obtained using the OpenAI (Open Artificial Intelligence) embedding model. The corpus consists of 380 document fragments (Chunks), and the corresponding user query (query) and the corresponding document index were manually annotated for each document. The hit rate (Hit Rate) and mean reciprocal ranking (MRR) indicators were used for evaluation. The hit rate refers to the recall text (true value) that will appear in the first k texts of the recall result. The higher the hit rate, the better the algorithm effect; the mean reciprocal ranking measures the average of the inverse of the average ranking of the relevant documents or information returned by the algorithm in a series of queries. The higher the value, the better the retrieval effect.
[0059] The formula for calculating the hit rate is: , Where: Hitrate@K is the hit rate indicator, which indicates the probability that the number of hits in the search results is K; NumberofHits@K is the number of hits, and GT is the total number of searches.
[0060] The mean reciprocal ranking (MRR) formula is: , Where: Q represents the total number of searches, rank i It indicates the position of the most relevant result in the i-th search, and its reciprocal is the quality of the returned results.
[0061] Combining Table 1, Table 2 and Figure 4 It can be seen that with the increase of k (k=1, 2, 3, 4, 5), both the hit rate and MRR indicators increase, but this method has a higher hit rate and higher accuracy than the traditional method.
[0062] Table 1 Evaluation results using this method , Table 2 Evaluation results using BM25 , In summary, compared with the traditional BM25 algorithm, this method has significantly improved indicators such as hit rate and average reciprocal ranking in the test environment, showing higher retrieval accuracy and efficiency. At the same time, it has good semantics and strong robustness to external interference, and can effectively meet the high-quality requirements of hydropower plant knowledge retrieval.
[0063] Embodiment 3 is the third embodiment of the present invention, which is different from the first two embodiments in that: If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0064] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0065] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk case (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0066] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0067] Embodiment 4 is the fourth embodiment of the present invention. This embodiment provides a full-text retrieval enhancement system for hydropower plant knowledge, including a construction module for constructing a triple structure encoding model and a target semantic recognition model, wherein the target semantic recognition model is obtained by introducing an enhanced contrastive learning algorithm training into the triple structure encoding model; A decomposition and expansion module, used to decompose the acquired query request into at least two query keywords, and perform expansion processing on the decomposed query keywords to obtain an expanded query request; The initial sorting module is used to convert the expanded query request into a target semantic vector through a triple structure encoding model, and perform the first similarity calculation between the target semantic vector and the semantic vector in the vector database to obtain the similarity between the target semantic vector and the semantic vector in the vector database, recall related documents according to the similarity, and preliminarily sort the recalled documents according to the similarity from large to small; The refined sorting module is used to perform a second similarity calculation on the recalled documents and the expanded query requests through the target semantic recognition model and to perform a second sorting on the recalled documents in descending order according to the calculated similarities to obtain the target retrieval sorting result, wherein the first similarity calculation and the second similarity calculation adopt different methods.
[0068] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A full-text retrieval enhancement method for hydropower plant knowledge, characterized by: include, A triple structure encoding model and a target semantic recognition model are constructed, wherein the target semantic recognition model is trained by introducing an enhanced contrastive learning algorithm into the triple structure encoding model; Decomposing the acquired query request into at least two query keywords, and performing expansion processing on the decomposed query keywords to obtain an expanded query request; The expanded query request is converted into a target semantic vector through a triple structure encoding model, and the first similarity calculation is performed between the target semantic vector and the semantic vector in the vector database to obtain the similarity between the target semantic vector and the semantic vector in the vector database. The relevant documents are recalled according to the similarity, and the recalled documents are preliminarily sorted from large to small according to the similarity; The target semantic recognition model is used to perform a second similarity calculation on the recalled documents and the expanded query requests, and the recalled documents are sorted secondarily in descending order according to the calculated similarities to obtain the target retrieval sorting result, in which the first similarity calculation and the second similarity calculation are performed in different ways.
2. A full-text retrieval enhancement method for hydropower plant knowledge as claimed in claim 1, characterized in that: The construction of the triple structure coding model includes: Preprocess the acquired hydropower plant knowledge documents to obtain target document data; Three pre-trained models with shared weights are used as the core network to construct the initial triple structure encoding model; Sentences are extracted from target document data to form training samples, and the initial triple structure encoding model is trained and optimized by using a contrastive learning method combined with a vector-based dense retrieval loss function to obtain the triple structure encoding model.
3. A full-text retrieval enhancement method for hydropower plant knowledge as claimed in claim 2, characterized in that: The semantic vector is obtained by encoding the target document data through a triple structure encoding model; The vector database is a database storing the semantic vectors.
4. A full-text retrieval enhancement method for hydropower plant knowledge as claimed in claim 3, characterized in that: Extended processing includes, According to the synonym dictionary and edit distance calculation, synonyms with similar semantics to the query keywords are selected to generate an expanded query matrix.
5. A full-text search enhancement method for hydropower plant knowledge as claimed in claim 4, characterized in that: The extracting sentences from the target document data to form training samples includes extracting multiple groups of sentences from the target document data to form the training samples, wherein each group of sentences includes: sentences used as anchors, positive sample sentences, and negative sample sentences.
6. A method for enhancing full-text retrieval of hydropower plant knowledge as claimed in claim 5, characterized in that: The formula for the first similarity calculation is: , Where: Similarity is the similarity score; q is the target semantic vector; d is the semantic vector in the vector database; Represents the modulus length of the target semantic vector q; Represents the modulus length of the semantic vector d in the vector database; is the value of the i-th dimension of the target semantic vector; is the square of the value of the i-th dimension of the target semantic vector; is the value of the i-th dimension of a semantic vector in the vector database; is the square of the value of the i-th dimension of a semantic vector in the vector database; n is the dimension of the semantic vector, .
7. A method for enhancing full-text retrieval of hydropower plant knowledge as claimed in claim 6, characterized in that: The enhanced contrastive learning algorithm loss function is: , Where: is the loss function of query q; q is the target semantic vector; It is a semantic vector representation that is positively correlated with q; is the semantic vector representation of q negative correlation; τ is a temperature hyperparameter, and K is the total number of negatively correlated samples; is a positively correlated sample pair Similarity score of ; are all negatively correlated sample pairs The sum of similarity scores.
8. A full-text search enhancement system for hydropower plant knowledge, using a full-text search enhancement method for hydropower plant knowledge as claimed in any one of claims 1 to 7, characterized in that: include: A construction module is used to construct a triple structure encoding model and a target semantic recognition model, wherein the target semantic recognition model is obtained by introducing an enhanced contrastive learning algorithm into the triple structure encoding model; A decomposition and expansion module, used to decompose the acquired query request into at least two query keywords, and perform expansion processing on the decomposed query keywords to obtain an expanded query request; The initial sorting module is used to convert the expanded query request into a target semantic vector through a triple structure encoding model, and perform the first similarity calculation between the target semantic vector and the semantic vector in the vector database to obtain the similarity between the target semantic vector and the semantic vector in the vector database, recall related documents according to the similarity, and preliminarily sort the recalled documents according to the similarity from large to small; The refined sorting module is used to perform a second similarity calculation on the recalled documents and the expanded query requests through the target semantic recognition model and to perform a second sorting on the recalled documents in descending order according to the calculated similarities to obtain the target retrieval sorting result, wherein the first similarity calculation and the second similarity calculation adopt different methods.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of a full-text retrieval enhancement method for hydropower plant knowledge described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a full-text retrieval enhancement method for hydropower plant knowledge described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Document retrieval method and system based on deep learning
CN115495555A
Method and system for recommending legal provisions, electronic equipment and storage medium
CN116662643A
Text classification model training method and device
CN118733761A
Archive information resource intelligent sharing method and system based on AI
CN119149704A
Enterprise-level knowledge management system based on large language model
CN119476460A
Cited By
Literature semantic search method and system based on elastic search
CN120429311A