Data search enhanced rearrangement method, system and equipment
By constructing a domain keyword database and generating multi-level feature vectors, combining the screening and rearrangement methods of transfer probability distance and spherical divergence distance, the problem of low retrieval accuracy and accuracy in RAG technology is solved, and more efficient semantic matching and the accuracy of generated content is achieved.
Patent Information
- Application Number
- CN202510552027.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing RAG technology is difficult to adapt to complex semantic scenarios and long-tail queries due to local matching, single representation and noise interference.
By constructing a domain keyword database, text blocks are divided and first and second high-dimensional feature vectors are generated, and secondary screening is performed by combining keyword matching and transfer probability distances, and candidate text blocks are rearranged through spherical divergence distances to improve the accuracy and adaptability of the search results.
It significantly improves the accuracy of professional fields and open domain searches, enhances the robustness of semantic matching in complex semantic scenarios, reduces the risks of noise interference and mismatch, and optimizes the factual consistency of the generated content.
Smart Images

Figure CN120067311A_ABST
Abstract
Description
Background Art
[0002] In recent years, large language models (LLMs) represented by ChatGPT have demonstrated capabilities approaching or even exceeding human levels in text understanding, generation, and logical reasoning tasks, driving the intelligent transformation of various industries. However, there are still significant deficiencies in the core technical framework of large models: Randomness of autoregressive prediction: The model generates content word by word based on historical tokens, and its probabilistic output mechanism leads to uncertainty in the generated results.
[0003] Noise in pre-training data and knowledge limitations: The timeliness and accuracy of the training corpus are insufficient, making the model prone to "hallucinations" in open-domain or professional domain tasks, that is, generating content that does not conform to facts or contains logical errors.
[0004] To alleviate the above problems, Retrieval-Augmented Generation (RAG) has become the mainstream solution. It retrieves external knowledge bases or Internet information and inputs relevant text fragments as context into the model to constrain the accuracy of the generated content. However, existing RAG solutions still face the following bottlenecks: Insufficient retrieval accuracy: Traditional methods rely on local similarity metrics such as cosine similarity for retrieval, making it difficult to distinguish content with similar semantics but factual conflicts.
[0005] Lack of distribution modeling: Existing solutions represent queries and text chunks with a single vector, ignoring the semantic distribution characteristics of the text, resulting in poor adaptability to long-tail queries or complex semantic scenarios.
[0006] Significant noise interference: Among the candidate text chunks retrieved initially, the misleading of low-correlation content in the generation stage is difficult to eliminate through simple threshold screening, and the actual application accuracy is usually only 70%-80%.
[0007] Based on this, the present invention proposes a data search enhancement and rearrangement method, system, and device. Summary of the Invention
[0008] To solve the above problems in the prior art, that is, the low accuracy of existing RAG technologies due to local matching, single representation, and noise interference, the present invention provides a data search enhancement and rearrangement method, system, and device.
[0009] In the first aspect of the present invention, a data search enhancement and rearrangement method is proposed, which includes: Collect professional knowledge and documents in the existing field and extract keywords to construct a keyword library corresponding to the field; Divide the professional knowledge and documents into multiple text blocks according to a fixed length. For each text block, extract keywords based on a keyword library respectively, and generate a first high-dimensional feature vector through a text representation model; Perform semantic segmentation on each text block to generate multiple semantic units, and extract high-dimensional feature vectors of each semantic unit through a text representation model to obtain a second high-dimensional feature vector; Extract keywords of the input query information based on the keyword library, and extract high-dimensional feature vectors corresponding to the keywords of the query information; obtain multiple semantically equivalent rewritten queries of the query information, and extract high-dimensional feature vectors corresponding to each rewritten query; Screen multiple text blocks that match the keywords of the query information as candidate text blocks; Match the high-dimensional feature vector corresponding to the keyword of the query information with the first high-dimensional feature vector, and perform secondary screening on the candidate text blocks based on the transfer probability distance; Calculate the spherical divergence distance between the high-dimensional feature vector corresponding to the rewritten query and the second high-dimensional feature vector after secondary screening, and rearrange the text blocks after secondary screening based on the spherical divergence distance.
[0010] Further, the method for performing secondary screening on the candidate text blocks based on the transfer probability distance is as follows: Combine the high-dimensional feature vector corresponding to the keyword of the query information with the first high-dimensional feature vectors of all candidate text blocks into a joint feature set; Calculate the cosine similarity between each feature vector in the joint feature set and other feature vectors; For each feature vector, retain the similarities corresponding to the top K neighbors with the highest cosine similarity, and set the remaining similarities to zero to obtain a similarity matrix, and perform normalization row by row to obtain a probability transfer matrix; Based on the probability transfer matrix, calculate the transfer probability distance between the high-dimensional feature vector corresponding to the keyword of the query information and the first high-dimensional feature vector of each candidate text block, and perform secondary screening on the candidate text blocks according to the sorting of the transfer probability distance and a preset distance threshold.
[0011] Further, the transfer probability distance , and its calculation method is: ; Among them, is the high-dimensional feature vector corresponding to the keyword of the query information, is the first high-dimensional feature vector of the i th candidate text block, N is the number of candidate text blocks, is transferring to thej The probability of a feature vector is the probability of transitioning to the j th feature vector.
[0012] Furthermore, the similarity matrix is calculated as follows: ; where is the j th feature vector in the joint feature set, is the similarity matrix between and
[0013] Furthermore, the spherical divergence distance is calculated as follows: Denote the set of high-dimensional feature vectors of each rewritten query as , and the set of second high-dimensional feature vectors after secondary screening as ; where is the number of feature vectors, m is the index of; Calculate and the Euclidean distance between any two feature vectors in ; where , ; Calculate the spherical divergence distance of each group of feature vectors based on the Euclidean distance.
[0014] Furthermore, calculate the spherical divergence distance of each group of feature vectors based on the Euclidean distance , and the method is as follows: ; where represents the within-group distance between the high-dimensional feature vectors of the rewritten query, is the within-group distance between the second high-dimensional feature vectors after secondary screening, is the loop index variable used to traverse the elements in the set, and its value range is , is the th feature vector in.
[0015] Furthermore, rearrange the text blocks after secondary screening based on the spherical divergence distance, and the method is as follows: Arrange the text blocks corresponding to each second high-dimensional feature vector after secondary screening in ascending order based on the spherical divergence distance to obtain the rearrangement result.
[0016] In a second aspect of the present invention, a data search enhancement and rearrangement system is proposed, based on a data search enhancement and rearrangement method. The system includes: A keyword library construction module configured to collect professional knowledge and documents in existing fields and extract keywords to construct a keyword library corresponding to the fields; A first high-dimensional feature vector generation module configured to divide the professional knowledge and documents into multiple text blocks according to a fixed length. For each text block, keywords are extracted based on the keyword library, and a first high-dimensional feature vector is generated through a text representation model; A second high-dimensional feature vector generation module configured to perform semantic segmentation on each text block to generate multiple semantic units, and extract high-dimensional feature vectors of each semantic unit through a text representation model to obtain a second high-dimensional feature vector; A query information processing module configured to extract keywords of the input query information based on the keyword library, and extract high-dimensional feature vectors corresponding to the keywords of the query information; obtain multiple semantically equivalent rewritten queries of the query information, and extract high-dimensional feature vectors corresponding to each rewritten query; A candidate text block initial screening module configured to screen multiple text blocks that match the keywords of the query information as candidate text blocks; A candidate text block secondary screening module configured to match the high-dimensional feature vector corresponding to the keyword of the query information with the first high-dimensional feature vector, and perform secondary screening on the candidate text blocks based on the transfer probability distance; A rearrangement module configured to calculate the spherical divergence distance between the high-dimensional feature vector corresponding to the rewritten query and the second high-dimensional feature vector after secondary screening, and rearrange the text blocks after secondary screening based on the spherical divergence distance.
[0017] In a third aspect of the present invention, an electronic device is proposed, including: At least one processor; and A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned data search enhancement and rearrangement method.
[0018] In a fourth aspect of the present invention, a computer-readable storage medium is proposed. The computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned data search enhancement and rearrangement method.
[0019] Advantages of the present invention: Enhancing domain adaptability and retrieval accuracy: Based on the text block screening mechanism of the domain keyword library, combining keyword matching and high-dimensional semantic feature matching, it can effectively distinguish text contents with similar semantics but conflicting facts, and significantly improve the accuracy of retrieval in professional domains and the open domain.
[0020] Enhancing complex semantic modeling capabilities: Through the semantic unit segmentation and multi-granularity feature representation of text blocks (the first and second high-dimensional feature vectors), it breaks through the limitations of traditional single-vector representation, accurately captures long-tail queries, polysemous expressions, and implicit semantic associations, and improves the robustness of semantic matching in complex scenarios.
[0021] Reducing noise interference and the risk of mis-matching: Adopting a three-level filtering mechanism of "initial keyword screening - secondary screening based on transition probability - rearrangement based on spherical divergence", modeling the context relevance through the transition probability distance, and combining the spherical divergence distance to quantify the semantic distribution difference, it suppresses the interference of low-correlation text blocks in the generation stage and reduces the "hallucination" phenomenon of the model.
[0022] Optimizing the factual consistency of generated content: Through the rearrangement optimization of retrieval results, it ensures that high-correlation text blocks are preferentially input into the large model, strengthens the anchoring effect of the generated content with the authoritative knowledge base, and improves the logical accuracy and factual credibility of the generated results.
[0023] System flexibility and scalability: The modular design supports the dynamic update of the keyword library and the independent optimization of the text representation model, adapts to different domains and multi-language scenarios, reduces the deployment cost, and provides an efficient technical foundation for the intelligent transformation of the industry. Description of the Drawings
[0024] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more obvious: Figure 1 It is a schematic flowchart of a data search enhancement rearrangement method of the present invention. Detailed Embodiments
[0025] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and do not limit the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0026] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0027] The present invention provides a data search enhancement rearrangement method, which includes: Step S10: Collect the professional knowledge and documents in the existing field, extract keywords, and construct a keyword library corresponding to the field. Step S20: Divide the professional knowledge and documents into multiple text blocks according to a fixed length. For each text block, extract keywords based on the keyword library respectively, and generate a first high-dimensional feature vector through a text representation model. Step S30: Perform semantic segmentation on each text block to generate multiple semantic units, and extract the high-dimensional feature vectors of each semantic unit through a text representation model to obtain a second high-dimensional feature vector. Step S40: Extract the keywords of the input query information based on the keyword library, and extract the high-dimensional feature vectors corresponding to the keywords of the query information; obtain multiple semantically equivalent rewritten queries of the query information, and extract the high-dimensional feature vectors corresponding to each rewritten query. Step S50: Screen multiple text blocks that match the keywords of the query information as candidate text blocks. Step S60: Match the high-dimensional feature vector corresponding to the keyword of the query information with the first high-dimensional feature vector, and perform secondary screening on the candidate text blocks based on the transfer probability distance. Step S70: Calculate the spherical divergence distance between the high-dimensional feature vector corresponding to the rewritten query and the second high-dimensional feature vector after secondary screening, and rearrange the text blocks after secondary screening based on the spherical divergence distance.
[0028] To more clearly illustrate a data search enhancement and rearrangement method of the present invention, the following combines Figure 1 Expand and detail steps S10 - S70 in the embodiments of the present invention, and the detailed descriptions of each step are as follows: Step S10: Collect the professional knowledge and documents in the existing field, extract keywords, and construct a keyword library corresponding to the field. The professional knowledge in this embodiment is the expert knowledge and experience in the field, and the documents are the relevant documents in the existing field. Based on the expert knowledge and experience in the field, combined with the relevant documents in the existing field, use a large language model (such as deepseek-R1) to summarize and extract corresponding keywords to establish a keyword library.
[0029] Step S20: Divide the professional knowledge and documents into multiple text blocks according to a fixed length. For each text block, extract keywords based on the keyword library respectively, and generate a first high-dimensional feature vector through a text representation model. In this embodiment, first divide the document into different text blocks based on a fixed length (such as 1024). For each text block, use a large language model (LLM) to extract keywords based on the keyword library and use a text representation model to extract high-dimensional representations respectively.
[0030] The text representation model can select the BERT model.
[0031] The high-dimensional feature vector in the present invention refers to converting a text block into a numerical vector with a fixed dimension (such as 768 dimensions, 1024 dimensions), and each dimension corresponds to a potential semantic or syntactic feature.
[0032] Step S30: Perform semantic segmentation on each text block to generate multiple semantic units, and extract the high-dimensional feature vectors of each semantic unit through the text representation model to obtain the second high-dimensional feature vector; In this embodiment, for each text block, it is segmented into several semantics by using a large language model (LLM), and then for each semantics, its high-dimensional representation is extracted by using the text representation model.
[0033] The present invention introduces the ability of a large language model (LLM), which can adaptively extract semantic information from the text blocks segmented by the knowledge base, making up for the representation defects caused by semantic loss / excessiveness in the existing fixed-length segmentation.
[0034] Step S40: Extract the keywords of the input query information based on the keyword library, and extract the high-dimensional feature vectors corresponding to the keywords of the query information; obtain multiple semantically equivalent rewritten queries of the query information, and extract the high-dimensional feature vectors corresponding to each rewritten query; In this embodiment, the large language model (LLM) is also used to extract the keywords of the input query information and perform several rewrites that maintain the semantics but have different expressions.
[0035] Step S50: Screen multiple text blocks that match the keywords of the query information as candidate text blocks; In this embodiment, the corresponding candidate text blocks are preliminarily screened according to the matching similarity technology (TF-IDF).
[0036] Step S60: Match the high-dimensional feature vector corresponding to the keywords of the query information with the first high-dimensional feature vector, and perform secondary screening on the candidate text blocks based on the transfer probability distance; In this embodiment, secondary screening is performed on the candidate text blocks based on the transfer probability distance, and the method is as follows: Combine the high-dimensional feature vector corresponding to the keywords of the query information with the first high-dimensional feature vectors of all candidate text blocks into a joint feature set; Calculate the cosine similarity between each feature vector in the joint feature set and other feature vectors; For each feature vector, retain the similarities corresponding to the top K neighbors with the highest cosine similarity, and set the remaining similarities to zero to obtain a similarity matrix, and perform normalization row by row to obtain a probability transfer matrix; Based on the probability transition matrix, calculate the transition probability distance between the high-dimensional feature vector corresponding to the keyword of the query information and the first high-dimensional feature vector of each candidate text block, and perform a secondary screening on the candidate text blocks according to the sorting of the transition probability distances and a preset distance threshold.
[0037] Assume that the high-dimensional feature vectors corresponding to the query information and the candidate text block in the corresponding second high-dimensional feature vectors are respectively and , then we can calculate a similarity matrix S , where it satisfies: ; where, is the j th feature vector in the joint feature set, is and the similarity matrix between.
[0038] Calculate , we can obtain the probability transition matrix P , where , so our transition probability distance , and its calculation method is: ; where, is the high-dimensional feature vector corresponding to the keyword of the query information, is the i th first high-dimensional feature vector of the candidate text block, N is the number of candidate text blocks, is the probability of transferring to the j th feature vector, is the probability of transferring to the j th feature vector.
[0039] Specifically, i represents the index of the feature vector, corresponding to the row number of the matrix.
[0040] i = 0, it means the high-dimensional feature vector corresponding to the keyword of the query information.
[0041] i = 1, 2,..., N when, it means the first high-dimensional feature vector .
[0042] j Indicates the index of the target feature vector, corresponding to the column number of the matrix.
[0043] j When it is 0, the target is the high-dimensional feature vector .
[0044] j When it is 1, 2, … N , the target is the j th candidate text block feature vector .
[0045] In the high-dimensional feature comparison and screening stage, different from the traditional cosine distance, we introduce the transition probability distance. This distance has very good properties: its value range is between [0, 1], and the distance is 0 only when two points have exactly the same neighbors; this distance is not affected by the data distribution, and there can be a unified distance standard for any data distribution. Compared with the cosine distance, this distance can greatly improve the robustness of screening.
[0046] Step S70, calculate the spherical divergence distance between the high-dimensional feature vector corresponding to the rewritten query and the second high-dimensional feature vector after secondary screening, and rearrange the text blocks after secondary screening based on the spherical divergence distance.
[0047] In this embodiment, for the spherical divergence distance, the method is as follows: Denote the high-dimensional feature vector set of each rewritten query as , and denote the second high-dimensional feature vector set after secondary screening as ; where is the number of feature vectors, m is the index of; Calculate and the Euclidean distance between any two feature vectors in ; where , ; Calculate the spherical divergence distance of each group of feature vectors based on the Euclidean distance.
[0048] Calculate the spherical divergence distance of each group of feature vectors based on the Euclidean distance , and the method is as follows: ; where represents the within-group distance between the high-dimensional feature vectors of the rewritten query, is the within-group distance between the second high-dimensional feature vectors after secondary screening, is a loop index variable used to traverse the elements in the set, and its value range is , is the th eigenvector in
[0049] In this embodiment, based on the spherical divergence distance, the text blocks after the secondary screening are rearranged, and the method is as follows: Arrange the text blocks corresponding to the second highest-dimensional eigenvectors after each secondary screening in ascending order based on the spherical divergence distance to obtain the rearrangement result.
[0050] Although the above steps are described in the above order in the above embodiments, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order, and they can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.
[0051] A data search enhanced rearrangement system according to the second embodiment of the present invention is based on a data search enhanced rearrangement method, and the system includes: A keyword library construction module configured to collect professional knowledge and documents in the existing field and extract keywords to construct a keyword library corresponding to the field; A first high-dimensional eigenvector generation module configured to divide the professional knowledge and documents into multiple text blocks according to a fixed length, for each text block, extract keywords based on the keyword library respectively, and generate first high-dimensional eigenvectors through a text representation model; A second high-dimensional eigenvector generation module configured to perform semantic segmentation on each text block to generate multiple semantic units, and extract high-dimensional eigenvectors of each semantic unit through a text representation model to obtain second high-dimensional eigenvectors; A query information processing module configured to extract keywords of the input query information based on the keyword library, and extract high-dimensional eigenvectors corresponding to the keywords of the query information; obtain multiple semantically equivalent rewritten queries of the query information, and extract high-dimensional eigenvectors corresponding to each rewritten query; A candidate text block primary screening module configured to screen multiple text blocks that match the keywords of the query information as candidate text blocks; A candidate text block secondary screening module configured to match the high-dimensional eigenvectors corresponding to the keywords of the query information with the first high-dimensional eigenvectors, and perform secondary screening on the candidate text blocks based on the transfer probability distance; A rearrangement module configured to calculate the spherical divergence distance between the high-dimensional feature vector corresponding to the rewritten query and the second high-dimensional feature vector after secondary screening, and rearrange the text blocks after secondary screening based on the spherical divergence distance. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and related descriptions of the above-described system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated herein.
[0052] It should be noted that the data search enhanced rearrangement system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiment can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. For the names of the modules and steps involved in the embodiments of the present invention, they are only used to distinguish each module or step and are not regarded as an improper limitation of the present invention.
[0053] An electronic device according to a third embodiment of the present invention includes: At least one processor; and A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above data search enhanced rearrangement method.
[0054] A computer-readable storage medium according to a fourth embodiment of the present invention, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above data search enhanced rearrangement method.
[0055] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and related descriptions of the above-described storage device and processing device can refer to the corresponding process in the foregoing method embodiments and will not be elaborated herein.
[0056] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0057] The terms "first", "second", etc. are used to distinguish similar objects, rather than to describe or represent a specific order or sequence.
[0058] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, so that a process, method, article, or device / apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or also elements inherent to these process, method, article, or device / apparatus.
[0059] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. A data search enhancement rearrangement method, characterized in that: The method includes: Collect professional knowledge and documents in existing fields and extract keywords to build a keyword library corresponding to the field; The professional knowledge and documents are divided into a plurality of text blocks according to a fixed length, and for each text block, keywords are extracted based on a keyword library, and a first high-dimensional feature vector is generated through a text representation model; Perform semantic segmentation on each text block to generate multiple semantic units, and extract the high-dimensional feature vector of each semantic unit through the text representation model to obtain the second high-dimensional feature vector; Extracting keywords of the input query information based on the keyword library, and extracting high-dimensional feature vectors corresponding to the keywords of the query information; obtaining multiple semantically equivalent rewritten queries of the query information, and extracting high-dimensional feature vectors corresponding to each rewritten query; Filtering multiple text blocks that match the keywords of the query information as candidate text blocks; Matching the high-dimensional feature vector corresponding to the keyword of the query information with the first high-dimensional feature vector, and performing secondary screening on the candidate text blocks based on the transition probability distance; The spherical divergence distance between the high-dimensional feature vector corresponding to the rewritten query and the second high-dimensional feature vector after secondary screening is calculated, and the text blocks after secondary screening are rearranged based on the spherical divergence distance.
2. A data search enhancement rearrangement method according to claim 1, characterized in that: The candidate text blocks are screened twice based on the transition probability distance, and the method is as follows: Combining the high-dimensional feature vector corresponding to the keyword of the query information and the first high-dimensional feature vectors of all candidate text blocks into a joint feature set; Calculating the cosine similarity between each feature vector and other feature vectors in the joint feature set; For each eigenvector, retain the similarities corresponding to the first K neighbors with the highest cosine similarity, set the remaining similarities to zero, obtain the similarity matrix, and normalize it row by row to obtain the probability transfer matrix; Based on the probability transfer matrix, the transfer probability distance between the high-dimensional feature vector corresponding to the keyword of the query information and the first high-dimensional feature vector of each candidate text block is calculated, and the candidate text blocks are secondary screened according to the order of the transfer probability distances and a preset distance threshold.
3. A data search enhancement rearrangement method according to claim 2, characterized in that: Transition probability distance , which is calculated as: ; in, is the high-dimensional feature vector corresponding to the keyword of the query information, For the i The first high-dimensional feature vector of the candidate text block, N is the number of candidate text blocks, for Transfer to j The probability of a feature vector, for Transfer to j The probability of a feature vector.
4. A data search enhancement rearrangement method according to claim 3, characterized in that: The similarity matrix is calculated as follows: ; in, is the first j feature vectors, for and The similarity matrix between .
5. A data search enhancement rearrangement method according to claim 1, characterized in that: The ball divergence distance, the method is: The high-dimensional feature vector set of each rewritten query is recorded as , the second highest dimensional feature vector set after secondary screening is recorded as ;in, is the number of eigenvectors, m for The index of calculate and The Euclidean distance between any two eigenvectors is ;in, , ; The spherical divergence distance of each set of feature vectors is calculated based on the Euclidean distance.
6. A data search enhancement rearrangement method according to claim 5, characterized in that: The spherical divergence distance of each set of feature vectors is calculated based on the Euclidean distance , the method is: ; in, represents the intra-group distance between high-dimensional feature vectors of the rewritten query, is the intra-group distance between the second highest dimensional feature vectors after secondary screening, Is a loop index variable used to traverse the elements in the collection. The value range is , for The feature vectors.
7. A data search enhancement rearrangement method according to claim 5, characterized in that: The text blocks after the secondary screening are rearranged based on the spherical divergence distance, and the method is as follows: The text blocks corresponding to the second highest dimensional feature vectors after each secondary screening are sorted in ascending order based on the spherical divergence distance to obtain a rearrangement result.
8. A data search enhancement and rearrangement system, based on a data search enhancement and rearrangement method according to any one of claims 1 to 7, characterized in that: The system includes: A keyword library construction module is configured to collect professional knowledge and documents in existing fields and extract keywords to build a keyword library corresponding to the field; A first high-dimensional feature vector generation module is configured to divide the professional knowledge and documents into a plurality of text blocks according to a fixed length, extract keywords from each text block based on a keyword library, and generate a first high-dimensional feature vector through a text representation model; A second high-dimensional feature vector generation module is configured to perform semantic segmentation on each text block to generate multiple semantic units, and extract a high-dimensional feature vector of each semantic unit through a text representation model to obtain a second high-dimensional feature vector; A query information processing module configured to extract keywords of the input query information based on the keyword library, and extract high-dimensional feature vectors corresponding to the keywords of the query information; obtain multiple semantically equivalent rewritten queries of the query information, and extract high-dimensional feature vectors corresponding to each rewritten query; A candidate text block preliminary screening module configured to screen multiple text blocks matching the keywords of the query information as candidate text blocks; a candidate text block secondary screening module, configured to match the high-dimensional feature vector corresponding to the keyword of the query information with the first high-dimensional feature vector, and perform secondary screening on the candidate text block based on the transition probability distance; The rearrangement module is configured to calculate the spherical divergence distance between the high-dimensional feature vector corresponding to the rewritten query and the second high-dimensional feature vector after secondary screening, and rearrange the text blocks after secondary screening based on the spherical divergence distance.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement a data search enhancement rearrangement method based on any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement a data search enhancement rearrangement method based on any one of claims 1-7.
Citation Information
Patent Citations
Language model training method, NLP task processing method and device
CN113420123A
Information retrieval method and device, electronic equipment and storage medium
CN118394916A
Nuclear power knowledge base retrieval method and system fusing semantic features and keywords
CN118626591A
Retrieval enhancement generation-based retrieval method, product, equipment and medium
CN119003795A
Structured Pruning of Vision Transformer
US20230073835A1