Similar text retrieval method and system based on multi-dimensional information reordering
By adopting a multi-dimensional information reordering method in similar text retrieval, using text key information extraction model and text summary information generation model, and integrating multiple representational search results, the problem of difficulty in long text representation is solved, and the search accuracy and quality are improved.
Patent Information
- Application Number
- CN202510036206.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
Existing similar text retrieval techniques are difficult to effectively characterize long texts, resulting in low retrieval accuracy.
The similar text search method based on multi-dimensional information reordering is adopted, and the text after the clause is processed through the pre-trained text key information extraction model and the text summary information generation model, and the search results of similar texts of multiple representations are integrated to optimize the search quality.
It improves the accuracy of similar text retrieval, effectively solves the problem of difficulty in long text representation, and optimizes the search quality.
Smart Images

Figure CN119938883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and in particular to a similar text retrieval method and system based on multi-dimensional information reordering. Background Art
[0002] With the development of informatization, information resources in all walks of life are becoming increasingly huge, which poses a challenge to the development of resource management and provides strong data support for decision-making assistance. In many industries such as law, medicine, and news, it is a common method to search for similar texts and match similar cases in a huge historical database to provide reference ideas for current decision-making. In the medical field, the search for similar texts is reflected in the recommendation of similar cases, and in the legal field, the search for similar texts is reflected in the search for similar cases. Taking similar case search as an example, it refers to the use of information retrieval technology to retrieve cases similar to a given legal issue or case from a legal database by analyzing the similarity and correlation between case texts. Specifically, similar case search technology refers to comparing a given query text with legal documents in a large number of search documents, and returning the most similar n document IDs (unique identification index of the search document). In this process, the degree of similarity between the query and search documents should be judged based on the three aspects of basic facts, dispute focus, and legal application issues.
[0003] However, existing similar text retrieval technologies mainly use similar text retrieval based on pre-trained model representation. Although this method can perform fast and efficient similar text recommendation, it still faces the technical problem of difficulty in representing long texts. Summary of the invention
[0004] 1. Technical issues to be solved
[0005] In view of the deficiencies of the prior art, the present invention provides a similar text retrieval method and system based on multi-dimensional information reordering, which solves the technical problem that the existing similar text retrieval technology is difficult to characterize long texts.
[0006] (II) Technical solution
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0008] In a first aspect, the present invention provides a similar text retrieval method based on multi-dimensional information reordering, comprising:
[0009] S1. Obtain the text to be processed and divide the text into sentences;
[0010] S2, processing the sentence-separated texts by using a pre-trained text key information extraction model and a text summary information generation model, respectively, to obtain a first representation-based similar text retrieval result and a second representation-based similar text retrieval result;
[0011] S3. Fusion the first representation formula similar text retrieval result and the second representation formula similar text retrieval result to obtain a final similar text retrieval result.
[0012] Preferably, the text to be processed is a legal document, the text key information extraction model is a case key information extraction model based on ERNIE, and the text summary information generation model is a case summary information generation model based on ChatglmV2.
[0013] Preferably, the sentence segmentation of the text includes: sentence segmentation of the text by terminators.
[0014] Preferably, the fusion of the first representational similar text search result and the second representational similar text search result includes:
[0015] The result set fusion method based on the position score fuses the first representation formula similar text retrieval result and the second representation formula similar text retrieval result.
[0016] Preferably, the position score-based result set fusion method fuses the first representation formula similar text retrieval result and the second representation formula similar text retrieval result, including:
[0017] The result set document format is as follows:
[0018] {
[0019] "QueryId1":[DocID1,DocID2,DocID3...DocID n ],
[0020] "QueryId2":[DocID1,DocID2,DocID3...DocIDn],
[0021] …
[0022] “QueryId m ":[DocID1,DocID2,DocID3...DocIDn],
[0023] }
[0024] Among them, n represents the number of the top n closest document sets, QueryID and DocID are the unique identification IDs of the query and the search set respectively:
[0025] The fusion steps include:
[0026] Step 1: Take out the search set list under the same QueryID of multiple result sets;
[0027] Step 2: Calculate the number of occurrences of each DocID in the result set list, and record it in the format of [DocID:Num];
[0028] Step 3: Record the position POS of each DocID in the result list and add them up, recording them as [DocID:POS]
[0029] Step 4: Sort the occurrence count list in descending order to get the sorted list as follows:
[0030] [Doc3:[Num:4,POS:5],Doc2:[Num:3,POS:7,Doc1:[Num:3,POS:8],Doc2:[Num:0,POS:0]];
[0031] Step 4: For Doc2:[Num:3,POS:7,Doc1:[Num:3,POS:8] with the same number of occurrences, sort them in ascending order according to their position scores to obtain the final similar text retrieval results.
[0032] In a second aspect, the present invention provides a similar text retrieval system based on multi-dimensional information re-ranking, comprising:
[0033] The text acquisition module is used to acquire the text to be processed and divide the text into sentences;
[0034] A text processing module, used to process the text after sentence segmentation through a pre-trained text key information extraction model and a text summary information generation model, respectively, to obtain a first representation type similar text retrieval result and a second representation type similar text retrieval result;
[0035] The fusion module is used to fuse the first representation formula similar text retrieval result and the second representation formula similar text retrieval result to obtain the final similar text retrieval result.
[0036] Preferably, the text to be processed is a legal document, the text key information extraction model is a case key information extraction model based on ERNIE, and the text summary information generation model is a case summary information generation model based on ChatglmV2.
[0037] Preferably, the sentence segmentation of the text includes: sentence segmentation of the text by terminators.
[0038] In a third aspect, the present invention provides a storage medium storing a computer program for similar text retrieval based on multi-dimensional information reordering, wherein the computer program enables a computer to execute the similar text retrieval method based on multi-dimensional information reordering as described above.
[0039] In a fourth aspect, the present invention provides an electronic device, comprising:
[0040] one or more processors;
[0041] Memory; and
[0042] One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including instructions for executing the similar text retrieval method based on multi-dimensional information re-ranking as described above.
[0043] (III) Beneficial effects
[0044] The present invention provides a similar text retrieval method and system based on multi-dimensional information reordering. Compared with the prior art, it has the following beneficial effects:
[0045] The present invention first obtains the text to be processed and divides the text into sentences, and then processes the text after sentence division through the pre-trained text key information extraction model and text summary information generation model, respectively, to obtain the first representation formula similar text retrieval result and the second representation formula similar text retrieval result; finally, the first representation formula similar text retrieval result and the second representation formula similar text retrieval result are integrated to obtain the final similar text retrieval result. The present invention solves the problem of long documents and low accuracy in the retrieval process by a re-ranking method based on multi-dimensional information. The multi-dimensional information includes text key information and text summary information, and the re-ranking is a position-based re-ranking method. The text key information is extracted by the text key information extraction model, and the text summary information is summarized and generated by the text summary information generation model, thereby optimizing the retrieval quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0047] Figure 1 This is a flowchart of the BM25 similar case retrieval method;
[0048] Figure 2A block diagram of a similar text retrieval method based on multi-dimensional information reordering according to an embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of the effect of the case key information extraction model;
[0050] Figure 4 A schematic diagram showing the effect of the model for generating case summary information;
[0051] Figure 5 This is an example diagram of the result set fusion method based on position score;
[0052] Figure 6 This is an example diagram of the result set fusion method based on logarithmic gain. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0054] The embodiments of the present application provide a similar text retrieval method and system based on multi-dimensional information reordering, thereby solving the technical problem that the existing similar text retrieval technology is difficult to characterize long texts, and utilizing structured information and summary information to solve the long document problem and accuracy problem in similar text retrieval, thereby improving the retrieval accuracy.
[0055] The technical solution in the embodiment of the present application is to solve the above technical problems, and the overall idea is as follows:
[0056] It should be noted that, in the embodiment of the present invention, similar case retrieval is taken as an example to explain similar text retrieval in detail.
[0057] The development of similar case retrieval technology has mainly gone through the following three stages:
[0058] Phase 1: Based on lexical features: represented by the Tf-idf and BM25 models of statistical word frequency features, the relationship network between document citations, and keyword extraction and matching methods. This method helps people use basic algorithms to quickly build judicial retrieval systems, but it cannot capture the semantic information in the case well.
[0059] Phase 2: Mainly based on traditional models: Represented by deep models of Transformer, CNN, and RNN architectures, many researchers have built similar case retrieval systems based on the characteristics of judicial texts, but the traditional architecture has limited input length and is relatively inefficient.
[0060] Phase 3: Pre-trained models in the judicial field: With the continuous development of pre-trained models in the judicial field, people began to use judicial representation models to represent legal texts and to build vector databases for fast and efficient case retrieval. However, they still faced the problem of difficulty in representing long texts.
[0061] The most commonly used technology in the current retrieval system is the similar case retrieval method based on BM25 matching (this method belongs to the technology in stage three). BM25 (Best Matching 25) is a classic information retrieval algorithm that sorts documents by calculating the degree of match between the query and the document. In similar case retrieval, the BM25 algorithm regards legal cases as documents and uses information such as word frequency and document frequency to evaluate the similarity between cases, thereby achieving efficient retrieval. Its advantage is that it uses statistical mathematical models to achieve fast retrieval without the need for training corpus and without limiting the length of the document. The flowchart of the BM25 similar case retrieval method is shown below. Figure 1 As shown in the figure, in the process of retrieval using this method, it relies on the term frequency-inverse document frequency (TF-IDF), which refers to the inverse number of the frequency of a word in the entire document collection. The calculation of the inverse document frequency is usually expressed as the logarithm of the inverse number of documents in the document collection that contain the word. This is done to measure the importance of a word, that is, the more common the word, the lower its inverse document frequency. BM25 will design a weight for each segmented term based on the term frequency-inverse document frequency, and ultimately use it to measure the relevance between the query and the document.
[0062] Although the BM25 similar case retrieval method is superior to many similar case retrieval methods, it still has some defects:
[0063] 1. Limited processing capability for long texts: The BM25 algorithm is mainly based on word frequency and document length, and its processing capability for long texts is relatively weak. In long documents, the distribution of word frequency may cause the algorithm to make inaccurate judgments on relevance, because the high or low word frequency cannot fully reflect the importance of a word.
[0064] 2. Insufficient processing of rare words: The BM25 algorithm has a weak processing capability for rare words. Rare words tend to have a high amount of information, but in BM25, due to its calculation method based on document frequency, the weight of words with fewer occurrences may not be sufficient to accurately reflect their importance.
[0065] 3. Unable to process semantic information: The BM25 algorithm is mainly based on statistical information such as word frequency and document length, and cannot capture the semantic relationship between words. In some tasks that require consideration of semantic information, the performance of the BM25 algorithm may not be as good as that of deep learning-based methods.
[0066] In summary, although the BM25 algorithm shows good performance in certain scenarios, it has certain limitations in processing long texts, rare words, and semantic information, and needs to be reasonably selected based on specific application scenarios.
[0067] In order to solve the above problems, an embodiment of the present invention proposes a similar text retrieval method based on multi-dimensional information reordering. The similar text retrieval method adopts a representational case-like retrieval method. In order to solve the problems of long documents and retrieval accuracy, a case key information extraction and case summary information generation module are introduced to input low-quality documents and focus on the endpoint content containing legal elements. At the same time, a result set fusion scheme based on position scoring is used to further optimize the retrieval quality.
[0068] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0069] The embodiment of the present invention provides a similar text retrieval method based on multi-dimensional information reordering, such as Figure 2 As shown, including:
[0070] S1. Obtain the text to be processed and divide the text into sentences.
[0071] S2, processing the sentence-separated texts by using a pre-trained text key information extraction model and a text summary information generation model, respectively, to obtain a first representation-based similar text retrieval result and a second representation-based similar text retrieval result;
[0072] S3. Fusion the first representation formula similar text retrieval result and the second representation formula similar text retrieval result to obtain a final similar text retrieval result.
[0073] In the embodiment of the present invention, a re-ranking method based on multi-dimensional information is used to solve the problem of long documents and low accuracy in the retrieval process. The multi-dimensional information includes text key information and text summary information. The re-ranking method is based on the position. The text key information is extracted by the text key information extraction model, and the text summary information is summarized and generated by the text summary information generation model, thereby optimizing the retrieval quality.
[0074] Taking similar case retrieval as an example, the similar text retrieval method based on multi-dimensional information reordering according to an embodiment of the present invention is described in detail below.
[0075] In step S1, the text to be processed is obtained and the text is divided into sentences. The specific implementation process is as follows:
[0076] Obtain the legal document to be processed, and separate the legal document into sentences using terminators.
[0077] In step S2, the text after sentence segmentation is processed by the pre-trained text key information extraction model and text summary information generation model to obtain the first representation similar text retrieval result and the second representation similar text retrieval result. The specific implementation process is as follows:
[0078] In similar case retrieval, the text key information extraction model is the case key information extraction model, and the text summary information generation model is the case summary information generation model.
[0079] The embodiment of the present invention uses ERNIE (Enhanced Representation through Knowledge Integration) as the key case information extraction model. ERNIE is a pre-trained language model proposed by Baidu. It is improved based on the BERT model and optimized in terms of knowledge fusion. It is suitable for tasks such as text classification and semantic matching. In the key case information extraction module, the data set with basic facts and dispute focus labels is first used as training data, and is input into a variety of deep models and pre-trained models as components of the extraction module. The extraction module is then used to predict the retrieved legal document sentences. With the powerful classification ability of the pre-trained model (ERNIE), sentences with key case information such as basic facts and dispute focus are classified as 1, and then organized into the format of case cause + basic facts + dispute focus + legal provisions + crime + punishment. A representational retrieval of similar cases is performed based on the data, thereby completing the retrieval of similar cases based on key case information. Figure 3 shown.
[0080] In the embodiment of the present invention, ChatglmV2 is selected as the case summary information generation model. ChatglmV2 is a language model based on the Transformer architecture and has 6B (6.2 billion parameters can be written). It uses a model structure called GLM, which has good parallel computing capabilities and can generate high-quality text at a faster speed. It is a large-scale dialogue model with rich knowledge. The case summary information generation model utilizes the rich knowledge and semantic understanding capabilities of the large model to summarize the lengthy legal documents. After many tests, the large model selected is ChatglmV2, which uses a prompt-based method to generate summary information content. The format of the prompt is "Assuming that you are a Chinese lawyer, please help me summarize the content of the paragraph in []: [content]", such as Figure 4 shown.
[0081] Thanks to ChatglmV2's ultra-long text input, you can focus on the entire content of almost all documents and summarize them within a length of 512 characters. By performing a representational similar case search on the summarized text, you can complete the similar case search input based on the case summary information.
[0082] In step S3, the first representation similar text retrieval result and the second representation similar text retrieval result are merged to obtain the final similar text retrieval result. The specific implementation process is as follows:
[0083] In the embodiment of the present invention, the first representational similar case (in the embodiment of the present invention, the similar case is the name of the similar text in the legal text) retrieval result and the second representational similar case retrieval result are fused through the result set fusion method based on the position score, or on the basis of the result set fusion method based on the position score, the first representational similar case (in the embodiment of the present invention, the similar case is the name of the similar text in the legal text) retrieval result and the second representational similar case retrieval result are fused by superimposing the result set fusion method based on the logarithmic gain. The two methods are described in detail below:
[0084] The result set fusion based on position score is as follows:
[0085] The format of the result set document is as follows, where n represents the number of the top n closest document sets, and QueryID and DocID are the unique identification IDs of the query and the search set, respectively:
[0086] {
[0087] "QueryId1":[DocID1,DocID2,DocID3...DocID n ],
[0088] "QueryId2":[DocID1,DocID2,DocID3...DocIDn],
[0089] “QueryId m ":[DocID1,DocID2,DocID3...DocIDn],
[0090] }
[0091] Here are the steps:
[0092] Step 1: Take out the search set list under the same QueryID of multiple result sets;
[0093] Step 2: Calculate the number of occurrences of each DocID in the result set list, and record it in the format of [DocID:Num];
[0094] Step 3: Record the position POS of each DocID in the result list and add them up, recording them as [DocID:POS]
[0095] Step 4: Sort the occurrence count list in descending order to get the sorted list as follows:
[0096] [Doc3:[Num:4,POS:5],Doc2:[Num:3,POS:7,Doc1:[Num:3,POS:8],Doc2:[Num:0,POS:0]];
[0097] Step 4: For Doc2:[Num:3,POS:7,Doc1:[Num:3,POS:8] with the same number of occurrences, sort them in ascending order according to their position scores. For example: Figure 5 shown.
[0098] exist Figure 5 Here, Rank_BM25, Rank_Model1, and Rank_Model2 represent the ranking results of three models for a query. In the query_id, the number of times it appears is its score. For example, id: 37960 appears three times, so its position score is 3, and it is ranked at the front. And id: 97663 also appears three times, so it should be sorted according to its position in the result set with 37960. The positions of 37960 are [1,1,2], and the positions of 97633 are [3,2,1]. Adding their position scores and arranging them in reverse order can get the order of the two, 37960:4, 97633:6. Sorting all document_ids in turn can get the final list.
[0099] The purpose of this is to merge the search results from multiple search lists. If multiple models believe that the searched document is similar to similar cases, it will be arranged at the front of the list; if the frequency scores are consistent, then they will be arranged in descending order according to the position in the corresponding list, because the position is also related to the similarity between the search set and the query content.
[0100] The result set fusion method based on logarithmic gain is as follows:
[0101] In the above, the result set ranking score based on the position score solves the problem of multi-model result set fusion very well. However, for the case where the overlap of individual ranking model result sets is not high (the difference between individual model result sets and other model result sets is too large and far below the average), the position score method will be affected by greater fluctuations.
[0102] The reason for this phenomenon is that the primary ranking basis of the position score is based on the number of occurrences. If the most similar document appears at the top of the result set in other models but does not appear in the result set with a large difference in effect, its ranking will drop sharply). In order to solve this problem, the present invention proposes a result set fusion solution based on logarithmic gain, such as Figure 6 shown.
[0103] For the above problem, it is necessary to introduce a weight factor that varies with position to fit the position information in the candidate set ranking. In the embodiment of the present invention, the log function is selected to fit the gain a:
[0104]
[0105] Among them, k is the length of the result set, and i is the current position. For example, for id: 37960, its position is 1, and the list length is 6. Its ranking in the Rank_BM25 list is 1, and its gain a is log26, which is about 2.59. It is also the same value in Rank_Model1, so its cumulative gain is about 5.18. The rest of the documentIDs are similar. Finally, all the sorting results can be measured by this value in terms of position and score to get the cumulative gain result set.
[0106] The effectiveness of the embodiments of the present invention is verified by comparative experiments below:
[0107] The most direct indicator for measuring the retrieval system is the NDCG value, which is related to two aspects: 1. Whether the document ID with high similarity is correctly found in the retrieval pool; 2. Whether the retrieval document is arranged in the corresponding position of the result set list according to the similarity between the retrieval document and the query content. NDCG@N represents whether the top N rankings are related to these two contents. In this experiment, NDCG@30 was used as a measurement indicator to evaluate the public case retrieval dataset LeCardV2 (55,192 judicial documents with an average length of 4,766, 160 queries and corresponding correct retrieval set labels). The results are shown in Table 1 (the representational retrieval model is unified as the SAILER model):
[0108] Table 1 Retrieval results of each method
[0109]
[0110] In the table, ERNIE+SAILER represents the similar case retrieval method based on case key information in the present invention, ChatglmV2+SAILER represents the similar case retrieval method based on case summary information generation, and POS represents the result set fusion method based on position score. Using the above method, the present invention finally improves the accuracy NDCG@30 value of similar case retrieval from 0.2716 of the traditional method to 0.5271, proving the beneficial effect of the method proposed in the embodiment of the present invention on the similar case retrieval task.
[0111] In fact, when the similar text retrieval method of the embodiment of the present invention is applied to similar case retrieval, compared with the existing interactive model similar case retrieval method, its advantages are as follows:
[0112] 1. Accuracy: The present invention proposes a similar case retrieval method based on structured information and summary information, which can focus on almost all the contents of the entire document, and proposes a fusion solution for the two types of retrieval results, thereby greatly improving the retrieval accuracy;
[0113] 2. Retrieval basis: The case key information extraction module and case summary information generation module in the present invention can generate structured information including the cause of the case, basic facts, dispute focus, legal provisions, penalties, and charges for the retrieval document in advance, as well as summary information summarized by the large model. In the process of retrieving documents, this information can provide users with a brief browsing and similar basis to a certain extent.
[0114] 3. Retrieval speed: By using the representation model and vector space technology to complete similar case retrieval, the vector space distance can be used to obtain a millisecond-level retrieval speed on a retrieval set of tens of thousands of articles, which is much higher than the effect of the interactive model.
[0115] 4. Retrieval scale: The representational model creates a representation vector when a document is entered into the database. Unlike the interactive model, it does not need to concatenate the query content with all the retrieval sets for similarity calculation when the retrieval task is initiated. Instead, it can be directly mapped into the vector space for spatial distance calculation, which is more suitable for large-scale retrieval tasks.
[0116] The embodiment of the present invention further provides a similar text retrieval system based on multi-dimensional information reordering, comprising:
[0117] The text acquisition module is used to acquire the text to be processed and divide the text into sentences;
[0118] A text processing module, used to process the text after sentence segmentation through a pre-trained text key information extraction model and a text summary information generation model, respectively, to obtain a first representation type similar text retrieval result and a second representation type similar text retrieval result;
[0119] The fusion module is used to fuse the first representation formula similar text retrieval result and the second representation formula similar text retrieval result to obtain the final similar text retrieval result.
[0120] It can be understood that the similar text retrieval system based on multi-dimensional information reordering provided in an embodiment of the present invention corresponds to the similar text retrieval method based on multi-dimensional information reordering mentioned above, and the explanations, examples, beneficial effects and other parts of the relevant contents can refer to the corresponding contents in the similar text retrieval method based on multi-dimensional information reordering, which will not be repeated here.
[0121] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program for similar text retrieval based on multi-dimensional information reordering, wherein the computer program enables a computer to execute the similar text retrieval method based on multi-dimensional information reordering as described above.
[0122] An embodiment of the present invention also provides an electronic device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include a method for executing the similar text retrieval method based on multi-dimensional information reordering as described above.
[0123] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0124] 1. In the embodiment of the present invention, a re-ranking method based on multi-dimensional information is used to solve the problem of long documents and low accuracy in the retrieval process. The multi-dimensional information includes text key information and text summary information. The re-ranking method is based on the position. The text key information is extracted by the text key information extraction model (ERNIE), and the text summary information is generated by the text summary information generation model, thereby optimizing the retrieval quality.
[0125] 2. The embodiment of the present invention provides two result set fusion methods. Compared with the existing voting scoring method, the result set fusion method based on position score in the embodiment of the present invention has the following advantages: (1) The model fusion method based on voting scoring is mostly used for the classification task of multiple models for one sample. Finally, different weights are given to the models according to the quality of the classification, thereby improving the overall effect. However, it is not applicable in this task. The result set fusion method based on position score and cumulative gain is suitable for the regression task of multiple models for multiple samples, and can be applied to the result set fusion of retrieval tasks. (2) The model fusion method based on voting scoring is difficult to combine the position information in the model results. For example, in application scenarios such as search engines and recommendation systems, not only the accuracy of the results is concerned, but also the ranking quality of the results needs to be taken into account. The method based on position score and logarithmic gain can combine the position information in the retrieval results, thereby better re-ranking the results of multiple models.
[0126] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0127] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A similar text retrieval method based on multi-dimensional information re-ranking, characterized in that: include: S1. Obtain the text to be processed and divide the text into sentences; S2, processing the sentence-separated texts by using a pre-trained text key information extraction model and a text summary information generation model, respectively, to obtain a first representation-based similar text retrieval result and a second representation-based similar text retrieval result; S3. Fusion the first representation formula similar text retrieval result and the second representation formula similar text retrieval result to obtain a final similar text retrieval result.
2. The similar text retrieval method based on multi-dimensional information reordering as claimed in claim 1, characterized in that: The text to be processed is a legal document, the text key information extraction model is a case key information extraction model based on ERNIE, and the text summary information generation model is a case summary information generation model based on ChatglmV2.
3. The similar text retrieval method based on multi-dimensional information reordering as claimed in claim 1, characterized in that: The sentence segmenting of the text includes: sentence segmenting of the text by terminators.
4. The similar text retrieval method based on multi-dimensional information re-ranking according to any one of claims 1 to 3, characterized in that: The fusion of the first representational similar text search result and the second representational similar text search result includes: The result set fusion method based on the position score fuses the first representation formula similar text retrieval result and the second representation formula similar text retrieval result.
5. The similar text retrieval method based on multi-dimensional information reordering as claimed in claim 4, characterized in that: The position score-based result set fusion method fuses the first representational similar text retrieval result and the second representational similar text retrieval result, including: The result set document format is as follows: Among them, n represents the number of the top n closest document sets, QueryID and DocID are the unique identification IDs of the query and the search set respectively: The fusion steps include: Step 1: Take out the search set list under the same QueryID of multiple result sets; Step 2: Calculate the number of occurrences of each DocID in the result set list, and record it in the format of [DocID:Num]; Step 3: Record the position POS of each DocID in the result list and add them up, recording them as [DocID:POS] Step 4: Sort the occurrence count list in descending order to get the sorted list as follows: [Doc3:[Num:4,POS:5],Doc2:[Num:3,POS:7,Doc1:[Num:3,POS:8],Doc2:[Num:0,POS:0]]; Step 4: For Doc2:[Num:3,POS:7,Doc1:[Num:3,POS:8] with the same number of occurrences, sort them from small to large according to their position scores to obtain the final similar text retrieval results.
6. A similar text retrieval system based on multi-dimensional information re-ranking, characterized in that: include: The text acquisition module is used to acquire the text to be processed and divide the text into sentences; A text processing module, used to process the text after sentence segmentation through a pre-trained text key information extraction model and a text summary information generation model, respectively, to obtain a first representation type similar text retrieval result and a second representation type similar text retrieval result; The fusion module is used to fuse the first representation formula similar text retrieval result and the second representation formula similar text retrieval result to obtain the final similar text retrieval result.
7. The similar text retrieval system based on multi-dimensional information re-ranking as claimed in claim 6, characterized in that: The text to be processed is a legal document, the text key information extraction model is a case key information extraction model based on ERNIE, and the text summary information generation model is a case summary information generation model based on ChatglmV2.
8. The similar text retrieval system based on multi-dimensional information re-ranking as claimed in claim 6, characterized in that: The sentence segmenting of the text includes: sentence segmenting of the text by terminators.
9. A storage medium, characterized in that: It stores a computer program for similar text retrieval based on multi-dimensional information re-ranking, wherein the computer program enables a computer to execute the similar text retrieval method based on multi-dimensional information re-ranking as claimed in any one of claims 1 to 5.
10. An electronic device, characterized in that: include: one or more processors; Memory; as well as One or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, the programs including a program for executing the similar text retrieval method based on multi-dimensional information reordering as described in any one of claims 1 to 5.