Text retrieval method and device based on law and regulation text blocks and medium
By taking each regulatory clause of the legal and regulatory document as the basic blocking unit, combining the adaptive blocking model and the FAISS database, the classification accuracy and cost problems in the blocking of legal and regulatory texts is solved, and efficient and accurate text retrieval and information provision are achieved.
Patent Information
- Application Number
- CN202510322514.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-11
AI Technical Summary
The existing text blocking technology cannot take into account both classification accuracy and cost in legal and regulatory application scenarios, and there are problems such as loss of context information, incomplete semantics, and high computing resource consumption.
Each regulatory clause in the legal and regulatory document is used as the basic chunking unit, chunking is performed according to the preset text length threshold, and long text is classified using an adaptive chunking model, combined with short text, and stored in the FAISS vector database, obtain the context content through similarity search and enter the LLM model to generate answers.
Improves retrieval accuracy, reduces computational overhead and model limitations, enhances context consistency and information integrity, and provides more accurate and comprehensive legal and regulatory information support.
Smart Images

Figure CN120296116A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text retrieval, and particularly to a text retrieval method, device, and medium based on chunking of legal and regulatory texts. Background Art
[0002] Retrieval-Augmented Generation (hereinafter referred to as RAG) is a deep learning technology that combines retrieval and generation models, mainly used to enhance the capabilities of large language models (LLMs), enabling them to incorporate information from external knowledge bases when generating responses, thereby obtaining more accurate and context-aware answers.
[0003] In the application of RAG technology, the chunking technique is a key processing technique aimed at reasonably splitting large-scale documents to improve the efficiency and accuracy of information retrieval. Existing chunking techniques include fixed chunking, rule-based chunking, and LLM-agent-based chunking methods. Fixed chunking truncates text according to a fixed number of characters, which may result in the loss of context information due to large segments of sentences or paragraphs, and the semantic incompleteness after chunking leads to a decrease in retrieval accuracy. Rule-based chunking identifies sentence boundaries and truncates according to unsupervised algorithms, but the lengths of the divided text chunks vary. Long text chunks have ambiguous semantics and high semantic complexity when vectorized, while short text chunks may lose core information and the fragmentation of short texts increases the complexity of retrieval and processing. There is also a way of LLM-agent-based chunking, which uses another LLM to analyze and chunk documents, and thus can obtain the best chunking results, but this chunking method is computationally expensive and there will be problems exceeding the model token limit. Summary of the Invention
[0004] The purpose of the present invention is to provide a text retrieval method, device, and medium based on chunking of legal and regulatory texts to solve the problem that existing text chunking techniques cannot balance classification accuracy and classification cost in the application scenario of laws and regulations.
[0005] The above object of this application is achieved by the following technical solutions:
[0006] S1: Obtain documents in the field of laws and regulations and their abstract summaries;
[0007] S2: Use each section of legal provisions text in each document as a basic chunking unit;
[0008] S3: According to a preset text length threshold L, divide each basic chunking unit into long texts and short texts; classify the long texts using an adaptive chunking model and combine with the short texts to obtain text chunks;
[0009] S4: Store the text block into the FAISS vector database afterwards;
[0010] S5: Obtain the user's question; retrieve a preset number of retrieved text blocks similar to the user's question in the FAISS database by means of similarity;
[0011] S6: Through the retrieved text blocks, reverse locate the corresponding positions in the document to obtain the context content related to the retrieved text blocks;
[0012] S7: Input the context content, the summary corresponding to the retrieved text block, and the user's question into the LLM model to generate a text retrieval answer.
[0013] Optionally, step S1 includes:
[0014] Use a pre-trained LLM model based on llama and a fixed prompt template to extract the summary of each document.
[0015] Optionally, step S3 includes:
[0016] The text block includes: the text block of short text and the text block of long text;
[0017] S31: Divide the paragraphs in the basic chunking unit that are longer than the text length threshold L into long texts;
[0018] S32: Divide the paragraphs in the basic chunking unit that are less than or equal to the text length threshold L into short texts;
[0019] S33: Use the short text as an independent text block;
[0020] S34: Use an adaptive chunking model to chunk the long text to obtain the text blocks of the long text.
[0021] Optionally, step S34 includes:
[0022] S34a: The adaptive chunking model is defined as follows: Suppose there is a long text T; tokenize the long text T and split it at the boundaries of the tokens according to the length of L / 10 to obtain n preliminary text blocks C i , then there is
[0023] T = C1 ∪ C2 ∪ … ∪ C n
[0024] where i = 1, 2, 3, …, n; L is the text length threshold;
[0025] S34b: Adjust the preliminary text blocks in a recursive manner to obtain the text blocks of the long text. The steps include: merging the preliminary text blocks based on the similarity between two adjacent preliminary text blocks until the length of the merged preliminary text block exceeds L or there are no similar preliminary text blocks, and determining the text blocks of the long text;
[0026] Define the chunking function as blocks(T), representing the result of chunking the long text T;
[0027] Suppose there are two adjacent chunks C i and C i+1 in the long text T, then the recursive process is expressed as:
[0028]
[0029] where T1 and T2 respectively represent the long text before C i and the long text after C i+1 ; | represents the chunking mark; sim(C i , C i+1 ) represents the weighted cosine similarity between two chunks.
[0030] Optionally, the weighted cosine similarity is calculated as follows:
[0031]
[0032] where S(C i ) represents the set of all words after C i is tokenized; we j is the weight of the word w i in the chunk C j , and the weight is determined by the frequency of the word in the long text T; e(w j ) represents the embedding representation of the word w j , that is, the embedding vector containing semantic information obtained by using the word embedding function; e(C i+1 ) is the embedding vector containing semantic information of the chunk C i+1 ; ‖v‖ represents the second norm of the vector v.
[0033] Optionally, the embedding representation e(C i ) of the chunk C i is obtained by weighted average fusion of the main embedding vector of each word constituting the text block and the sub-word embedding , that is:
[0034]
[0035] where S(w j ) represents the word wj The corresponding sub - word set, which is the set of all words composed of the word with a length not exceeding n characters; the main embedding vector represents the word w j 's embedding vector; the sub - word embedding represents the sub - word s k 's embedding vector.
[0036] An electronic device, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that the electronic device executes a text retrieval method based on the chunking of legal and regulatory texts.
[0037] A computer - readable storage medium stores instructions. When the instructions are executed, a text retrieval method based on the chunking of legal and regulatory texts is executed.
[0038] The beneficial effects brought by the technical solution provided in this application are:
[0039] 1. Improve retrieval accuracy: Each section of legal provisions text is used as the basic chunking unit, and reasonable chunking is carried out according to the text length of the legal articles, avoiding the problems of context loss and semantic incompleteness that may be caused by fixed chunking and rule - based chunking. Short texts retain the complete meaning of laws and regulations, while long texts are segmented by an adaptive chunking model, ensuring that the semantics of each text block are clear and easy to process, thus improving the final retrieval accuracy.
[0040] 2. Reduce computational overhead and model limitations: Compared with the LLM - agent - based chunking method, by using the structure of legal articles for chunking, the high computational cost and token limitations of using the LLM model for large - scale document segmentation are avoided. Through pre - processing and reasonable chunking of legal provisions, the dependence on large models is greatly reduced, ensuring the chunking effect while reducing the consumption of computing resources.
[0041] 3. Enhance context consistency and information integrity: Use the LLM to extract the abstract summary of the document, and provide the most similar text blocks together with the context and the summary as background knowledge, which helps to improve the context consistency of the retrieval. By combining the summary information, the essence of legal provisions can be better captured, thus providing more comprehensive and accurate legal and regulatory information support during retrieval. Description of the Drawings
[0042] The following will further illustrate the present application in conjunction with the drawings and embodiments. In the drawings:
[0043] Figure 1 is the step diagram in the embodiment of the present application;
[0044] Figure 2 It is the question type diagram in the embodiment of the present application;
[0045] Figure 3 It is the result diagram in the embodiment of the present application;
[0046] Figure 4 It is the basic block unit diagram in the embodiment of the present application;
[0047] Figure 5 It is the preliminary text block diagram in the embodiment of the present application;
[0048] Figure 6 It is the text block segmentation result diagram in the embodiment of the present application;
[0049] Figure 7 It is the schematic diagram of the electronic device structure in the embodiment of the present application. Detailed implementation manners
[0050] For a clearer understanding of the technical features, objectives, and effects of the present application, the detailed implementation manners of the present application will now be described with reference to the accompanying drawings.
[0051] The embodiment of the present application provides a text retrieval method based on legal and regulatory text segmentation.
[0052] Please refer to Figure 1 , Figure 1 , which is the step diagram of a text retrieval method based on legal and regulatory text segmentation in the embodiment of the present application, including:
[0053] S1: Obtain the documents in the legal and regulatory field and their abstract summaries;
[0054] S2: Take each section of the legal provisions text of each document as a basic block unit;
[0055] S3: According to the preset text length threshold L, divide each basic block unit into long texts and short texts; use an adaptive block model to classify the long texts and combine with the short texts to obtain text blocks;
[0056] S4: Store the text blocks into the FAISS vector database;
[0057] As an embodiment, after obtaining the text blocks, use a pre-trained embedding model based on BERT to encode the text blocks to generate an embedding E P (p) in vector form, and store it in the FAISS vector database. After storage, index the FAISS to speed up the retrieval speed.
[0058] S5: Obtain the user's question; retrieve a preset number of retrieved text chunks similar to the user's question in the FAISS database by means of similarity.
[0059] As an embodiment, after the user asks a question, the user's question q is encoded using another BERT-based embedding model to generate the embedding E Q (q). Use similarity retrieval in the FAISS database to find the most relevant text chunks. The similarity between the question and the text chunks is represented by the dot product between the two embedding vectors.
[0060] S6: Through the retrieved text chunks, reverse-locate the corresponding positions in the document to obtain the context content related to the retrieved text chunks.
[0061] As an embodiment, perform regular expression prefix matching on the document positions to extract the hierarchy, and extract all relevant content at this level, such as sub-clauses, explanations, etc.
[0062] S7: Input the context content, the summary corresponding to the retrieved text chunks, and the user's question into the LLM model to generate the text retrieval answer.
[0063] As an embodiment, the present invention proposes a text chunking and retrieval algorithm in the field of laws and regulations. Based on the characteristics of laws and regulations documents, first, each article of the regulations is regarded as an element. If the text length contained in the element is short, it is directly used as a text chunk. Otherwise, the chunking prediction model proposed by the present invention is used for chunking. At the same time, use the LLM to extract the summary of each document. After the retrieval is completed, use the most similar embedding to find the nearest text chunk, and use the context and summary of this text chunk together as background knowledge. Input the background knowledge and the user's question into the LLM to obtain the reply and return it to the user.
[0064] Step S1 includes:
[0065] Use a pre-trained LLM model based on llama and a fixed prompt template to extract the summary of each document.
[0066] As an embodiment, the documents of the present application can adopt the Chinese Q&A set in the field of laws and regulations (JEC-QA). Each question in the JEC-QA dataset consists of a question description and four candidate options and includes single-choice questions and multiple-choice questions. The questions in the dataset are divided into two categories: knowledge-driven questions (KD questions) and case-analysis questions (CA questions), such as Figure 2As shown, the correct options are marked in red. KD questions mainly focus on the definition and interpretation of legal and regulatory concepts, while CA questions require the analysis of actual cases. Answering both types of questions requires relevant document knowledge and strong reasoning abilities.
[0067] As an example, the retrieval task is performed on the JEC-QA dataset using the method of the present invention, and the results are as Figure 3 shown.
[0068] Step S3 includes:
[0069] The text blocks include: text blocks of short texts and text blocks of long texts;
[0070] S31: Divide the paragraphs in the basic chunking unit that are longer than the text length threshold L into long texts;
[0071] S32: Divide the paragraphs in the basic chunking unit that are less than or equal to the text length threshold L into short texts;
[0072] S33: Take the short text as an independent text block;
[0073] S34: Use an adaptive chunking model to chunk the long text to obtain text blocks of the long text.
[0074] As an example, the text length threshold L is selected as the token limit length of the embedding model, or the optimal value can be obtained through multiple tests. To ensure the semantic coherence of the long text and prevent semantic loss caused by truncation of professional domain vocabulary. Use an adaptive chunking model to classify the long text.
[0075] As an example, the chunking process is shown as Figure 4 、 Figure 5 shown. Set the long text threshold L = 100 words. Except for the first paragraph, each subsequent paragraph does not exceed the long text threshold length. Therefore, each subsequent paragraph is treated as a short text, and the first paragraph is treated as a long text. Then, with θ = 0.65 as the recursive merging threshold, recursive merging is performed, and the merging completion result is as Figure 6 shown.
[0076] Step S34 includes:
[0077] S34a: The adaptive chunking model is defined as follows: Given a long text T; tokenize the long text T and divide it at the boundaries of the tokens according to a length of L / 10 to obtain n preliminary text blocks C i , then there is
[0078] T = C1 ∪ C2 ∪ … ∪ C n
[0079] Among them, i = 1, 2, 3, …, n; L is the text length threshold;
[0080] S34b: Adjust the preliminary text blocks in a recursive manner to obtain the text blocks of the long text. The steps include: Based on the similarity between two adjacent preliminary text blocks, merge the preliminary text blocks until the length of the merged preliminary text block exceeds L or there is no similar preliminary text block, and determine the text blocks of the long text;
[0081] Define the chunking function as blocks(T), which represents the result of chunking the long text T;
[0082] Suppose there are two adjacent chunks C i and C i+1 in the long text T, then the recursive process is expressed as:
[0083]
[0084] where T1 and T2 respectively represent the long text before C i and the long text after C i+1 ; | represents the chunking mark; sim(C i , C i+1 ) represents the weighted cosine similarity between two chunks.
[0085] The weighted cosine similarity is calculated as follows:
[0086]
[0087] where S(C i ) represents the set of all words after word segmentation of C i ; we j is the weight of the word w i in the chunk C j , and the weight is determined by the frequency of the word in the long text T; e(w j ) represents the embedding representation of the word w j , that is, the embedding vector containing semantic information obtained by using the word embedding function; e(C i+1 ) is the embedding vector containing semantic information of the chunk C i+1 ; ‖v‖ represents the second norm of the vector v.
[0088] The embedding representation e(C i ) of the chunk C i is obtained by weighted average fusion of the main embedding vector of each word constituting the text block and the sub-word embedding , that is:
[0089]
[0090] Among them, S(w j ) represents the set of sub-words corresponding to the word w j . The set of sub-words is the set composed of all words of this word that are no longer than n characters; the main embedding vector represents the embedding vector of the word w j ; the sub-word embedding v sk represents the embedding vector of the sub-word s k .
[0091] As an example, the main embedding vector and the sub-word embedding vector v sk are optimized by maximizing the conditional probability that the given context is the target word, as follows: Given the target word w and its context C w , maximize where v w is the word vector of the target word w, is the word vector of each word in the context; v w′ represents the word vector of the word w′; w′ represents the word in the context C w .
[0092] This application also discloses an electronic device. Referring to Figure 7 , Figure 7 is a schematic structural diagram of an electronic device disclosed in an embodiment of this application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.
[0093] Among them, the communication bus 502 is used to realize the connection and communication between these components.
[0094] Among them, the user interface 503 may include a display screen. Optionally, the user interface 503 may further include a standard wired interface and a wireless interface.
[0095] Among them, the network interface 504 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0096] This application also discloses a computer-readable storage medium. The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the above-mentioned text retrieval method based on legal regulation text chunking.
[0097] The above are only exemplary embodiments of the present disclosure and cannot limit the scope of the present disclosure. That is, all equivalent changes and modifications made according to the teachings of the present disclosure still fall within the scope covered by the present disclosure.
[0098] This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or customary technical means in the technical field not recorded in the present disclosure. The specification and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A text retrieval method based on text chunking of laws and regulations texts, characterized in that, The method includes the following steps: S1: Obtain documents in the field of laws and regulations and their abstract summaries; S2: Use each section of the legal clause text of each document as a basic chunk unit; S3: According to the preset text length threshold L, divide each basic chunk unit into long texts and short texts; Use an adaptive chunking model to classify long texts and combine short texts to obtain text chunks; S4: Store the text chunks into the FAISS vector database afterwards; S5: Obtain the user's question; Through the method of similarity, retrieve a preset number of retrieved text chunks similar to the user's question in the FAISS database; S6: Through the retrieved text chunks, reverse-locate the corresponding positions in the document to obtain the context content related to the retrieved text chunks; S7: Input the context content, the abstract summary corresponding to the retrieved text chunks, and the user's question into the LLM model to generate a text retrieval answer.
2. The text retrieval method based on text chunking of laws and regulations texts as claimed in claim 1, wherein Step S1 includes: Use a pre-trained LLM model based on llama and a fixed prompt template to extract the abstract summary of each document.
3. A text retrieval method based on text chunking of laws and regulations texts according to claim 1, characterized in that, Step S3 includes: The text chunks include: text chunks of short texts and text chunks of long texts; S31: Divide the paragraphs in the basic chunk unit that are longer than the text length threshold L into long texts; S32: Divide the paragraphs in the basic chunk unit that are less than or equal to the text length threshold L into short texts; S33: Use the short text as an independent text chunk; S34: Use an adaptive chunking model to chunk the long text to obtain text chunks of the long text.
4. The text retrieval method based on the chunking of legal and regulatory texts according to claim 3, characterized in that, Step S34 includes: S34a: The adaptive chunking model is defined as follows: Given a long text T; the long text T is tokenized and segmented at the token boundaries according to a length of L / 10 to obtain n preliminary text chunks C i , then there is T = C1 ∪ C2 ∪ … ∪ C n Among them, L is the text length threshold; S34b: Use a recursive method to adjust the preliminary text chunks to obtain text chunks of the long text. The steps include: Based on the similarity between two adjacent preliminary text chunks, merge the preliminary text chunks until the length of the merged preliminary text chunks exceeds L or there are no similar preliminary text chunks, and determine the text chunks of the long text; Define the chunking function as blocks(T), representing the result of chunking the long text T; Let there be two adjacent blocks C i and C i+1 in the long text T, then the recursive process is expressed as: Among them, T1 and T2 respectively represent the long text before C i and the long text after C; | represents a chunk marker; sim(C i+1 , C i , C i+1 ) represents the weighted cosine similarity between two chunks.
5. The text retrieval method based on the text chunking of laws and regulations texts according to claim 4, characterized in that, The weighted cosine similarity, the calculation formula is as follows: Among them, S(C i ) represents the set of all words after i word segmentation of C; we j is the weight of the word w i in the block C j , and the weight is determined by the frequency of the word in the long text T; e(w j ) represents the embedding representation of the word w j , that is, the embedding vector containing semantic information obtained by using the word embedding function; e(C i+1 ) is the embedding vector containing semantic information of the block C i+1 ; ‖v‖ represents the second norm of the vector v.
6. The text retrieval method based on the chunking of legal and regulatory texts according to claim 5, characterized in that, The chunk C i 's embedded representation e(C i ) is obtained by the weighted average fusion of the main embedding vectors of each word constituting the text chunk and the sub-word embeddings , that is: Among them, S(w j ) represents the set of sub-words corresponding to the word w j . The set of sub-words is the set of all words composed of no more than n characters of this word; the main embedding vector represents the embedding vector of the word w j . The sub-word embedding represents the embedding vector of the sub-word s k .
7. An electronic device, characterized in that, It includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes the method described in any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions. When the instructions are executed by a computer, the method described in any one of claims 1-6 is executed.