Method and system for constructing guarantee knowledge think tank based on LLM and RAG technologies

Through word segmentation and vector embedding processing based on LLM and RAG technology, the problems of data errors and information loss in the guarantee knowledge think tank were solved, more accurate document retrieval and intelligent assistant services were achieved, and the efficiency of guarantee business and customer satisfaction were improved.

CN120706525AActive Publication Date: 2025-09-26ANHUI CREDIT FINANCING GUARANTEE GROUP CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511167460.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-09-26
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

The data in the knowledge think tank in traditional guarantee business contains errors or inaccurate information, which leads to biased retrieval results, and serious information loss during the data vectorization process, affecting the accuracy of document retrieval.

Method used

By adopting a method based on LLM and RAG technology, by collecting guarantee knowledge information, using optimized word segmentation strategy and optimal embedding model, the text data is divided into large text blocks and small text blocks, and vector embedding processing is performed to establish a vector database and form a guarantee knowledge think tank.

Benefits of technology

It improves the data accuracy and retrieval recall rate of the knowledge think tank, can capture user query intentions from multiple angles, provide efficient knowledge services and intelligent assistant tools, and promote the digital transformation of the guarantee industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706525A_ABST
    Figure CN120706525A_ABST
Patent Text Reader

Abstract

The invention provides a guarantee knowledge think tank construction method and system based on LLM and RAG technologies, and the method comprises the steps: S1, collecting a guarantee knowledge data file, and importing the guarantee knowledge data file into a retrieval enhancement generation RAG system, so as to form a primary edition knowledge base; s2, loading and analyzing a guarantee knowledge data file in the first edition knowledge base, and storing text content obtained by analysis in a memory to form complete knowledge text data; s3, performing word segmentation on the knowledge text data by using an optimized word segmentation strategy to obtain a plurality of large text blocks and corresponding small text blocks; and S4, determining and using an optimal embedding model, performing vector embedding processing on the knowledge blocks formed by the large text blocks and the corresponding small text blocks, converting the knowledge blocks into vector representations, and storing the vector representations into a vector database to form a guarantee knowledge think tank. According to the method, the guarantee knowledge think tank can be constructed based on LLM and RAG technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language model (LLM) and retrieval enhancement generation (RAG), and in particular to a method and system for constructing a guarantee knowledge think tank based on LLM and RAG technologies. Background Art

[0002] As an integral part of the financial system, the guarantee industry plays a vital role in economic activities. In recent years, with the continuous expansion of financial markets and the increasing complexity of transactions, the guarantee business has faced numerous challenges, including increased difficulty in risk management, inefficient business processes, and a need to improve customer service quality.

[0003] To address these challenges, the industry is actively introducing various digital technologies. Traditional guarantee operations rely on manual experience and simple information systems, resulting in significant deficiencies in knowledge management. With the application of big data technology, guarantee institutions are beginning to collect and analyze large amounts of business data, establishing guarantee knowledge think tanks. These think tanks are then used for subsequent applications such as automated question-and-answering on guarantee knowledge.

[0004] However, if there are errors or inaccurate information in the data of the guarantee knowledge think tank, the generated results will be biased. Even if a lot of optimization is performed in the retrieval stage, it is difficult to effectively solve this problem. In addition, there is information loss in the data vectorization process. The complexity of text data makes it difficult for limited vectors to fully express all its details and features, affecting the accuracy of subsequent document retrieval. Therefore, there is an urgent need for an effective method to construct a guarantee knowledge think tank. Summary of the Invention

[0005] In order to solve the technical problems existing in the above-mentioned prior art, the present invention provides a method and system for constructing a guarantee knowledge think tank based on LLM and RAG technologies. The technical solution is as follows:

[0006] On the one hand, a method for constructing a guarantee knowledge think tank based on LLM and RAG technology is provided, which includes:

[0007] S1. Collect guarantee knowledge data files and import them into the retrieval enhancement generation RAG system to form the initial version of the knowledge base;

[0008] S2. Load and parse the guarantee knowledge data file in the first edition knowledge base, and store the parsed text content in the memory to form complete knowledge text data;

[0009] S3. Segmenting the knowledge text data using an optimized segmentation strategy to obtain a plurality of large text blocks and corresponding small text blocks;

[0010] S4. Determine and use the optimal embedding model to perform vector embedding processing on the knowledge blocks composed of large text blocks and corresponding small text blocks, convert them into vector representations and store them in the vector database to form a guarantee knowledge think tank.

[0011] Optionally, the S3 specifically includes:

[0012] The knowledge text is segmented using a word segmentation model based on the pre-trained BERT architecture. According to the logical structure and semantic integrity of the knowledge content, the formula is adopted:

[0013] ;

[0014] Calculating block quality scores ,determine the division criteria of large text blocks and obtain multiple large text blocks with reasonable semantic units;

[0015] in, represents the semantic coherence aggregation weight, represents the position dispersion compensation coefficient, represents the punctuation density suppression coefficient, Representative The mean semantic similarity of adjacent blocks, Represents the starting position index of the current candidate block, Representative The base position index of the semantic anchor point, Represents the number of punctuation marks in the current block. Represents the total number of characters in the current block. Represents the number of adjacent blocks;

[0016] After the large text block segmentation is completed, the length of each large text block is judged. If the size of the large text block exceeds the pre-set text block length threshold, the length segmentation method is used to perform a secondary segmentation on the large text block, subdividing it into multiple small text blocks that meet the length requirements, and establishing an association between the small text blocks and the large text block. If the length of the large text block is less than or equal to the pre-set text block length threshold, the large text block is directly stored as a complete unit.

[0017] Optionally, establishing an association relationship between the small text block and the large text block specifically includes:

[0018] The information of the large text block to which the small text block belongs is recorded by adding a "parent_block_id" field in the data structure.

[0019] Optionally, the length segmentation method performs automatic segmentation processing by setting the segmentation length and overlap length, specifically including:

[0020] Preset word segmentation length and overlap length;

[0021] Preset word segmentation length and overlap length , according to the formula:

[0022] ;

[0023] Calculating dynamic offset , intercepted from the beginning of the text The length segment is used as the segmentation result, and the adjacent segmentation segments are retained. For the part with the same length, we get the word segmentation result set;

[0024] in, Represents the sequence number of the current word segment, Represents the preset word segmentation length, Represents the preset overlap length, Representative The dynamic offset of the starting position of each segment, Represents a dynamic balancing factor based on indexes and preset parameters;

[0025] To ensure the coherence and accuracy of word segmentation, there is a certain overlap between adjacent word segmentation segments according to the pre-set overlap length.

[0026] Optionally, the S3 further includes:

[0027] For some knowledge text data, manual tagging is used to achieve word segmentation by manually adding specific identifiers to the text content, including:

[0028] Based on the semantic structure of the text and business needs, annotators insert specific identifiers into the text to mark the start and end positions of knowledge blocks and clearly define a complete knowledge unit.

[0029] When loading text, these tags are automatically identified and the text is segmented according to the tags, dividing it into corresponding large text blocks and small text blocks.

[0030] Optionally, the method further includes:

[0031] Determine an optimal large language model (LLM), submit each large text block to the LLM, and pre-generate a set of relevant questions and a summary based on the content of the large text block;

[0032] The optimal embedding model performs vector embedding processing on the relevant questions and summary abstracts, converts them into vector representations, and stores them together with the large text blocks and the corresponding small text blocks in the vector database, and establishes an index connection between them to form the guarantee knowledge think tank.

[0033] Optionally, establishing an index relationship between them specifically includes:

[0034] Create an index table for each large text block, recording the vector IDs of the large text block, pre-generated related questions and summary summaries, and the corresponding relationships between them.

[0035] On the other hand, a guarantee knowledge think tank construction system based on LLM and RAG technology is provided, the system comprising:

[0036] The collection and import module is used to collect guarantee knowledge data files and import them into the retrieval enhancement generation RAG system to form the initial version of the knowledge base;

[0037] A loading and parsing module is used to load and parse the guarantee knowledge data file in the first edition knowledge base, store the parsed text content in the memory, and form complete knowledge text data;

[0038] A word segmentation module, configured to segment the knowledge text data using an optimized word segmentation strategy to obtain a plurality of large text blocks and corresponding small text blocks;

[0039] The embedding module is used to determine and use the optimal embedding model, perform vector embedding processing on the knowledge block composed of large text blocks and corresponding small text blocks using the optimal embedding model, convert it into a vector representation and store it in the vector database to form a guarantee knowledge think tank.

[0040] On the other hand, an electronic device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned method for constructing a guarantee knowledge think tank based on LLM and RAG technology.

[0041] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the above-mentioned method for constructing a guarantee knowledge think tank based on LLM and RAG technology.

[0042] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0043] The construction method of the guarantee knowledge think tank of the present invention not only collects a large amount of guarantee knowledge information, but also uses an optimized word segmentation strategy to obtain multiple large text blocks and corresponding small text blocks. It also pre-generates a set of related questions and a summary summary based on the large text block during the data vectorization process, and performs vector embedding processing on the related questions and summary summary. After converting them into vector representations, they are stored in the vector database together with the large text block and the corresponding small text block. The knowledge think tank constructed in this way not only has more accurate data content, but also can capture the user's query intention from different angles in subsequent user retrieval, increase the probability of recalling relevant documents, and obtain more accurate answers, providing the guarantee group with efficient knowledge services and intelligent assistant tools, such as case analysis, industry research, business assistants, etc., to promote the digital transformation of the guarantee industry, improve business efficiency, risk management capabilities and customer satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0045] Figure 1 This is a flow chart of a method for constructing a guarantee knowledge think tank based on LLM and RAG technologies provided by an embodiment of the present invention;

[0046] Figure 2 This is a block diagram of a guarantee knowledge think tank construction system based on LLM and RAG technologies provided by an embodiment of the present invention;

[0047] Figure 3 is a structural diagram of an electronic device provided by an embodiment of the present invention;

[0048] Figure 4 Construct a flow chart for the invention guarantee knowledge think tank;

[0049] Figure 5 This is the text block diagram after cleaning in the present invention. DETAILED DESCRIPTION

[0050] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0051] See also Figure 1-Figure 5 An embodiment of the present invention provides a method for constructing a guarantee knowledge think tank based on LLM and RAG technologies. The method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of the method is shown, and the processing flow may include the following steps:

[0052] S1. Collect guarantee knowledge data files and import them into the retrieval enhancement generation RAG system to form the initial version of the knowledge base;

[0053] The embodiment of the present invention collects a large number of guarantee knowledge documents such as bonds and bond guarantees, compliance management, financial management and informatization. The document material formats support word, pdf, txt, and excel. After collection, they are imported into the retrieval enhancement generation RAG system to form a preliminary knowledge base.

[0054] Among them, the Retrieval-Augmented Generation (RAG) system combines information retrieval technology with language generation models to attempt to solve the problem that language models lack real-time and accurate knowledge when answering questions. It retrieves relevant information from external knowledge bases to provide supplementary knowledge for the language model, enabling the model to answer questions more accurately.

[0055] S2. Load and parse the guarantee knowledge data file in the first edition knowledge base, and store the parsed text content in the memory to form complete knowledge text data;

[0056] This embodiment of the present invention supports loading knowledge text in a variety of common formats, such as TXT, DOCX, and PDF. During the loading process, the file is verified for validity. If the file format is incorrect, damaged, or unreadable, an error message is logged and the user is prompted to re-upload. Furthermore, to improve loading efficiency, multi-threading technology can be used to process multiple files in parallel.

[0057] The present invention uses a parsing strategy tailored to the file format: for TXT files, the content is directly read line by line; for DOCX files, a specialized document parsing library is used to extract the text content; for PDF files, OCR technology is used to recognize and extract text. After parsing, the text content is stored in memory, forming complete knowledge text data.

[0058] S3. Segmenting the knowledge text data using an optimized segmentation strategy to obtain a plurality of large text blocks and corresponding small text blocks;

[0059] Optionally, the S3 specifically includes:

[0060] The knowledge text is segmented using a word segmentation model based on the pre-trained BERT architecture. According to the logical structure and semantic integrity of the knowledge content, the formula is adopted:

[0061] ;

[0062] Calculating block quality scores ,determine the division criteria of large text blocks and obtain multiple large text blocks with reasonable semantic units;

[0063] in, represents the semantic coherence aggregation weight, represents the position dispersion compensation coefficient, represents the punctuation density suppression coefficient, Representative The mean semantic similarity of adjacent blocks, Represents the starting position index of the current candidate block, Representative The base position index of the semantic anchor point, Represents the number of punctuation marks in the current block. Represents the total number of characters in the current block. Represents the number of adjacent blocks;

[0064] The knowledge text is segmented using a word segmentation model based on the pre-trained BERT architecture. Specifically, the knowledge text to be processed is first preprocessed, including removing irrelevant characters and standardizing the format. Then, the system uses sentence terminators (such as periods, question marks, and exclamation marks) as initial potential segmentation points to process a text paragraph containing three sentences: "Guarantee refers to a credit guarantee provided by a third party to a debtor to a creditor. Guarantee usually occurs in lending or trading activities and aims to enhance the security and feasibility of the creditor's rights. Throughout the process, the guarantor promotes capital flow and maintains market credit order by assessing risks and charging fees." The system will evaluate the quality of the segmentation between the first and second sentences. At this time, the first candidate text block is "Guarantee refers to a credit guarantee provided by a third party to a debtor to a creditor." The number of adjacent blocks is Set to 1 (i.e. only consider the semantic association of the previous block). Since it is the first block, there is no previous block, and the average semantic similarity of its adjacent blocks is The initial value is set to 0.8, and a grid search and cross-validation experiment is conducted on a dataset of 5,000 technical documents to determine the semantic coherence aggregation weight. is 0.58, position dispersion compensation coefficient is 0.25, the punctuation density suppression coefficient is 0.12. These numerical combinations achieve the highest ROUGE-L score of 0.89 on the test set. The starting position index of the candidate block is 0, the base position index of the preset semantic anchor point in the text (usually the chapter title or topic sentence) Also 0, the total number of characters in this block is 28, including the number of punctuation marks is 1, substitute it into the formula to calculate the block quality score ,Then, the system evaluates the quality of merging the first two sentences into one block. ,At this time, the candidate block is "Guarantee usually occurs in lending or trading activities, ,aims to enhance the security and realization of debt. Throughout the process, the guarantor promotes ,capital flow and maintains market credit order by assessing risks and ,collecting fees.", and the average semantic similarity of its adjacent blocks ,is The cosine similarity of the BERT sentence vectors of the first and second sentences is calculated, and its value is 0.92. The total number of characters in this block is 65, the number of punctuation marks 2, by iteratively calculating the block quality scores of all potential split point combinations, the average of all final blocks is selected. The segmentation scheme with the highest value, for example, if the first sentence is segmented separately The score is 0.75, and the first two sentences combined If the score is 0.88, the system tends to the latter. In this way, the division criteria of large text blocks are determined according to the logical structure and semantic integrity of the knowledge content, and multiple large text blocks with reasonable semantic units are obtained.

[0065] The formula is beneficial in that it introduces the semantic coherence Item, representing the importance of the structure Items and content fluency Item, combined with an adjustable weight coefficient 、 、 , realizing a multi-dimensional comprehensive evaluation of the quality of text block segmentation.

[0066] Table 1: Example of block quality score calculation

[0067]

[0068] As shown in Table 1, the candidate blocks generated by different segmentation methods are scored and calculated. The block quality score of the first candidate block "Guarantee refers to the credit guarantee provided by a third party for the debtor to the creditor" is The calculation process is: The second candidate block is "Guarantee usually occurs in lending or trading activities, and is intended to enhance the security and realization of debt. Throughout the process, the guarantor promotes capital flow and maintains market credit order by assessing risks and charging fees." The calculation process is: , in the formula, the summation operation Used to accumulate the semantic similarity of all adjacent blocks and divide it by the number of adjacent blocks To obtain the average similarity, it reflects the overall semantic coherence of the block, the square root and absolute value operation It is used to calculate the distance between the starting position of the current block and the core semantic anchor point. The absolute value of the distance ensures that the offset is positive. The square root operation smoothes the effect of the distance growth. The subtraction operation A penalty term is introduced to reduce the weight of blocks that use punctuation too densely and may cause content fragmentation. The multiplication operation is used to combine the weight coefficients 、 、 , integrating features of different dimensions into the final block quality score The preset block quality score threshold is 0.5. The score of the second candidate block is 0.53, which is higher than the threshold, indicating that it is more reasonable to merge the first two sentences into one text block. This numerical result will be used as a decision basis to guide the system to select a better text segmentation point, and finally form a large text block sequence containing complete knowledge fragments.

[0069] After the large text block segmentation is completed, the length of each large text block is judged. If the size of the large text block exceeds the pre-set text block length threshold, the length segmentation method is used to perform a secondary segmentation on the large text block, subdividing it into multiple small text blocks that meet the length requirements, and establishing an association between the small text blocks and the large text block. If the length of the large text block is less than or equal to the pre-set text block length threshold, the large text block is directly stored as a complete unit.

[0070] The embodiment of the present invention not only improves the word segmentation accuracy and effectively reduces the knowledge block segmentation errors through the optimized word segmentation strategy, but also recalls not only the small text blocks but also the related large text blocks during subsequent retrieval and recall through the association relationship between the large and small text blocks, making the recalled documents more comprehensive and accurate.

[0071] Optionally, establishing an association relationship between the small text block and the large text block specifically includes:

[0072] The information of the large text block (parent block) to which the small text block belongs is recorded by adding a "parent_block_id" field in the data structure.

[0073] Optionally, the length segmentation method performs automatic segmentation processing by setting the segmentation length and overlap length, specifically including:

[0074] Preset word segmentation length and overlap length;

[0075] Preset word segmentation length and overlap length , according to the formula:

[0076] ;

[0077] Calculating dynamic offset , intercepted from the beginning of the text The length segment is used as the segmentation result, and the adjacent segmentation segments are retained. For the part with the same length, we get the word segmentation result set;

[0078] in, Represents the sequence number of the current word segment, Represents the preset word segmentation length, Represents the preset overlap length, Representative The dynamic offset of the starting position of each segment, Represents a dynamic balancing factor based on indexes and preset parameters;

[0079] Preset word segmentation length Character and overlap length Character, dynamic offset calculated according to the formula ,when hour, , intercept the fragment of characters 0 to 8 from the starting position of the text, when hour, , intercept the segment from 5.25 to 13.25 characters at the starting position (rounded to 5 to 13 characters), and there is a 3-character overlap between adjacent segments in the 5-8 character range. hour, , truncate the segments of 10.5 to 18.5 characters (rounded to 10 to 18 characters), and overlap with the previous segment in the range of 10-13 characters by 3 characters. Through dynamic offset calculation, ensure that the adjacent segments retain the preset overlap length while compensating for the index offset error through square root and absolute value operations, and finally obtain a continuous and complete text segmentation result set.

[0080] Table 2: Dynamic offset calculation example

[0081]

[0082] As shown in Table 2, the dynamic offset It is obtained by calculating the formula, where the square root term exist At that time ,exist At that time , the compensation item value is dynamically adjusted as the index increases. When the calculation result contains decimals, the truncation interval is rounded down to ensure that the character index is an integer. The overlapping length of adjacent segments always maintains the preset 3 characters. Through the dynamic compensation mechanism, the actual overlapping area fluctuates between 3-4 characters, improving the fault tolerance of word segmentation boundaries.

[0083] Parameters in the formula The average semantic unit length experiment of the text was determined. Statistics of 100 news texts showed that 8 characters long covered 80% of the independent semantic units. The parameter Set to 37.5% (i.e. 3 characters), through cross-validation testing, this ratio strikes a balance between recall and precision, the square root term Used to compensate for the linear offset cumulative error caused by index growth. The absolute value operation ensures that the compensation amount is non-negative. The denominator Normalize the compensation amount to prevent over-compensation, for example, when When , the compensation term is The dynamic offset increment is stabilized at 5.25 / fragment. The formula realizes nonlinear offset compensation through a combination of four operations: multiplication, subtraction, square root, and absolute value. The final word segmentation result set covers the complete text and the error in the overlapping area is controlled within ±1 character.

[0084] To ensure the consistency and accuracy of word segmentation, there will be a certain amount of overlap between adjacent word segmentation segments according to the pre-set overlapping length (for example, if the word segmentation length is set to 20 words and the overlapping length is 5 words, the 1st to 20th words will be cut off as the first word segmentation segment, and then 20 words starting from the 16th word (that is, the 16th to 35th words) will be cut off as the second word segmentation segment, and so on, until the word segmentation operation of the entire text is completed).

[0085] Optionally, the S3 further includes:

[0086] For some knowledge text data, manual tagging is used to achieve word segmentation by manually adding specific identifiers to the text content, including:

[0087] Based on the semantic structure of the text and business needs, annotators insert specific tags into the text to mark the start and end positions of knowledge blocks and clearly define a complete knowledge unit (for example, in a legal document, annotators can mark different legal clauses as large text blocks and specific paragraphs in the clauses as small text blocks). <kcs>: Large text block, <tcs>Small text blocks) to achieve accurate word segmentation);

[0088] When loading text, these tags are automatically identified and the text is segmented according to the tags, dividing it into corresponding large text blocks and small text blocks.

[0089] The biggest advantage of manual word segmentation is its extremely high accuracy. Since it is manually marked according to semantics and business logic, it can best meet actual needs and ensure that the word segmentation results meet expectations. However, the workload is huge, so the embodiment of the present invention adopts a parallel strategy of multiple word segmentation schemes to give priority to the use of word segmentation models for word segmentation of large and small texts. For situations where the recall effect of some knowledge texts is not ideal, it can be considered to use manual word segmentation appropriately.

[0090] S4. Determine and use the optimal embedding model to perform vector embedding processing on the knowledge blocks composed of large text blocks and corresponding small text blocks, convert them into vector representations and store them in the vector database to form a guarantee knowledge think tank.

[0091] Based on the guarantee knowledge data file, the embodiment of the present invention determined bge-large-zh-v1.5 as the optimal embedding model through replacement testing, and performed vector embedding processing on the knowledge blocks composed of large text blocks and corresponding small text blocks, converted them into vector representations and stored them in the vector database to form a guarantee knowledge think tank.

[0092] Optionally, the method further includes:

[0093] Determine the optimal large language model (LLM) (through replacement testing, this embodiment of the present invention determined that the large chatbot model glm4-9b-chat developed by Zhipu AI was the optimal large language model). Submit each large text block to the LLM. The LLM pre-generates a set of relevant questions and a summary based on the content of the large text block (for example, for a knowledge block related to the guarantee business process, it generates questions such as "What are the application conditions for the guarantee business?" and "What are the key steps in the guarantee approval process?" as well as a brief summary of the business process).

[0094] The optimal embedding model performs vector embedding processing on the relevant questions and summary abstracts, converts them into vector representations, and stores them together with the large text blocks and the corresponding small text blocks in the vector database, and establishes an index connection between them to form the guarantee knowledge think tank.

[0095] When a user asks a question subsequently, the embodiment of the present invention will not only generate a corresponding vector based on the query statement entered by the user, but also match it from multiple dimensions. In addition to matching with the vector of the original knowledge block data, it will also match with the pre-generated related question vector and summary vector. Through this multi-dimensional matching method, even if there is a certain amount of information loss in the data vectorization process, it can capture the user's query intention from different angles, increase the probability of retrieving and recalling relevant documents, and thus effectively reduce the problem of inaccurate retrieval caused by information loss.

[0096] Optionally, establishing an index relationship between them specifically includes:

[0097] Create an index table for each large text block, recording the vector IDs of the large text block, pre-generated related questions and summary summaries, and the corresponding relationships between them.

[0098] like Figure 2 As shown, an embodiment of the present invention further provides a guarantee knowledge think tank construction system based on LLM and RAG technology, the system comprising:

[0099] The collection and import module 210 is used to collect guarantee knowledge data files and import them into the retrieval enhancement generation RAG system to form a preliminary knowledge base;

[0100] The loading and parsing module 220 is used to load and parse the guarantee knowledge data file in the first edition knowledge base, store the parsed text content in the memory, and form complete knowledge text data;

[0101] A word segmentation module 230 is configured to segment the knowledge text data using an optimized word segmentation strategy to obtain a plurality of large text blocks and corresponding small text blocks;

[0102] The embedding module 240 is used to determine and use the optimal embedding model, perform vector embedding processing on the knowledge block composed of the large text block and the corresponding small text block using the optimal embedding model, convert it into a vector representation and store it in the vector database to form a guarantee knowledge think tank.

[0103] An embodiment of the present invention provides a guarantee knowledge think tank construction system based on LLM and RAG technology, and its functional structure corresponds to a guarantee knowledge think tank construction method based on LLM and RAG technology provided by an embodiment of the present invention, which will not be repeated here.

[0104] Figure 3 It is a structural diagram of an electronic device 300 provided in an embodiment of the present invention. The electronic device 300 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 301 and one or more memories 302, wherein the memory 302 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 301 to implement the steps of the above-mentioned method for constructing a guarantee knowledge think tank based on LLM and RAG technology.

[0105] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory device containing instructions. These instructions are executable by a processor in a terminal to implement the aforementioned method for constructing a guarantee knowledge think tank based on LLM and RAG technologies. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0106] Figure 4 It is the process of building a guarantee knowledge think tank.

[0107] Data collection: Data related to guarantee business is collected through multiple channels, including various text-based materials such as guarantee business case documents and industry research reports, as the basic data source for building a guarantee knowledge think tank.

[0108] Text cleaning: Perform text cleaning operations on the collected data to remove noise data in the text (such as irrelevant characters, incorrect encoding, repeated content, etc.), standardize the text format, improve the quality of the data, and provide clean and effective text content for subsequent processing. The cleaned text is stored in the guarantee knowledge base.

[0109] Text segmentation: The text in the guarantee knowledge base is segmented into blocks. Based on the semantic association, content theme and other rules of the text, the long text is divided into multiple guarantee knowledge blocks, making the knowledge presentation more granular and facilitating subsequent targeted processing and knowledge mining.

[0110] Text vectorization: Perform text vectorization conversion on the segmented guarantee knowledge blocks, and use word embedding, vector encoding and other technologies in natural language processing to convert the text into a vector form that is computer-recognizable, easy to calculate and retrieve. The generated vector data is stored in the guarantee knowledge vector library, realizing the structured and numerical representation of knowledge, and facilitating the rapid matching, retrieval and intelligent application of knowledge.

[0111] Figure 5 This is a diagram of the text block processing results, showing the structured organization of the cleaned text:

[0112] Two types of text blocks are distinguished: KCS (large text block): This serves as a top-level container and can nest multiple TCSs. The figure shows a multi-level KCS nesting structure, reflecting the macro-grouping of text. TCS (small text block): This is the basic unit, carrying specific text content (such as items numbered 1-6) and serving as the fundamental carrier of text semantics.

[0113] Through hierarchical nesting (KCS includes TCS, and KCS can nest KCS), the organizational structure of the cleaned text from the overall (large module) to the local (small unit) is clearly displayed, reflecting the structured splitting and classification logic of the content by the text segmentation algorithm. It is often used in NLP processes such as text mining and information extraction to assist subsequent semantic analysis and data processing.

[0114] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0115] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.< / tcs> < / kcs>

Claims

1. A method for constructing a guarantee knowledge think tank based on LLM and RAG technology, characterized by: The method comprises: S1. Collect guarantee knowledge data files and import them into the retrieval enhancement generation RAG system to form the initial version of the knowledge base; S2. Load and parse the guarantee knowledge data file in the first edition knowledge base, and store the parsed text content in the memory to form complete knowledge text data; S3. Segmenting the knowledge text data using an optimized segmentation strategy to obtain a plurality of large text blocks and corresponding small text blocks; S4. Determine and use the optimal embedding model to perform vector embedding processing on the knowledge blocks composed of large text blocks and corresponding small text blocks, convert them into vector representations and store them in the vector database to form a guarantee knowledge think tank.

2. The method according to claim 1, characterized in that Said S3 specifically includes: The knowledge text is segmented using a word segmentation model based on the pre-trained BERT architecture. According to the logical structure and semantic integrity of the knowledge content, the formula is adopted: ; Calculating block quality scores ,determine the division criteria of large text blocks and obtain multiple large text blocks with reasonable semantic units; in, represents the semantic coherence aggregation weight, represents the position dispersion compensation coefficient, represents the punctuation density suppression coefficient, Representative The mean semantic similarity of adjacent blocks, Represents the starting position index of the current candidate block, Representative The base position index of the semantic anchor point, Represents the number of punctuation marks in the current block. Represents the total number of characters in the current block. Represents the number of adjacent blocks; After the large text block segmentation is completed, the length of each large text block is judged. If the size of the large text block exceeds the pre-set text block length threshold, the length segmentation method is used to perform a secondary segmentation on the large text block, subdividing it into multiple small text blocks that meet the length requirements, and establishing an association between the small text blocks and the large text block. If the length of the large text block is less than or equal to the pre-set text block length threshold, the large text block is directly stored as a complete unit.

3. The method according to claim 2, characterized in that The establishing of the association relationship between the small text block and the large text block specifically includes: The information of the large text block to which the small text block belongs is recorded by adding a "parent_block_id" field in the data structure.

4. The method according to claim 2, characterized in that The length segmentation method performs automatic segmentation processing by setting the segmentation length and overlap length, specifically including: Preset word segmentation length and overlap length; Preset word segmentation length and overlap length , according to the formula: ; Calculating dynamic offset , intercepted from the beginning of the text The length segment is used as the segmentation result, and the adjacent segmentation segments are retained. For the part with the same length, we get the word segmentation result set; in, Represents the sequence number of the current word segment, Represents the preset word segmentation length, Represents the preset overlap length, Representative The dynamic offset of the starting position of each segment, Represents a dynamic balancing factor based on indexes and preset parameters; To ensure the coherence and accuracy of word segmentation, there is a certain overlap between adjacent word segmentation segments according to the pre-set overlap length.

5. The method according to claim 2, characterized in that Said S3 further includes: For some knowledge text data, manual tagging is used to achieve word segmentation by manually adding specific identifiers to the text content, including: Based on the semantic structure of the text and business needs, annotators insert specific identifiers into the text to mark the start and end positions of knowledge blocks and clearly define a complete knowledge unit. When loading text, these tags are automatically identified and the text is segmented according to the tags, dividing it into corresponding large text blocks and small text blocks.

6. The method according to claim 1, wherein The method further comprises: Determine an optimal large language model (LLM), submit each large text block to the LLM, and pre-generate a set of relevant questions and a summary based on the content of the large text block; The optimal embedding model performs vector embedding processing on the relevant questions and summary abstracts, converts them into vector representations, and stores them together with the large text blocks and the corresponding small text blocks in the vector database, and establishes an index connection between them to form the guarantee knowledge think tank.

7. The method according to claim 6, characterized in that The establishment of index connections between them specifically includes: Create an index table for each large text block, recording the vector IDs of the large text block, pre-generated related questions and summary summaries, and the corresponding relationships between them.

8. A guarantee knowledge think tank construction system based on LLM and RAG technology, the system is used to implement the method according to any one of claims 1 to 7, characterized in that: The system comprises: The collection and import module is used to collect guarantee knowledge data files and import them into the retrieval enhancement generation RAG system to form the initial version of the knowledge base; A loading and parsing module is used to load and parse the guarantee knowledge data file in the first edition knowledge base, store the parsed text content in the memory, and form complete knowledge text data; A word segmentation module, configured to segment the knowledge text data using an optimized word segmentation strategy to obtain a plurality of large text blocks and corresponding small text blocks; The embedding module is used to determine and use the optimal embedding model, perform vector embedding processing on the knowledge block composed of large text blocks and corresponding small text blocks using the optimal embedding model, convert it into a vector representation and store it in the vector database to form a guarantee knowledge think tank.

9. An electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, characterized in that: The at least one instruction is loaded and executed by the processor to implement the guarantee knowledge think tank construction method based on LLM and RAG technology as described in any one of claims 1-7.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The at least one instruction is loaded and executed by the processor to implement the guarantee knowledge think tank construction method based on LLM and RAG technology as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Construction method, system and device of retrieval enhancement generation system and medium

    CN118797060A

  • Method for constructing vectorized knowledge base based on dynamic fields for RAG system

    CN119377418A

  • Automatic RAG knowledge base construction method and system based on large model

    CN119558393A

  • Agricultural knowledge base construction method based on large language model

    CN119962649A

  • Structured document knowledge base construction method, system and equipment based on tree structure and semantic perception and medium

    CN120179676A