Method and device for accurately decomposing AI text into large model cue words
By dividing the text into titles, directories and paragraphs for vectorization, the problem of difficulty in distinguishing similar content in the document construction knowledge base is solved, and text query with high accuracy and high search rate is achieved.
Patent Information
- Application Number
- CN202410274783.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-11
- Publication Date
- 2025-07-22
AI Technical Summary
Existing document construction knowledge base technology cannot effectively distinguish similar content between multiple documents, resulting in too small granularity and difficulty in accurately finding data, especially when answering in government or legal documents, the error rate is high.
By dividing the text into three containers, storage, and using the SentenceTransformers tool for vectorization, merge each container vector and compare the semantic closest vector with the user query statement, and return the corpus for use in the big model.
It improves the accuracy and search rate of text query, ensures the retention of subject information and the synchronization of information after segmentation, and realizes accurate text query at the title, directory and paragraph levels.
Smart Images

Figure CN120353883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and particularly to a method and device for accurately decomposing AI text into large model prompts. Background Art
[0002] In the existing technology of building a knowledge base from documents, there is no discussion on how to distinguish the similar content among various documents. If the text segmentation technology of OCR is uniformly used, the granularity will be too small, making it difficult to distinguish the differences between various subjects. Even accurate data cannot be found - because when constructing the vector library, clustering will occur, resulting in the loss of information of other subjects.
[0003] For example, in the knowledge base-based dialogue systems in patents such as those with the publication numbers CN117093698A, CN116992005A, and CN117009492A, when it comes to building a knowledge base from documents, basically, the OCR tool is uniformly used for document segmentation to generate word vectors, and then FAISS is used to generate a word vector library. When a user queries, the question words input by the user are vectorized and compared with the knowledge base vector library (generally, the cosine distance or Euclidean distance is calculated) to retrieve relevant corpus information. Then, the large model is used to answer questions based on this corpus information. It does not explain how to effectively solve the problems of ambiguity and overlap in the presence of multiple documents. For example, the NLP module is not clear and does not focus on this. Currently, it is difficult to solve the problem that the large model gives incorrect answers when answering questions from the local knowledge base, especially when dealing with government or legal documents where high precision is required for time, place, person, and data, and the error rate is often high.
[0004] For example, for the "Industrial Transfer Main Platform Policy of a Certain Province", relevant policies will be introduced in each of its cities, such as the "Policy for a Certain City to Undertake the Industrial Main Platform". Different implementation methods exist for the same policy in different prefecture-level cities. When asking about a certain policy, both the provincial policy and the policies of each city will be listed. However, due to the limitations of the large model's answers, there may be a situation where not all policies are fully listed, and even the policy situations of each city may be confused.
[0005] In the existing technology of building a knowledge base from documents, there is no discussion on how to distinguish the similar content among various documents. If the text segmentation technology of OCR is uniformly used, the granularity will be too small, making it difficult to distinguish the differences between various subjects. Even accurate data cannot be found - because when constructing the vector library, clustering will occur, resulting in the loss of information of other subjects. Summary of the Invention
[0006] The main objective of the present invention is to provide a method and device for accurately decomposing AI text into large model prompts, aiming to solve the technical problem of low accuracy in existing text generation.
[0007] To achieve the above object, the present invention provides a method for accurately decomposing AI text into large model prompts. A method for accurately decomposing AI text into large model prompts includes the following steps:
[0008] S1: Read the text through a text reading tool;
[0009] S2: Store the read text in three containers: title, table of contents, and paragraphs;
[0010] S3: Vectorize the content in the three containers based on a text library tool;
[0011] S4: Merge the vectors of the three containers;
[0012] S5: Merge the vectors of all documents of the text;
[0013] S6: Vectorize the user query statement, compare it with the three-dimensional vectors in the document, find the vector with the closest semantics, and return the user query statement as a corpus for use by the large model.
[0014] In the method for accurately decomposing AI text into large model prompts provided in this application, reading the text through a text reading tool includes: reading the text through the text reading tool textLoader.
[0015] In the method for accurately decomposing AI text into large model prompts provided in this application, the three containers respectively include: T, M, P, where T stores the title, M stores the table of contents, and P stores the paragraphs.
[0016] In the method for accurately decomposing AI text into large model prompts provided in this application, one topic corresponds to multiple tables of contents, one table of contents corresponds to multiple paragraphs, and one paragraph is divided into multiple words.
[0017] In the method for accurately decomposing AI text into large model prompts provided in this application, the vector of one document is N*2304, and the vectors of M articles are three-dimensional vectors of M*N*2304.
[0018] In the method for accurately decomposing AI text into large model prompts provided in this application, vectorizing the user query statement includes: decomposing the user query statement and performing vectorization using a text library tool.
[0019] In the method for accurately decomposing AI text into large model prompts provided in this application, the text library tool includes the SentenceTransformers tool.
[0020] In addition, to achieve the above object, the present invention further provides a device for accurately decomposing AI text into large model prompts. The device for accurately decomposing AI text into large model prompts includes:
[0021] A text reading module for reading text through a text reading tool;
[0022] A storage module for storing the read text in three containers: title, table of contents, and paragraphs;
[0023] A vectorization module for vectorizing the contents in the three containers based on a text library tool;
[0024] A first merging module for merging the vectors of the three containers;
[0025] A second merging module for merging the vectors of all documents of the text;
[0026] A statement query module for vectorizing the user query statement, comparing it with the three-dimensional vectors in the document, finding the vector with the closest semantics, and returning the user query statement as a corpus for use by the large model.
[0027] In addition, to achieve the above object, the present invention further provides a device for accurately decomposing AI text into large model prompts. The device for accurately decomposing AI text into large model prompts includes a processor, a memory, and a program for accurately decomposing AI text into large model prompts stored on the memory and executable by the processor. When the program for accurately decomposing AI text into large model prompts is executed by the processor, the steps of the method for accurately decomposing AI text into large model prompts as described above are implemented.
[0028] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium. A program for accurately decomposing AI text into large model prompts is stored on the computer-readable storage medium. When the program for accurately decomposing AI text into large model prompts is executed by a processor, the steps of the method for accurately decomposing AI text into large model prompts as described above are implemented.
[0029] The present invention provides a method for accurately decomposing AI text into large model prompts. The method reads the text through a text reading tool, stores the read text in three containers, vectorizes the content in the containers based on a text library tool, merges the vectors of each container, merges all documents, vectorizes the user query statement, compares it with the three-dimensional vectors in the documents, finds the vector with the closest semantics, and returns the user query statement as a corpus for use by the large model. This article proposes a text segmentation technology that can not only retain the main information but also synchronize the segmented technical information with the main information. The segmentation in this article is carried out according to three levels: title, table of contents, and paragraphs, so as to achieve the integration of the main body and the local part, and achieve accurate text query in three dimensions: the main body (title decomposition), the local part (table of contents subtitle decomposition), and the details (paragraph decomposition).
[0030] This article mainly changes the NLP module and achieves accurate segmentation of the corpus by subdividing the article title, subtitle, and paragraphs. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic flowchart of the first embodiment of the method for accurately decomposing AI text into large model prompts according to the present invention;
[0032] Figure 2 It is a schematic diagram of the method for accurately decomposing AI text into large model prompts in the method for accurately decomposing AI text into large model prompts according to the present invention;
[0033] Figure 3 It is a schematic diagram of the effect of testing and comparing the method for accurately decomposing AI text into large model prompts according to the present invention with the prior art in terms of accuracy rate and recall rate;
[0034] Figure 4 It is a schematic diagram of the hardware structure of the device for accurately decomposing AI text into large model prompts involved in the embodiment solution of the present invention.
[0035] The realization, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] This application proposes a text segmentation technology that can not only retain the main information but also synchronize the segmented technical information with the main information. The segmentation in this article is carried out according to three levels: title, table of contents, and paragraphs, so as to achieve the integration of the main body and the local part, and achieve accurate text query in three dimensions: the main body (title decomposition), the local part (table of contents decomposition), and the details (paragraph decomposition).
[0038] This article mainly changes the NLP module. By subdividing headings, tables of contents, and paragraphs, it achieves precise semantic segmentation of the corpus.
[0039] The embodiment of the present invention provides a method for accurately decomposing AI text into large model prompts.
[0040] Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the method for accurately decomposing AI text into large model prompts according to the present invention.
[0041] In this embodiment, the method for accurately decomposing AI text into large model prompts includes the following steps:
[0042] A method for accurately decomposing AI text into large model prompts includes the following steps:
[0043] S1: Read the text through a text reading tool;
[0044] Specifically, read the text through the text reading tool textLoader.
[0045] S2: Store the read text in three containers: headings, table of contents, and paragraphs;
[0046] Store the read text in three containers. The T_container stores headings, the M_container stores the table of contents, and the P_container stores paragraphs.
[0047] S3: Vectorize the content in the three containers based on a text library tool;
[0048] Use the SentenceTransformers tool to vectorize the content in the containers of step 2. (For example, when a sentence is put into SentenceTransformers, it will output a 1*768 vector). For example, a single-sentence heading is a 1*768 vector, and another single-sentence heading is another 1*768-dimensional vector. If each paragraph has N sentences, there will be N*768 vectors.
[0049] S4: Merge the vectors of the three containers;
[0050] Merge the vectors generated in step S3. One topic corresponds to multiple tables of contents, one table of contents corresponds to multiple paragraphs, and one paragraph is further segmented into multiple words.
[0051] For example, a certain document has a title T, two tables of contents M1 and M2. There are two paragraphs P1 and P2 under M1; there are two paragraphs Q1 and Q2 under M2. Then the entire document becomes four combined vectors: TM1P1, TM2P2, TM2Q1, and TM2Q2. TM1P1 is composed of three vectors: 768 + 768 + 768 = 2304. Therefore, the entire document is a two-dimensional vector of 4 * 2304.
[0052] S5: Combine the vectors of all documents of the text;
[0053] Combine all the documents of the text. If a document has N large paragraphs, there will be N * 2304 vectors. If there are M documents, it will become a three-dimensional vector of M * N * 2304.
[0054] S6: Vectorize the user query statement, compare it with the three-dimensional vectors in the document, find the vector with the closest semantics, and return the user query statement as a corpus for the large model to use.
[0055] Decompose the user query statement into vectors using the SentenceTransformers algorithm in step S3, compare it with the three-dimensional vectors in step S5 (perform cosine vectors), find the vector with the closest semantics, and the corresponding M and N correspond to that part of the content of that article. Return the statement as a corpus for the large model to use.
[0056] The present invention provides a method for accurately decomposing AI text into large model prompt words. Read the text through a text reading tool; store the read text in three containers; vectorize the content in the containers based on a text library tool; combine the vectors of each container; combine all the documents; vectorize the user query statement, compare it with the three-dimensional vectors in the document, find the vector with the closest semantics, and return the user query statement as a corpus for the large model to use.
[0057] This article proposes a text segmentation technology that can both retain the main information and synchronize the technical information after segmentation with the main information. The segmentation in this article is carried out according to three levels: title, table of contents, and paragraph, to achieve the integration of the main body and the local part, and to achieve accurate text query in three dimensions: the main body (title decomposition), the local part (table of contents decomposition), and the details (paragraph decomposition). This article mainly changes the NLP module, and achieves accurate segmentation of the corpus by subdividing the title, table of contents, and paragraph.
[0058] The NLP (Natural Language Processing) module is a software tool or library used to process and analyze human language. It can be used in various applications, such as speech recognition, machine translation, sentiment analysis, text classification, information retrieval, etc.
[0059] Specifically, the NLP module usually includes the following functions:
[0060] Lexical analysis: decomposing the text into words, punctuation marks, and other language elements.
[0061] Syntactic analysis: analyzing the structure and grammar of sentences, such as determining the subject, predicate, object, etc.
[0062] Semantic analysis: understanding the meaning and context of the text, such as identifying synonyms, antonyms, concept relationships, etc.
[0063] Text classification: assigning the text to predefined categories, such as sentiment classification, topic classification, etc.
[0064] Information retrieval: searching and retrieving relevant information in a text database.
[0065] Machine translation: translating the text of one language into another language.
[0066] Speech recognition: converting spoken language into text.
[0067] Text generation: generating new text, such as abstracts, articles, dialogues, etc.
[0068] The NLP module can use various technologies and algorithms to implement these functions, such as statistical machine learning, deep learning, rules and grammar, semantic networks, etc. They are usually provided in the form of libraries or APIs of programming languages so that developers can integrate them into their own applications.
[0069] Some common NLP modules include StanfordNLP, OpenNLP, SpaCy, NLTK (Natural Language Toolkit), etc. These modules provide various functions and tools to help developers process and analyze natural language text.
[0070] In the method for accurately decomposing AI text into large model prompts provided in this application, reading the text through a text reading tool includes: reading the text through the text reading tool textLoader.
[0071] In the method for accurately decomposing AI text into large model prompts provided in this application, the three containers respectively include: T, M, P, where T stores the title, M stores the table of contents, and P stores the paragraphs.
[0072] In the method for accurately decomposing AI text into large model prompts provided in this application, one title corresponds to multiple tables of contents, one table of contents corresponds to multiple paragraphs, and one paragraph is divided into multiple words. For example Figure 2As shown, an article generally has a title, M sub - titles, and P paragraphs. Generally, one paragraph corresponds to one sub - title, that is, P >= M.
[0073] Suppose the length of the token embedding vector of the large model is 768. Our storage method is to concatenate a title + a sub - title + a paragraph, which becomes P * (768 * 3), that is, P vectors of 2304. This P * 2304 represents the encoding vector of this article. If there are N articles, there will be N * P * 2304 vectors. Thus, we use the above - mentioned encoding method for the title, sub - title, and paragraph to accurately distinguish the vectors with the same meaning in different articles.
[0074] In the method for accurately decomposing AI text into large - model prompts provided in this application, a document vector is N * 2304, and for M articles, it is a three - dimensional vector of M * N * 2304.
[0075] In the method for accurately decomposing AI text into large - model prompts provided in this application, the vectorization of the user query statement includes: decomposing the user query statement into vectorization using a text library tool.
[0076] In the method for accurately decomposing AI text into large - model prompts provided in this application, the text library tool includes the SentenceTransformers tool.
[0077] Sentence Transformers is a Python library based on PyTorch and Transformers. It can be used for sentence, text, and image embeddings. It can calculate text embeddings for more than 100 languages and can easily be used for common tasks such as semantic text similarity, semantic search, and synonym mining.
[0078] This framework is based on PyTorch and Transformers and provides a large number of pre - trained models for various tasks. It can also be easily fine - tuned according to its own models.
[0079] Through testing, this application has a significant improvement in both accuracy and recall rate. The effect comparison is as Figure 3 shown. In terms of accuracy, the accuracy of implementing this application can exceed 80%, and the accuracy before implementing this application is 70% +; in terms of recall rate, the accuracy of implementing this application can exceed 50%, and the accuracy before implementing this application is 30% +.
[0080] In addition, to achieve the above - mentioned purpose, the present invention also provides a device for accurately decomposing AI text into large - model prompts. The device for accurately decomposing AI text into large - model prompts includes:
[0081] A text reading module for reading text through a text reading tool;
[0082] A storage module that stores the read text in three containers: title, table of contents, and paragraphs;
[0083] A vectorization module that vectorizes the content in the three containers based on a text library tool;
[0084] A first merging module that merges the vectors of the three containers;
[0085] A second merging module that merges the vectors of all the documents of the text;
[0086] A statement query module that vectorizes the user's query statement, compares it with the three-dimensional vectors in the document, finds the vector with the closest semantics, and returns the user's query statement as a corpus for use by the large model.
[0087] In addition, to achieve the above object, the present invention also provides a device for accurately decomposing AI text into large model prompt words. The device for accurately decomposing AI text into large model prompt words includes a processor, a memory, and a program for accurately decomposing AI text into large model prompt words stored on the memory and executable by the processor. When the program for accurately decomposing AI text into large model prompt words is executed by the processor, the steps of the method for accurately decomposing AI text into large model prompt words as described above are implemented.
[0088] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium. A program for accurately decomposing AI text into large model prompt words is stored on the computer-readable storage medium. When the program for accurately decomposing AI text into large model prompt words is executed by a processor, the steps of the method for accurately decomposing AI text into large model prompt words as described above are implemented.
[0089] Among them, the method implemented when the program for accurately decomposing AI text into large model prompt words is executed can refer to the various embodiments of the method for accurately decomposing AI text into large model prompt words of the present invention, which will not be elaborated here.
[0090] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0091] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority order of the embodiments.
[0092] The method for accurately decomposing AI text into large model prompts involved in the embodiments of the present invention is mainly applied to a device for accurately decomposing AI text into large model prompts. The device for accurately decomposing AI text into large model prompts can be a device with display and processing functions such as a PC, a portable computer, a mobile terminal, etc.
[0093] Refer to Figure 4 , Figure 4 which is a schematic diagram of the hardware structure of the device for accurately decomposing AI text into large model prompts involved in the solution of the embodiments of the present invention. In the embodiments of the present invention, the device for accurately decomposing AI text into large model prompts may include a processor 1001 (such as a CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components; the user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard); the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface); the memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory, and the memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.
[0094] Those skilled in the art can understand that Figure 2 the hardware structure shown in
[0095] does not constitute a limitation on the device for accurately decomposing AI text into large model prompts, and may include more or fewer components than shown in the figure, or combine some components, or arrange different components. Figure 2 , Figure 2 the memory 1005 as a computer-readable storage medium in
[0096] In Figure 2 , the network communication module is mainly used to connect to the server and communicate with the server for data; while the processor 1001 can call the program for accurately decomposing AI text into large model prompts stored in the memory 1005 and execute the method for accurately decomposing AI text into large model prompts provided by the embodiments of the present invention.
[0097] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0099] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for accurately decomposing AI text into large model prompts, characterized in that, The method for accurately decomposing AI text into large model prompts includes the following steps: S1: Read the text through a text reading tool; S2: Store the read text in three containers: title, table of contents, and paragraphs; S3: Vectorize the content in the three containers based on a text library tool; S4: Merge the vectors of the three containers; S5: Merge the vectors of all documents of the text; S6: Vectorize the user query statement, compare it with the three-dimensional vectors in the document, find the vector with the closest semantics, and return the user query statement as a corpus for the large model to use.
2. The method for accurately decomposing AI text into large model prompts according to claim 1, characterized in that, The step of reading the text through a text reading tool includes: reading the text through the text reading tool textLoader.
3. The method for accurately decomposing AI text into large model prompts as claimed in claim 1, wherein, The three containers respectively include: T, M, P, where T stores the title, M stores the table of contents, and P stores the paragraphs.
4. The method for accurately decomposing AI text into large model prompts as claimed in claim 1, wherein, One title corresponds to multiple tables of contents, one table of contents corresponds to multiple paragraphs, and one paragraph is divided into multiple words.
5. The method for accurately decomposing AI text into large model prompts according to claim 1, characterized in that, The vector of a single document is N*2304, and the vectors of M articles form a three-dimensional vector of M*N*2304.
6. The method for accurately decomposing AI text into large model prompts as described in claim 1, characterized in that, The step of vectorizing the user query statement includes: decomposing the user query statement and using a text library tool for vectorization.
7. The method for accurately decomposing AI text into large model prompts as described in claim 1, characterized in that, The text library tool includes the SentenceTransformers tool.
8. The method for accurately decomposing AI text into large model prompts as described in claim 1, wherein The text reading tool includes the textLoader tool.
9. An apparatus for accurately decomposing AI text into large model prompts, characterized in that, The device for accurately decomposing AI text into large model prompts includes: A text reading module for reading the text through a text reading tool; A storage module for storing the read text in three containers: title, table of contents, and paragraphs; A vectorization module for vectorizing the content in the three containers based on a text library tool; A first merging module for merging the vectors of the three containers; A second merging module for merging the vectors of all documents of the text; A statement query module for vectorizing the user query statement, comparing it with the three-dimensional vectors in the document, finding the vector with the closest semantics, and returning the user query statement as a corpus for the large model to use.
10. The device for accurately decomposing AI text into large model prompts as described in claim 9, wherein One title corresponds to multiple tables of contents, one table of contents corresponds to multiple paragraphs, and one paragraph is divided into multiple words.
Citation Information
Patent Citations
Engineering consultation report retrieval method combining semantics and association matching
CN115858813A
Financial statement attached event extraction method and system and storage medium
CN116150361A
Document processing and response generation system
CN116157790A
Text question and answer processing method and device, electronic equipment and storage medium
CN117171328A
Text generation method and apparatus, storage medium, and electronic device
WO2020258948A1