Complex text analyzing and processing system and method based on large language model

Through a complex text analysis and processing system based on a large language model, the problem of accuracy and inefficiency of complex text analysis in the prior art is solved, efficient analysis and processing of non-standard format files is achieved, the accuracy and structural integrity of text content is ensured, and the value of information retrieval and analysis is enhanced.

CN120106045APending Publication Date: 2025-06-06数字宁波科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510264687.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively parse and process complex text, especially non-standard format files, such as scanned documents, which leads to inaccuracy and inefficiency in parsing, and garbled codes are prone to appear in the text and loses the original structure.

Method used

A complex text analysis and processing system based on large language models is adopted, including preprocessing modules, multimodal analysis modules, large language model sorting modules, entity extraction modules and database storage modules. The system converts non-encoded text into picture files through image processing technology, uses multimodal large model to identify and extract text information, and performs semantic analysis and error correction through large language models, and finally stores the text in the database.

Benefits of technology

It improves the electronic processing efficiency and accuracy of complex texts, ensures the readability and correctness of file content, and enhances the retrieval and analytical value of text information through entity extraction and database storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106045A_ABST
    Figure CN120106045A_ABST
Patent Text Reader

Abstract

The invention discloses a complex text analysis and processing system and method based on a large language model, and relates to the technical field of text analysis and processing, and the system comprises a preprocessing module, a multi-modal analysis module, a large language model arrangement module, an entity extraction module and a database storage module. Wherein the preprocessing module receives a complex text, performs format detection, determines whether the complex text is in a coding format, and converts the complex text in a non-coding format into a picture file; the multi-modal analysis module uses a multi-modal large model to identify and extract text information in the picture file, and converts the picture file into text data in a coding format; the large language model sorting module uses a large language model to perform semantic analysis and error correction to obtain an output text, and sorts the format and layout of the output text; the entity extraction module extracts entity information from the output text; and the database storage module stores the entity information into a database to realize structured storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text parsing and processing, and in particular to a complex text parsing and processing system and method based on a large language model. Background Art

[0002] With the advent of the information age, the electronic processing of complex text information has become the key to improving work efficiency and information transparency. Traditional text processing relies on the analysis of coding formats, which is effective when processing electronic documents in standard formats. However, for non-standard format files such as scanned copies, traditional coding format analysis methods often cannot effectively read the text content. In addition, even if the file can be parsed, garbled characters often appear due to problems such as format conversion, which seriously affects the readability and accuracy of the file.

[0003] In the existing technology, some solutions try to parse the text in the scanned documents through optical character recognition (OCR) technology. However, the recognition accuracy of OCR technology will drop significantly when facing complex backgrounds, content obstruction or low-resolution scans. In addition, OCR technology is usually unable to process the format and layout information in the file, resulting in the loss of the original structure of the parsed text, which brings difficulties to the subsequent information collation and analysis.

[0004] In order to improve the efficiency and accuracy of electronic processing of non-standard format files, some researchers have tried to apply deep learning technology to the field of document parsing. These methods have improved the recognition accuracy to a certain extent, but there are still some limitations. First, these methods usually require a large amount of annotated data for training, and the annotation of complex texts is time-consuming and expensive. Second, these methods often focus on text recognition and ignore the extraction and processing of other important information in the file (such as watermarks, annotations, etc.).

[0005] Therefore, technicians in this field are committed to developing a new text parsing and processing system and method to solve the above-mentioned defects in the prior art. Summary of the invention

[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to more effectively parse and process complex texts in various formats, improve the efficiency and accuracy of electronic processing of complex texts, and at the same time ensure the readability and correctness of the file content.

[0007] To achieve the above-mentioned purpose, the present invention provides a complex text parsing and processing system based on a large language model, comprising a preprocessing module, a multimodal parsing module, a large language model sorting module, an entity extraction module and a database storage module; The preprocessing module receives the complex text and performs format detection to determine whether the complex text is in a codable format. For the complex text in a non-codable format, the preprocessing module converts the complex text into a picture file through image processing technology; The multimodal analysis module is connected to the preprocessing module, receives the image file, uses the multimodal large model to identify and extract text information in the image file, and converts the image file into text data in an encodable format; The large language model arrangement module is connected to the preprocessing module and the multimodal analysis module, receives the complex text in the encodable format in the preprocessing module and the text data in the encodable format converted from the image file in the multimodal analysis module, uses the pre-trained large language model to perform semantic analysis and error correction to obtain output text, and arranges the format and layout of the output text; The entity extraction module is connected to the large language model sorting module to extract entity information from the output text; The database storage module is connected to the entity extraction module to store the entity information into a database to achieve structured storage.

[0008] Furthermore, the preprocessing module also preprocesses the image file, and the preprocessing methods adopted include denoising, contrast enhancement or binarization.

[0009] Furthermore, the multimodal large model used in the multimodal parsing module is the Qwen-VL-Chat multimodal large model, and the Qwen-VL-Chat multimodal large model identifies and extracts the text information in the image file by analyzing the text layout and semantic content in the image file.

[0010] Furthermore, the pre-trained large language model used in the large language model sorting module is BERT or GPT.

[0011] Furthermore, the entity extraction module uses NER technology to extract the entity information.

[0012] The present invention also provides a complex text parsing and processing method based on a large language model, the method comprising the following steps: Step 1: receiving a complex text and performing format detection to determine whether the complex text is in a codable format. For the complex text in a non-codable format, converting it into a picture file through image processing technology; Step 2: receiving the image file, using the multimodal large model to identify and extract text information in the image file, and converting the image file into text data in an encodable format; Step 3, receiving the complex text in the encodable format in step 1 and the text data in the encodable format converted from the image file in step 2, performing semantic analysis and error correction using a pre-trained large language model to obtain output text, and arranging the format and layout of the output text; Step 4: extracting entity information from the output text; Step 5: Storing the entity information in a database to achieve structured storage.

[0013] Furthermore, the step 1 includes the following sub-steps: Step 1.1, receiving the complex text, which is a file from different channels, including a scanned copy of an electronic document or a paper document; Step 1.2: Perform format detection to identify the format of the file. For the complex text in a non-encoded format, convert it into the picture file through image processing technology.

[0014] Furthermore, the multimodal large model used in step 2 is the Qwen-VL-Chat multimodal large model, and the Qwen-VL-Chat multimodal large model identifies and extracts the text information in the image file by analyzing the text layout and semantic content in the image file.

[0015] Furthermore, the pre-trained large language model used in step 3 is BERT or GPT.

[0016] Furthermore, the step 4 uses NER technology to extract the entity information.

[0017] The complex text parsing and processing system and method based on a large language model provided by the present invention has at least the following technical effects: 1. The technical solution provided by the present invention first uses the traditional solution to try to parse. If the parsing fails or serious garbled characters appear, the file is converted into a picture format, and then the Qwen-VL-Chat multimodal large model is used for content parsing. This method not only improves the parsing success rate of non-standard format files, but also can better maintain the integrity and accuracy of the file content. At the same time, the large language model is used to organize and correct the parsed content, further improving the readability and correctness of the text. Traditional text parsing methods often rely on the parsing of encoding formats. When facing non-standard format files, such as stamped scans, they often cannot effectively read the text content, resulting in parsing failures or a large amount of garbled characters, which seriously affects the readability and accuracy of the file. In addition, even if the file can be parsed, the original structure of the text is often lost due to problems such as format conversion, which brings difficulties to subsequent information collation and analysis. The complex text parsing and processing method based on a large language model proposed in the present invention can effectively solve these problems.

[0018] 2. The technical solution provided by the present invention extracts each entity in the file through the entity extraction module, which enhances the retrieval and analysis value of the file content, and stores all parsed and extracted data in the database instead of simply outputting it as a file, which greatly facilitates the long-term management and utilization of complex texts. Users can retrieve, analyze and update the data in the database at any time, which greatly improves the utilization efficiency and value of text information.

[0019] In summary, the technical solution provided by the present invention not only improves the accuracy and efficiency of complex text parsing, but also provides strong support for in-depth analysis and long-term management of text information through entity extraction and database storage, which is unmatched by traditional parsing methods.

[0020] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of system modules and connection relationships of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0022] The following describes several preferred embodiments of the present invention with reference to the drawings in the specification, so that the technical content is clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0023] The embodiment of the present invention provides a system and method for parsing and processing complex and professional text information, which combines traditional parsing technology, multimodal large models and large language models, aiming to improve the efficiency and accuracy of electronic processing of complex texts while ensuring the readability and correctness of the file content. The technical route provided by the embodiment of the present invention can effectively parse and process complex texts in various formats, including non-standard format files such as stamped scans, providing strong technical support for electronic management and efficient query of information.

[0024] Example 1 like Figure 1 As shown, an embodiment of the present invention provides a complex text parsing and processing system based on a large language model, including a preprocessing module, a multimodal parsing module, a large language model sorting module, an entity extraction module and a database storage module; The preprocessing module receives complex text and performs format detection to determine whether the complex text is in a codable format. For complex text in non-encoded format, it is converted into a picture file through image processing technology. In particular, the preprocessing module receives complex text files from different sources. For complex text files in non-encoded format, such as scanned copies, this module will use advanced image processing technology to convert complex text files in non-encoded format into high-resolution picture format for subsequent processing; The multimodal parsing module is connected to the preprocessing module, receives the image file, uses the multimodal large model to identify and extract the text information in the image file, and converts the image file into text data in an encodable format; The large language model arrangement module is connected to the preprocessing module and the multimodal parsing module, receives the complex text in the encodable format in the preprocessing module and the text data in the encodable format converted from the image file in the multimodal parsing module, uses the pre-trained large language model to perform semantic analysis and error correction to obtain output text, and arranges the format and layout of the output text; The entity extraction module is connected to the large language model sorting module to identify and classify entities in different fields and extract entity information from the output text, including names of people, places, and institutions. The extracted entity information is then stored in a structured manner, which not only helps improve the retrieval of the file content, but also facilitates subsequent data analysis and information extraction. The database storage module is connected to the entity extraction module to store entity information in the database to achieve structured storage. This module not only provides data addition, deletion, modification and query functions to meet the user's data management needs, but also ensures data security and consistency. By storing the parsed data in the database, users can easily retrieve and analyze it, and it also provides a solid foundation for the long-term management and utilization of complex texts.

[0025] Example 2 On the basis of Example 1, the preprocessing module also preprocesses the image file, and the preprocessing methods adopted include denoising, contrast enhancement or binarization. These steps are all to improve the accuracy and efficiency of the subsequent analysis stage. Through the above preprocessing steps, the system can lay a solid foundation for subsequent multimodal analysis.

[0026] In particular, the multimodal large model used in the multimodal parsing module is the Qwen-VL-Chat multimodal large model, in which the Qwen-VL-Chat multimodal large model identifies and extracts text information from image files by analyzing the text layout and semantic content in the image files. The Qwen-VL-Chat multimodal large model is specially designed to process input in image format. It can not only understand the content of the image, but also convert it into structured text data. The Qwen-VL-Chat multimodal large model can accurately identify and extract text information by analyzing the text layout and semantic content in the image, and can maintain high accuracy even in the case of poor image quality or complex background.

[0027] In particular, the multimodal parsing module can also perform preliminary semantic analysis on the parsed text, determine the structure and hierarchy of the text, and provide support for the subsequent large language model compilation.

[0028] In particular, the pre-trained large language model used in the large language model finishing module is BERT or GPT. The pre-trained large language model can understand the deep meaning of the text and identify grammatical errors and typos in the text for correction. In addition, the large language model finishing module is also responsible for arranging the format and layout of the text to restore the original layout of the file as much as possible, ensuring that the output text content is both accurate and easy to read.

[0029] In particular, the entity extraction module uses NER technology to extract entity information, using named entity recognition (NER) technology combined with pre-trained models to identify and classify entities in different fields. The extracted entity information is then stored in a structured manner, which not only helps to improve the retrieval of file content, but also facilitates subsequent data analysis and information extraction.

[0030] Example 3 The embodiment of the present invention also provides a complex text parsing and processing method based on a large language model, comprising the following steps: Step 1: Receive complex text and perform format detection to determine whether the complex text is in a codable format. For complex text in a non-codable format, convert it into a picture file through image processing technology; Step 2: Receive the image file, use the multimodal large model to identify and extract the text information in the image file, and convert the image file into text data in an encodable format; Step 3: receiving the complex text in the encodable format in step 1 and the text data in the encodable format converted from the image file in step 2, performing semantic analysis and error correction using the pre-trained large language model to obtain output text, and organizing the format and layout of the output text; Step 4: Extract entity information from the output text; Step 5: Store the entity information into the database to achieve structured storage.

[0031] Example 4 Based on Example 3, step 1 includes the following sub-steps: Step 1.1, receiving complex texts, which are files from different channels, including scanned copies of electronic documents or paper documents; Step 1.2: Perform format detection, identify the file format, determine the file format and perform necessary format conversion; for complex texts in non-encoded formats, such as scanned copies, use efficient image processing technology to convert the file into a high-resolution image file suitable for further analysis.

[0032] In particular, these images are also preprocessed by denoising, contrast enhancement or binarization to improve the quality of the images and provide clear input for the subsequent multimodal parsing module.

[0033] Example 5 Based on Example 3 or 4, the multimodal large model used in step 2 is the Qwen-VL-Chat multimodal large model. The Qwen-VL-Chat multimodal large model can understand the visual content in the picture and convert it into structured text data. By analyzing the text layout and semantic content in the picture, it can accurately identify and extract text information, even in the case of poor picture quality or complex background. It can maintain high accuracy. The output of this step is the preliminary parsed text content, which will be passed to the next step for further processing and organization.

[0034] In particular, the pre-trained large language model used in step 3 is BERT or GPT. The purpose of this step is to correct grammatical errors and typos in the text, and to organize the format and layout of the text to restore the original layout of the file as much as possible. Through these meticulous processing, the output text content will be more accurate and easier to read.

[0035] In particular, step 4 uses NER technology to extract entity information, using named entity recognition (NER) technology combined with pre-trained models to identify and classify entities in different fields. The extracted entity information is then stored in a structured manner, which not only helps to improve the retrieval of text content, but also facilitates subsequent data analysis and information extraction.

[0036] In particular, step 5 stores all the obtained entities, document contents, etc. into the database, which not only provides the data addition, deletion, modification and query functions to meet the user's data management needs, but also ensures the security and consistency of the data. By storing the parsed data in the database, users can easily retrieve and analyze it, and it also provides a solid foundation for the long-term management and utilization of complex texts.

[0037] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A complex text parsing and processing system based on a large language model, characterized in that: It includes preprocessing module, multimodal parsing module, large language model sorting module, entity extraction module and database storage module; The preprocessing module receives the complex text and performs format detection to determine whether the complex text is in a codable format. For the complex text in a non-codable format, the preprocessing module converts the complex text into a picture file through image processing technology; The multimodal analysis module is connected to the preprocessing module, receives the image file, uses the multimodal large model to identify and extract text information in the image file, and converts the image file into text data in an encodable format; The large language model arrangement module is connected to the preprocessing module and the multimodal analysis module, receives the complex text in the encodable format in the preprocessing module and the text data in the encodable format converted from the image file in the multimodal analysis module, uses the pre-trained large language model to perform semantic analysis and error correction to obtain output text, and arranges the format and layout of the output text; The entity extraction module is connected to the large language model sorting module to extract entity information from the output text; The database storage module is connected to the entity extraction module to store the entity information into a database to achieve structured storage.

2. The complex text parsing and processing system based on a large language model as claimed in claim 1, characterized in that: The preprocessing module also preprocesses the image file, and the preprocessing methods adopted include denoising, contrast enhancement or binarization.

3. The complex text parsing and processing system based on a large language model as claimed in claim 1, characterized in that: The multimodal large model used in the multimodal analysis module is the Qwen-VL-Chat multimodal large model. The Qwen-VL-Chat multimodal large model identifies and extracts the text information in the image file by analyzing the text layout and semantic content in the image file.

4. The complex text parsing and processing system based on a large language model as claimed in claim 1, characterized in that: The pre-trained large language model used in the large language model sorting module is BERT or GPT.

5. The complex text parsing and processing system based on a large language model as claimed in claim 1, characterized in that: The entity extraction module extracts the entity information using NER technology.

6. A complex text parsing and processing method based on a large language model, characterized in that: The method comprises the following steps: Step 1: receiving a complex text and performing format detection to determine whether the complex text is in a codable format. For the complex text in a non-codable format, converting it into a picture file through image processing technology; Step 2: receiving the image file, using the multimodal large model to identify and extract text information in the image file, and converting the image file into text data in an encodable format; Step 3, receiving the complex text in the encodable format in step 1 and the text data in the encodable format converted from the image file in step 2, performing semantic analysis and error correction using a pre-trained large language model to obtain output text, and arranging the format and layout of the output text; Step 4: extracting entity information from the output text; Step 5: Storing the entity information in a database to achieve structured storage.

7. The complex text parsing and processing method based on a large language model as claimed in claim 6, characterized in that: The step 1 includes the following sub-steps: Step 1.1, receiving the complex text, which is a file from different channels, including a scanned copy of an electronic document or a paper document; Step 1.2: Perform format detection to identify the format of the file. For the complex text in a non-encoded format, convert it into the picture file through image processing technology.

8. The complex text parsing and processing method based on a large language model as claimed in claim 6, characterized in that: The multimodal large model used in step 2 is the Qwen-VL-Chat multimodal large model, and the Qwen-VL-Chat multimodal large model identifies and extracts the text information in the image file by analyzing the text layout and semantic content in the image file.

9. The complex text parsing and processing method based on a large language model as claimed in claim 6, characterized in that: The pre-trained large language model used in step 3 is BERT or GPT.

10. The complex text parsing and processing method based on a large language model as claimed in claim 6, characterized in that: The step 4 uses NER technology to extract the entity information.

Citation Information

Cited By

  • Multi-format text analysis method and system based on large language model

    CN122113896A