Document image information extraction method and device based on large language model
By reconstructing the layout information of the document image and generating prompt words, combined with the large language model, the problems of weak generalization ability and limited understanding ability of document information extraction are solved, and efficient and accurate document image information extraction is achieved.
Patent Information
- Application Number
- CN202510597634.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the document information extraction method based on OCR template files has weak generalization ability, high development and maintenance costs, and the large language model has limited understanding of digital image documents, resulting in low accuracy of information extraction.
By obtaining text information and coordinate information of the document image, reconstructing the layout information of the document image, generating layout-aware documents, combining large language models to generate prompt words, and determining the coordinate position of the answer based on text line index information to generate final result information.
It realizes high-accuracy information extraction for various layout documents images, reduces development and maintenance costs, improves development efficiency, and is suitable for image information extraction in unseen or complex scenarios.
Smart Images

Figure CN120472475A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a document image information extraction method and device based on a large language model. Background Art
[0002] In today's digital age, the recognition and extraction of document image information is becoming increasingly important. Document image information extraction is the process of extracting target fields from digitized document images. In order to accurately locate the target field (answer), it is usually necessary to set the corresponding keyword (question). Taking the second-generation ID card as an example, the specific value of the ID card number is the target field, and the corresponding keyword is "citizen identity number." With the help of keywords, the specific location of the ID card number can be inferred, and then a piece of text at the target location can be used as the target field value. This set of keywords and target field feature description rule sets constitute the OCR template file, and the inference engine extracts information from the image by parsing the template file.
[0003] In actual scenarios, due to the diversity and complexity of document image formats, document information extraction methods based on OCR template files rely heavily on manually defined rules and have weak generalization capabilities. In engineering, you may encounter endless "abnormal" samples, and the workload of development and maintenance is extremely large. The large language model has extremely strong feature extraction and semantic understanding capabilities, and its zero-sample reasoning ability performs well in natural language understanding and generation tasks. However, the large language model has limited understanding of digital image documents, and the accuracy of problem information extraction is not high, which cannot meet actual business needs. Summary of the Invention
[0004] The present invention provides a method and device for extracting document image information based on a large language model, which is used to overcome the shortcomings of the document information extraction method based on OCR template files in the existing technology, such as weak generalization ability, high development and maintenance costs, and limited understanding ability of the large language model for digitized image documents, and realize accurate, efficient and highly generalized recognition and extraction of document image information.
[0005] The present invention provides a document image information extraction method based on a large language model, comprising:
[0006] Acquire a document image to be identified, perform detection and identification on the document image, and obtain text information in the document image and text coordinate information corresponding to the text information;
[0007] reconstructing the layout information of the document image using preset symbols based on the text information and the text coordinate information to generate a layout-aware document;
[0008] Obtaining question information, and generating prompt words based on the layout-aware document and the question information according to a preset question generation rule;
[0009] Inputting the prompt word into a large language model to obtain preliminary result information output by the large language model, wherein the preliminary result information includes a question, answer text information corresponding to the question, and text line index information corresponding to the answer text information;
[0010] The coordinate position information of the answer text information is determined according to the text line index information, and the question, the answer text information corresponding to the question and the coordinate position information are combined to generate final result information according to a preset answer generation rule.
[0011] According to a document image information extraction method based on a large language model provided by the present invention, based on the text information and the text coordinate information, the layout information of the document image is reconstructed using preset symbols to generate a layout-aware document, including:
[0012] Clustering all text lines in the text information into rows;
[0013] Calculating the average character width and average character height of the text line;
[0014] For the text lines in the same row, a corresponding number of first preset symbols are filled between adjacent text lines according to the horizontal distance between adjacent text lines and the average character width; for the text lines in different rows, the vertical distance between the text lines is calculated, and a corresponding number of second preset symbols are filled between the text lines in the vertical direction according to the vertical distance and the average character height;
[0015] According to the arrangement order of the text lines, the text line index information is allocated to the text coordinate information corresponding to each text line, and the text line index information is inserted at the beginning of the corresponding text line.
[0016] According to a document image information extraction method based on a large language model provided by the present invention, all text lines in the text information are clustered by line, including:
[0017] Taking the horizontal coordinate of the left edge of the text line as the basis, all the text lines are arranged from left to right on the horizontal coordinate axis in ascending order of the horizontal coordinate;
[0018] Calculating the vertical overlap of each of the text lines, and clustering the text lines with an overlap greater than a preset threshold into the same line;
[0019] For each clustered text line, the vertical coordinate of the upper boundary of the text line is used as the basis, and all the text lines are arranged from top to bottom on the vertical coordinate axis in the order of the vertical coordinate from small to large.
[0020] According to a document image information extraction method based on a large language model provided by the present invention, for text lines in the same row, a corresponding number of first preset symbols are filled between adjacent text lines according to the horizontal distance between adjacent text lines and the average character width, including:
[0021] For the text lines in the same row, the horizontal distance is calculated based on the difference between the horizontal coordinates of two adjacent text lines, the ratio between the horizontal distance and the average character width is calculated, and a corresponding number of spaces are filled between the adjacent text lines based on the ratio.
[0022] According to a document image information extraction method based on a large language model provided by the present invention, for different text lines, a vertical distance between each of the text lines is calculated, and a corresponding number of second preset symbols are filled in the vertical direction between each of the text lines based on the vertical distance and the average character height, including:
[0023] For the text lines of different rows, the vertical distance is calculated based on the difference in the vertical coordinates between the upper and lower text lines, the ratio between the vertical distance and the average character height is calculated, and a corresponding number of line breaks are filled between the upper and lower text lines based on the ratio.
[0024] According to a document image information extraction method based on a large language model provided by the present invention, the prompt word is represented as follows:
[0025] "The text information is: \n" + the layout-aware document + "\nPlease follow the format according to the text information: \n" + question information 1, question information 2, ..., question information N + "\nAnswer: \n"; where \n is a line break character.
[0026] According to a document image information extraction method based on a large language model provided by the present invention, the prompt word also includes example information;
[0027] In the case where the prompt word includes example information, the prompt word is represented as follows:
[0028] "The text example is: \n" + the perceived layout document of the example document image + "Please follow the text example in the format: \n" + example question information 1, example question information 2, ..., example question information N + "Answer: \n" + example answer text information 1, example answer text information 2, ..., example answer text information N + "\nThe above is a document question and answer example, please refer to the above content and conduct document question and answer for the following text\n" + "The text information is: \n" + the layout perceived document of the document image to be identified + "\nPlease follow the text information in the format: \n" + question information 1, question information 2, ..., question information N + "\nAnswer: \n".
[0029] The present invention also provides a document image information extraction device based on a large language model, comprising:
[0030] A document image recognition module is used to obtain a document image to be recognized, detect and recognize the document image, and obtain text information in the document image and text coordinate information corresponding to the text information;
[0031] a layout-aware document generation module, configured to reconstruct the layout information of the document image using preset symbols based on the text information and the text coordinate information, thereby generating a layout-aware document;
[0032] A prompt word generation module is used to obtain question information and generate prompt words based on the layout-aware document and the question information according to preset question generation rules;
[0033] an information extraction module, configured to input the prompt word into a large language model and obtain preliminary result information output by the large language model, wherein the preliminary result information includes a question, text information of an answer corresponding to the question, and text line index information corresponding to the answer text information;
[0034] A post-processing module is used to determine the coordinate position information of the answer text information based on the text line index information, and generate final result information according to the preset answer generation rules by combining the question, the answer text information corresponding to the question and the coordinate position information.
[0035] The present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for extracting document image information based on a large language model as described above is implemented.
[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for extracting document image information based on a large language model.
[0037] The present invention provides a document image information extraction method and device based on a large language model. The method and device obtain a document image to be identified, detect and identify the document image, and obtain text information in the document image and text coordinate information corresponding to the text information. Based on the text information and text coordinate information, the method reconstructs the layout information of the document image using preset symbols to generate a layout-aware document. The method obtains question information and generates prompt words based on the layout-aware document and the question information according to preset question generation rules. The prompt words are input into the large language model to obtain preliminary result information output by the large language model. The preliminary result information includes the question, the answer text information corresponding to the question, and the text line index information corresponding to the answer text information. The coordinate position information of the answer text information is determined based on the text line index information, and the question, the answer text information corresponding to the question, and the coordinate position information are converted into final result information according to the preset answer generation rules. Through the application of this method and device, the image analysis and processing capabilities of the OCR recognition engine and the semantic understanding capabilities of the large language model are combined. The method and device can recognize and analyze document images of various formats, have strong generalization capabilities, and high information extraction accuracy. For simple scenarios, users do not need any configuration and can use it out of the box. For images of unprecedented or complex scenes, there is no need to train a large language model to meet user needs, which improves development efficiency and reduces development and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 It is a flowchart of the document image information extraction method based on the large language model provided by the present invention;
[0040] Figure 2 is a schematic diagram of layout-aware document generation in an embodiment of the present invention;
[0041] Figure 3 Schematic diagram of prompt word generation and information extraction in an embodiment of the present invention;
[0042] Figure 4 It is a structural diagram of a document image information extraction device based on a large language model provided by the present invention;
[0043] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0044] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0045] The following combination Figure 1-Figure 3 The document image information extraction method based on a large language model of the present invention is described.
[0046] like Figure 1 As shown, the document image information extraction method based on the large language model provided by the present invention includes the following steps:
[0047] S1. Obtain a document image to be identified, detect and identify the document image, and obtain text information in the document image and text coordinate information corresponding to the text information.
[0048] The document image to be identified mentioned in the present invention is a digitized document image, such as an ID card image, a driver's license image, a financial bill image, etc. Specifically, the present invention uses an OCR recognition module based on an OCR (Optical Character Recognition) recognition engine to extract information from document images. The OCR engine is a technical software development kit (SDK) that analyzes and processes image files and automatically identifies and obtains text information and layout information. The OCR recognition module can directly extract structured document information, including text information and corresponding coordinate information in the document image, wherein the text information is presented in the form of text lines, and the coordinate information is represented by the coordinates of the minimum circumscribed rectangle that can contain the content of the text line. Figure 2 For example, taking the upper left corner of the ID card image as the origin, downward as the positive direction of the y-axis, and rightward as the positive direction of the x-axis, the coordinate information of the text line "Li Si" in the ID card image is expressed as [327,195,385,231], where [327,195] represents the coordinate of the upper left corner of the minimum enclosing rectangle corresponding to the text line "Li Si", and [385,231] represents the coordinate of the lower right corner of the minimum enclosing rectangle corresponding to the text line "Li Si".
[0049] S2. Based on the text information and text coordinate information, use preset symbols to reconstruct the layout information of the document image to generate a layout-aware document.
[0050] Specifically, the steps to generate a layout-aware document are as follows:
[0051] S21. Cluster all text lines in the text information by line.
[0052] Specifically, the clustering steps are as follows:
[0053] S211. Based on the horizontal coordinate of the left edge of the text line, all text lines are arranged from left to right on the horizontal coordinate axis in ascending order of the horizontal coordinate.
[0054] by Figure 2 For example, Figure 2 The coordinate information of the text line "Li Si" in the ID card image shown is [327,195,385,231], the unit is pixel (px), that is, the horizontal coordinate of the left boundary is 327; the coordinate information of the text line "21-06-2018" is [220,262,309,281], that is, the horizontal coordinate of the left boundary is 220, then the text line "21-06-2018" is arranged to the left of the text line "Li Si" on the horizontal axis.
[0055] S212: Calculate the vertical overlap of each text line, and cluster the text lines with an overlap greater than a preset threshold into the same line.
[0056] Specifically, the vertical overlap of each text line is calculated to determine whether the text lines belong to the same line, and the intersection-over-union (IOU) or intersection-over-aspect (IOA) of the text line coordinates in the y-axis direction is calculated to determine whether the text lines belong to the same line. Figure 2 For example, Figure 2 The coordinate information of the text line "Li Si" in the ID card image shown is [327,195,385,231], and the calculated line height is 231-195=36px; the coordinate information of the text line "11-03-1969" is [217,218,309,232], and the calculated line height is 232-218=14px; the overlapping height of the two text lines in the y direction is 231-218=13px; then the IOU of the two text lines is 13 / (36+14-13)=0.35. If the preset threshold is 0.3, the two text lines are judged to be the same line.
[0057] S213 , for each clustered text line, taking the vertical coordinate of the upper boundary of the text line as the standard, and arranging all text lines from top to bottom on the vertical coordinate axis in the order of the vertical coordinate from small to large.
[0058] by Figure 2 For example, Figure 2The coordinate information of the text line "Li Si" in the shown ID card image is [327, 195, 385, 231], that is, the ordinate of the upper boundary is 195; the coordinate information of the text line "21-06-2018" is [220, 262, 309, 281], that is, the ordinate of the upper boundary is 262. Then, arrange the text line "21-06-2018" below the text line "Li Si".
[0059] S22. Calculate the average character width and average character height of the text line.
[0060] Specifically, the OCR engine can detect each character one by one and return its font height. Add up the character heights of all characters and divide by the number of characters to obtain the average character height. Calculate the width of the text line based on the difference between the left boundary and the right boundary in the coordinate information of the text line. By calculating the widths of each text line, add up all the width values and divide by the number of all characters to obtain the average character width.
[0061] S23. For text lines in the same row, fill in the corresponding number of first preset symbols between adjacent text lines according to the horizontal distance and average character width between adjacent text lines; for text lines in different rows, calculate the vertical distance between each text line, and fill in the corresponding number of second preset symbols in the vertical direction between each text line according to the vertical distance and average character height.
[0062] Among them, for text lines in the same row, fill in the corresponding number of first preset symbols between adjacent text lines according to the horizontal distance and average character width between adjacent text lines, as follows:
[0063] For text lines in the same row, calculate the horizontal distance based on the difference in abscissa between two adjacent text lines, calculate the ratio between the horizontal distance and the average character width, and fill in the corresponding number of spaces between adjacent text lines according to the ratio.
[0064] Take Figure 2 as an example. Figure 2 The coordinate information of the text line "Li Si" in the shown ID card image is [327, 195, 385, 231]; the coordinate information of the text line "11-03-1969" is [217, 218, 309, 232]. In step S212, these two text lines have been determined to be in the same row. The calculated horizontal distance between the two text lines is 385 - 309 = 76px; assume Figure 2 the average character width of the characters is 19px, 76 / 19 = 4, then fill in 4 spaces between the text line "Li Si" and the text line "11-03-1969".
[0065] For text lines in different rows, calculate the vertical distance between each pair of text lines. According to the vertical distance and the average character height, fill in the corresponding number of second preset symbols between each pair of text lines in the vertical direction, specifically as follows:
[0066] For text lines in different rows, calculate the vertical distance based on the difference in the vertical coordinates between the upper and lower text lines. Calculate the ratio between the vertical distance and the average character height, and fill in the corresponding number of line break symbols between the upper and lower text lines according to the ratio.
[0067] Take Figure 2 as an example, Figure 2 In the ID card image shown, the coordinate information of the text line "Li, Si" is [327, 195, 385, 231]; the coordinate information of the text line "1, 76" is [217, 232, 255, 251]. According to the method in step S213, it can be known that the text line "1, 76" is arranged below the text line "Li, Si". The vertical coordinate of the lower boundary of the text line "Li, Si" is 231, and the vertical coordinate of the lower boundary of the text line "1, 76" is 251. The vertical distance between the two text lines is 251 - 231 = 20px. Assume Figure 2 the average character height of Chinese characters is 19px, 20 / 19 = 1.05, which is rounded to 1 after rounding. Then, fill in a line break symbol between the text line "Li, Si" and the text line "1, 76".
[0068] S24. According to the arrangement order of the text lines, assign text line index information to the text coordinate information corresponding to each text line, and insert the text line index information at the beginning of the corresponding text line.
[0069] In an optional embodiment of the present invention, the text line index information is represented as a text line index number, and the text line index number is inserted at the beginning of the text line content, with the format <idx + index number> + text line content. As Figure 2 shown, according to the arrangement order of the text lines, insert index numbers for the text lines in the order from left to right and from top to bottom. After inserting the index number into the text line "Li, Si", it becomes " <idx0>Li, Si", the text line "248665824785" becomes " <idx1>248665824785"...The text line "3274568(4)" becomes " <idx12>3274568(4)”.
[0070] S3. Obtain question information, and generate prompt words based on the layout-aware document and the question information according to preset question generation rules.
[0071] In the AI big model, the main function of the prompt is to prompt the AI model with the context of the input information and the parameter information of the input model. In an optional embodiment of the present invention, the prompt is expressed as follows:
[0072] "The text information is: \n" + layout-aware document + "\nPlease follow the format according to the text information: \n" + question information 1, question information 2, ..., question information N + "\nAnswer: \n"; where \n is a line break character.
[0073] For example, Figure 2 For example, if the question is "What is your name?", an optional prompt word is: "The text information is:
[0074] <idx0>Li, Si
[0075] <idx1> 248665824785
[0076] <idx2>CHEN,TING
[0077] <idx4> 11-03-1969 <idx3>Li Si <idx5>CSF
[0078] <idx6> 1,76
[0079] <idx7> 2145882
[0080] <idx8> 21-06-2018 <idx9> 19-07-1985 <idx10> 11-03-19693F
[0081] <idx11> 21-06-2008 <idx12>3274568
[0082] Please follow the format according to the text information:
[0083] What is your name?
[0084] answer:
[0085] ”
[0086] In an optional embodiment of the present invention, in order to assist the large language model to more easily understand the user's questions and the answers the user wants, and to further improve the accuracy of document information extraction, the prompt word also includes example information. The example information includes the layout-aware document information of the example image, the example question information, and the example answer information, that is, it includes the input and output information of a complete large language model. Among them, the example image is a digitized document image, which is the same type of image as the image to be tested. For example, the image to be tested is an ID card photo image, and the image to be tested is also an ID card photo image. In the case where the prompt word includes example information, the prompt word is expressed as follows:
[0087] "The text example is: \n" + the perceived layout document of the example document image + "Please follow the format according to the text example: \n" + example question information 1, example question information 2, ..., example question information N + "Answer: \n" + example answer text information 1, example answer text information 2, ..., example answer text information N + "\nThe above is a document question and answer example. Please refer to the above content to conduct document question and answer for the following text\n" + "The text information is: \n" + the layout perceived document of the document image to be recognized + "\nPlease follow the format according to the text information: \n" + question information 1, question information 2, ..., question information N + "\nAnswer: \n".
[0088] S4. Input the prompt word into the large language model to obtain preliminary result information output by the large language model.
[0089] The preliminary result information includes the question, the answer text information corresponding to the question, and the text line index information corresponding to the answer text information. In an optional embodiment of the present invention, the preliminary result information is a string in json format. Figure 3 For example, the preliminary result information is represented as:
[0090] {"Name":[{"text":"Li Si","idx":[3]}]}.
[0091] S5. Determine the coordinate position information of the answer text information according to the line index information of the text, and generate final result information according to the preset answer generation rules by combining the question, the answer text information corresponding to the question, and the coordinate position information.
[0092] In step S1, the coordinate position information of each text line has been determined. Therefore, the coordinate position information can be directly determined based on the text line corresponding to the index information. Taking the above preliminary result information as an example, the text line index information (idx) is 3, and the corresponding text line is "Li Si", and its position coordinates are [327, 195, 385, 231]. Therefore, the final result information is expressed as:
[0093] {"Name": [{"text":"Li Si","bbox":[327,195,385,231}]}, where bbox represents the coordinate position.
[0094] In summary, the present invention provides a document image information extraction method based on a large language model. The method comprises: obtaining a document image to be identified, detecting and identifying the document image, obtaining text information in the document image and text coordinate information corresponding to the text information; reconstructing the layout information of the document image using preset symbols based on the text information and text coordinate information, and generating a layout-aware document; obtaining question information, and generating prompt words based on the layout-aware document and the question information according to preset question generation rules; inputting the prompt words into the large language model, and obtaining preliminary result information output by the large language model. The preliminary result information includes the question, the answer text information corresponding to the question, and the text line index information corresponding to the answer text information; determining the coordinate position information of the answer text information based on the text line index information, and generating final result information according to the preset answer generation rules based on the question, the answer text information corresponding to the question, and the coordinate position information. This method combines the respective advantages of an OCR recognition engine and a large language model, and can directly extract structured document information. In combination with text line indexing technology, it has the ability to extract text coordinate information. For simple scenarios, users can use it out of the box without any configuration. For complex scenarios, users only need to annotate a sample image to significantly improve the accuracy of document information extraction. For images that have never been seen or are complex scenes, no training is required to meet user needs, thereby improving development efficiency and reducing development and maintenance costs.
[0095] Based on the same inventive concept, the present invention also provides a document image information extraction device based on a large language model. The document image information extraction device based on a large language model provided by the present invention is described below. The document image information extraction device based on a large language model described below and the document image information extraction method based on a large language model described above can be referenced to each other.
[0096] like Figure 4 As shown, the document image information extraction device based on the large language model provided by the present invention includes a document image recognition module 41, a layout-aware document generation module 42, a prompt word generation module 43, an information extraction module 44 and a post-processing module 45.
[0097] The document image recognition module 41 is used to obtain a document image to be recognized, detect and recognize the document image, and obtain text information in the document image and text coordinate information corresponding to the text information.
[0098] The layout-aware document generation module 42 is configured to reconstruct the layout information of the document image using preset symbols based on the text information and the text coordinate information, and generate a layout-aware document.
[0099] The prompt word generation module 43 is used to obtain question information and generate prompt words based on the layout-aware document and the question information according to preset question generation rules.
[0100] The information extraction module 44 is used to input the prompt word into the large language model to obtain preliminary result information output by the large language model, wherein the preliminary result information includes the question, the answer text information corresponding to the question, and the text line index information corresponding to the answer text information.
[0101] The post-processing module 45 is used to determine the coordinate position information of the answer text information based on the text line index information, and generate final result information based on the question, the answer text information corresponding to the question and the coordinate position information according to the preset answer generation rules.
[0102] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the document image information extraction method based on the large language model provided by the above methods, which includes:
[0103] Acquire a document image to be identified, perform detection and identification on the document image, and obtain text information in the document image and text coordinate information corresponding to the text information;
[0104] reconstructing the layout information of the document image using preset symbols based on the text information and the text coordinate information to generate a layout-aware document;
[0105] Obtaining question information, and generating prompt words based on the layout-aware document and the question information according to a preset question generation rule;
[0106] Inputting the prompt word into a large language model to obtain preliminary result information output by the large language model, wherein the preliminary result information includes a question, answer text information corresponding to the question, and text line index information corresponding to the answer text information;
[0107] The coordinate position information of the answer text information is determined according to the text line index information, and the question, the answer text information corresponding to the question and the coordinate position information are combined to generate final result information according to a preset answer generation rule.
[0108] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0109] On the other hand, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is capable of performing the document image information extraction method based on a large language model provided by the above methods, the method comprising:
[0110] Acquire a document image to be identified, perform detection and identification on the document image, and obtain text information in the document image and text coordinate information corresponding to the text information;
[0111] reconstructing the layout information of the document image using preset symbols based on the text information and the text coordinate information to generate a layout-aware document;
[0112] Obtaining question information, and generating prompt words based on the layout-aware document and the question information according to a preset question generation rule;
[0113] Inputting the prompt word into a large language model to obtain preliminary result information output by the large language model, wherein the preliminary result information includes a question, answer text information corresponding to the question, and text line index information corresponding to the answer text information;
[0114] The coordinate position information of the answer text information is determined according to the text line index information, and the question, the answer text information corresponding to the question and the coordinate position information are combined to generate final result information according to a preset answer generation rule.
[0115] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the document image information extraction method based on a large language model provided by the above methods, the method comprising:
[0116] Acquire a document image to be identified, perform detection and identification on the document image, and obtain text information in the document image and text coordinate information corresponding to the text information;
[0117] reconstructing the layout information of the document image using preset symbols based on the text information and the text coordinate information to generate a layout-aware document;
[0118] Obtaining question information, and generating prompt words based on the layout-aware document and the question information according to a preset question generation rule;
[0119] Inputting the prompt word into a large language model to obtain preliminary result information output by the large language model, wherein the preliminary result information includes a question, answer text information corresponding to the question, and text line index information corresponding to the answer text information;
[0120] The coordinate position information of the answer text information is determined according to the text line index information, and the question, the answer text information corresponding to the question and the coordinate position information are combined to generate final result information according to a preset answer generation rule.
[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0122] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention. < / idx11> < / idx10> < / idx9> < / idx8> < / idx7> < / idx6> < / idx4> < / idx1>
Claims
1. A document image information extraction method based on a large language model, characterized in that: include: Acquire a document image to be identified, perform detection and identification on the document image, and obtain text information in the document image and text coordinate information corresponding to the text information; reconstructing the layout information of the document image using preset symbols based on the text information and the text coordinate information to generate a layout-aware document; Obtaining question information, and generating prompt words based on the layout-aware document and the question information according to a preset question generation rule; Inputting the prompt word into a large language model to obtain preliminary result information output by the large language model, wherein the preliminary result information includes a question, answer text information corresponding to the question, and text line index information corresponding to the answer text information; The coordinate position information of the answer text information is determined according to the text line index information, and the question, the answer text information corresponding to the question and the coordinate position information are combined to generate final result information according to a preset answer generation rule.
2. The document image information extraction method based on a large language model according to claim 1, characterized in that: Reconstructing the layout information of the document image using preset symbols based on the text information and the text coordinate information to generate a layout-aware document, including: Clustering all text lines in the text information into rows; Calculating the average character width and average character height of the text line; For the text lines in the same row, a corresponding number of first preset symbols are filled between adjacent text lines according to the horizontal distance between adjacent text lines and the average character width; for the text lines in different rows, the vertical distance between the text lines is calculated, and a corresponding number of second preset symbols are filled between the text lines in the vertical direction according to the vertical distance and the average character height; According to the arrangement order of the text lines, the text line index information is allocated to the text coordinate information corresponding to each text line, and the text line index information is inserted at the beginning of the corresponding text line.
3. The document image information extraction method based on a large language model according to claim 2, characterized in that: Clustering all text lines in the text information by line, including: Taking the horizontal coordinate of the left edge of the text line as the basis, all the text lines are arranged from left to right on the horizontal coordinate axis in ascending order of the horizontal coordinate; Calculating the vertical overlap of each of the text lines, and clustering the text lines with an overlap greater than a preset threshold into the same line; For each clustered text line, the vertical coordinate of the upper boundary of the text line is used as the basis, and all the text lines are arranged from top to bottom on the vertical coordinate axis in the order of the vertical coordinate from small to large.
4. The document image information extraction method based on a large language model according to claim 3 is characterized in that: For the text lines in the same row, filling a corresponding number of first preset symbols between adjacent text lines according to the horizontal distance between adjacent text lines and the average character width, including: For the text lines in the same row, the horizontal distance is calculated based on the difference between the horizontal coordinates of two adjacent text lines, the ratio between the horizontal distance and the average character width is calculated, and a corresponding number of spaces are filled between the adjacent text lines based on the ratio.
5. The document image information extraction method based on a large language model according to claim 3 is characterized in that: For different text lines, calculating the vertical distance between the text lines, and filling a corresponding number of second preset symbols between the text lines in the vertical direction according to the vertical distance and the average character height, including: For the text lines of different rows, the vertical distance is calculated based on the difference in the vertical coordinates between the upper and lower text lines, the ratio between the vertical distance and the average character height is calculated, and a corresponding number of line breaks are filled between the upper and lower text lines based on the ratio.
6. The document image information extraction method based on a large language model according to claim 1, characterized in that: The prompt words are as follows: "The text information is: \n" + the layout-aware document + "\nPlease follow the text information in the format: \n" + question information 1, question information 2, ..., question information N + "\nAnswer: \n"; where \n is a line break character.
7. The document image information extraction method based on a large language model according to claim 6, characterized in that: The prompt word also includes example information; In the case where the prompt word includes example information, the prompt word is represented as follows: "The text example is: \n" + the described perceived layout document of the example document image + "Please follow the format according to the text example: \n" + example question information 1, example question information 2, ..., example question information N + "Answer: \n" + example answer text information 1, example answer text information 2, ..., example answer text information N + "\nThe above is a document question and answer example. Please refer to the above content to conduct document question and answer for the following text\n" + "The text information is: \n" + the described layout perceived document of the document image to be identified + "\nPlease follow the format according to the text information: \n" + question information 1, question information 2, ..., question information N + "\nAnswer: \n".
8. A document image information extraction device based on a large language model, characterized in that: include: A document image recognition module is used to obtain a document image to be recognized, detect and recognize the document image, and obtain text information in the document image and text coordinate information corresponding to the text information; a layout-aware document generation module, configured to reconstruct the layout information of the document image using preset symbols based on the text information and the text coordinate information, thereby generating a layout-aware document; A prompt word generation module is used to obtain question information and generate prompt words based on the layout-aware document and the question information according to preset question generation rules; an information extraction module, configured to input the prompt word into a large language model and obtain preliminary result information output by the large language model, wherein the preliminary result information includes a question, text information of an answer corresponding to the question, and text line index information corresponding to the answer text information; A post-processing module is used to determine the coordinate position information of the answer text information based on the text line index information, and generate final result information according to the preset answer generation rules by combining the question, the answer text information corresponding to the question and the coordinate position information.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the document image information extraction method based on a large language model as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the document image information extraction method based on a large language model as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Zero sample template inference and document structured recognition method and device
CN120877300A