Method, system, medium and computer device for determining coordinates of characters in a document
By parsing the source document to obtain text coordinates and matching them with the processing results of the large model, the problem that the large model cannot provide coordinate information is solved, and accurate labeling of document review or analysis results on the front end is achieved, which improves the convenience and visualization of document processing.
Patent Information
- Application Number
- CN202510787109.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Large models cannot provide the specific coordinate information of the text in the source document during document review or analysis, resulting in the inability to accurately mark and display the review or analysis results on the front end, affecting the visualization and ease of use of the document review and analysis results.
By parsing the source document, the coordinates of the text are obtained and accurately matched with the processing results of the large model. The sliding window algorithm and the matching algorithm of the maximum number of consecutive identical characters are used to determine the corresponding coordinates of the processing results of the large model in the source document.
It achieves precise matching of large model output content with source document coordinates, improves the feasibility and intuitiveness of front-end annotation display of document review or analysis results, significantly improves annotation efficiency and visualization, and enhances user convenience and work efficiency.
Smart Images

Figure CN120316276B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method, system, medium and computer equipment for determining the coordinates of characters in a document. Background Art
[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.
[0003] With the rapid development of artificial intelligence technology, large models, with their powerful language understanding and analysis capabilities, are widely used in areas such as document review and content analysis, improving document processing efficiency for both businesses and users. Numerous researchers are also actively exploring the application of large models in document processing, integrating them into office software, document management systems, and other products to enable intelligent review of document content and risk identification.
[0004] However, existing technologies that use large models to review or analyze documents generally face a key problem: large models can only output the specific text content after analysis, but cannot provide the specific coordinate information of these texts in the source document. The root cause of this flaw is that the design logic of large models focuses primarily on semantic understanding and content generation, and their output mechanism does not include the positioning data of the text in the document's physical layout. Due to the lack of coordinate information, the review or analysis results cannot be accurately annotated and displayed on the front end, making it difficult for users to intuitively associate the analysis conclusions with the source document, seriously affecting the visualization and ease of use of the document review and analysis results.
[0005] Although some technologies attempt to use optical character recognition (OCR) technology for coordinate positioning, this method has obvious shortcomings. The recognition process requires multiple complex steps such as image preprocessing, character segmentation, feature extraction and recognition, and the amount of calculation is huge, resulting in slow coordinate positioning speed and unable to meet the needs of actual application scenarios for real-time annotation and display. In addition, simple text retrieval or manual annotation methods also have problems of low efficiency and difficulty in accurate matching, and neither can meet the needs of actual application scenarios for efficient and accurate annotation and display. Summary of the Invention
[0006] In order to address the deficiencies of the prior art, the present invention provides a method, system, medium and computer device for determining the coordinates of text in a document. By parsing the source document, the coordinates of the text in the source document are obtained, and they are accurately matched with the processing results of a large model, thereby obtaining the corresponding coordinates of the processing results of the large model in the source document, solving the technical problem that the processing results of the large model cannot be annotated and displayed on the front end, so that the source document review or analysis results can be intuitively annotated on the front end.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a method for determining coordinates of text in a document.
[0009] A method for determining the coordinates of text in a document, comprising the following steps:
[0010] Extract the coordinates of each character in the source document and merge the coordinates with the same vertical coordinates on the same page;
[0011] Obtaining a large model processing result of the source document, and sliding a sliding window simultaneously between the large model processing result and the source document, wherein the large model processing result is a text sequence;
[0012] Calculate the number of consecutive identical characters between two text segments within the sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates;
[0013] When the obtained matching coordinates are continuous, the continuous matching coordinates are merged to finally determine the corresponding coordinates of the large model processing result in the source document.
[0014] In an implementation of the first aspect of the present invention, extracting the coordinates of each character in the source document includes:
[0015] The characters in the source document are extracted one by one, and the coordinates of each character in the page and the page number are recorded. The coordinates of each character in the page include the upper left corner horizontal coordinate, upper left corner vertical coordinate, lower right corner horizontal coordinate and lower right corner vertical coordinate of the character box.
[0016] In an implementation of the first aspect of the present invention, merging coordinates with consistent vertical coordinates on the same page includes:
[0017] All coordinates on the same page are traversed through a double loop. The outer loop selects a coordinate as a reference, and the inner loop compares the remaining coordinate data with it. When the vertical coordinate difference is within the first set error range, it is determined to be the same row;
[0018] For the coordinates of the same row, all the coordinates of the row are merged into a new coordinate according to the maximum value of the horizontal coordinate of the lower right corner of the text box and the maximum value of the vertical coordinate of the lower right corner of the text box.
[0019] In an implementation of the first aspect of the present invention, using a sliding window to slide simultaneously between the large model processing result and the source document includes:
[0020] A sliding window of fixed length is set to slide simultaneously on the large model processing result and the text sequence corresponding to the source document.
[0021] In an implementation of the first aspect of the present invention, the maximum number of consecutive identical characters and the corresponding matching coordinates are recorded. When the maximum number of consecutive identical characters corresponding to multiple positions in the source document is the same, the matching coordinates corresponding to each position are recorded.
[0022] In an implementation of the first aspect of the present invention, determining whether the obtained matching coordinates are continuous includes:
[0023] On the same page, after the coordinates of the same line of text are merged, a coordinate list is obtained. Each element of the coordinate list corresponds to the coordinate data of a line of text and has a unique index.
[0024] When the indexes of the obtained multiple matching coordinates are continuous, the page numbers are the same, and the horizontal coordinate errors are within a second set threshold range, the multiple matching coordinates are determined to be continuous matching coordinates.
[0025] As a further qualification, merge consecutive matching coordinates, including:
[0026] From the multiple continuous matching coordinates, take the horizontal coordinate of the smallest upper left corner of the text box as the starting horizontal coordinate of the merged coordinate, the horizontal coordinate of the largest lower right corner of the text box as the ending horizontal coordinate, the vertical coordinate of the smallest upper left corner of the text box as the starting vertical coordinate, and the vertical coordinate of the largest lower right corner of the text box as the ending vertical coordinate, retain the common page number value, and obtain the merged matching coordinates.
[0027] In a second aspect, the present invention provides a system for determining coordinates of text in a document.
[0028] A system for determining coordinates of text in a document, comprising:
[0029] The coordinate extraction unit is configured to: extract the coordinates of each character in the source document and merge the coordinates with the same vertical coordinates on the same page;
[0030] a sliding matching unit configured to: obtain a large model processing result of the source document, and slide a sliding window simultaneously between the large model processing result and the source document, wherein the large model processing result is a text sequence;
[0031] a coordinate matching unit configured to: calculate the number of consecutive identical characters between two text segments within a sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates;
[0032] The coordinate merging unit is configured to: when the obtained matching coordinates are continuous, merge the continuous matching coordinates, and finally determine the corresponding coordinates of the large model processing result in the source document.
[0033] In a third aspect, the present invention provides a computer device comprising: a processor and a computer-readable storage medium;
[0034] a processor adapted to execute a computer program;
[0035] A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the method for determining the coordinates of characters in a document as described in the first aspect of the present invention is implemented.
[0036] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program is suitable for being loaded by a processor and executing the method for determining the coordinates of text in a document as described in the first aspect of the present invention.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. The present invention innovatively proposes a method for determining the coordinates of characters in a document. By parsing the source document, the coordinates of the characters in the source document are obtained, and the coordinates are accurately matched with the processing results of the large model, thereby obtaining the corresponding coordinates of the processing results of the large model in the source document, achieving accurate matching of the output content of the large model with the coordinates of the source document, and solving the technical problem that the processing results of the large model cannot be annotated and displayed on the front end. The source document review or analysis results can be intuitively annotated on the front end, which greatly improves the feasibility and intuitiveness of the front-end annotation display of the document review or analysis results, provides users with a more convenient and efficient document review and analysis interactive experience, and significantly improves the visualization and practicality of the document review and analysis results.
[0039] 2. Compared with traditional OCR coordinate positioning technology, this invention avoids time-consuming steps such as complex image preprocessing and character segmentation, supports direct acquisition of text coordinates, and combines with efficient matching algorithms to greatly shorten the coordinate positioning time, thereby increasing the annotation efficiency by several times or even higher; at the same time, compared with manual annotation and simple text retrieval, the automated processing flow greatly reduces manual intervention and significantly improves the efficiency of document review and analysis result annotation.
[0040] 3. Based on the matching algorithm of the maximum number of consecutive identical characters and the coordinate continuity judgment mechanism, the present invention can accurately identify the position of the output content of the large model in the document. Even if there are repeated expressions or complex layouts in the document, it can also accurately locate them, providing a reliable data foundation for front-end annotation.
[0041] 4. The present invention accurately matches the large model processing results with the source document coordinates and optimizes the annotations. Users can intuitively see the specific location of the audit analysis results in the source document at the front end without manual search or comparison, which greatly improves the user's visualization experience of the document audit analysis results, makes document processing more convenient and efficient, and helps to improve user work efficiency and satisfaction.
[0042] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0044] Figure 1 A schematic diagram of a method for determining the coordinates of text in a document provided by an exemplary embodiment of the present invention;
[0045] Figure 2 A flowchart of a method for determining the coordinates of text in a document provided by an exemplary embodiment of the present invention;
[0046] Figure 3 A schematic diagram of a content matching method based on the maximum number of consecutive identical characters provided by an exemplary embodiment of the present invention;
[0047] Figure 4 A schematic diagram of a multi-row coordinate continuity determination and merging optimization process provided by an exemplary embodiment of the present invention;
[0048] Figure 5 A schematic diagram of a system for determining coordinates of text in a document provided by an exemplary embodiment of the present invention;
[0049] Figure 6 A schematic diagram of a computer device provided for an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0050] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0051] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0052] As mentioned in the background technology, when using a large model to review or analyze the specific content of a source document, the large model can only output specific text, but cannot provide the specific coordinates of the text in the source document, resulting in the inability to annotate and display the review or analysis results on the front end. In view of this, this implementation proposes a method for determining the coordinates of text in a document, which fundamentally makes up for the defect of the large model outputting content that lacks coordinate information. Specifically, Figure 1 As shown in the figure, taking a source document in PDF format as an example, the pdfplumber library is used to parse the document to obtain coordinate information, establishing a physical location connection between the large model output content and the document; coordinates with consistent vertical coordinates on the same page are merged. The coordinate merging technology simplifies the coordinate data structure, making it more consistent with document typesetting rules and subsequent matching requirements; the large model detection results and the coordinates obtained by parsing are obtained. Based on the matching algorithm of the maximum number of consecutive identical characters, the text similarity measurement is used to accurately find the corresponding position of the large model output content in the document, calculate the maximum number of identical characters, and output the corresponding coordinates (if multiple exist, all are output); multi-line coordinate continuity judgment and merging optimization are carried out to determine whether multiple matching coordinates of the source document are continuous, merge the continuous matching coordinates, further improve the coordinate information, and ensure the accuracy and completeness of the front-end annotation. These technical solutions are closely linked and closely focus on the core issue of annotating the coordinates of the large model output content, forming a complete solution. This effectively solves the problem of the existing technology that cannot display the large model audit and analysis results on the front-end annotation.
[0053] More specifically, Figure 2 As shown, the following process is included:
[0054] S101: extracting the coordinates of each character in the source document, and merging the coordinates with the same vertical coordinates on the same page;
[0055] S102: Obtaining a large model processing result of the source document, and sliding a sliding window simultaneously between the large model processing result and the source document, wherein the large model processing result is a text sequence;
[0056] S103: Calculate the number of consecutive identical characters between two text segments within the sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates;
[0057] S104: When the obtained matching coordinates are continuous, the continuous matching coordinates are merged to finally determine the corresponding coordinates of the large model processing result in the source document.
[0058] In step S101 of the present invention, the source document is taken as PDF, and the coordinates of the source document are parsed based on the pdfplumber library.
[0059] The pdfplumber library is a powerful Python-based PDF parsing tool that can read the geometric information of text in PDF documents. Its core principle is to parse the page content stream of a PDF file and use the description of text objects in the PDF specification (such as character encoding, font attributes, and position coordinates) to extract each character from the document one by one. It then records the precise coordinates of each character on the page, including the horizontal (X-axis), vertical (Y-axis), and the page number (for example, with the lower left corner of the page as the origin). In the actual implementation process, the pdfplumber.open() function is used to open the PDF document. By traversing each page of the document, the extract_words() method of the page is called to obtain the text and its coordinate data on each page, and finally a list containing all the text coordinate information is formed, for example: [{'text':'Specific text', 'x0': 100, 'y0': 200,'x1':150, 'y1': 220, 'page': 1},···], where x0 and y0 represent the coordinates of the upper left corner of the text box, and x1 and y1 represent the coordinates of the lower right corner.
[0060] In this implementation, coordinates with the same vertical coordinate on the same page are merged according to the largest coordinate. Since there may be text fragments of different lengths in the horizontal direction in the same line of text, merging with the largest coordinate can completely cover the display area of all text in the same line, and more accurately fit the document layout. During implementation, a double-layer loop is used to traverse the coordinate list of the same page. The outer loop selects a coordinate data as a benchmark, and the inner loop compares the remaining coordinate data with it. When the vertical coordinate difference is within the tolerable error range (taking into account the slight errors that may exist in document rendering and parsing, the error range can be set to ±2 pixels), it is considered to be the same line. For the coordinates of the same line, the maximum value of the horizontal coordinate x1 (the horizontal coordinate of the lower right corner of the text box) and y1 (the vertical coordinate of the lower right corner of the text box) is selected, and all the coordinates of the line are merged into a new coordinate data. For example, the original coordinate list for the same line is [{'x0': 100, 'y0': 200, 'x1': 150, 'y1': 220},{'x0': 160, 'y0': 200, 'x1': 200, 'y1': 220}, {'x0': 210, 'y0': 200, 'x1':250, 'y1': 220}], and after merging, the result is {'x0': 100, 'y0': 200, 'x1': 250, 'y1': 220}. This processing method not only reduces the subsequent matching calculation amount but also more accurately reflects the actual display range of the same line of text in the document, providing a more precise coordinate basis for subsequent matching with the output content of the large model.
[0061] In step S102 and step S103 of this implementation, Figure 3 As shown, the matching relationship between the large model output and the text corresponding to the document coordinates is determined by calculating the maximum number of consecutive identical characters. The core of this approach is to compare the two text sequences (the large model output and the text corresponding to the coordinates) and find the subsequence with the longest consecutive identical characters. The position and length of this subsequence reflect the degree of match between the two. The sliding window algorithm plays a key role in this process. It sets a fixed-length window (the window length can be adjusted based on actual needs and experience, such as an initial setting of 10 characters) and slides it simultaneously across the two text sequences.
[0062] Each time the window slides, the number of consecutive identical characters between the two text segments is calculated, and the maximum value and its corresponding coordinate information are recorded. For example, if the large model outputs "This is a sample text" and the coordinates correspond to the text "This is an important sample text message," when the window length is 5, the first comparison of the characters in the window between "This is a sample" and "This is a heavy" shows that the number of identical characters is 4. Continuing to slide, when the window contains "This is a sample," the number of identical characters reaches 5, and the coordinate information of this position is recorded. This process is repeated until the entire text sequence is traversed, and the maximum number of consecutive identical characters and their corresponding coordinates are finally determined.
[0063] In practice, the coordinate information corresponding to multiple different locations may match the maximum number of consecutive identical characters in the output of the large model. This is due to the possibility of repeated statements or similar paragraphs in the document. To ensure that no possible matches are missed, this solution outputs all matching coordinate information. For example, if there are two similar descriptions in a document, and the maximum number of consecutive identical characters between the output of the large model and both descriptions is 8, the coordinate information corresponding to both descriptions will be recorded, ensuring the completeness and accuracy of the match and providing data support for the subsequent accurate labeling of all relevant areas on the front end.
[0064] In step S104 of this implementation, a mechanism for determining the continuity of coordinates of multiple lines of text is proposed, namely, in the same document page, after the coordinates of the same line of text are merged, a list containing the merged coordinate information will be obtained. Each list element corresponds to the coordinate data of a line of text and has a unique index. The present invention determines whether two or more lines of text above and below are continuous by determining the continuity of the list element indexes.
[0065] Specific process, such as Figure 4As shown, the method includes: first obtaining a list of coordinate information after merging the coordinates of the same page and the same row, setting the list as coordinate_list, traversing the list coordinate_list, and checking whether the indexes of adjacent elements are continuous (that is, the index value of the subsequent element is equal to the index value of the previous element plus 1). If the indexes of all adjacent elements are continuous, it is preliminarily determined that the text in these rows is continuous in terms of layout; if the indexes of adjacent elements are discontinuous, such as when the index values jump, the text in the corresponding rows is determined to be discontinuous.
[0066] For example, suppose there are 5 elements in coordinate_list, with indexes 0, 1, 2, 3, and 4, and the difference between adjacent indexes is 1, which means that the corresponding 5 lines of text in the list are continuous in layout. If the list element indexes are 0, 1, 3, and 4, then since there is a jump between indexes 1 and 3, the two lines of text with indexes 1 and 3 are not continuous in layout.
[0067] After confirming index continuity, further verification of coordinate validity is required. For coordinate data corresponding to consecutive indexes, check that their page numbers are identical (ensure they are on the same page), and that the horizontal coordinates of text in the same column are roughly the same (a ±3-pixel tolerance can be set to account for document parsing errors). Only when the indexes are continuous and both the page numbers and horizontal coordinates meet these requirements can the coordinates of the text rows be considered continuous.
[0068] For example, the coordinates of the elements indexed 0, 1, and 2 in coordinate_list are:
[0069] {'x0': 80, 'y0': 180, 'x1': 120, 'y1': 200, 'page': 1} (index 0);
[0070] {'x0': 80, 'y0': 220, 'x1': 120, 'y1': 240, 'page': 1} (index 1);
[0071] {'x0': 80, 'y0': 260, 'x1': 120, 'y1': 280, 'page': 1} (index 2).
[0072] The page numbers of these coordinates are all 1, and the differences in the horizontal coordinates of the corresponding positions are within the error tolerance range. Combined with the judgment of index continuity, it can be finally determined that the coordinates of these three lines of text are continuous.
[0073] Optimized merging of multi-line text coordinates: For multi-line text coordinates that are determined to be continuous through the above judgment mechanism, in order to fully present the original text area when annotating on the front end, coordinate merging is required. When merging, from the coordinate data corresponding to the continuous index, take the smallest x0 value (i.e., the smallest horizontal coordinate of the upper left corner of the text box) of all coordinates as the starting horizontal coordinate of the merged coordinate, and the largest x1 value (i.e., the largest horizontal coordinate of the lower right corner of the text box) as the ending horizontal coordinate; the smallest y0 value (i.e., the smallest vertical coordinate of the upper left corner of the text box) as the starting vertical coordinate, and the largest y1 value (i.e., the largest vertical coordinate of the lower right corner of the text box) as the ending vertical coordinate; at the same time, retain the common page number value to form a new coordinate data unit.
[0074] Assume that the elements indexed 0, 1, 2, and 3 in coordinate_list correspond to four consecutive rows of coordinates:
[0075] {'x0': 50, 'y0': 150, 'x1': 90, 'y1': 170, 'page': 1} (index 0);
[0076] {'x0': 50, 'y0': 190, 'x1': 90, 'y1': 210, 'page': 1} (index 1);
[0077] {'x0': 50, 'y0': 230, 'x1': 90, 'y1': 250, 'page': 1} (index 2);
[0078] {'x0': 50, 'y0': 270, 'x1': 90, 'y1': 290, 'page': 1} (index 3).
[0079] After merging, we get {'x0': 50,'y0':150,'x1': 90, 'y1':290, 'page':1}. In the new coordinate data unit, x0=50 is the smallest starting horizontal coordinate of all continuous coordinates, x1=90 is the largest ending horizontal coordinate, y0=150 is the smallest starting vertical coordinate, and y1=290 is the largest ending vertical coordinate. It completely covers the display area of the four lines of text in the document. When annotating on the front end, it can accurately present the complete position of the source document in multi-line typesetting, significantly improving the accuracy and visualization of annotations.
[0080] It should be pointed out that the present invention is based on the common PDF document format and the widely used Python library (pdfplumber), and has good versatility and extensibility. It is not only suitable for PDF documents, but can also be applied to other similar electronic document formats after appropriate adjustments (for example, using other coordinate extraction methods). It provides a universal solution for document review, analysis, and annotation display in different scenarios, and has high commercial value and application prospects.
[0081] The method according to the embodiment of the present invention is described in detail above. To facilitate better implementation of the method according to the embodiment of the present invention, a system according to the embodiment of the present invention is provided below.
[0082] Figure 5 A system for determining coordinates of text in a document is shown, comprising:
[0083] The coordinate extraction unit 501 is configured to: extract the coordinates of each character in the source document and merge the coordinates with the same vertical coordinates on the same page;
[0084] The sliding matching unit 502 is configured to: obtain a large model processing result of the source document, and slide a sliding window between the large model processing result and the source document simultaneously, wherein the large model processing result is a text sequence;
[0085] The coordinate matching unit 503 is configured to: calculate the number of consecutive identical characters between two text segments within the sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates;
[0086] The coordinate merging unit 504 is configured to: when the obtained matching coordinates are continuous, merge the continuous matching coordinates, and finally determine the corresponding coordinates of the large model processing result in the source document.
[0087] In the coordinate extraction unit 501, the coordinates of each character in the source document are extracted, specifically including:
[0088] The characters in the source document are extracted one by one, and the coordinates of each character in the page and the page number are recorded, wherein the coordinates include the upper left corner horizontal coordinate, the upper left corner vertical coordinate, the lower right corner horizontal coordinate and the lower right corner vertical coordinate of the character box.
[0089] In the coordinate extraction unit 501, coordinates with the same vertical coordinates on the same page are merged, specifically including:
[0090] All coordinates on the same page are traversed through a double loop. The outer loop selects a coordinate as a reference, and the inner loop compares the remaining coordinate data with it. When the vertical coordinate difference is within the first set error range, it is determined to be the same row;
[0091] For the coordinates of the same row, all the coordinates of the row are merged into a new coordinate according to the maximum value of the horizontal coordinate of the lower right corner of the text box and the maximum value of the vertical coordinate of the lower right corner of the text box.
[0092] In the sliding matching unit 502, a sliding window is used to slide simultaneously between the large model processing result and the source document, specifically including:
[0093] A sliding window of fixed length is set to slide simultaneously on the large model processing result and the text sequence corresponding to the source document.
[0094] In the coordinate matching unit 503 , the maximum number of consecutive identical characters and the corresponding matching coordinates are recorded. When the maximum number of consecutive identical characters corresponding to multiple positions in the source document is the same, the matching coordinates corresponding to each position are recorded.
[0095] In the coordinate merging unit 504, whether the obtained matching coordinates are continuous is determined, specifically including:
[0096] On the same page, after the coordinates of the same line of text are merged, a coordinate list is obtained. Each element of the coordinate list corresponds to the coordinate data of a line of text and has a unique index.
[0097] When the indexes of the obtained multiple matching coordinates are continuous, the page numbers are the same, and the horizontal coordinate errors are within a second set threshold range, the multiple matching coordinates are determined to be continuous matching coordinates.
[0098] In the coordinate merging unit 504, continuous matching coordinates are merged, more specifically, including:
[0099] From the multiple continuous matching coordinates, take the horizontal coordinate of the smallest upper left corner of the text box as the starting horizontal coordinate of the merged coordinate, the horizontal coordinate of the largest lower right corner of the text box as the ending horizontal coordinate, the vertical coordinate of the smallest upper left corner of the text box as the starting vertical coordinate, and the vertical coordinate of the largest lower right corner of the text box as the ending vertical coordinate, retain the common page number value, and obtain the merged matching coordinates.
[0100] It is understandable that each of the above-mentioned units can be separately or completely combined into one or several other units to form a unit, or one (or some) of the units can be further divided into multiple functionally smaller units to form a unit, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the system may also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units.
[0101] According to another embodiment of the present application, the system described in this embodiment can be constructed by running a computer program (including program code) capable of executing the steps involved in the corresponding method described in Example 1 on a general-purpose computing device such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.
[0102] Figure 6 A computer device is shown, which includes a processor 601, a communication interface 602, and a computer-readable storage medium 603. The processor 601, the communication interface 602, and the computer-readable storage medium 603 may be connected via a bus or other means.
[0103] Among them, the communication interface 602 is used to receive and send data, the computer-readable storage medium 603 can be stored in the memory of the electronic device, the computer-readable storage medium 603 is used to store computer programs, the computer programs include program instructions, and the processor 601 is used to execute the program instructions stored in the computer-readable storage medium 603.
[0104] The processor 601 is the computing core and control core of the electronic device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement corresponding method processes or corresponding functions.
[0105] The processor 601 is configured to perform the following process:
[0106] Extract the coordinates of each character in the source document and merge the coordinates with the same vertical coordinates on the same page;
[0107] Obtaining a large model processing result of the source document, and sliding a sliding window simultaneously between the large model processing result and the source document, wherein the large model processing result is a text sequence;
[0108] Calculate the number of consecutive identical characters between two text segments within the sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates;
[0109] When the obtained matching coordinates are continuous, the continuous matching coordinates are merged to finally determine the corresponding coordinates of the large model processing result in the source document.
[0110] The present invention also provides a computer-readable storage medium, which is a memory device in an electronic device for storing programs and data. It is understood that the computer-readable storage medium herein may include both built-in storage media in the electronic device and, of course, extended storage media supported by the electronic device. The computer-readable storage medium provides storage space that stores the processing system of the electronic device.
[0111] Furthermore, the storage space also stores one or more instructions suitable for being loaded and executed by the processor. These instructions may be one or more computer programs (including program code). It should be noted that the computer-readable storage medium herein may be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device; alternatively, it may be at least one computer-readable storage medium located remotely from the processor.
[0112] In one embodiment, the computer-readable storage medium stores one or more instructions; the processor loads and executes the one or more instructions stored in the computer-readable storage medium to implement the following process:
[0113] Extract the coordinates of each character in the source document and merge the coordinates with the same vertical coordinates on the same page;
[0114] Obtaining a large model processing result of the source document, and sliding a sliding window simultaneously between the large model processing result and the source document, wherein the large model processing result is a text sequence;
[0115] Calculate the number of consecutive identical characters between two text segments within the sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates;
[0116] When the obtained matching coordinates are continuous, the continuous matching coordinates are merged to finally determine the corresponding coordinates of the large model processing result in the source document.
[0117] The present invention also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the following process:
[0118] Extract the coordinates of each character in the source document and merge the coordinates with the same vertical coordinates on the same page;
[0119] Obtaining a large model processing result of the source document, and sliding a sliding window simultaneously between the large model processing result and the source document, wherein the large model processing result is a text sequence;
[0120] Calculate the number of consecutive identical characters between two text segments within the sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates;
[0121] When the obtained matching coordinates are continuous, the continuous matching coordinates are merged to finally determine the corresponding coordinates of the large model processing result in the source document.
[0122] Those skilled in the art will appreciate that the units and algorithmic steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0123] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data processing device such as a server or data center that integrates one or more available media. Available media can include magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives).
[0124] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for determining the coordinates of characters in a document, characterized in that: The following processes are included: Extract the coordinates of each character in the source document and merge the coordinates with the same vertical coordinates on the same page; Obtaining a large model processing result of the source document, and sliding a sliding window simultaneously between the large model processing result and the source document, wherein the large model processing result is a text sequence; Calculate the number of consecutive identical characters between two text segments within the sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates; When the obtained matching coordinates are continuous, the continuous matching coordinates are merged to finally determine the corresponding coordinates of the large model processing result in the source document; Determine whether the obtained matching coordinates are continuous, including: On the same page, after the coordinates of the same line of text are merged, a coordinate list is obtained. Each element of the coordinate list corresponds to the coordinate data of a line of text and has a unique index. When the indexes of the obtained multiple matching coordinates are continuous, the page numbers are the same, and the horizontal coordinate errors are within a second set threshold range, the multiple matching coordinates are determined to be continuous matching coordinates; Merge consecutive matching coordinates, including: From the multiple continuous matching coordinates, take the horizontal coordinate of the smallest upper left corner of the text box as the starting horizontal coordinate of the merged coordinate, the horizontal coordinate of the largest lower right corner of the text box as the ending horizontal coordinate, the vertical coordinate of the smallest upper left corner of the text box as the starting vertical coordinate, and the vertical coordinate of the largest lower right corner of the text box as the ending vertical coordinate, retain the common page number value, and obtain the merged matching coordinates.
2. The method for determining the coordinates of characters in a document according to claim 1, wherein: Extract the coordinates of each character in the source document, including: The characters in the source document are extracted one by one, and the coordinates of each character in the page and the page number are recorded. The coordinates of each character in the page include the upper left corner horizontal coordinate, upper left corner vertical coordinate, lower right corner horizontal coordinate and lower right corner vertical coordinate of the character box.
3. The method for determining the coordinates of characters in a document according to claim 1, wherein: Merge coordinates with the same vertical coordinates on the same page, including: All coordinates on the same page are traversed through a double loop. The outer loop selects a coordinate as a reference, and the inner loop compares the remaining coordinate data with it. When the vertical coordinate difference is within the first set error range, it is determined to be the same row; For the coordinates of the same row, all the coordinates of the row are merged into a new coordinate according to the maximum value of the horizontal coordinate of the lower right corner of the text box and the maximum value of the vertical coordinate of the lower right corner of the text box.
4. The method for determining the coordinates of characters in a document according to claim 1, wherein: Using a sliding window to slide simultaneously between the large model processing result and the source document includes: A sliding window of fixed length is set to slide simultaneously on the large model processing result and the text sequence corresponding to the source document.
5. The method for determining the coordinates of characters in a document according to claim 1, wherein: The maximum number of consecutive identical characters and the corresponding matching coordinates are recorded. When the maximum number of consecutive identical characters corresponding to multiple positions in the source document is the same, the matching coordinates corresponding to each position are recorded.
6. A system for determining the coordinates of characters in a document, characterized in that: include: The coordinate extraction unit is configured to: extract the coordinates of each character in the source document and merge the coordinates with the same vertical coordinates on the same page; a sliding matching unit configured to: obtain a large model processing result of the source document, and slide a sliding window simultaneously between the large model processing result and the source document, wherein the large model processing result is a text sequence; a coordinate matching unit configured to: calculate the number of consecutive identical characters between two text segments within a sliding window, record the maximum number of consecutive identical characters and the corresponding matching coordinates, and traverse the entire source document to obtain all matching coordinates; a coordinate merging unit configured to: when the obtained matching coordinates are continuous, merge the continuous matching coordinates, and finally determine the corresponding coordinates of the large model processing result in the source document; Determine whether the obtained matching coordinates are continuous, including: On the same page, after the coordinates of the same line of text are merged, a coordinate list is obtained. Each element of the coordinate list corresponds to the coordinate data of a line of text and has a unique index. When the indexes of the obtained multiple matching coordinates are continuous, the page numbers are the same, and the horizontal coordinate errors are within a second set threshold range, the multiple matching coordinates are determined to be continuous matching coordinates; Merge consecutive matching coordinates, including: From the multiple continuous matching coordinates, take the horizontal coordinate of the smallest upper left corner of the text box as the starting horizontal coordinate of the merged coordinate, the horizontal coordinate of the largest lower right corner of the text box as the ending horizontal coordinate, the vertical coordinate of the smallest upper left corner of the text box as the starting vertical coordinate, and the vertical coordinate of the largest lower right corner of the text box as the ending vertical coordinate, retain the common page number value, and obtain the merged matching coordinates.
7. A computer device, characterized in that: include: a processor and a computer-readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the method for determining the coordinates of text in a document according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the method for determining coordinates of text in a document according to any one of claims 1 to 5.
Citation Information
Patent Citations
Privacy information shielding method and device based on version identification and key field positioning
CN119227138A