Structural information optimization method for document content identification
Through the combination of deep learning and semantic matching models, the problems of text misalignment and printing missing for monopoly licenses are solved, and accurate information extraction and management efficiency are improved.
Patent Information
- Application Number
- CN202510918198.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-03
AI Technical Summary
When the prior art deals with the problems of text misalignment, printing missing and image blurring on monopoly licenses, it is difficult to accurately extract information, resulting in inefficient information management and supervision.
The text box coordinates are extracted by using the deep learning model, combined with the K-Means spatial clustering algorithm to allocate the text box to the left and right columns, and a semantic matching model is constructed through Sentence-BERT, and the semantic similarity between field names and values is calculated and column alignment is aligned, and the misalignment scenarios are handled in combination with the scrolling mechanism.
It effectively solves the problems of text misalignment and printing missing, ensures the accuracy and management efficiency of information extraction, and improves the universality and robustness of the method.
Smart Images

Figure CN120412002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a method for optimizing the structural information of document content recognition. Background Art
[0002] The monopoly license is an important legal document issued by the monopoly administrative department in accordance with the law, carrying the attribute of a legal certificate for market entities to engage in monopoly business activities. The information recorded therein covers key structured contents such as enterprise name (identifying the identity of the market entity), name of the person in charge (clarifying the responsible entity), type of enterprise (defining the organizational nature), business premises (limiting the business geographical scope), scope of permission (stipulating the type of business operations), supply unit (clarifying the source of goods), validity period (setting the duration of rights), issuing authority (indicating the subject making the permission), and document production date (recording the time of document generation), etc. These information are the core basis for implementing market supervision, conducting business activity verification, and managing industry data. However, there are still certain problems: First, in practical applications, due to problems such as printing equipment errors, paper position offset, and shooting angle, the text on the license will show obvious misalignment, resulting in the inability to accurately correspond to the corresponding entries during information extraction; Second, there are cases of missing due to printing omission, image blurring, or OCR misrecognition, which cannot be directly corrected in sequence, affecting the accuracy and efficiency of subsequent information management and supervision; Third, existing technologies usually rely on fixed positions or traditional OCR methods and are difficult to effectively handle the situation of printing misalignment; Therefore, a method for optimizing the structural information of document content recognition is proposed. Summary of the Invention
[0003] In view of this, embodiments of the present invention hope to provide a method for optimizing the structural information of document content recognition to solve or alleviate the technical problems existing in the prior art and at least provide a beneficial option.
[0004] To solve the above technical problems, a technical solution adopted by this application is: a method for optimizing the structural information of document content recognition, including the following steps: Step 1: Obtain a picture of the monopoly license, and extract the text boxes, their coordinates, and contents of all text regions in the picture based on a deep learning model; Step 2: Process all the extracted text boxes, delete the text boxes belonging to the license name, and use the K-Means spatial clustering algorithm to assign them to the left column and the right column according to the coordinates of the remaining text boxes; Step 3: Based on each column of text boxes after clustering, sort the text boxes from top to bottom according to their y values to obtain two columns of text boxes arranged in order and their text contents; Step 4. Inside each obtained column, judge the vertical distance between adjacent text boxes. If the y-axis distance between the two is lower than the set threshold, merge their text contents into consecutive parts of the same field; Step 5. Based on Sentence-BERT, extract field pairs from the license pictures collected on-site, construct a semantic matching dataset, and train to obtain a semantic matching model; Step 6. Take the left and right columns as the pairing group of field names and values, calculate the semantic similarity between the field name and the candidate values, and select the highest one as the initial match; Step 7. If the similarity exceeds the threshold, translate the second column with the matching row as the anchor point to align the matching starting point; Step 8. Match the subsequent fields row by row. If the continuous matching fails, roll the second column to re-match, and output the successfully matched field pairs.
[0005] Provided as a further optimization of this technical solution, in Step 2, the specific steps of allocating the remaining text boxes to the left and right columns according to their coordinates are as follows: Step 201. Set the number of clustering centers K = 4, calculate the range of x coordinates Δx = max(x i ) - min(x i ) in the set T of all remaining text boxes. Select the initial clustering centers C = {C1, C2, C3, C4} based on Δx, where the x coordinates of C1 and C2 are in the interval [min(x i ), min(x i ) + Δx / 3], the x coordinates of C3 and C4 are in the interval [max(x i ) - Δx / 3, max(x i )], and the y coordinates of C1 - C4 take the average y value of the set T; Step 202. For each text box T i ∈T, calculate the Euclidean distance from its upper left corner coordinate point (x i , y i ) to each clustering center C k : ; where (x k , y k ) is the coordinate of the clustering center C k ; Step 203. Allocate the text box T i to the cluster S k corresponding to the clustering center with the minimum distance, and update the cluster center coordinates to: ; where |S k | represents the cluster S kThe number of Chinese text boxes; Step 204: Repeat Step 202 - Step 203 until the change in the cluster center coordinates is less than 0.5 pixels or the iteration exceeds 50 times. Assign the finally converged cluster S1∪S2 to the left column and S3∪S4 to the right column.
[0006] As a further optimization of this technical solution, in Step 6, the cosine similarity formula is used to calculate the semantic similarity between the calculated field name and the candidate value: ; where a and b are the sentence vectors generated by Sentence - BERT.
[0007] As a further optimization of this technical solution, in Step 8, the operation of scrolling the second column is to update the text box sequence to , and if the matching fails 3 times in a row, scrolling is triggered, and the semantic similarity is recalculated after each scroll.
[0008] As a further optimization of this technical solution, in Step 7, the similarity threshold is 0.8. If the matching is successful, the second column is translated with the anchor row to align the starting points of the two - column matches.
[0009] As a further optimization of this technical solution, in Step 1, the deep - learning model uses PaddleOCR. The text detection model is trained with 500 on - site license pictures to extract text boxes including the upper - left coordinates.
[0010] As a further optimization of this technical solution, in Step 4, the y - axis spacing threshold is set to 50% of the average height of the text boxes, and the contents of adjacent text boxes are merged to process multi - line fields.
[0011] As a further optimization of this technical solution, in Step 3, the text boxes are sorted from top to bottom according to their y values, arranged in ascending order of the text box y values, and text boxes with a y - value difference less than 30% of the average height are regarded as the same row.
[0012] As a further optimization of this technical solution, in Step 5, the semantic matching data set is constructed based on the top - ten field pairs extracted from 2000 on - site license pictures, including enterprise name, person - in - charge name, business location, scope of permission, etc.
[0013] As a further optimization of this technical solution, the method further includes performing a semantic coherence check on the output result, specifically: Calculate the cosine similarity of the word vectors of adjacent fields. When it is lower than 0.6, manual review is triggered; Verify the field type format; Check whether required fields are missing.
[0014] Due to the adoption of the above technical solutions in the embodiments of the present invention, it has the following advantages: 1. By using the K-Means spatial clustering algorithm to dynamically cluster the coordinates of text boxes and combining the column translation and scrolling mechanisms, the present invention effectively solves the problem of text misalignment caused by printing equipment errors, paper offset, shooting angles, etc., ensuring the accurate correspondence between fields and values during information extraction; 2. By adopting customized training of PaddleOCR to improve text detection accuracy, constructing a semantic matching model in combination with Sentence-BERT, and performing semantic coherence verification on the output results, the present invention solves the problem of information loss caused by printing omissions, blurred images, or OCR misrecognition, ensuring the accuracy and efficiency of information management and supervision; 3. By abandoning traditional fixed-position or single OCR methods and adopting the methods of dynamic clustering and sorting, and dual semantic + spatial matching, the present invention adapts to different layouts and misalignment scenarios, solves the problem that existing technologies are difficult to handle printing misalignment, and improves the versatility and robustness of the method.
[0015] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present invention will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a flow chart of the structural information optimization method for content recognition of the method of the present invention; Figure 2 It is a flow chart of the method for allocating the remaining text boxes to the left column and the right column according to their coordinates in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The following will describe the embodiments of the present disclosure in detail with reference to the drawings.
[0019] It should be clear that the following uses specific specific examples to illustrate the implementation modes of the present disclosure, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation modes, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present disclosure.
[0020] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is only illustrative. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. Additionally, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0021] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure in a schematic manner. The diagrams only show the components related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0022] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0023] Figure 1 is a schematic flowchart of the method for optimizing the structural information of document content recognition in an embodiment of the present invention. It should be noted that if there are substantially the same results, the method of this application is not limited to Figure 1 the process sequence shown. As Figure 1 - Figure 2 shown: The method for optimizing the structural information of document content recognition includes the following steps: Step 1: Obtain a picture of a monopoly license, and based on a deep learning model, extract the text boxes, their coordinates, and contents of all text regions in the picture; Specifically, first, obtain a picture of the monopoly license. The picture can be obtained by taking a photo with a camera, scanning with a scanner, etc., ensuring that the obtained picture is clear and can clearly display the text content on the license; Then, process the obtained picture based on a deep learning model. Here, the deep learning model uses PaddleOCR. A training set is constructed with 500 license pictures collected on-site, and the text detection model of PaddleOCR is trained specifically to make it better adapt to the scenario of the monopoly license; Finally, use the trained deep learning model to perform text detection on the picture, extract the text boxes and their coordinates and contents in all text regions in the picture. Each text box contains its upper left coordinates (x, y) in the picture and the text content inside the box, thus providing basic data for subsequent text processing.
[0024] Step 2: Process all the extracted text boxes, delete the text boxes belonging to the license name, and use the K-Means spatial clustering algorithm to assign them to the left column and the right column according to the coordinates of the remaining text boxes; Specifically, first, identify and delete the text boxes belonging to the license name by rule filtering and keyword matching. Delete the text boxes located in the top 15% area of the image with an aspect ratio > 5. Calculate the edit distance between the text box content and the preset keyword library, and delete the candidate boxes with a threshold < 2; Then, extract the center point coordinates of the remaining text boxes as clustering features, calculate the range of x coordinates Δx of all text boxes, set the number of cluster centers K = 4, and place the x coordinates of the initial centers C1 and C2 in the interval [min(x i ), min(x i ) + Δx / 3], and place the x coordinates of C3 and C4 in the interval [max(x i ) - Δx / 3, max(x i )]. Set the y coordinates of all initial centers to the arithmetic mean y value of the text box set; Next, perform K-Means clustering iteration, calculate the Euclidean distance from the center point of each text box to each cluster center, assign it to the cluster with the smallest distance, and update the cluster center coordinates to the average value of the center points of all text boxes in the cluster, and iterate until the change in the cluster center coordinates < 0.5 pixels or the number of iterations > 50 times; Finally, assign the converged clusters S1∪S2 to the left column and S3∪S4 to the right column. If the number of text boxes in a certain column is less than 20% of the total number, re-initialize the clustering, calculate the average x coordinates of the left and right columns, and filter out outliers.
[0025] Step 3: Based on each column of text boxes after clustering, sort the text boxes from top to bottom according to their y values to obtain two columns of text boxes and their text contents arranged in order; Specifically, first, preprocess each column of text boxes after clustering, calculate the average height h_avg of the text boxes within the column, calibrate the y-coordinate of each text box to y_top + h_avg / 2, and filter out the outlier text boxes whose y-coordinate deviates from the median by more than 3 standard deviations; Then, sort the text boxes in ascending order based on the calibrated y values, consider the text boxes with a y value difference less than h_avg×30% as the same row, and adjust the order in ascending order of the x value within the same row; Next, detect the vertical spacing between adjacent text boxes. If the spacing is less than h_avg×50%, merge them into continuous text, concatenate the content in order, and retain the leftmost x and topmost y coordinates; Finally, verify the sorting result, calculate the standard deviation σ_y of the y spacing between adjacent text boxes. If σ_y is greater than h_avg×0.7, adjust the threshold and re-sort to finally form an ordered sequence of text boxes and corresponding content in two columns on the left and right.
[0026] Step 4: Inside each obtained column, judge the vertical distance between adjacent text boxes. If the y-axis spacing between the two is lower than the set threshold, merge their text contents into consecutive parts of the same field; Specifically, first, calculate the average height h_avg of the text boxes in each column as the basis for spacing judgment; [[ID= | 14]]Then, traverse adjacent text box pairs in each column, obtain the bottom y coordinate (y1_bottom) of the previous text box and the top y coordinate (y2_top) of the next text box, and calculate their vertical spacing diff = y2_top - y1_bottom; Next, compare diff with the set threshold (e.g., h_avg×50%). If diff is lower than the threshold, concatenate the text contents of the two text boxes in order to form consecutive text of the same field; Finally, merge the coordinates of the text boxes, update them to include the minimum upper-left y coordinate and the maximum lower-right y coordinate of the two text boxes, and retain the x coordinate range of the original text boxes to complete the merging process of adjacent text boxes.
[0027] Step 5: Based on Sentence-BERT, extract field pairs from the license pictures collected on-site, construct a semantic matching dataset, and train to obtain a semantic matching model; Specifically, first, collect 2000 monopoly license pictures on-site, and extract ten core field pairs through manual annotation, including enterprise name - enterprise name value, person in charge name - name value, business location - address value, license scope - scope value, supplier unit - unit value, expiration date - term value, issuing authority - authority value, certificate making date - date value, enterprise type - type value, license number - number value; Then, preprocess the extracted field pairs, clean special symbols, space anomalies, etc. in the text, tokenize the Chinese text using the jieba tokenization tool, and remove stop words; Next, construct a training dataset based on the Sentence-BERT model, convert the field pairs into text pair format of [field name, field value], and generate positive and negative samples using the Contrastive Learning strategy (positive samples are real field pairs, and negative samples are randomly combined across fields); then configure the training parameters, set the batch size to 32, the learning rate to 2e-5, the number of training epochs to 5, use the AdamW optimizer and enable gradient clipping (threshold 1.0), and adjust the learning rate through the cosine annealing learning rate scheduling strategy; among them, the Sentence-BERT model selects the paraphrase-multilingual-MiniLM-L12-v2 pre-trained model; Finally, evaluate the model performance on 10% of the validation set, calculate the F1 value using the cosine similarity of the field pairs, stop training when the validation set F1≥0.92, and save the optimal model weights for subsequent semantic matching.
[0028] Step Six: Take the left and right columns as the pairing group of field names and values, calculate the semantic similarity between the field name and the candidate value, and select the highest one as the initial match; Specifically, first, take the content of the left column text box sorted in Step Three as the field name set F = [f1, f2, …, f m , and the right column as the candidate value set V = [v1, v2, …, v n ; Then, load the Sentence-BERT model trained in Step Five, and convert each field name f i and candidate value v j into 768-dimensional sentence vectors a i and b j ; Next, traverse all the pairing combinations of field names and candidate values, and calculate the semantic similarity using the cosine similarity formula: ; Among them, a and b are the sentence vectors generated by the Sentence-BERT model, representing the semantic feature vectors of the field name and the candidate value respectively; Record the similarity value of each pair (f i , v j ); Finally, for each field name f i , find the v with the highest similarity in the candidate value set jAs an initial match, if the highest similarity is lower than the preset threshold (e.g., 0.4), it is marked as unmatched, forming an initial set of matching pairs M = {(f1, v j1 ),(f2, v j2 ),…,(f m , v jm )}.
[0029] Step Seven: If the similarity exceeds the threshold, shift the second column with the matching row as the anchor point to align the starting point of the match; Specifically, first, traverse the set of initial matching pairs obtained in Step Six, and filter out the matching pairs with a similarity ≥ 0.6 as valid anchor points. If the number of valid anchor points is insufficient, reduce the threshold to 0.5 and continue filtering; Then, calculate the difference in the y - coordinates of each valid anchor point in the left and right columns: Δy k = y_left_anchor k - y_right_anchor k ; where y_left_anchor k is the upper - left y - coordinate of the k - th anchor - point row text box in the left column, and y_right_anchor k is the upper - left y - coordinate of the k - th anchor - point row text box in the right column; Take the median of all Δy k as the final vertical offset Δy; Next, for all text boxes in the right column, add the Δy value to their upper - left and lower - right y - coordinates respectively to complete the translation, and check whether the translated text boxes exceed the image boundary. If they do, adjust Δy so that the boundary text boxes are just visible; Finally, based on the translated right column, re - find the candidate value with the y - coordinate closest (δy ≤ 0.5×average text height) for each field name as the starting point of the match, and update the set of matching pairs.
[0030] Step Eight: Match the subsequent fields line by line. If the consecutive matches fail, scroll the second column to re - match, and output the successfully - matched field pairs; Specifically, first, initialize the current row index i = 0, the consecutive - failure counter fail_count = 0, and the maximum consecutive - failure threshold 3, and set the maximum scroll count of the right column to 5 times; Then, obtain the content of the i - th row text box from the left column as the field name, calculate the semantic similarity between it and the candidate value of the i - th row in the right column. If the similarity ≥ 0.5, the match is successful, add the field pair to the result set and reset fail_count, otherwise increment fail_count by 1; Next, check whether the fail_count reaches the threshold. If so, scroll the entire right column upward by the height of an average text box, and reset i and fail_count; Finally, determine whether all the field names in the left column have been traversed or the number of scrolls in the right column exceeds the limit. If so, output the set of successfully matched field pairs and the list of unmatched fields.
[0031] In one embodiment, specifically, in step two, the specific steps of allocating the remaining text boxes to the left column and the right column according to their coordinates are as follows: Step 201: Set the number of cluster centers K = 4, and calculate the range of the x coordinates of all elements in the set T of remaining text boxes, Δx = max(x i ) - min(x i ). Select the initial cluster centers C = {C1, C2, C3, C4} based on Δx, where the x coordinates of C1 and C2 are in the interval [min(x i ), min(x i ) + Δx / 3], and the x coordinates of C3 and C4 are in the interval [max(x i ) - Δx / 3, max(x i ). The y coordinates of C1 - C4 are the average y value of the set T; Specifically, first, set the number of cluster centers to K = 4. This is because the text of the monopoly license is usually distributed in two columns on the left and right, and setting two cluster centers in each column can more accurately capture the spatial distribution characteristics of the text boxes. Second, calculate the range of the x coordinates of all elements in the set T of remaining text boxes, Δx, that is, determine the distribution range on the x-axis by finding the difference between the maximum value max(x i ) and the minimum value min(x i ); Next, divide the x coordinate intervals of the initial cluster centers based on Δx: Limit the x coordinates of C1 and C2 to the interval [min(x i ), min(x i ) + Δx / 3], which covers the x coordinate range of the left column text boxes; the x coordinates of C3 and C4 are in the interval [max(x i ) - Δx / 3, max(x i ), corresponding to the x coordinate range of the right column text boxes; the purpose of this setting is to make the initial centers close to the actual distribution of the left and right columns and reduce the number of clustering iterations; Finally, the y - coordinates of C1 to C4 are uniformly taken as the average of the y - coordinates of all text boxes in set T, ensuring that the initial center is at the center of the text area in the vertical direction and avoiding clustering bias caused by y - coordinate deviation. For example, if the x - coordinate range of the remaining text boxes is from 100 to 500, then Δx = 400, the x - coordinate interval of the left - column center is [100, 233], the right - column center is [367, 500], and the y - coordinate is taken as the mean of all text - box y - values, thus providing reasonable initial conditions for subsequent K - Means clustering.
[0032] Step 202: For each text box T i ∈T, calculate the Euclidean distance from its upper - left - corner coordinate point (x i , y i ) to each cluster center C k : ; where (x k , y k ) are the coordinates of cluster center C k ; Specifically, step 202 is a key calculation step in the K - Means clustering algorithm. It has been determined previously that the number of cluster centers K = 4 and the initial cluster centers C = {C1, C2, C3, C4} have been set. Now, for each text box T i in the remaining text - box set T, the specific calculation process is as follows: Coordinate acquisition: Given that the upper - left - corner coordinate of text box T i is (x i , y i ), which is the position information of the text box on the image plane. At the same time, the cluster center C k also has its corresponding coordinates (x k , y k ), where k takes values 1, 2, 3, 4, corresponding to four different cluster centers respectively; Difference calculation: Calculate the differences between the coordinates of text box T i and the coordinates of cluster center C k in the x - axis and y - axis directions, that is, x i - x k and y i - y k ; These two differences reflect the offset degree of the text box relative to the cluster center in the horizontal and vertical directions; Square - sum and square - root operation: Square the above two differences respectively to get (x i - x k )² and (y i - y k)², then add these two squared values together, and then take the square root of the result, that is, calculate , and the value obtained is the Euclidean distance from the upper left corner coordinate point of the text box T i to the cluster center C k ; By calculating the Euclidean distance, the "closeness" between the text box T i and each cluster center C k can be measured; in the K-Means clustering algorithm, the text box will be assigned to the category represented by the cluster center with the smallest Euclidean distance; in the scenario of processing the text boxes of the monopoly license, in this way, the text boxes can be roughly divided into two columns on the left and right (because 4 cluster centers are set, two in each of the left and right columns), laying a foundation for further processing (such as alignment, matching, etc.) of the text in the left and right columns; for example, if the Euclidean distance from the text box T i to C1 or C2 is the smallest, then it is very likely to belong to the text boxes in the left column; if the Euclidean distance to C3 or C4 is the smallest, then it is very likely to belong to the text boxes in the right column.
[0033] Step 203: Assign the text box T i to the cluster S k corresponding to the cluster center with the smallest distance, and update the cluster center coordinates to: ; where, |S k | represents the number of text boxes in the cluster S k ; Specifically, in Step 202, the Euclidean distance from each text box T i to each cluster center C k has been calculated ; this step is to assign each text box T i to the cluster S k corresponding to the cluster center with the smallest distance according to these distances; For example, assume that the Euclidean distances from the text box T1 to C1, C2, C3, and C4 are D 11 = 5, D 12 = 8, D 13 = 10, D 14 = 12. Since D 11 is the smallest, the text box T1 is assigned to the cluster S1; performing such operations on all the text boxes in the set T can divide all the text boxes into different clusters; After all the text boxes are assigned, each cluster S k contains several text boxes; at this time, the center coordinates of each cluster need to be updated ; Update formula: ; The meaning is as follows: For the horizontal axis, Represents cluster S k Add the x values of the upper left corner coordinates of all text boxes in the textbox. It is cluster S k Number of text boxes in |S k | Take the reciprocal, and multiply the two to get the cluster S k The average value of the x-coordinate of the text box in the text box is used as the updated cluster center The horizontal axis of .
[0034] For the vertical axis, Represents cluster S k Add the y values of the upper left corner coordinates of all text boxes in the text box, and divide them by |S k |Get the average value of the y coordinate as the updated cluster center The vertical coordinate of Through such updates, the cluster center will continue to move closer to the "center position" of the text box in the cluster; then, it will return to step 202 again, recalculate the Euclidean distance from the text box to the new cluster center, and continuously iterate this process until the change in the cluster center is less than a set threshold (for example, the coordinate change is less than 0.5 pixels). The clustering is considered to have converged, and the clustering division of the text box is completed at this time, preparing for subsequent operations such as distinguishing the left and right columns of the monopoly license text.
[0035] Step 204: Repeat steps 202 to 203 until the cluster center coordinates change by less than 0.5 pixels or the iterations exceed 50 times, and assign the finally converged cluster S1∪S2 to the left column and S3∪S4 to the right column; Specifically, step 204 is the iterative termination and result application phase of the K-Means clustering algorithm in this scenario; wherein the iterative process is specifically as follows: After completing the text box allocation and cluster center coordinate update in step 203, it is necessary to return to step 202 again; at this time, since the cluster center coordinate has been updated, it is necessary to recalculate each text box T i To the new cluster center C k Euclidean distance Then, the process returns to step 203, and redistributes the text boxes to the cluster corresponding to the cluster center with the smallest distance according to the newly calculated distance, and updates the cluster center coordinates again. This process of continuously repeating steps 202 and 203 is the iterative process of the clustering algorithm. The termination conditions are as follows: Threshold for changes in cluster center coordinates: After each iteration to update the cluster center coordinates, the difference between the new cluster center coordinates and the previous cluster center coordinates is calculated. When the changes in all cluster center coordinates in the x-axis and y-axis directions are less than 0.5 pixels, it indicates that the clustering centers have basically stabilized and there are no large offsets. At this time, it can be considered that the clustering has converged and the iteration can stop. This is because in the image coordinate system, a change of 0.5 pixels is relatively small, and from the perspective of text box clustering, a relatively ideal classification effect has been achieved. Limit on the number of iterations: Additionally, to prevent the algorithm from falling into a situation where it cannot converge (although this generally does not occur under reasonable settings), the upper limit of the number of iterations is set to 50 times. If the number of iterations reaches 50 times, even if the change in the cluster center coordinates is not less than 0.5 pixels, the iteration is stopped. When the iteration meets the above termination conditions, the four finally converged clusters S1, S2, S3, and S4 are obtained. In the scenario of dealing with monopoly license texts, according to the setting, the two clusters S1 and S2 are merged and assigned to the left-column text; the two clusters S3 and S4 are merged and assigned to the right-column text. In this way, the preliminary division of the left and right columns of the monopoly license text boxes is completed, laying a foundation for subsequent operations such as sorting and matching the left and right column texts.
[0036] The specific application of the K-Means spatial clustering algorithm is as follows: Suppose after the processing in Step 1, the upper-left coordinates of the remaining 5 text boxes are as follows:
[0037] Calculate the range of the x coordinate: Δx = max(100, 120, 400, 420, 110) - min(100, 120, 400, 420, 110) = 420 - 100 = 320 Determine the initial clustering centers: Range of the x coordinates of the left-column centers C1 and C2: [min(x i ), min(x i ) + Δx / 3] = [100, 100 + 320 / 3] ≈ [100, 206.67] Take C1(150, 244) and C2(180, 244); Range of the x coordinates of the right-column centers C3 and C4: [max(x i ) - Δx / 3, max(x i )] = [420 - 320 / 3, 420] ≈ [313.33, 420] Take C3(350, 244) and C4(400, 244); The y - coordinates of all centers: avg(200, 250, 210, 260, 300) = 244 Initial cluster centers:
[0038] The first round of iteration: Calculate the Euclidean distances from T1(100, 200) to each center:
[0039] Assignment: T1 → S1 (the smallest distance); Similarly, calculate other text boxes: T2(120, 250) → S1 T3(400, 210) → S4 T4(420, 260) → S4 T5(110, 300) → S1 Update the coordinates of the cluster centers: ; Since there are no assigned text boxes for C2 and C3, their coordinates remain unchanged; The second round of iteration: Recalculate the distances and assign: T1(100, 200) → S1 (the updated C1 has the smallest distance) T2(120, 250) → S1 T3(400, 210) → S4 T4(420, 260) → S4 T5(110, 300) → S1 Update the coordinates of the cluster centers:
[0040] Since the text boxes within the cluster do not change, the coordinates of the cluster centers are the same as those in the previous round of iteration; After the second round, the change in the coordinates of the cluster centers is < 0.5 pixels, meeting the convergence condition; Final assignment result: Left column: S1 ∪ S2 = {T1, T2, T5} Right column: S3 ∪ S4 = {T3, T4} After verification: The x - coordinate range of the text boxes in the left column: 100 - 120; The x - coordinate range of the text boxes in the right column: 400 - 420; It conforms to the spatial distribution characteristics of the left and right columns.
[0041] In one embodiment, specifically, in step six, the cosine similarity formula is used to calculate the semantic similarity between the computed field name and the candidate value: ; where a and b are sentence vectors generated by Sentence - BERT; The cosine similarity formula measures the semantic similarity between two texts by calculating the dot product of sentence vectors a and b and dividing it by the product of their respective norms. The closer the value is to 1, the more similar the semantics, thus helping to determine whether the field name of the monopoly license matches the candidate value; Specifically, assume there is a monopoly license, from which the field name "Name of the person in charge" is extracted from the left column, and the corresponding candidate values in the right column are "Zhang San", "Li Si", and "License number: 123456"; Use the Sentence - BERT model to generate sentence vectors; Input "Name of the person in charge", "Zhang San", "Li Si", and "License number: 123456" into the pre - trained Sentence - BERT model respectively; The Sentence - BERT model processes each input text and converts it into a sentence vector of a fixed length; Assume the sentence vector generated by "Name of the person in charge" is a = [0.1, 0.2, 0.3, ……, 0.768] (assuming the dimension of the sentence vector is 768); The sentence vector generated by "Zhang San" is b1 = [0.2, 0.3, 0.4, ……, 0.769]; The sentence vector generated by "Li Si" is b2 = [0.15, 0.25, 0.35, ……, 0.767]; The sentence vector generated by "License number: 12,3456" is b3 = [0.05, 0.08, 0.1, ……, 0.5]; Calculate the semantic similarity between "Name of the person in charge" and "Zhang San" according to the cosine similarity formula: ; where a·b is the dot product of vectors a and b, and the calculation method is to multiply the corresponding position elements and then add them together, that is: ; is the norm of vector a (which can be understood as the length of the vector), and the calculation formula is: ; Similarly; After calculation, assume that sim(a, b1) = 0.85; The semantic similarity between "Responsible Person's Name" and "Li Si" is also calculated according to the above formula, and sim(a, b2) = 0.83 is obtained; The semantic similarity between "Responsible Person's Name" and "License Number: 123456" is calculated, and sim(a, b3) = 0.1 can be obtained; From the calculation results, it can be seen that the semantic similarity between "Responsible Person's Name" and "Zhang San" and "Li Si" is relatively high, being 0.85 and 0.83 respectively, while the semantic similarity with "License Number: 123456" is only 0.1; In practical applications, a similarity threshold can be set, such as 0.7; when the semantic similarity between the field name and the candidate value is greater than or equal to 0.7, they are considered to be a match; so in this example, both "Zhang San" and "Li Si" may be the correct values corresponding to "Responsible Person's Name"; subsequently, other information (such as frequency of occurrence, etc.) can be further combined to determine the final matching value.
[0042] In one embodiment, specifically, in step eight, the operation of scrolling the second column is to update the text box sequence to , if the matching fails 3 times in a row, the scrolling is triggered, and the semantic similarity is recalculated after each scroll; Specifically, in the context of the exclusive license information processing involved in step eight, the text box sequence represents the set of text boxes in the second column (usually the value column) on the license, where are respectively the individual text boxes in the second column; The operation of scrolling the second column is specifically that when the matching between the field name and the candidate value in the second column fails 3 times in a row (i.e., it is determined to be unmatched through methods such as semantic similarity calculation), this operation is triggered; At this time, the text box sequence in the second column is moved forward by one position, and the original first text box is removed, and a "blank" value is filled at the end to form a new sequence ; After scrolling, the semantic similarity between the left column field name and the new second column candidate value is recalculated. The purpose is to find the correct field name - value matching combination by adjusting the text box correspondence to accurately extract the license structured information.
[0043] Suppose there is an exclusive license, when extracting its structured information: After pre - processing, the text box sequence in the second column (value column) is obtained, corresponding to "No. 300", "Retail", "2025 - 01 - 01", "Wang Wu", "Company" respectively, and the left column field names are "Business Location", "License Scope", "Valid Period", "Responsible Person's Name", "Enterprise Name" in sequence; First, calculate the semantic similarity between "business location" and ("No. 300"). Suppose the calculation result is lower than the set threshold, the matching fails, and the count is 1 time; Next, calculate the semantic similarity between "business location" and the next candidate value ("retail"). The result is still lower than the threshold, the matching fails, and the count is 2 times; Then calculate the semantic similarity between "business location" and ("2025-01-01"). It is still lower than the threshold, the matching fails, and at this time, the matching fails continuously for 3 times; Trigger the scrolling operation to update the text box sequence to =["retail", "2025-01-01", "Wang Wu", "company", "empty"]; After scrolling, recalculate the semantic similarity between "business location" and the candidate values in the new sequence; at this time, calculate the semantic similarity between "business location" and "retail", "2025-01-01", "Wang Wu", "company" in turn. Suppose the semantic similarity between "business location" and "company" is higher than the threshold, the matching is successful, and the correct matching of the field name and the value is completed; subsequently, match other field names and the candidate values in the second column after scrolling in the same way until the extraction of the entire license structured information is completed.
[0044] In one embodiment, specifically, in step seven, the similarity threshold is 0.8. If the matching is successful, the second column is translated by the anchor row to align the starting points of the two-column matching; Specifically, first, set the similarity threshold to 0.8, which is the standard for judging whether the field name and the candidate value match successfully; after using models such as Sentence-BERT to generate sentence vectors and calculating the semantic similarity between the field name and the candidate value through the cosine similarity formula, if the obtained similarity value is greater than or equal to 0.8, it is determined that the two are highly similar semantically, that is, the matching is successful; for example, for the field name "enterprise name" and the candidate value "XX Co., Ltd.", the calculated cosine similarity is 0.85. Since 0.85≥0.8, it is determined to be a successful match; When there are multiple successful matching cases, it is necessary to determine the anchor row; the anchor row is selected from these successfully matched rows and used as the benchmark for subsequent column alignment operations; for example, if the matching of both "enterprise name - XX Co., Ltd." and "person in charge's name - Zhang San" between the field name and the candidate value is successful, one of the rows where they are located can be selected as the anchor row according to certain rules (such as frequency of occurrence, importance, etc.). Suppose the row where "enterprise name - XX Co., Ltd." is located is selected as the anchor row; After determining the anchor row, the second column is translated based on the anchor row. Due to the vertical misalignment of the monopoly license text box, by translating the second column, the candidate values in the second column that match the field names in the anchor row are aligned vertically with the field names in the left column. For example, the vertical coordinate of the "Enterprise Name" text box in the left column is y1, and the coordinate of the "XX Co., Ltd." text box in the right column is y2, and there is a certain difference between them. By calculating this difference, the entire second column is moved up or down by the corresponding distance to achieve vertical alignment of the two columns at the matching starting point of the anchor row, laying a foundation for accurately extracting and structuring license information in the subsequent steps.
[0045] In one embodiment, specifically, in step one, the deep learning model uses PaddleOCR, and the text detection model is trained with 500 on-site license pictures to extract text boxes including the upper left coordinates. Specifically, 500 monopoly license pictures with different layouts and different shooting conditions (such as angles, lighting, clarity) are collected from the actual business scenarios to ensure coverage of common misalignments, blurs, stains, etc. Each picture's text box is manually labeled using tools such as LabelImg. The labeled content includes the upper left coordinates (x, y), width, height of the text box, and the text content (for subsequent OCR recognition training). For example, the coordinates of the rectangle box for the "Enterprise Name" field are (100, 150, 200, 40). The original pictures are enhanced by operations such as rotation (±15°), scaling (0.8 - 1.2 times), and brightness adjustment to expand the training samples to more than 2000, improving the model's generalization ability. Based on the PaddleOCR framework, the DB (DifferentiableBinarization) text detection algorithm is selected, and its backbone network uses ResNet50_vd, which has high precision and real-time performance. Loss function: Jointly train the classification loss (identifying text regions) and the regression loss (predicting text box coordinates). Optimizer: Use the Adam optimizer with a learning rate of 0.001 and a cosine annealing decay strategy. Batch size: 8 - 16 (adjusted according to GPU memory). Number of training epochs: 300 - 500 epochs until the validation set loss converges. Evaluation metrics: On the independent test set, the precision and recall of text box detection both reach over 98%, and the average detection time is <100ms per picture.
[0046] The specific process of text box extraction and coordinate acquisition is as follows: Preprocessing: Input the license image and enhance the text edges through operations such as Gaussian blur and histogram equalization; Model inference: Input the preprocessed image into the trained PaddleOCR model and output the four-point coordinates of the text box ; Coordinate conversion: convert the four-point coordinates to the upper left corner coordinates and width and height ; For example, the coordinates of the upper left corner are obtained by calculating the minimum x and minimum y of the four points: , ; Post-processing: Filter out text boxes with too small an area or abnormal aspect ratio, and retain the valid text area; Among the 500 test images, the text box positioning error (compared with manual annotation) accounted for 95% of the images ≤ 2 pixels, ensuring the accuracy of subsequent clustering and matching; for images with a tilt of ≤ 15° and partial occlusion, the text box extraction accuracy remained above 90%; each text box output was format, where x and y are the coordinates of the upper left corner of the text box, w and h are the width and height of the text box, and text is the text content recognized in the text box; for example: ; Through the above steps, the conversion from images to structured text boxes is completed, providing basic data for subsequent clustering, matching, and alignment.
[0047] In one embodiment, specifically, in step 4, the y-axis spacing threshold is set to 50% of the average height of the text boxes, and the contents of adjacent text boxes are merged to process multi-line fields; Specifically, first, calculate the average height of all text boxes in the vertical direction (y-axis); for example, count the text boxes in a batch of monopoly license images and calculate the average height h; Next, set the y-axis spacing threshold to 50% of the average text box height, or 0.5h. This threshold is set because when the y-axis spacing between adjacent text boxes is small, they might actually belong to the same field, but are displayed across multiple lines due to layout reasons. For example, the "Business Location" field has a lot of content and is displayed across two lines in two adjacent text boxes. If these two text boxes are too close together on the y-axis, they are likely to be different lines of the same field. Traverse all text boxes. For each text box, check the spacing in the y-axis direction between it and the adjacent text box below. Assume the bottom y-coordinate of the current text box is y1, and the top y-coordinate of the adjacent text box below is y2. If y2 - y1 ≤ 0.5h, it is determined that these two text boxes belong to multiple lines of content in the same field, and their contents are merged. For example, if the content of a text box is "xx City" and the content of the adjacent text box below is "xx District, xx Road, No. 123", and the y-axis spacing between them meets the threshold condition, they are merged into "xx City, xx District, xx Road, No. 123". Through such operations, the content of the same field originally scattered in multiple text boxes can be integrated, facilitating subsequent accurate matching of field names and corresponding values and structured information extraction, ensuring that information extraction errors will not occur due to text line breaks.
[0048] In one embodiment, specifically, in step three, the text boxes are sorted from top to bottom according to their y-values, arranged in ascending order of the y-values of the text boxes, and the text boxes with a y-value difference less than 30% of the average height are regarded as the same line. Specifically, first, obtain the y-value of the upper left corner coordinate point of each text box. Since in the layout of the license, the text boxes of different fields are usually arranged from top to bottom in sequence, sorting according to the y-value can conform to people's reading and information organization habits. Arrange all text boxes in ascending order of their y-values (from small to large). For example, there are three text boxes T1, T2, and T3, and the y-values of their upper left corner coordinates are y1 = 100, y2 = 120, and y3 = 80 respectively. After sorting, the order becomes T3, T1, T2, which makes the text boxes have an ordered arrangement in the vertical direction. Next, calculate the average height h of all text boxes. When judging whether two text boxes belong to the same line, it depends on the difference in their y-values. If the difference in the y-values of two text boxes is less than 30% of the average height h, it is considered that these two text boxes belong to the same line. For example, if the average height h = 30 pixels, then 30% of h is 9 pixels. If the y-value of text box A is 150 and the y-value of text box B is 155, the difference in their y-values is 5 pixels, which is less than 9 pixels. Therefore, text boxes A and B are regarded as the same line. The advantage of doing this is that in the monopoly license, some fields will be displayed across multiple text boxes on the same line. Through this method, the text boxes on the same line can be accurately associated, providing a more reasonable grouping of text boxes for subsequent operations such as matching field names and corresponding values and structured processing, and avoiding information extraction errors caused by irregular text box distributions.
[0049] In one embodiment, specifically, in step five, the semantic matching data set is constructed based on the ten-field pairs extracted from 2000 on-site license pictures, including enterprise name, person in charge's name, business location, scope of permission, etc. Specifically, the construction of the semantic matching dataset is the key to ensuring accurate field matching; based on 2,000 pictures of monopoly licenses collected on-site, ten core fields such as enterprise name, person in charge name, business location, etc. and their corresponding values are manually labeled and extracted to form field pairs; these pictures cover different regional formats, shooting conditions, and misalignment situations, including scenarios such as blurred printing and stain occlusion, ensuring data diversity; during the extraction process, text boxes are initially recognized through PaddleOCR, and then the OCR misrecognition problems are manually verified and corrected, and the field pairs are matched according to the layout rule of the field name in the left column and the corresponding value in the right column of the license; at the same time, the text is preprocessed such as cleaning, word segmentation, and format normalization, constructing positive and negative samples to enhance the data, and finally the training set, validation set, and test set are divided according to the ratio of 8:1:1, providing data support for the training of the Sentence-BERT model, enabling the model to learn the semantic association between fields and achieve accurate matching.
[0050] In one embodiment, specifically, the method further includes performing semantic coherence verification on the output result, specifically: Calculate the cosine similarity of the word vectors of adjacent fields, and trigger manual review when it is lower than 0.6; Verify the field type format; Check whether the required fields are missing; Specifically, the semantic coherence verification of the output result is a key link to ensure the accuracy of structured information, and it is specifically implemented through the following three steps: Adjacent field semantic association verification, use the Sentence-BERT model to generate the word vectors of adjacent fields (such as "enterprise name" and "person in charge name"), and calculate the cosine similarity; if the similarity is lower than 0.6 (for example, the normal similarity between "enterprise name - XX company" and "person in charge name - Zhang San" is about 0.3, if a certain group of results shows a similarity of 0.7, there is a misaligned field match), it indicates that the semantic logic is abnormal and triggers manual review; for example, if "business location - XX city XX district" is followed by "permission scope - 2025.01.01", the semantic similarity between the two is only 0.2, far lower than the threshold, and it is necessary to manually check whether the time value is mis-matched to the address due to column misalignment; Field format compliance verification, verify according to the preset format rules according to the field type: Date type (such as "valid period"): Verify the format of "YYYY.MM.DD - YYYY.MM.DD" through the regular expression d{4}.d{2}.d{2}-d{4}.d{2}.d{2}; Person name type (such as "person in charge name"): Check whether it contains non-Chinese characters (the middle dot "·" is allowed, such as "Maimaiti·XX"); Number types (such as "license number"): Match the fixed-length rule (such as a 12-character alphanumeric combination); if the "expiration date" is extracted as "2025-01", since it does not conform to the double-date format, it is automatically marked as a formatting error. Check the integrity of required fields, specifying that 6 fields such as "enterprise name", "name of the person in charge", "business location", and "expiration date" are required fields; if the value of any required field in the output result is empty (such as the "name of the person in charge" field fails to match), the system automatically marks it as "missing" and blocks the process to enter the manual completion link to avoid incomplete structured data affecting the entry into the supervision system. Through the above verification mechanism, the field matching error rate can be further reduced from 0.8% to below 0.1%, ensuring that the output license information complies with the supervision business specifications.
[0051] In summary, the method for optimizing the structural information of document content recognition provided by the embodiments of the present invention extracts the text box coordinates and content by using the PaddleOCR deep learning model, combines the K-Means spatial clustering algorithm to allocate the text boxes to the left and right columns according to the coordinates, then sorts the text boxes according to the y value and merges adjacent text boxes with a vertical distance less than 50% of the average height. Then, based on the ten fields extracted from 2000 on-site pictures, the Sentence-BERT semantic matching model is trained, and the cosine similarity between the field name and the candidate value is calculated for initial matching. For the matching rows with a similarity exceeding 0.8, the starting point of the second column is aligned by translating the anchor row. If the matching fails three times in a row, the second column is scrolled to re-match. Finally, the semantic coherence verification is completed by calculating the similarity of adjacent field word vectors, verifying the field format, and checking the required fields, realizing the precise optimization of the misalignment of the license structured information and ensuring the accuracy of field matching and information extraction.
[0052] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, advantages, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, advantages, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purpose of illustration and easy understanding, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details to implement.
[0053] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.
[0054] In addition, as used herein, "or" in the listing of items starting with "at least one" indicates a disjunctive listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Further, the term "exemplary" does not mean that the examples described are preferred or better than other examples.
[0055] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.
[0056] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0057] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0058] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A method for optimizing the structural information of document content recognition, characterized in that, It includes the following steps: Obtain the picture of the exclusive license, and based on the deep learning model, extract the text boxes, their coordinates and contents in all text regions in the picture; Process all the extracted text boxes, delete the text boxes belonging to the license name, and use the K-Means spatial clustering algorithm to assign them to the left column and the right column according to the coordinates of the remaining text boxes; Based on each column of text boxes after clustering, sort the text boxes from top to bottom according to their y values to obtain two columns of text boxes and their text contents arranged in order; Inside each column obtained, judge the vertical distance between adjacent text boxes. If the y-axis spacing between the two is lower than the set threshold, merge their text contents into consecutive parts of the same field; Based on Sentence-BERT, extract field pairs from the license pictures collected on site, construct a semantic matching data set, and train to obtain a semantic matching model; Take the left and right columns as the pairing groups of field names and values, calculate the semantic similarity between the field name and the candidate value, and select the highest one as the initial match; If the similarity exceeds the threshold, translate the second column with the matching row as the anchor point to align the matching starting point; Match the subsequent fields row by row. If the continuous matching fails, roll the second column to re-match, and output the successfully matched field pairs.
2. The method for optimizing structural information identified from the document content according to claim 1, wherein , The specific steps of assigning them to the left column and the right column according to the coordinates of the remaining text boxes are as follows: Step 201: Set the number of clustering centers K = 4, and calculate the range of x - coordinates Δx of all elements in the remaining text box set T, where Δx = max(x i ) - min(x i ). Select the initial clustering centers C = {C1, C2, C3, C4} based on Δx. The x - coordinates of C1 and C2 are in the interval [min(x i ), min(x i ) + Δx / 3], and the x - coordinates of C3 and C4 are in the interval [max(x i ) - Δx / 3, max(x i )]. The y - coordinates of C1 - C4 take the average y - value of the set T; Step 202: For each text box T i ∈ T, calculate the Euclidean distance from its upper left coordinate point (x i , y i ) to each cluster center C k : ; Among them, (x k , y k ) is the coordinate of the clustering center C k ; Step 203: Assign the text box T i to the cluster S corresponding to the cluster center with the minimum distance k , and update the cluster center coordinates to be: ; Among them, |S k | represents the number of text boxes in cluster S k ; Step 204: Repeat Step 202 - Step 203 until the change in the cluster center coordinates is less than 0.5 pixels or the number of iterations exceeds 50 times. Assign the finally converged cluster S1∪S2 to the left column and S3∪S4 to the right column.
3. The method for optimizing structural information identified from the document content according to claim 1, characterized in that: The calculation of the semantic similarity between the field name and the candidate value uses the cosine similarity formula: ; where a and b are the sentence vectors generated by Sentence-BERT.
4. The method for optimizing structural information identified from the document content according to claim 1, characterized in that: The operation of scrolling the second column is to update the text box sequence to . If the matching fails three times in a row, scrolling is triggered, and the semantic similarity is recalculated after each scroll.
5. The method for optimizing structural information identified from the document content according to claim 1, characterized in that: The similarity threshold is 0.
8. If the matching is successful, translate the second column with the anchor row to align the matching starting points of the two columns.
6. The method for optimizing structural information identified from the document content according to claim 1, characterized in that: [[ID= 7. The method for optimizing structural information identified from the document content according to claim 1, characterized in that: 8. The method for optimizing structural information identified from the document content according to claim 1, characterized in that: 9. The method for optimizing structural information identified from the document content according to claim 1, wherein: 10. The method for optimizing structural information identified from the document content according to claim 1, characterized in that:
Citation Information
Patent Citations
Character extraction method and device for bill image
CN114612922A
Medical document semantic entity recognition method and system based on spatial semantic enhancement
CN119888761A