Cross-border trade document image intelligent identification and data extraction method and system
By adaptively adjusting the gradient magnitude and information entropy difference in cross-border trade document images, the instability of binarization threshold recognition in complex backgrounds is solved, improving the reliability of text recognition and the accuracy of data extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO HENGXING UNIV OF SCI & TECH
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies in cross-border trade document image processing lack effective adaptive adjustment of the binarization threshold, making it difficult to obtain reliable text recognition results under complex interference conditions, which affects the automated application of document data in actual business scenarios.
Interference field regions are identified by calculating the statistical variance of the gradient magnitude of the target field region, and the binarization threshold is adaptively adjusted based on the information entropy difference value. A closed-loop iterative process is constructed to optimize the threshold, ensuring the stability and reliability of the identification results.
It improves the reliability and determinism of document image text recognition, reduces the interference of complex backgrounds on recognition results, and achieves efficient recognition and data extraction in complex environments.
Smart Images

Figure CN121921785A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document image recognition and extraction technology, and more specifically, this application relates to a method and system for intelligent recognition and data extraction of cross-border trade document images. Background Technology
[0002] In cross-border trade, document images serve as crucial carriers of cargo information, transaction details, and transit documentation. The accuracy of text recognition directly impacts data extraction and processing efficiency. Current technologies for text recognition in document images typically rely on pre-defined image segmentation parameters for uniform image processing. The binarization threshold, a key factor influencing text-to-background separation, is often set using fixed rules or empirical methods. However, significant differences exist in the generation, scanning, and transmission of cross-border trade documents, making it difficult to maintain consistent character distribution, grayscale variations, and background complexity in document images. This makes fixed or weakly adaptive binarization methods unsuitable for ensuring recognition stability across various document scenarios in practical applications.
[0003] In complex business scenarios, document images often contain overlapping seals, background textures, and table lines that obscure the image. These factors can cause grayscale competition between characters and the background in local areas. Existing technologies often struggle to effectively assess the validity and reliability of recognition results, and invalid character fragments caused by noise or background textures are easily mixed into the recognition output. Due to the lack of quantitative criteria for judging the rationality of the internal structure of the recognition results, the system often cannot distinguish between real and valid text and random interference information. This leads to a lack of clear optimization direction in the binarization threshold selection process, resulting in significant uncertainty in the threshold adjustment process and substantial fluctuations in recognition quality depending on image conditions.
[0004] Furthermore, existing technologies, when optimizing recognition results, typically focus only on the effect of a single processing step, lacking the ability to analyze the overall parameter adjustment process itself. When the binarization threshold changes, different regions exhibit significant differences in their response to the threshold change. Existing methods struggle to reflect the complexity of the image content through parameter changes, leading to the continued use of the original region for processing even when characters are not fully rendered or the region coverage is insufficient, thus further amplifying recognition errors. When significant anomalies appear in the recognition results, existing technologies also lack a clear judgment mechanism to distinguish between "situations that can be improved through further adjustments" and "situations that cannot be eliminated through automatic processing," making it difficult for the system to make timely and reasonable decisions.
[0005] To address the aforementioned issues, there is an urgent need in this field for a technical solution that can adaptively adjust around the binarization threshold and combine the inherent stability and consistency characteristics of the recognition results to effectively judge and process interference areas, so as to improve the reliability and determinism of text recognition and processing in complex image environments for cross-border trade documents.
[0006] In cross-border trade document image processing, existing technologies lack effective adaptive adjustment criteria for binarization thresholds, making it difficult to obtain stable and reliable text recognition results under complex interference conditions. This, in turn, affects the automated application of document data in actual business scenarios. Summary of the Invention
[0007] To address the aforementioned technical issues, this paper provides a method and system for intelligent recognition and data extraction of cross-border trade document images. This technical solution resolves the problems mentioned in the background section.
[0008] In a first aspect, embodiments of this application provide a method for intelligent recognition and data extraction of cross-border trade document images, comprising the following steps: S1, acquiring a target document image, and using an optical character recognition algorithm to recognize at least one target field region's grayscale image and first text information based on a first binarization threshold; S2, calculating the statistical variance of the gradient magnitude in the grayscale image of the target field region, and if the statistical variance is greater than a preset variance threshold, then the target field region is recorded as an interference field region; S3, based on the difference between the statistical variance of the interference field region and the preset variance threshold, generating a first adjustment amount and adjusting the first binarization threshold accordingly to obtain a second binarization threshold; S4, using an optical character recognition algorithm to extract data based on the second binarization threshold... S5. The information entropy difference between the first and second text information is calculated. If the information entropy difference is less than zero and its absolute value is greater than a preset negative threshold, the difference between the first adjustment amount and the preset adjustment step size is subtracted and the sign is reversed to obtain the second adjustment amount. The first binarization threshold is then adjusted accordingly to obtain a new second binarization threshold. S6. Steps S4 to S5 are repeated until the information entropy difference obtained after adjustment is greater than zero or the number of iterations reaches the maximum threshold. The iteration is then terminated, and the corresponding second binarization threshold is obtained and recorded as the optimal binarization threshold. S7. The third text information is obtained by recognizing the target document image using an optical character recognition algorithm based on the optimal binarization threshold and then output.
[0009] Secondly, embodiments of this application provide a cross-border trade document image intelligent recognition and data extraction system, including: a data acquisition module: used to acquire a target document image, and to recognize at least one target field region grayscale image and first text information by using an optical character recognition algorithm according to a first binarization threshold; an interference field region acquisition module: used to calculate the statistical variance of the gradient magnitude in the grayscale image of the target field region, and if the statistical variance is greater than a preset variance threshold, the target field region is recorded as an interference field region; a binarization threshold first adjustment module: used to generate a first adjustment amount based on the difference between the statistical variance and the preset variance threshold, and to perform a first adjustment on the first binarization threshold accordingly to obtain a second binarization threshold; and a text processing module: used to process the interference field region according to the second binarization threshold by using an optical character recognition algorithm. The second text information is obtained through domain recognition; the second binarization threshold adjustment module is used to calculate the information entropy difference between the first and second text information. If the information entropy difference is less than zero and its absolute value is greater than a preset negative threshold, the difference between the first adjustment amount and the preset adjustment step size is subtracted and the sign is reversed to obtain the second adjustment amount. Based on this, the first binarization threshold is adjusted to obtain a new second binarization threshold; the threshold iteration module is used to repeatedly execute the text information processing module and the second binarization threshold adjustment module until the information entropy difference obtained after adjustment is greater than zero or the number of iterations reaches the maximum threshold, then the iteration is terminated, and the corresponding second binarization threshold is obtained, which is recorded as the optimal binarization threshold; the text output module is used to recognize the target document image according to the optimal binarization threshold through the optical character recognition algorithm to obtain the third text information and output it.
[0010] Thirdly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for intelligent recognition and data extraction of cross-border trade document images.
[0011] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0012] 1. By calculating the gradient magnitude statistical variance of the target field region and identifying the interference field region, the binarization threshold adjustment is no longer based on a uniform parameter across the entire image, but focuses on local areas with significant grayscale fluctuations. This reduces the interference of complex backgrounds such as seals, backgrounds, and table lines on the optical character recognition results and improves the effectiveness of the first text information.
[0013] 2. By constructing the information entropy difference value between the first text information and the second text information, and using this information entropy difference value as the feedback basis for binarization threshold adjustment, the iterative process of binarization threshold has a clear optimization direction, avoiding invalid oscillations during threshold adjustment, thereby obtaining the optimal binarization threshold within a limited number of iterations, and improving the determinism and convergence stability of the recognition process.
[0014] 3. By analyzing the inflection points of the changes in the second binarization threshold iteration sequence and the information entropy difference iteration sequence, and generating different region adjustment coefficients based on the differences in inflection points, different judgment extractions are performed respectively. This can distinguish between the normal situation where parameter changes and recognition results change synchronously and the abnormal situation where parameter changes and recognition results change asynchronously. This enables accurate localization and classification of potential abnormal regions in document images, improving the credibility and interpretability of the overall recognition results. Attached Figure Description
[0015] Figure 1 A schematic diagram illustrating the steps of the intelligent recognition and data extraction method for cross-border trade document images provided in this application embodiment;
[0016] Figure 2 A schematic diagram of the logical flow for threshold and information entropy change trend analysis provided in the embodiments of this application;
[0017] Figure 3 This is a schematic diagram of the structure of the intelligent recognition and data extraction system for cross-border trade documents provided in the embodiments of this application. Detailed Implementation
[0018] This application embodiment solves the technical problem of insufficient reliability in separating and extracting characters from interference backgrounds in cross-border trade document image processing by using a method and system for intelligent recognition and data extraction of cross-border trade document images.
[0019] In the automated processing of cross-border trade documents, the instability of text recognition results stems not from the recognition model itself, but from the difficulty in reliably separating characters from the background in an image. Due to the high degree of inconsistency in document origin, generation method, and image quality, uniformly set image segmentation parameters often only apply to certain scenarios. When there is stamp overlay, background interference, or drastic fluctuations in grayscale in local areas, the boundary between characters and background becomes significantly blurred, directly affecting the reliability of the recognition results. Therefore, the key to document image processing is not obtaining a recognition result in one go, but rather establishing a processing logic around image segmentation parameters that can dynamically adjust according to changes in image content.
[0020] Based on the aforementioned issues, this solution first focuses on the varying complexity of different regions within the document image. By analyzing the image features of the target field region, it identifies areas more susceptible to background interference. These regions typically exhibit significant instability in grayscale changes. If the same segmentation parameters as other regions are used, it can easily lead to fragmented character structures or the incorrect introduction of background information. Therefore, it is necessary to process these regions separately and use them as the primary focus for subsequent parameter adjustments.
[0021] When processing interference regions, relying solely on experience or fixed rules to adjust image segmentation parameters is insufficient to guarantee the correctness of the adjustment direction. This scheme further introduces an evaluation of the stability of the recognition results themselves. By analyzing the changes in the recognized text under different parameter conditions, it determines whether the current parameter adjustments are moving in a direction that is conducive to the presentation of effective characters.
[0022] When the character distribution in the recognition result is more reasonable and the structure is more complete, it indicates that the parameter adjustment direction is effective; otherwise, the adjustment direction and magnitude need to be corrected. In this way, the parameters are no longer passively set, but gradually converge based on the feedback from the recognition results, thereby avoiding disordered or blind repeated attempts.
[0023] During parameter adjustment, different regions showed varying degrees of response to changes in segmentation parameters. Some regions exhibited significant fluctuations even with minor parameter changes, reflecting a strong competitive relationship between the characters and the background. This variation itself contains important information about the image structure.
[0024] This solution identifies highly sensitive regions by analyzing the trends of relevant data during parameter adjustment, and determines whether the current region has completely covered the interfered character content. When the analysis results indicate that the original region is insufficient to support stable recognition, this solution adjusts the region appropriately to better match the actual character distribution, thereby avoiding amplification of recognition errors due to improper region selection.
[0025] After regional adjustments, this solution does not simply accept the new identification results, but further focuses on the consistency between the identified content before and after the adjustment. When the optimized identification results can completely cover the original valid information and the overall stability is improved, it indicates that the current adjustment is effective, and the processing flow can continue. If the identification results still show significant inconsistencies or instability after multiple adjustments, it means that the area has exceeded the range that automatic processing can reliably cover. In this case, continuing to adjust parameters or regions will not only fail to improve the results but may also introduce more uncertainty. This solution explicitly identifies this state, precisely pinpointing the problem area, thereby achieving a smooth transition from automatic optimization to explicit decision output.
[0026] Based on the above logic, this solution organically combines the adjustment of image segmentation parameters, the stability assessment of recognition results, and the dynamic correction of the region range in complex document image environments. This enables the entire data processing process to continuously self-correct and gradually converge around the actual image content, thereby improving the reliability and controllability of cross-border trade document text recognition results while ensuring processing efficiency.
[0027] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0028] like Figure 1 The diagram shown is a structural schematic of the intelligent recognition and data extraction method for cross-border trade document images provided in this application embodiment. The intelligent recognition and data extraction method for cross-border trade document images includes the following steps: S1, acquiring a target document image, and using an optical character recognition algorithm to recognize at least one target field region's grayscale image and first text information based on a first binarization threshold; S2, calculating the statistical variance of the gradient magnitude in the grayscale image of the target field region, and if the statistical variance is greater than a preset variance threshold, then the target field region is recorded as an interference field region; S3, based on the difference between the statistical variance of the interference field region and the preset variance threshold, generating a first adjustment amount and adjusting the first binarization threshold accordingly to obtain a second binarization threshold; S4, using an optical character recognition algorithm to... The recognition algorithm identifies the second text information by recognizing the interference field region based on the second binarization threshold; S5, calculate the information entropy difference between the first and second text information. If the information entropy difference is less than zero and its absolute value is greater than the preset negative threshold, then subtract the preset adjustment step size from the first adjustment amount and reverse the sign to obtain the second adjustment amount. Based on this, the first binarization threshold is adjusted to obtain a new second binarization threshold; S6, repeat steps S4 to S5 until the information entropy difference obtained after adjustment is greater than zero or the number of iterations reaches the maximum threshold, then terminate the iteration and obtain the corresponding second binarization threshold, which is recorded as the optimal binarization threshold; S7, the optical character recognition algorithm recognizes the target document image based on the optimal binarization threshold to obtain the third text information and outputs it.
[0029] Step S1 is used to establish the initial input conditions for the entire recognition process. By performing grayscale and initial binarization processing on the target document image, the optical character recognition algorithm can output the first text information and the corresponding grayscale image of the target field area, thereby providing a basis for subsequent judgment of field quality.
[0030] The first binarization threshold can be obtained based on the statistical distribution of brightness of common paper types in historical samples.
[0031] Step S2 is used to identify field regions that may contain noise, ghosting, or background interference. The statistical variance of the gradient magnitude in the grayscale image is calculated to reflect the severity of image edge changes. A larger statistical variance indicates higher texture complexity in the region, making it more likely to affect character segmentation. For example, when a field region has a red stamp superimposed on it, its gradient magnitude statistical variance is significantly higher than that of a plain text region. If it exceeds a preset variance threshold obtained through sample statistics, the region is marked as an interference field region.
[0032] Step S3 is used to directionally correct the binarization threshold according to the degree of interference, thereby improving the separation effect between the character foreground and background. The first adjustment value reflects the degree to which the current statistical variance deviates from the normal level, and its magnitude and sign determine the direction of adjustment of the binarization threshold. For example, when the statistical variance is significantly greater than the preset variance threshold, the first adjustment value is positive, indicating that the binarization threshold is appropriately increased to suppress background noise.
[0033] Step S4 is used to re-identify the interference field region under the updated binarization threshold, thereby obtaining second text information to evaluate the impact of threshold adjustment on the recognition result. For example, in the stamp interference region, after re-identification by increasing the binarization threshold, the character outline is clearer, but some strokes may be missing.
[0034] Step S5 quantifies the change in information uncertainty between the two recognition results using the information entropy difference value. A decrease in information entropy usually indicates a more stable text structure. When the information entropy difference value is less than zero and the magnitude exceeds a preset negative threshold, it indicates that the current adjustment direction may lead to information loss. Therefore, it is necessary to reverse the adjustment direction and introduce a preset adjustment step size for finer adjustment. For example, if raising the threshold for the first time causes stroke loss and increases information entropy, the threshold is lowered by reversing the adjustment to restore character integrity.
[0035] Step S6 constructs a closed-loop adjustment process based on information entropy feedback, gradually approaching the binarization threshold that optimizes the stability of text information through multiple iterations. The maximum threshold is used to prevent infinite iteration. For example, after multiple fine-tunings, when the information entropy difference value turns positive for the first time, it indicates that the information structure of the recognition result is better than the initial state, and the current threshold can be determined as the optimal binarization threshold.
[0036] Step S7 applies the optimal binarization threshold globally to obtain a consistent recognition result, ensuring that the output third-party text information is superior to the initial recognition result in both completeness and stability. For example, applying the optimal binarization threshold to the entire invoice image can simultaneously improve the recognition quality of multiple fields.
[0037] Furthermore, step S6 also includes: constructing a binary threshold iteration sequence from the second binarized threshold corresponding to each iteration during the iteration process, and constructing an information entropy difference iteration sequence from the corresponding information entropy difference values; this is used to record the entire process data of threshold adjustment and information entropy change, providing a basis for subsequent trend analysis. For example, a binary threshold iteration sequence and an information entropy difference iteration sequence of length five are formed in five iterations.
[0038] Curve fitting is performed on the binarized threshold iteration sequence and the information entropy difference iteration sequence respectively to obtain the binarized threshold change curve and the information entropy difference change curve. Curve fitting smooths the discrete iterative data, facilitating the analysis of changing trends. Existing techniques such as polynomial fitting or spline fitting can be used. For example, quadratic polynomial fitting can eliminate the influence of single abnormal fluctuations.
[0039] Calculate the first derivatives of the binarization threshold change curve and the information entropy difference change curve respectively, and extract the number of iterations corresponding to the maximum value of the absolute value of the first derivative. These are used as the number of iterations for the inflection point of the binarization threshold change and the number of iterations for the inflection point of the information entropy difference change, respectively, and are denoted as the first inflection point and the second inflection point. This is used to identify the position where the change is most drastic during the adjustment process. The inflection point reflects the moment when the system state undergoes a significant change, such as from the rapid convergence stage to the steady stage.
[0040] Determine whether the difference between the first inflection point and the second inflection point is within the preset inflection point difference threshold range; this is used to assess the synchronicity between threshold changes and information entropy changes. The preset inflection point difference threshold can be determined by the average deviation of the two types of inflection points in historical samples.
[0041] If so, a first region adjustment coefficient is generated based on the ratio of the average of the first and second inflection points to the total number of iterations, and a first extended region is generated accordingly. The first extended region analysis and judgment are then performed to obtain the first extended region judgment result. This result is used to verify the rationality of field boundaries by employing a directional region expansion strategy when the threshold and information entropy changes are highly consistent. For example, when two inflection points appear almost simultaneously, it indicates that there may be a slight truncation at the field boundary.
[0042] If not, based on the ratio of the absolute difference between the first and second inflection points to the total number of iterations, a second region adjustment coefficient is generated, and a second extended region is generated accordingly. The second extended region analysis and judgment are then performed to obtain the second extended region judgment result. This is used to employ a center-expanding strategy to capture potentially missed character regions when the threshold and information entropy changes asynchronously.
[0043] In this embodiment, Figure 2 This is a schematic diagram of the logical flow for threshold and information entropy change trend analysis provided in the embodiments of this application.
[0044] Based on the dynamic changes in threshold and entropy during the iteration process, the potential nature of the interference is inferred, and subsequent regional processing strategies are intelligently selected.
[0045] The inflection points of the binarization threshold and information entropy appear in similar iteration intervals, indicating that the grayscale distribution of the interference area is relatively uniform and continuous. When the threshold changes, the information entropy of the entire area responds synchronously. This usually corresponds to overall degradation, such as slight blurring, uneven global illumination, or uniform light-colored stains, without introducing strong heterogeneous noise.
[0046] A preset threshold range for inflection point differences, such as [0,1], is used to determine whether two inflection points appear essentially synchronously. For example, considering the three iterations of data in the "total amount" region mentioned above, polynomial fitting and differentiation analysis revealed that the threshold curve exhibits the largest rate of change in the second iteration, with the first inflection point Iter_T=2. The entropy difference curve exhibits the largest rate of change in the third iteration, with the second inflection point Iter_H=3. The absolute difference between the two is |2-3|=1, falling within the preset range [0,1], therefore it is determined that the inflection points are "essentially synchronous."
[0047] Since the problem may stem from inaccurate delineation of the region boundary, a "one-sided expansion detection" approach is adopted. After expanding in one direction, if the recognition quality significantly improves, it proves that the expansion direction is correct and the original region is too small. At this point, the strategy of delineating the new region in reverse, using the original boundary as a benchmark, is an efficient binary search process. This aims to gradually approach and isolate the incorrectly identified sub-regions that are truly causing the decrease in information entropy, until the "hard" problem region that cannot be repaired by adjusting the threshold is located.
[0048] Inflection point separation indicates that the inflection points of the binarization threshold and information entropy occur at different iteration stages, suggesting the existence of significantly heterogeneous grayscale components within the interference region. For example, a dark stamp covers part of the characters. In the early stages of iteration, threshold changes primarily affect the separation of the stamp from the background; only as the threshold continues to be adjusted does it begin to significantly affect the visibility of the covered characters. The separation degree between the two inflection points quantifies the "distance" between the interference object and the main character in the grayscale space.
[0049] If the two inflection points are far apart, for example, Iter_T=5 and Iter_H=10, and the difference of 5 is not within the preset range, it means that the drastic change in the threshold and the drastic change in the entropy value occur asynchronously, which may indicate the existence of isolated outliers in the region.
[0050] At this point, the initial region may simultaneously contain normal characters, distractors, and interfered characters. Simple unilateral expansion might include both distractors and subsequent characters at once. Therefore, a "global expansion followed by fine-tuning" approach is adopted. First, the coefficients for the second region are adjusted, their size being positively correlated with the inflection point separation, to expand the region and ensure complete context capture. Then, the region is gridded, and by detecting the "low points" of local information entropy and utilizing their spatial continuity, character strips that are contaminated by the stamp and cannot be reliably identified are accurately selected without mistakenly reporting normal character regions as well.
[0051] Furthermore, the specific execution process of the first extended region analysis and judgment is as follows: taking the first side of the outer rectangle of the interference field region as the reference side, the outer rectangle is extended along a preset direction perpendicular to the reference side, according to the first region adjustment coefficient, to obtain the first extended region; this is used to perform directional extension in the main extension direction of the field to reduce the introduction of irrelevant regions. For example, a slight vertical extension is performed on horizontally arranged fields.
[0052] The first extended region is identified using an optical character recognition algorithm based on the optimal binarization threshold to obtain the fourth text information, which is used to verify whether there is any valuable supplementary text within the extended region.
[0053] Determine whether the fourth text information completely contains the third text information and whether the difference in information entropy between the two is less than a preset threshold; determine the rationality of the extension by considering both the text inclusion relationship and the stability of information entropy.
[0054] If so, the first extended region is determined to be a normal region, and the reference edge is translated in the opposite direction of the preset direction according to the preset translation step size, and the newly defined region is used as the new interference field region, and the process returns to step S4; this is used to fine-tune the field positioning after confirming that there is no abnormality in the extension, so as to further optimize the boundary.
[0055] If not, the first extended region is determined to be an abnormal region.
[0056] In this embodiment, pseudo-"interference" caused by slight deviations in the initial region detection box positioning is specifically addressed. The reference edge can be any of the four sides of a rectangle, with a preset direction being the outward direction of its normal. For example, for the "Total Amount" region, its upper edge is selected as the reference edge, and the upward extension distance is the length of the upper edge. The first expanded region is then identified using the determined optimal threshold optical character recognition algorithm = 130, yielding the fourth text information, possibly "TotalAmount: 1,250.00". Next, it is determined that the fourth text information "TotalAmount: 1,250.00" completely contains the third text information "1,250.00" extracted from the final recognition result of the entire image, and the difference in information entropy between the two is within a preset small threshold. This indicates that the expanded region is continuous and normal text, and the original region positioning may have segmented complete semantic units due to deviations. Therefore, the first expanded region is determined to be a normal region. Subsequently, the reference edge is shifted in the opposite direction of the preset direction by a preset step size to obtain a new rectangular region that is slightly larger than the original region but still smaller than the first extended region. This new rectangular region is then used as the new interference field region, and the process returns to step S4 of claim 1, where it is identified and verified using the optimal binarization threshold. This is essentially a fine-tuning calibration of the region boundary.
[0057] Furthermore, the specific execution process of the second extended region analysis and extraction is as follows: taking the center of the outer rectangle of the interference field region as the origin, the length and width of the outer rectangle are simultaneously extended outward according to the second region adjustment coefficient to obtain the second extended region; this is used for symmetrical expansion when the direction of interference is uncertain.
[0058] The second extended region is identified using an optical character recognition algorithm based on the optimal binarization threshold to obtain the fifth text information; the character set difference between the third and fifth text information is extracted; and it is determined whether there are characters in the extended region that were not included in the original recognition result.
[0059] If the difference in the character set is not empty, the second extended region is used as the current search region, and search segmentation and extraction are performed on the current search region to obtain the judgment result of the current search region; otherwise, the adjustment coefficient of the second region continues to increase until the difference in the character set is not empty, or the adjustment coefficient of the second region increases to the preset upper limit, or the extended region reaches any boundary of the target document image, at which point the extension is terminated; the search range is controlled by step-by-step extension to avoid over-extension.
[0060] If the character set difference remains empty when the adjustment coefficient of the second region increases to the preset upper limit or when the expanded region reaches any boundary of the target document image, then the second expanded region is determined to be an abnormal region. This is used to give a clear abnormal conclusion when no new characters can be found.
[0061] In this embodiment, the aim is to detect and capture discrete interference elements that may be located at the edge or inside the original region. The expansion is performed uniformly outwards; for example, in another "product description" field, there might be a small ink dot inside. After expansion, the fifth text information is identified using an optical character recognition algorithm. The character set difference between the third and fifth text information is calculated—that is, the set of characters that appear in the fifth text but not in the third text; if the result is empty, it indicates that the expansion did not capture any new valid or interfering characters. At this point, expansion is performed incrementally with a fixed step size, and recognition and character set difference calculation are performed after each expansion. Suppose that the expanded area contains sporadic characters from another adjacent field, causing the character set difference between the third and fifth text information to be non-empty. In this case, expansion is terminated, and the largest expanded area is marked as the current search area, entering the subsequent search, segmentation, and extraction process to accurately locate the interference source. If the character set difference remains empty even after expansion to the image boundary, it may mean that the original interference is uniform or that recognition has reached its limit; this is then identified as an abnormal region and recorded.
[0062] Furthermore, after returning to step S4, if the analysis and judgment of the first extended region are all entered and all are determined to be normal regions after a preset number of consecutive loops, and a normal information report is output, then step S4 is not returned and the redefined region is recorded as the final region. The text information is obtained by recognizing the final region using the optical character recognition algorithm based on the optimal binarization threshold and then output.
[0063] In this embodiment, a termination condition is set for the boundary fine-tuning process to prevent infinite loops caused by boundary judgment logic. A preset number of loops is used as a safety threshold. When such a "fine-tuning-verification-normal determination" cycle occurs continuously for the preset number of loops, it is considered that a stable and reliable optimal region range has been found. At this point, the iteration stops, and the region defined after the last translation is marked as the final region. Subsequently, the final region is identified using the optimal binarization threshold, and the accurate text obtained is output as part of the normal information report. This ensures the highest recognition accuracy even when field labels and content are extracted as a whole.
[0064] Furthermore, the specific execution process of search segmentation and extraction is as follows: the current search region is divided into equal-width sub-region strips along the length of its circumscribed rectangle; the text content of each sub-region strip is identified, and its information entropy is calculated; sub-region strips with information entropy lower than a preset low-entropy threshold and whose adjacent sub-region strips have normal information entropy are marked as suspected problem strips; morphological closing operations are performed on all suspected problem strips to obtain suspected problem regions, which are then determined to be abnormal regions, and an abnormal information report is output; the remaining regions of the current search region other than suspected problem regions are determined to be normal regions, and a normal information report is output.
[0065] In this embodiment, the preset number of segments can be dynamically determined based on the region width, for example, one segment every 20 pixels.
[0066] A preset low-entropy threshold is used to identify stripes with abnormal information content. For example, if the current search area is 200 pixels wide, it is horizontally divided into 10 sub-strips, each 20 pixels wide. Each strip is identified, and its information entropy is calculated. Assume stripes 1-4 have normal entropy values, stripes 5-6 have significantly lower entropy values corresponding to ink dots, and stripes 7-10 have abnormal entropy values corresponding to blank spaces or table lines. According to the rules, stripes 5 and 6 are marked as suspected problem stripes because their entropy values are below the low-entropy threshold and their adjacent stripe 4 has a normal entropy value. Then, a morphological closing operation is performed on these discrete suspected stripe images to connect them into a connected component, i.e., the suspected problem region, which precisely corresponds to the ink dots and the range of misidentified text images they cause.
[0067] The abnormal area is removed from the current search area, and the remaining part is judged as a normal area. A normal information report containing accurate text is output, thus achieving precise splitting and cleaning of mixed content areas.
[0068] Furthermore, the anomaly report specifically includes: the precise coordinates of the anomaly region, the original image slice corresponding to the anomaly region, and textual information and the optimal binarization threshold from all iterations. This defines the specific content of the anomaly report to ensure that the problem is traceable and analyzable.
[0069] Furthermore, the normal information report specifically includes: text information obtained by recognizing normal regions using an optical character recognition algorithm based on an optimal binarization threshold. This limits the simplicity of the normal information report, meaning it only outputs the final, verified, and accurate results.
[0070] Figure 3This is a schematic diagram of the structure of the intelligent recognition and data extraction system for cross-border trade document images provided in this application embodiment. The intelligent recognition and data extraction system for cross-border trade document images includes: a data acquisition module: used to acquire a target document image, and to recognize at least one grayscale image of a target field region and first text information by using an optical character recognition algorithm based on a first binarization threshold; an interference field region acquisition module: used to calculate the statistical variance of the gradient magnitude in the grayscale image of the target field region, and if the statistical variance is greater than a preset variance threshold, the target field region is recorded as an interference field region; a binarization threshold first adjustment module: used to generate a first adjustment amount based on the difference between the statistical variance and the preset variance threshold, and to perform a first adjustment on the first binarization threshold accordingly to obtain a second binarization threshold; and a text processing module: used to process the text by using an optical character recognition algorithm based on the second... The binarization threshold is used to identify the interference field region and obtain the second text information; the second binarization threshold adjustment module is used to calculate the information entropy difference between the first and second text information. If the information entropy difference is less than zero and its absolute value is greater than a preset negative threshold, the difference between the first adjustment amount and the preset adjustment step size is subtracted and the sign is reversed to obtain the second adjustment amount. Based on this, the first binarization threshold is adjusted to obtain a new second binarization threshold; the threshold iteration module is used to repeatedly execute the text information processing module and the second binarization threshold adjustment module until the information entropy difference obtained after adjustment is greater than zero or the number of iterations reaches the maximum threshold, then the iteration is terminated, and the corresponding second binarization threshold is obtained, which is recorded as the optimal binarization threshold; the text output module is used to identify the target document image according to the optimal binarization threshold using an optical character recognition algorithm to obtain the third text information and output it.
[0071] This application also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements a method for intelligent recognition and data extraction of cross-border trade document images.
[0072] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0073] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0074] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0075] These computer program instructions can also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0076] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.
[0077] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for intelligent recognition and data extraction of cross-border trade document images, characterized in that, Includes the following steps: S1. Obtain the target document image, and use an optical character recognition algorithm to recognize the target document image according to the first binarization threshold to obtain a grayscale image of at least one target field region and the first text information; S2. Calculate the statistical variance of the gradient magnitude in the grayscale image of the target field region. If the statistical variance is greater than the preset variance threshold, the target field region is recorded as the interference field region. S3. Based on the difference between the statistical variance of the interference field region and the preset variance threshold, generate a first adjustment amount and adjust the first binarization threshold accordingly to obtain a second binarization threshold. S4. The second text information is obtained by identifying the interference field region using an optical character recognition algorithm based on the second binarization threshold. S5. Calculate the information entropy difference between the first text information and the second text information. If the information entropy difference is less than zero and its absolute value is greater than the preset negative threshold, then subtract the difference of the preset adjustment step size from the first adjustment amount and reverse the positive and negative values to obtain the second adjustment amount. Based on this, the first binarization threshold is adjusted in the second way to obtain a new second binarization threshold. S6. Repeat steps S4 to S5 until the information entropy difference value obtained after adjustment is greater than zero or the number of iterations reaches the maximum threshold, then terminate the iteration and obtain the corresponding second binarization threshold, which is recorded as the optimal binarization threshold. S7. The third text information is obtained by recognizing the target document image using an optical character recognition algorithm based on the optimal binarization threshold and then output.
2. The method for intelligent recognition and data extraction of cross-border trade document images according to claim 1, characterized in that, Step S6 further includes: The second binarization threshold corresponding to each iteration in the iteration process is constructed into a binarization threshold iteration sequence, and the corresponding information entropy difference value is constructed into an information entropy difference iteration sequence. Curve fitting was performed on the binarization threshold iteration sequence and the information entropy difference iteration sequence respectively to obtain the binarization threshold change curve and the information entropy difference change curve; Calculate the first derivatives of the binarization threshold change curve and the information entropy difference change curve respectively, and extract the number of iterations corresponding to the maximum value of the absolute value of the first derivative. These are used as the number of iterations for the inflection point of the binarization threshold change and the number of iterations for the inflection point of the information entropy difference change, and are denoted as the first inflection point and the second inflection point respectively. Determine whether the difference between the first inflection point and the second inflection point is within the preset inflection point difference threshold range; If so, then based on the ratio of the average of the first inflection point and the second inflection point to the total number of iterations, a first region adjustment coefficient is generated and a first extended region is generated accordingly. The first extended region analysis and judgment are then performed to obtain the first extended region judgment result. If not, based on the ratio of the absolute difference between the first inflection point and the second inflection point to the total number of iterations, a second region adjustment coefficient is generated, and a second extended region is generated accordingly. The second extended region analysis and judgment are then performed to obtain the second extended region judgment result.
3. The method for intelligent recognition and data extraction of cross-border trade document images according to claim 2, characterized in that, The specific execution process of the analysis and judgment of the first extended region is as follows: Using the first side of the outer rectangle of the interference field region as the reference side, the outer rectangle is extended along a preset direction perpendicular to the reference side, and the first extended region is obtained by adjusting the coefficient of the first region. The first extended region is identified using an optical character recognition algorithm based on the optimal binarization threshold to obtain the fourth text information; Determine whether the fourth text information completely contains the third text information and whether the difference in information entropy between the two is less than a preset threshold. If so, the first extended region is determined to be a normal region, and the reference edge is translated in the opposite direction of the preset direction according to the preset translation step size, and the newly defined region is used as the new interference field region, and the process returns to step S4. If not, the first extended region is determined to be an abnormal region.
4. The method for intelligent recognition and data extraction of cross-border trade document images according to claim 2, characterized in that, The specific execution process of the second extended region analysis and extraction is as follows: Using the center of the outer rectangle of the interference field region as the origin, the length and width of the outer rectangle are simultaneously extended outward according to the second region adjustment coefficient to obtain the second extended region. The second extended region is identified using an optical character recognition algorithm based on the optimal binarization threshold to obtain the fifth text information; Extract the character set difference between the third and fifth text information; If the difference in the character set is not empty, the second extended region is used as the current search region, and search segmentation and extraction are performed on the current search region to obtain the judgment result of the current search region; otherwise, the adjustment coefficient of the second region continues to increase until the difference in the character set is not empty or the adjustment coefficient of the second region increases to the preset upper limit or the extended region reaches any boundary of the target document image, at which point the extension is terminated. If the character set difference is still empty when the adjustment coefficient of the second region increases to the preset upper limit or when the extended region reaches any boundary of the target document image, then the second extended region is determined to be an abnormal region.
5. The method for intelligent recognition and data extraction of cross-border trade document images according to claim 3, characterized in that, After returning to step S4, if the analysis and judgment of the first extended region are all entered and all are determined to be normal regions after a preset number of consecutive loops, and a normal information report is output, then step S4 will not be returned and the redefined region will be recorded as the final region. The text information will be obtained by recognizing the final region using the optical character recognition algorithm based on the optimal binarization threshold and then output.
6. The method for intelligent recognition and data extraction of cross-border trade document images according to claim 4, characterized in that, The specific execution process of the search segmentation and extraction is as follows: Divide the current search region into equal-width sub-region strips along the length of its bounding rectangle, according to a preset number of segments; Identify the text content of each sub-region strip and calculate its information entropy; Sub-region strips whose information entropy is lower than a preset low entropy threshold and whose information entropy of adjacent sub-region strips is normal are marked as suspected problem strips; Perform morphological closing operations on all suspected problematic strips to obtain suspected problematic regions, identify them as abnormal regions, and output an abnormal information report; For the remaining areas in the current search area, excluding the suspected problem areas, they are determined to be normal areas, and a normal information report is output.
7. The method for intelligent recognition and data extraction of cross-border trade document images according to claim 6, characterized in that, The anomaly information report specifically includes: the precise coordinates of the anomaly region, the original image slice corresponding to the anomaly region, and text information and the optimal binarization threshold for all iterations.
8. The method for intelligent recognition and data extraction of cross-border trade document images according to claim 6, characterized in that, The normal information report specifically includes: text information obtained by recognizing normal regions using an optical character recognition algorithm based on an optimal binarization threshold.
9. A cross-border trade document image intelligent recognition and data extraction system, characterized in that, include: Data acquisition module: used to acquire the target document image, and to identify the target document image by an optical character recognition algorithm according to a first binarization threshold to obtain a grayscale image of at least one target field region and the first text information; Interference field region acquisition module: used to calculate the statistical variance of the gradient magnitude in the grayscale image of the target field region. If the statistical variance is greater than the preset variance threshold, the target field region is recorded as the interference field region. The first binarization threshold adjustment module is used to generate a first adjustment amount based on the difference between the statistical variance and the preset variance threshold, and adjust the first binarization threshold accordingly to obtain the second binarization threshold. Text processing module: used to identify the second text information by using an optical character recognition algorithm based on a second binarization threshold to identify the interference field region; The second binarization threshold adjustment module is used to calculate the information entropy difference between the first text information and the second text information. If the information entropy difference is less than zero and its absolute value is greater than the preset negative threshold, the difference between the first adjustment amount and the preset adjustment step size is subtracted and the sign is reversed to obtain the second adjustment amount. Based on this, the first binarization threshold is adjusted to obtain a new second binarization threshold. Threshold Iteration Module: This module is used to repeatedly execute the text information processing module and the second binarization threshold adjustment module until the information entropy difference value obtained after adjustment is greater than zero or the number of iterations reaches the maximum threshold. Then, the iteration is terminated, and the corresponding second binarization threshold is obtained, which is recorded as the optimal binarization threshold. Text output module: Used to identify the target document image using an optical character recognition algorithm based on the optimal binarization threshold, obtain the third text information, and output it.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.