Invoice relation extraction method based on natural language analysis

By using a natural language analysis-based approach, combining the spatial fusion features and semantic dependency structure sequence of invoice image and text data, the system dynamically identifies elements such as amount and payer in invoices, solving the problems of delays and risk omissions in invoice data recognition in existing technologies, and achieving intelligent recognition and risk assessment.

CN120744116BActive Publication Date: 2025-11-25GANSU SHINING SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511196206.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-25
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify amounts and transaction elements in invoices when faced with diverse invoice structures and complex scenarios, leading to data recognition delays and omissions of risk items, which in turn affects the accurate flow of financial information and the ability to manage risks in advance.

Method used

By using natural language analysis-based methods, combined with spatial fusion features of invoice image and text data, matching consistency identifiers, semantic dependency structure sequences, and linked field clustering groups, elements such as amount and payer are dynamically identified and aggregated to achieve automatic identification of complex transaction relationships and intelligent judgment of abnormal risks.

Benefits of technology

It improves the accuracy of invoice data processing and the proactive risk control capabilities of financial management, enabling intelligent identification of invoice data and real-time generation of risk labels in complex transaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744116B_ABST
    Figure CN120744116B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of relation extraction, in particular to an invoice relation extraction method based on natural language analysis, comprising the following steps: through analyzing multiple features such as image seal, two-dimensional code and handwritten notes, matching text and image space information, checking the consistency of two-dimensional code and tax control code, identifying field combination and amount and payee abnormality. In the present application, through multi-level fusion analysis of seal, two-dimensional code and handwritten marks and other detail features in the invoice image, spatial joint discrimination of image and text content is realized, the credibility of information verification is improved by using character security comparison and coding consistency, with the help of context semantic dependency structure, the deep connection between payment behavior and entities is strengthened, elements such as amount and payee are dynamically classified and aggregated, automatic recognition of complex transaction relations is realized, further through linkage clustering and difference tracking, intelligent judgment and real-time label generation of abnormal risk behavior are realized, and the active risk control ability of financial management is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of relation extraction technology, and in particular to a method for extracting invoice relations based on natural language processing. Background Technology

[0002] Relation extraction refers to the use of natural language processing (NLP) techniques to extract entities with specific semantic relationships and the relationships between them from text. It is an important branch of information extraction, primarily used to identify and extract domain-specific knowledge from large amounts of unstructured text data. Traditional invoice relation extraction methods utilize NLP to extract relationships involving information such as amount, date, and transacting parties from invoice data. These methods typically employ rule-based matching and template-based matching, preprocessing the invoice text to extract specific fields or keywords to identify relationships.

[0003] When dealing with diverse invoice structures and mixed writing content, existing technologies are prone to omissions in field extraction due to templated rules, and manual patterns are difficult to handle complex situations such as invoice image distortion and annotation interference. Fixed relationship extraction methods limit the depth of semantic understanding, and the potential relationship between amount and transaction elements cannot be dynamically captured. When facing batch invoices or multi-scenario integration needs, data recognition delays and omissions of risk items are likely to occur, affecting the accurate flow of financial information and early risk control. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing an invoice relationship extraction method based on natural language analysis.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: an invoice relationship extraction method based on natural language analysis, comprising the following steps:

[0006] S1: Based on the extracted invoice image and text data, the pixel structure of the partitioned QR code is analyzed, the geometric shape of the handwritten annotation is analyzed, and the spatial coordinates are combined to match the OCR-recognized text with the image area. The matching situation between the two is judged to obtain the spatial fusion features.

[0007] S2: Based on the spatial fusion features, the QR code recognition character sequence and the tax control code field are matched item by item, the difference types are counted, samples that are completely consistent are filtered, and a unique identifier is assigned to the filtering results to obtain a matching consistency identifier.

[0008] S3: Based on the matching consistency identifier, analyze the semantic fields related to payment in the invoice text, sequentially search for words before and after the target field, determine the part of speech and verb structure collocation in the short sentence, filter the structure containing payment behavior words, encode the dependency relationship, and obtain the semantic dependency structure sequence;

[0009] S4: Based on the semantic dependency structure sequence, compare the combination of invoice number and invoice time fields, retrieve the co-occurrence of amount and time-related information, determine whether there is a payment verb in the phrase, and classify the field combinations that meet the conditions to obtain the linked field cluster group.

[0010] The present invention improves upon this invention by including the following features: the spatial fusion feature includes a field mapping identifier, a region association attribute, and a fusion confidence factor; the matching consistency identifier includes an identification code, a matching index, and a standard alignment number; the semantic dependency structure sequence includes a syntax dependency tag, a semantic association number, and a behavior chain structure; and the linked field clustering group includes a temporal grouping identifier, a field linkage parameter, and an event classification number.

[0011] The present invention is improved in that the step of obtaining the spatial fusion feature is specifically as follows:

[0012] S111: Based on the extracted invoice image and text data, identify the red seal area, determine the consistency between the area edge contour and the image pixel distribution, filter morphological features and summarize the area shape to obtain the seal contour morphological features.

[0013] S112: Based on the outline morphological features of the seal, optimize the pixel region segmentation of the QR code, analyze the geometric structure of the handwritten annotations, compare the spatial position distribution of each annotation, the QR code, and the seal, organize the relative coordinates between regions, and obtain the spatial coordinate matching distribution.

[0014] S113: Based on the spatial coordinate matching distribution, determine the correspondence between OCR text and image regions, filter each pairing result, and obtain spatial fusion features.

[0015] The present invention is improved in that the step of obtaining the matching consistency identifier is specifically as follows:

[0016] S211: Based on the spatial fusion features, analyze the character sequence identified in the QR code area, compare the content, position and arrangement of each character in the character group of the tax control code field, determine the insertion, omission and misalignment differences between corresponding characters, identify the character combinations with differences, and obtain the number of character difference combinations;

[0017] S212: Based on the number of character difference combinations, determine the differences between each group of characters, filter character combinations that are completely consistent in number, order and content, exclude character sequences containing differences, and organize the positioning and identification information of character combinations to obtain a consistent character combination index sequence.

[0018] S213: Based on the consistent character combination index sequence, optimize the character location content, source image number and QR code field label, determine the association between each number and the original field mapping, and obtain the matching consistency identifier.

[0019] The present invention is improved in that the step of obtaining the semantic dependency structure sequence is specifically as follows:

[0020] S311: Based on the matching consistency identifier, analyze the field fragments in the invoice text, filter out action words that contain payment, transfer and issuance, determine the grammatical relationship between the target field and its adjacent words, and obtain the action grammatical structure group;

[0021] S312: Based on the action grammar structure group, determine the arrangement of verbs, nouns and time words in each group, compare the logical relationship between word orders, calculate the verb-dominant structure type in the word order combination, classify the combinations with action semantic features, and establish a semantic linkage structure group.

[0022] S313: Based on the semantic linkage structure group, determine the dependency direction between verbs and subordinate phrases in each group, analyze semantic relationships, select verb fragments with stable grammatical dependency structures, optimize the aggregation method of structural differences, and obtain a semantic dependency structure sequence.

[0023] The present invention is improved in that the step of obtaining the clustering group of the linkage field is specifically as follows:

[0024] S411: Based on the semantic dependency structure sequence, determine the positional relationship between the invoice number field and the invoice time field in the short text sentence, the dependency path and time expression combination, compare the occurrence of the structure in the text, and calculate its distribution characteristics in the context to obtain the field combination frequency.

[0025] S412: Based on the frequency of the field combination, filter short sentences that simultaneously contain the amount field and time-related content, determine whether payment verbs appear in them, analyze the grammatical structure, dependency relationship and collocation of the verbs in the short sentences, and obtain the verb dependency matching quantity.

[0026] S413: Based on the verb dependency matching quantity, compare the field combination frequency quantity with the verb dependency matching quantity to obtain the structural linkage index, identify the field combinations that fall within the linkage interval range, and obtain the linkage field cluster group.

[0027] The present invention is improved in that the steps further include:

[0028] S5: Based on the clustering group of the linked fields, filter continuous invoice data, analyze the range of changes in the amount of adjacent invoices, determine the differences in the names of payers for different invoices, screen combinations with significant amount fluctuations and inconsistent payers, and statistically analyze the identification results to obtain risk trigger data.

[0029] The risk triggering data includes anomaly type codes, risk association tags, and sensitive combination identifiers.

[0030] The present invention is improved in that the step of obtaining the risk trigger data is specifically as follows:

[0031] S511: Based on the clustering group of the linked fields, analyze the time sorting information, and by comparing the data order of consecutive invoices with the corresponding payer names, filter invoices that are closely connected in time and have different payer content to obtain the number of payer change combinations.

[0032] S512: Based on the number of payer change combinations, determine the amount field of each pair of invoices, calculate the changes in the amounts of adjacent invoices, screen the pairs with key changes, and classify the pairs with amount changes into the number of amount difference pairs.

[0033] S513: Based on the number of pairs with the amount difference, and combined with the performance of the payer's abnormality and the amount change, analyze the distribution of each pair in the sequence, filter the pairs that meet the abnormal characteristics, and mark them as abnormal groups to obtain risk trigger data.

[0034] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0035] In this invention, by multi-level fusion analysis of detailed features such as seals, QR codes, and handwritten marks in invoice images, spatial joint discrimination of text and image content is achieved. Character security comparison and encoding consistency are used to improve the credibility of information verification. With the help of contextual semantic dependency structure, the deep connection between payment behavior and entities is strengthened. Elements such as amount and payer are dynamically classified and aggregated to achieve automatic identification of complex transaction relationships. Furthermore, through linkage clustering and difference tracking, intelligent judgment of abnormal risk behavior and real-time tag generation are achieved, optimizing the invoice data processing flow and improving the proactive risk control capability of financial management. Attached Figure Description

[0036] Figure 1 This is a flowchart of the main steps of the present invention;

[0037] Figure 2 This is a flowchart illustrating the acquisition of spatial fusion features in this invention.

[0038] Figure 3 This is a flowchart of the process for obtaining the consistency identifier in this invention;

[0039] Figure 4 This is a flowchart illustrating the process of obtaining the semantic dependency structure sequence in this invention.

[0040] Figure 5 This is a flowchart illustrating the process of obtaining cluster groups for linked fields in this invention.

[0041] Figure 6 This is a flowchart illustrating the process of obtaining risk trigger data in this invention. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0043] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0044] Example

[0045] Please see Figure 1 This invention provides a technical solution: an invoice relationship extraction method based on natural language analysis, comprising the following steps:

[0046] S1: Based on the extracted invoice image and text data, analyze the outline of the red stamp area in the invoice image, partition the QR code pixel structure, combine the geometric features of the handwritten annotation handwriting shape, match the OCR text results with the image area according to the spatial coordinates, determine the spatial matching between the image and text data, and summarize the related matching items to obtain the spatial fusion features.

[0047] S2: Based on spatial fusion features, determine whether the character sequence identified in the QR code area matches the character set of the tax control code field one by one, count the differences between characters, filter samples that meet the complete matching conditions, and perform unique identification processing on the filtering results to obtain the matching consistency identifier.

[0048] S3: Based on the matching consistency identifier, analyze the semantic fields related to payment in the invoice text. By sequentially searching the words before and after the target field, determine the collocation relationship between each part of speech and the verb structure combination in the short sentence, filter the payment, transfer and issuance behavior words that appear in the short sentence, and encode the difference type dependency relationship to obtain the semantic dependency structure sequence.

[0049] S4: Based on the semantic dependency structure sequence, compare the combination of the invoice number field and the invoice time field. During the text context analysis, retrieve the simultaneous occurrence of the amount field and time-related information, determine whether there is a payment verb in the current phrase structure, and classify the field combinations that meet the conditions to obtain the linked field cluster group.

[0050] S5: Based on the clustering of linked fields, filter consecutive combinations of invoices, calculate the range of change of the amount field of adjacent invoices, determine the corresponding differences of the payer names of each invoice, retrieve invoice groups with large amount changes and different payers, and statistically analyze the judgment results to obtain risk trigger data.

[0051] Spatial fusion features include field mapping identifiers, regional association attributes, and fusion confidence factors; matching consistency identifiers include identification codes, matching indexes, and standard alignment numbers; semantic dependency structure sequences include syntax dependency tags, semantic association numbers, and behavioral chain structures; linked field clustering groups include time-series grouping identifiers, field linkage parameters, and event classification numbers; and risk triggering data includes anomaly type codes, risk association labels, and sensitive combination identifiers.

[0052] In S1, pixel structure partitioning refers to dividing different areas in the invoice image, such as QR codes, seals, and handwritten annotations, into independent modules at the pixel level, preparing for subsequent feature extraction and comparison; geometric features refer to the shape, outline, area, angle, and other information of seals or handwritten handwriting in the invoice image, which are core parameters for judging specific elements (such as the authenticity of seals and the location of annotations) in image recognition; spatial coordinates refer to the specific two-dimensional coordinates (x, y) of each text or graphic element on the invoice image, used to establish a correlation between the OCR-recognized text and the specific physical location on the invoice; spatial matching refers to determining whether the two belong to the same business field or physical area by comparing the position of the OCR text and the position of the image area (such as the text "invoice code" and the corresponding QR code area); correlation matching refers to the combination of graphic elements that can be matched one-to-one after the above spatial and content comparisons, for example: the content of the "invoice code" OCR text is consistent with the content in the QR code and is in the same area.

[0053] In S2, the character sequence refers to a string of characters output by OCR or QR code recognition technology (such as a set of invoice numbers or tax control codes); the tax control code field refers to the unique number on the invoice used for anti-counterfeiting and management (such as the anti-counterfeiting code on the invoice issued by the tax authority or the data hidden in the QR code); the character difference type refers to the type of inconsistency found when comparing characters one by one (such as omissions, misalignments, and incorrect characters); the matching condition refers to the pre-set judgment criteria for a set of characters to be completely consistent (without difference); and the unique identifier processing refers to assigning a unique number or label to each sample that is confirmed to be correct after comparison for subsequent tracking and association.

[0054] In S3, associated semantic fields refer to information entries in the invoice text that are directly related to payment, such as the transaction entity, amount, and date (e.g., payer, amount, payment time, etc.); adjacent words refer to other words that are immediately next to the target field (e.g., "payer") in the text analysis, used to determine the contextual relationship; collocation refers to the structural and logical combination of words of different parts of speech (nouns, verbs, time words, etc.) in a sentence or phrase, such as "payer transfer date"; encoding differential type dependencies refers to using a unified code to mark various dependencies (e.g., "action-object", "time-event", etc.) by analyzing the grammatical or semantic dependencies between different words, which facilitates subsequent relationship analysis and modeling.

[0055] In S4, the Invoice Number field and the Invoice Date field refer to the unique number and date of issue on the invoice, respectively, and are often used for business process tracing and document time sequence analysis. Simultaneous appearance means that multiple target information (such as amount and time) are mentioned together in the same sentence or text fragment, which helps to identify the relationship between information. Payment verbs refer to payment-related action words that appear in the invoice text, such as "payment", "remittance", "issue", etc., which are key trigger words for automatic relationship identification. Conditional field combinations refer to the field groups that are determined to meet the analysis target (such as payment relationship) after semantic analysis, word collocation, verb detection and other judgments.

[0056] In S5, consecutive invoice combinations refer to a group of invoice data that are adjacent and consecutive in time or sequence (e.g., multiple invoices issued by the same customer for three consecutive days), which facilitates trend and risk analysis; corresponding differences refer to inconsistencies or changes found when comparing key fields (such as amount and payer) in invoices; different invoice groups refer to several sets of invoices that have been divided after searching based on the aforementioned differences (e.g., invoices with large amount fluctuations and changes in payer are grouped into one group), which serves risk identification and anomaly warning services.

[0057] Please see Figure 2 The specific steps for obtaining spatial fusion features are as follows:

[0058] S111: Based on the extracted invoice image and text data, identify the red seal area, determine the consistency between the area edge contour and the image pixel distribution, filter morphological features and summarize the area shape to obtain the seal contour morphological features.

[0059] When performing pixel processing on the original document image, channel separation is first performed to extract the red channel from the image and compare its pixel intensity with the green and blue channels. The high-saturation features of the red region are then enhanced. Subsequently, pixel blocks with higher brightness values ​​are located in the enhanced red image, and a baseline intensity value is set, for example, 210 based on common red seal samples. A set of regions with pixel values ​​higher than this baseline is selected, and these regions are initially considered as seal candidate areas. Based on this, contour extraction is performed on these regions, and the edge boundaries are detected through pixel gradient changes. The boundary contours are further binarized to obtain a clear edge structure, and the closure of the edge lines is determined. The continuity of the boundary is statistically analyzed through the connectivity between edge points. The degree of complete connection of the boundary is considered to be closed when the overall edge closure reaches more than 80%. It is then retained as a valid candidate seal area. Subsequently, its geometric features are statistically analyzed, including the ratio between the major axis and the minor axis of the area. If the ratio is close to 1, it indicates that the shape is close to a circle or ellipse. At the same time, it is statistically analyzed whether the edge direction change angle of the area is concentrated in a certain range. For example, in the samples, the edge angle of common seal outlines is concentrated between 25° and 60°. The proportion of the overall area of ​​the area to the area of ​​the whole ticket image is calculated. If the area proportion falls within a reasonable range, such as between 1% and 8%, the area is confirmed as a standard seal area. Its outline shape information is extracted and organized into a set of features for subsequent matching.

[0060] S112: Based on the outline morphological features of the seal, optimize the pixel region segmentation of the QR code, analyze the geometric structure of the handwritten annotations, compare the spatial position distribution of each annotation, the QR code, and the seal, organize the relative coordinates between regions, and obtain the spatial coordinate matching distribution.

[0061] The document image is divided into a gridded area, with each unit representing a set of spatial coordinates. Using the center coordinates of the area containing the seal as a reference point, the area is offset downwards by a set distance, such as approximately 15% of the image height. The area is then scanned for dense black-and-white block structures, similar to the clearly defined black-and-white block arrangement in printed QR codes. The ratio of black pixels to total pixels in this local area is calculated. If this ratio is between 60% and 90%, the area is considered a candidate for a QR code. Further analysis is conducted to determine if the edge lines of this area maintain a regular orientation, specifically whether the edge lines are predominantly horizontal or vertical. If this directional feature is concentrated within a horizontal or vertical linear range, the area is considered to have a complete outline and can be read by the QR code structure recognizer. This area is designated as the QR code area. Simultaneously, the handwriting area is identified from the image. By detecting indicators such as the average grayscale value of the ink color, the range of line curvature variation of the handwriting, and the number of strokes per unit area, the annotation area is located and it is determined whether it spatially overlaps with the seal or QR code area. When the handwriting area intersects or is adjacent to the boundary of the QR code or seal, the spatial center distance between the handwriting area and the above two key areas is further calculated. If the distance between a certain annotation area and the QR code area is significantly shorter than its distance to the seal area, it is classified as an annotation "close to the QR code", and its relative positional relationship is recorded. Finally, the horizontal and vertical coordinate values ​​of the center points of each area are classified and organized into a complete set of spatial coordinate distribution mapping information tables, which serve as the basic data for subsequent spatial fusion analysis.

[0062] S113: Based on spatial coordinate matching distribution, determine the correspondence between OCR text and image regions using the following formula:

[0063] ;

[0064] Filter each pairing result to obtain spatial fusion features. ,in, Represents the number of paired items. Representing the The degree of overlap between the OCR text and the image region coordinates. Representing the Spatial distance parameters between item annotations and target areas Representing the The area of ​​the QR code pixel partition. Representing the The outline and shape characteristics of the seal.

[0065] The spatial fusion feature is a set of results obtained by normalizing, comparing, superimposing, and transforming multiple sets of spatial parameters based on the above formula. It can reflect the spatial and structural coupling relationship between invoice graphic data. This feature can be used to determine whether the key fields of the invoice are accurately associated, and whether there are spatial conflicts or anomalies among multiple elements, providing support for subsequent invoice relationship identification and risk screening.

[0066] The pixel regions of the recognized OCR text are located within the image structure. Their initial boundary coordinates and width / height range in the image's two-dimensional coordinate system are determined, and a coordinate reference frame with the top-left corner of the image as the origin is established. Subsequently, the center point coordinates of each image region are extracted. By calculating the Euclidean distance, overlap boundary length, and angle between the direction vectors between the text and image regions, the structural mapping information of each pairing group is obtained. For each group of data, the overlap degree between the OCR text and image region coordinates is extracted. Spatial distance of annotation area QR code partition area 1. Characteristics of the outline of the seal The original data for one of the paired groups is as follows: , , , Its normalized corresponding value is , , , Substitute it into the formula, when At that time, let the other three sets of paired data be:

[0067] Group 2 (original): , , , After normalization: , , , ;

[0068] Group 3 (original): , , , After normalization: , , , ;

[0069] Group 4 (original): , , , After normalization: , , , ;

[0070] Calculate group by group first item:

[0071] Group 1 is ;

[0072] Group 2 is ;

[0073] Group 3 is ;

[0074] Group 4 is ;

[0075] Substitute into the absolute value calculation:

[0076] Group 1: ;

[0077] Group 2: ;

[0078] Group 3: ;

[0079] Group 4: ;

[0080] Taking the average, we get:

[0081] ;

[0082] This result indicates that spatial fusion features The corresponding values, under a unified normalization dimension, reflect the overall coordination between OCR text and key regions (QR codes, stamps, annotations) in the image structure in terms of spatial coordinates, visual morphology, and relative geometric distribution. The degree of fusion lies between perfect matching (approaching 0) and high offset (approaching 1), reflecting that each pairing group has a certain degree of matching consistency in local space but still has distinguishable differences. This numerical result, as a collective expression of multiple matching results, directly serves as the spatial fusion feature required for the next step, used to identify the coupling level of the image-text pairing relationship in multiple dimensions. The smaller the value, the tighter the fusion. By organizing the fusion output sequences corresponding to each pairing group, the basic field set for generating group labels, spatial matching levels, and subsequent linked field structure analysis can be further derived.

[0083] Please see Figure 3 The specific steps for obtaining the consistency identifier are as follows:

[0084] S211: Based on spatial fusion features, analyze the character sequence identified in the QR code area, compare the content, position and arrangement of each character in the character group of the tax control code field, determine the insertion, omission and misalignment differences between corresponding characters, identify the character combinations with differences, and obtain the number of character difference combinations;

[0085] Extract the character sequence from the QR code area. Retrieve the spatial binding information between the completed image area and OCR text to extract the characters within the QR code pixel area into a one-dimensional sequence. Then, based on the invoice structure field template, call the standard character group of the known tax control code field. In this example, the standard tax control code is set to "03456789123456789ABCD1234567890". Compare this standard sequence with the QR code parsing sequence character by character, checking if the character content at each corresponding position is consistent. For inconsistent character pairs, record their index position in the original sequence. Then, perform an insertion difference judgment, i.e., check if the QR code character sequence has one more character than the standard sequence in a certain segment. If there is one more character in the QR code sequence but the content after the position of the extra character remains aligned, record it as "insertion difference"; otherwise, if the QR code characters skip a standard position in a certain segment... Alignment misalignment is recorded as "omission difference". Further inspection reveals that if the character content at the same position is inconsistent but the position matches, it is marked as "misalignment difference". For example, if the 8th digit of the standard code is "1" but the QR code recognition result is "7", it is misalignment. In a set of invoices, the QR code sequence "034567891234X6789ABCD1234567890" was compared with the standard tax control code, and the 13th digit was found to be "X" instead of "4", which was recorded as a misalignment difference. In another set of data, the QR code recognition sequence was "03456789123456789AABCD1234567890". Comparison revealed an extra "A" in the 18th digit, confirming it as an insertion difference. By counting the number of occurrences of the above three types of differences in the current character sequence, and marking them according to character position, all the positions of characters with the above differences are packaged and organized to form the number of character difference combinations.

[0086] S212: Based on the number of character difference combinations, determine the differences between each group of characters, filter character combinations that are completely consistent in number, order and content, exclude character sequences containing differences, and organize the positioning and identification information of character combinations to obtain a consistent character combination index sequence.

[0087] The compared character sequences are screened, and each character combination is checked against the standard tax control code in terms of character quantity, order, and content. First, the total number of characters in the QR code recognition character sequence is extracted. If this number differs from the total number of characters in the standard tax control code (e.g., the standard is 32 bits while the QR code parsing result is 31 or 33 bits), the sequence is excluded as inconsistent. Next, the character order is checked by comparing each character position to ensure complete correspondence. For example, if the 10th digit in the standard tax control code is "5" but the 10th digit in the recognition sequence is "E", it is marked as inconsistent. If all digits match, the sequence is retained. Finally, the character content is checked for a one-to-one match, i.e., the QR code characters are equated to the standard characters. If all characters meet the content consistency requirement, the sequence is recorded as a matching character combination. Its starting coordinates in the image and the associated field name are also marked to form a unique identifier. In a real-world scenario, if the QR code sequence "03456789123456789ABCD1234567890" is found to be completely consistent with the standard tax control code, and the starting coordinates of the corresponding area in the image are (360, 480), with the corresponding field being "Value-Added Tax Special Invoice Anti-counterfeiting Code", then this combination is included in the consistency sequence and assigned the index number "ID_00012". All character combinations that meet the above conditions are numbered and sorted sequentially to form a consistent character combination index sequence.

[0088] S213: Based on a consistent character combination index sequence, optimize the character location content, source image number, and QR code field labeling, using the following formula:

[0089] ;

[0090] Determine the association between each number and the original field mapping to obtain a matching consistency indicator, where, Indicates the first The consistent character combination and the first A consistency identifier generated by combining QR code fields. The index of the consistent character combination set is the first one. The character index number of the item. Represents the first region in the image. The index tag number of each QR code field. This indicates the total number of records involved in the scan record and acquisition time merging operation under the current combination. Indicates the first The encoded number of each QR code scan record. Indicates the first Each QR code scan record corresponds to a collection time code. Indicates the first Layer identifier encoding associated with a combination of characters, Indicates the first The encoding of the source path of each QR code field.

[0091] The consistency matching identifier refers to the unique identification generated by combining the spatial, temporal, and structural features of the invoice image when the content of the QR code and the tax control code field are completely consistent at the character level. It is a unique encoding result that integrates "text content consistency" and "image data traceability features" and serves as the basic index for subsequent operations such as invoice relationship collection, data association, and risk control.

[0092] Based on the character location data extracted from the consistent character combination index set, the image number, layer structure, and QR code field attribution markers corresponding to the characters are optimized, and the character index numbers of the character combinations are uniformly converted. Field index number of the QR code field And extract the layer identifier code. Encoding of source path Used to calculate the structural dimension correlation of combined pairs, collect the scan record number associated with each group of QR code fields. and its acquisition time code Based on a unified numbering structure, normalization is performed to ensure consistent scale across different data dimensions in numerical calculations. Then, all fields are substituted into the consistency identifier calculation formula, using consistent character combinations. Corresponding field index Layer identifier length Source path encoding After normalization, they are respectively , , , QR code scanning record Collection time encoding After normalization , The consistency flag is calculated as follows:

[0093] Sequence number difference squared:

[0094] ;

[0095] Sum of the products of scanning and acquisition times:

[0096] ;

[0097] Square root of denominator for structural dimension:

[0098] ;

[0099] calculate:

[0100] ;

[0101] This result indicates that the first Group character combination and the first The consistency indicator between the two QR code fields is close to 1, with a value of 0.9982, indicating that the two elements have extremely high consistency in multidimensional features such as spatial location, character content, and QR code scanning behavior. Specifically, this value reflects a high degree of matching between the two in terms of character position, character content, scan record, and acquisition time. That is, the character sequences they correspond to have an almost completely identical structural relationship with the QR code fields. The formula couples character index deviation, behavioral record features, and structural path information through a unified normalization scale, which can realize a quantitative expression of the mapping relationship between character combinations and image fields, enhancing the accuracy and tracking capability of character relationship extraction.

[0102] Please see Figure 4 The specific steps for obtaining the semantic dependency structure sequence are as follows:

[0103] S311: Based on the matching consistency identifier, analyze field fragments in the invoice text, filter out statements containing action words such as payment, transfer and issuance, determine the grammatical relationship between the target field and its adjacent words, and obtain action grammatical structure groups;

[0104] Based on the identified identical invoice field groups, the context of each field is determined. Sentence fragments containing adjacent fields are extracted, and an index position table of each word within the sentence is created. Using noun fields as anchors, keyword searches are performed extending no more than three word positions before and after them. The searches sequentially check for the existence of any one of the words "payment," "transfer," or "issue" as an action word. Simultaneously, the part-of-speech tag of this action word is confirmed. If it is a verb and located near the target field, its position in the sentence is recorded. Then, it is determined whether the word arrangement between the action word and the target field constitutes a basic grammatical structure such as subject-predicate, verb-object, or modifier-head. For example, in the sentence "The company issued an invoice in May,"... The target field "company" is the subject, "issue" is the verb, and "invoice" is the object, forming a subject-verb-object structure. Based on this, it is determined whether the target field and adjacent words form a valid grammatical collocation relationship. If its word order conforms to the general Chinese written expression pattern and the parts of speech are reasonably matched, it is classified into the action structure group, and its word order structure, verb position, field position, surrounding vocabulary structure, and other parameters are recorded. In the actual invoice statement "The payer completed the transfer on May 20", "transfer" is the action word, preceded by "payer" as the subject, and followed by "amount" as the object. The structures on both sides of the verb are complete and semantically fluent, satisfying the judgment condition. Syntax groups with clear action logic structures are integrated, and the output is the action grammatical structure group.

[0105] S312: Based on the action grammar structure group, determine the arrangement of verbs, nouns and time words in each group, compare the logical relationship between word order, calculate the verb-dominant structure type in the word order combination, classify the combination with action semantic features, and establish semantic linkage structure group;

[0106] The word order information in each structural group was extracted, and the sequence numbers of verbs, nouns, and time words were marked. The frequency of different combinations in the invoice text was counted. A reasonable sample size of 100 word order combinations was set for analysis. For each structure, the position of the verb was recorded, and it was determined whether it was in a dominant position in the sentence. If there were dependent phrases before and after the verb, and the verb was in the third word position or after in the sentence, it was considered the core of the dominant structure. The direction of dependence between the verb and the noun was then analyzed, i.e., whether the noun was the subject or object of the verb. If the noun was immediately before the verb, it was marked as a subject structure; if the noun appeared after the verb, it was marked as an object structure. Finally, the time words were analyzed... For positional analysis, if the time word appears within two word positions before the verb, it is recorded as a "pre-modifying time structure"; if it appears after the verb, it is recorded as a "post-modifying time structure". For example, in the sentence "The company made payment on April 15", the time word "April 15" precedes the verb "payment", and the noun "company" is at the beginning to form the subject. Based on the position of the verb structure, it is determined to be a "subject-time-verb" type structure. According to this, all structural combinations are classified. Combinations in which the verb has at least one noun and one time word modifying the structure and the order is stable are classified as combinations with action semantic features, forming a semantic linkage structure group with the verb as the main body and time and noun logical modification components.

[0107] S313: Based on semantic linkage structure groups, determine the dependency direction between verbs and subordinate phrases in each group, analyze semantic relationships, select verb fragments with stable grammatical dependency structures, optimize the aggregation method of structural differences, and obtain semantic dependency structure sequences.

[0108] A directional relationship marker table is established for the verbs in each structure and their dependent nouns, time words, and other subordinate phrases. Each record in the table uses the verb as the core identifier, indicating the lexical number and dependency direction of its preceding and following dependent objects. If a verb is simultaneously modified by a preceding noun and limited by a following time word, it is recorded as a bidirectional dependency structure. Then, semantic relationship analysis is performed on all structure groups to extract whether there is semantic ambiguity caused by word order changes. For example, the verb "receive invoice" in the sentence "After payment is completed, the company receives the invoice" and "The company receives the invoice after payment is completed" have different subjects and different semantic dependency directions. The process involves determining the primary-subordinate relationship between verbs and their subordinate phrases based on their dependency logic, then searching for stable dependency marker combinations within verb fragments. If a verb fragment appears in the same position and has a stable dependency direction in over 90% of the samples, it is considered to have a stable grammatical dependency structure and is retained as a core semantic unit. Simultaneously, it summarizes situations such as word order changes and missing modifiers in the structure, reorganizes the differences between similar structures, and merges and classifies structural fragments with consistent dependency paths to generate a set of unified structure encoding identifiers, thus constructing a complete semantic dependency structure sequence.

[0109] Please see Figure 5 The specific steps for obtaining the cluster groups of linked fields are as follows:

[0110] S411: Based on the semantic dependency structure sequence, determine the positional relationship between the invoice number field and the invoice time field in the short text sentence, the combination of dependency path and time expression, compare the occurrence of the structure in the text, and calculate its distribution characteristics in the context to obtain the field combination frequency.

[0111] All short sentences marked with semantic dependency structures in the invoice text are indexed and extracted. Sentence fragments containing the invoice number and invoice date fields are selected. The relative position numbers of the two fields within the same short sentence are extracted, and their order of appearance in the sentence is recorded. Then, based on the syntactic analysis results, it is determined whether there is a dependency path connection between the number field and the time field. If the two fields form a dependency chain through verbs, conjunctions, prepositions, etc., it is marked as structurally valid. Next, time expressions related to the invoice date field are extracted. Two word positions are searched before and after this field to confirm whether there are word combinations conforming to the date and time format, such as standard time expressions like "June 30, 2025" or "Q4, 2024". If a sentence contains at least one time expression and the dependency relationship is clearly defined, it is considered to constitute a valid combination structure. The frequency of this type of structure in all short sentences in the text is counted and the frequency value is recorded. Then, the number of short sentences that have appeared with this field combination structure is divided by the total number of short sentences to calculate the contextual distribution ratio of this structure. If the ratio is above 0.15, it is classified as a frequent structure; otherwise, it is classified as a sparse structure. In a certain text set containing 126 short sentences, 32 structures with the fields of invoice number and invoice time appeared. Among them, 26 had valid dependency paths and time expression combinations. The structure coverage rate of this combination in the context was calculated to be 20.6%. The frequency result of this combination structure is output as the field combination frequency quantity through the above statistics.

[0112] S412: Based on the frequency of field combinations, filter short sentences that simultaneously contain monetary and time-related content, determine whether payment verbs appear in them, analyze the grammatical structure, dependency relationships, and collocation of verbs with adjacent words in the short sentences, and obtain the verb dependency matching quantity.

[0113] Call short sentence samples where the combined frequency value of all fields is greater than 0.1, and judge sentence by sentence whether they contain both an amount field and time-related words. Identify the key expressions of the amount field, such as the identifiers "¥", "yuan", "amount", etc., and at the same time identify the standard format of time expressions and then perform cross-matching. Screen out sentences that meet both the amount and time conditions, and then retrieve whether there are verbs in the sentence. Call the词性标注表 (lexical category tagging table) to confirm the lexical category of each word. If there is a lexical item with the lexical category of verb and the verb belongs to the payment semantic category, such as "pay", "remit", "make a payment", etc., then the sentence is included in the candidate set. Perform syntactic structure analysis on the structure where the verb is located to judge whether it forms a subject-predicate or verb-object structure. At the same time, retrieve the adjacent words in the two positions before and after the verb, and judge the lexical category collocation relationship between the words and the verb. If a combination of "noun-verb" or "time word-verb" appears, record it as a verb dependency matching instance. For example, in the sentence "In December 2023, the company paid 50,000 yuan for the goods", the time word "December 2023" is before the verb "pay", the noun "company" is the subject, "goods payment" is the object, the verb is in the core position in the sentence, the verb dependency relationship is clear and the lexical category matches, meeting the dependency matching standard. Record this structure in the dependency matching quantity. Count the number of sentences that meet the above conditions in all candidate short sentences, and then divide by the total number of sentences that meet the field combination frequency conditions to obtain the verb dependency matching ratio in this batch. Convert this ratio into the verb dependency matching quantity.

[0114] S413: Based on the verb dependency matching quantity, compare the field combination frequency quantity with the verb dependency matching quantity, using the formula:

[0115]

[0116] Obtain the structure linkage index , identify the field combinations that fall within the linkage range, and obtain the linkage field clustering group, where represents the position vector of the invoice number field represents the position vector of the invoice issuing time field represents the context similarity of the amount field represents the semantic dependency of the payment verb represents the field structure association strength. ​​​​​​and the position vector of the invoice time field Then, the absolute value of the difference is calculated, and the contextual similarity of the amount field is analyzed. Semantic dependency of the verb "payment" and the correlation strength of field structure By combining the above parameters, the calculation formula yields the structural linkage index. .

[0119] The original values ​​are as follows:

[0120] (Invoice number field position vector) (Normalized invoice number field position vector);

[0121] (Invoice time field position vector) (Normalized invoice time field position vector);

[0122] (Context similarity of amount field) (Context similarity of the amount field, normalized);

[0123] (Semantic dependency of payment verbs) (The semantic dependency of the payment verb has been normalized);

[0124] (Field structure association strength) (Field structure association strength has been normalized).

[0125] Substitute into the formula to calculate:

[0126] ;

[0127] This result indicates that the structural linkage index This reflects the strength of the linkage between the invoice number field, the invoice date field, the amount field, and the payment verb. The higher the value, the stronger the semantic and structural correlation between the fields, which means that they have a high degree of consistency and connection in the text. Based on this value, field combinations that meet the linkage conditions can be further filtered out and classified into the same linkage field cluster group, thereby providing a basis for subsequent tasks such as invoice relationship analysis and risk detection.

[0128] Please see Figure 6 The specific steps for obtaining risk trigger data are as follows:

[0129] S511: Based on the clustering group of linked fields, analyze the time sorting information, and by comparing the data order of consecutive invoices with the corresponding payer names, filter invoices that are closely connected in time and have different payer content to obtain the number of payer change combinations.

[0130] Extract invoice records containing invoice number and invoice date fields from each pair. Sort the records in ascending order by invoice date field and construct a time series index linked list. Pair adjacent record numbers in this linked list sequentially. For each pair, read the corresponding payer name field and use string similarity analysis to determine if the two payer names are identical. If the two strings differ by two or more characters in the beginning, end, or middle, and the overall similarity is less than 0.8, they are considered different payers. Further, perform time difference calculation on the time field content, extract the invoice date from each pair of records, and calculate the difference in days between the two dates. If the time interval is within 3 days, it is considered a close temporal connection. Next, determine whether the combination simultaneously meets the two conditions of different payers and close timing. If both are met, mark the pair as an abnormal combination instance. In an instance, if two consecutive invoices have dates of June 1, 2024 and June 2, 2024, respectively, with numbers FP20240601001 and FP20240602002, and the corresponding payers are "XXXX Co., Ltd." and "XXXX Intelligent Equipment" respectively, the two strings differ by 7 characters, and the time difference is 1 day, which meets the set conditions. The data pair is then classified as an abnormal combination. Repeat the above process to batch judge and filter all sorted combinations, count the total number of pairs that meet the conditions, and output the number of payer abnormal combinations.

[0131] S512: Based on the number of payer change combinations, determine the amount field of each pair of invoices, calculate the changes in the amounts of adjacent invoices, screen the pairs with key changes, and classify the pairs with amount changes into the number of amount difference pairs.

[0132] For each pair of marked payer change combinations, read the amount field. If the field contains amount identifiers such as "¥", "RMB", or "yuan", extract the corresponding numerical part as the standard amount value. Perform a difference calculation on the amounts of the previous and next invoices in the pair to calculate the amount change range between them. Use the absolute value of the result of subtracting the previous invoice from the next one to judge the fluctuation. Set the threshold for recognizing the amount change range to 5000 yuan. If the difference is equal to or exceeds this threshold, the amount difference is considered significant, and the pair is recorded as an amount difference pair. In actual data, if invoice A is 12800 yuan and invoice B is 5800 yuan, the difference between the two is 7000 yuan, which is greater than the set threshold. Therefore, this pair is recorded as an amount difference. If the invoice amounts in a pair are 5300 yuan and 5400 yuan respectively, the difference is only 100 yuan, so it is not included in the group. Repeat this judgment logic to process all payer change combinations, accumulate the number of pairs that meet the amount change condition, and output it as the number of amount difference pairs.

[0133] S513: Based on the number of pairs with amount differences, combined with the performance of payer anomalies and amount changes, analyze the distribution of each pair in the sequence, screen pairs that meet the abnormal characteristics, and mark them as abnormal groups to obtain risk trigger data;

[0134] Each pair of invoices with changes in amount and payer anomalous markers is read sequentially. The position information of the pair is recorded in the invoice time series, and the pair's position in the time sequence is generated. The analysis is performed to see if there are duplicate payer, amount, or invoice date fields before and after the pair. If the same payer name appears alternately in multiple positions, or the amount value is repeated on different dates, the pair is determined to be a periodic fluctuation type. Conversely, if all fields are unique and appear concentrated in a short period of time, it is recorded as a concentrated anomalous type. All pair distribution types are then statistically classified, and pairs are established based on three indicators: concentration, frequency, and structural regularity. The structural scoring is defined as follows: combinations that meet any two of the three abnormal conditions are considered abnormal groups. Concentration is judged by whether the time interval between the occurrence of the combination is less than 2 days; frequency is judged by whether the number of combinations exceeds 20% of the total number of combinations; and structural regularity is judged by whether there are repeated changes in multiple fields. If, in a certain set of invoices, combinations ID#0023, ID#0024, ID#0026, and ID#0027 are found to appear within 4 consecutive days, with different payers and amounts varying by more than 8,000 yuan, and similar abbreviations of payer names, they are judged as concentrated abnormal groups. All combinations that meet the above conditions are packaged and output as risk trigger data.

[0135] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for extracting invoice relationships based on natural language processing, characterized in that, Includes the following steps: S1: Based on the extracted invoice image and text data, the pixel structure of the partitioned QR code is analyzed, the geometric shape of the handwritten annotation is analyzed, and the spatial coordinates are combined to match the OCR-recognized text with the image area. The matching situation between the two is judged to obtain the spatial fusion features. S2: Based on the spatial fusion features, the QR code recognition character sequence and the tax control code field are matched item by item, the difference types are counted, samples that are completely consistent are filtered, and a unique identifier is assigned to the filtering results to obtain a matching consistency identifier. S3: Based on the matching consistency identifier, analyze the semantic fields related to payment in the invoice text, sequentially search for words before and after the target field, determine the part of speech and verb structure collocation in the short sentence, filter the structure containing payment behavior words, encode the dependency relationship, and obtain the semantic dependency structure sequence; The steps for obtaining the semantic dependency structure sequence are as follows: S311: Based on the matching consistency identifier, analyze the field fragments in the invoice text, filter out action words that contain payment, transfer and issuance, determine the grammatical relationship between the target field and its adjacent words, and obtain the action grammatical structure group; S312: Based on the action grammar structure group, determine the arrangement of verbs, nouns and time words in each group, compare the logical relationship between word orders, calculate the verb-dominant structure type in the word order combination, classify the combinations with action semantic features, and establish a semantic linkage structure group. S313: Based on the semantic linkage structure group, determine the dependency direction between verbs and subordinate phrases in each group, analyze semantic relationships, select verb fragments with stable grammatical dependency structures, optimize the aggregation method of structural differences, and obtain a semantic dependency structure sequence. S4: Based on the semantic dependency structure sequence, compare the combination of invoice number and invoice time field, retrieve the co-occurrence of amount and time-related information, determine whether there is a payment verb in the phrase, and classify the field combinations that meet the conditions to obtain the linked field cluster group. The specific steps for obtaining the linked field clustering group are as follows: S411: Based on the semantic dependency structure sequence, determine the positional relationship between the invoice number field and the invoice time field in the short text sentence, the dependency path and time expression combination, compare the occurrence of the structure in the text, and calculate its distribution characteristics in the context to obtain the field combination frequency. S412: Based on the frequency of the field combination, filter short sentences that simultaneously contain the amount field and time-related content, determine whether payment verbs appear in them, analyze the grammatical structure, dependency relationship and collocation of the verbs in the short sentences, and obtain the verb dependency matching quantity. S413: Based on the verb dependency matching quantity, compare the field combination frequency quantity with the verb dependency matching quantity to obtain the structural linkage index, identify the field combinations that fall within the linkage interval range, and obtain the linkage field cluster group.

2. The invoice relationship extraction method based on natural language analysis according to claim 1, characterized in that, The spatial fusion features include field mapping identifiers, regional association attributes, and fusion confidence factors; the matching consistency identifiers include identification codes, matching indexes, and standard alignment numbers; the semantic dependency structure sequence includes syntax dependency tags, semantic association numbers, and behavioral chain structures; and the linked field clustering groups include temporal grouping identifiers, field linkage parameters, and event classification numbers.

3. The invoice relationship extraction method based on natural language analysis according to claim 1, characterized in that, The specific steps for obtaining the spatial fusion features are as follows: S111: Based on the extracted invoice image and text data, identify the red seal area, determine the consistency between the area edge contour and the image pixel distribution, filter morphological features and summarize the area shape to obtain the seal contour morphological features. S112: Based on the outline morphological features of the seal, optimize the pixel region segmentation of the QR code, analyze the geometric structure of the handwritten annotations, compare the spatial position distribution of each annotation, the QR code, and the seal, organize the relative coordinates between regions, and obtain the spatial coordinate matching distribution. S113: Based on the spatial coordinate matching distribution, determine the correspondence between OCR text and image regions, filter each pairing result, and obtain spatial fusion features.

4. The invoice relationship extraction method based on natural language analysis according to claim 1, characterized in that, The specific steps for obtaining the matching consistency identifier are as follows: S211: Based on the spatial fusion features, analyze the character sequence identified in the QR code area, compare the content, position and arrangement of each character in the character group of the tax control code field, determine the insertion, omission and misalignment differences between corresponding characters, identify the character combinations with differences, and obtain the number of character difference combinations; S212: Based on the number of character difference combinations, determine the differences between each group of characters, filter character combinations that are completely consistent in number, order and content, exclude character sequences containing differences, and organize the positioning and identification information of character combinations to obtain a consistent character combination index sequence. S213: Based on the consistent character combination index sequence, optimize the character location content, source image number and QR code field label, determine the association between each number and the original field mapping, and obtain the matching consistency identifier.

5. The invoice relationship extraction method based on natural language analysis according to claim 1, characterized in that, The steps also include: S5: Based on the clustering group of the linked fields, filter continuous invoice data, analyze the range of changes in the amount of adjacent invoices, determine the differences in the names of payers for different invoices, screen combinations with significant amount fluctuations and inconsistent payers, and statistically analyze the identification results to obtain risk trigger data. The risk triggering data includes anomaly type codes, risk association tags, and sensitive combination identifiers.

6. The invoice relationship extraction method based on natural language analysis according to claim 5, characterized in that, The specific steps for obtaining the risk trigger data are as follows: S511: Based on the clustering group of the linked fields, analyze the time sorting information, and by comparing the data order of consecutive invoices with the corresponding payer names, filter invoices that are closely connected in time and have different payer content to obtain the number of payer change combinations. S512: Based on the number of payer change combinations, determine the amount field of each pair of invoices, calculate the changes in the amounts of adjacent invoices, screen the pairs with key changes, and classify the pairs with amount changes into the number of amount difference pairs. S513: Based on the number of pairs with the amount difference, and combined with the performance of the payer's abnormality and the amount change, analyze the distribution of each pair in the sequence, filter the pairs that meet the abnormal characteristics, and mark them as abnormal groups to obtain risk trigger data.

Citation Information

Patent Citations

  • Multi-mode invoice automatic classification identification method, verification method and system

    CN116052186A