Bill element intelligent extraction and verification method and system

By linking the invoice image and encoding information, the system identifies and analyzes invoice format differences, generates specific correction guidelines, solves the problem of low efficiency in invoice information processing in existing technologies, and achieves efficient and reliable format correction and compliance verification.

CN121938019APending Publication Date: 2026-04-28PURANG SERVICES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PURANG SERVICES
Filing Date
2025-12-29
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies cannot accurately analyze format differences in invoice information processing, making it difficult for users to quickly and accurately locate and correct problems, thus reducing processing efficiency.

Method used

By linking the invoice image, extracted element information, and coding information, the system identifies the current filling format and matches it with the standard format, analyzes the type of difference, triggers corresponding processing strategies, and generates specific correction guidelines.

Benefits of technology

It improves the efficiency of bill information processing, can automatically determine the type of discrepancy and generate actionable correction guidelines, enhances overall reliability and compliance, and lowers the operational threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938019A_ABST
    Figure CN121938019A_ABST
Patent Text Reader

Abstract

The invention relates to a bill element intelligent extraction and verification method and system, and relates to the field of artificial intelligence, and the method comprises the steps: collecting bill image information; extracting bill element information based on the bill image information; obtaining a current information filling format according to the bill image information and the bill element information; obtaining bill code information according to the bill element information; matching a standard information filling format through the bill coding information; if the current information filling format is inconsistent with the standard information filling format, combining the current information filling format with the standard information filling format to generate a format difference condition; and determining a difference processing method based on the format difference condition, and performing prompt processing according to the difference processing method. The method and the device have the effect of improving the bill information processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to a method and system for intelligent extraction and verification of invoice elements. Background Technology

[0002] Intelligent extraction and verification of invoice elements refers to the process of using artificial intelligence technology to automatically identify key business information in invoice images and verify their compliance, logic, and authenticity.

[0003] In scenarios such as bill discounting, circulation, and auditing, information entry and verification of various bills are required. Currently, optical character recognition (OCR) technology is commonly used to extract text from bill images. Subsequently, the extracted information is compared with a preset bill format template using rule matching to determine whether the information is filled out correctly.

[0004] Current technologies, after recognizing and comparing the formats of invoice images, can only generate a general prompt indicating a difference when inconsistencies are found. They cannot specifically analyze whether the information is incorrect or incomplete, nor can they provide differentiated processing and guidance for different types of format differences. This makes it difficult for users to quickly and accurately locate and correct format errors, thus reducing the efficiency of invoice information processing and requiring improvement. Summary of the Invention

[0005] To improve the efficiency of bill information processing, this invention provides a method and system for intelligent extraction and verification of bill elements.

[0006] Firstly, the present invention provides a method for intelligent extraction and verification of document elements, employing the following technical solution: A method for intelligent extraction and verification of invoice elements, comprising: Collect ticket image information; Extracting essential information from invoice images; The current information filling format is obtained based on the invoice image information and invoice element information; Based on the information of the elements of the bill, the bill code information is obtained; The standard information filling format is matched using the invoice code information; If the current information filling format is inconsistent with the standard information filling format, the format difference will be generated by combining the current information filling format and the standard information filling format. Based on the format differences, determine the method for handling the differences, and then process the prompts according to the method for handling the differences.

[0007] By adopting the above technical solution, and linking the document image, extracted element information, and coding information, the system can not only identify the current filling format but also accurately match the standard format that should be used for that document type. When inconsistencies are detected, the system can further analyze the specific manifestations of the differences and automatically determine the type of difference, triggering corresponding processing strategies. This transforms the originally vague error messages into specific and actionable correction guidelines, thereby improving the efficiency of document information processing.

[0008] Optionally, format differences include incorrect information and incomplete information. The methods for handling differences in incorrect information include: When the format difference situation indicates an information filling error, the specific filling error content is obtained based on the format difference situation; Determine the location of the error based on the invoice code information and the specific incorrect content; Based on the standard information filling format and the location of incorrect content, you can know the type of content that should be filled in correctly; The individual error content type is obtained based on the specific incorrect content; When a single incorrect entry falls into the correct entry type, a content position correction plan is generated by combining the single incorrect entry with the standard information entry format. The content location correction scheme corrects the specific errors in the content.

[0009] Optional methods for handling discrepancies due to incomplete information include: When the format difference situation results in incomplete information, the specific missing location is determined based on the format difference situation. Determine the specific missing content by combining the specific location of the missing information with the standard information filling format; generate a content supplementation request based on the specific missing content; In response to a content supplementation request, generate a content supplementation interface and collect the user's supplemented missing information; The missing information provided by the user is integrated with the invoice element information to form a complete invoice dataset.

[0010] Optional, methods for verifying the authenticity of invoices may also be included: Extracting digital watermark features based on ticket image information; The basic watermark features are determined based on the invoice coding information; The watermark matching degree is calculated by combining digital watermark features and baseline watermark features. When the watermark matching degree is lower than the preset standard watermark matching degree, an abnormal ticket notification will be reported.

[0011] Optionally, the method for verifying the authenticity of invoices further includes: When the watermark matching degree is not lower than the preset standard watermark matching degree, the historical circulation status record is determined based on the ticket code information. The current date and current holder information of the negotiable instrument are determined based on its key information. The current status statement is determined by combining the current date of the negotiable instrument with the current holder's information; Perform a logical continuity comparison between the current state declaration and the historical transition state records; If the logical continuity comparison fails, the bill status is determined to be abnormal, and an abnormal bill notification containing specific continuity conflicts is reported.

[0012] Optionally, a method for verifying the logical consistency of invoices may also be included: Key logical fields are obtained based on the information of the invoice elements; Based on the invoice code information, the relevant business rule base is matched; Constraint relationships of key logical fields are validated by associating them with the business rule base. If a key logical field violates a preset constraint, a logical conflict report will be generated by combining the key logical field with the associated business rule base. Logical conflict reports are used to determine the level of logical anomalies, and graded early warning prompts are issued according to the level of logical anomalies.

[0013] Optional, user behavior risk assessment methods may also be included: Use logical conflict reports to pinpoint the core conflict fields; Based on the format differences, filter out the related format difference items that are associated with the core conflicting fields; By comparing the format of invoice element information with the standard information, compliance simulation and correction are performed on the differences in related formats. Based on the simulated and corrected bill element information, the constraint relationship verification in the logical consistency verification method is executed again. If the verification shows that the conflict in the logical conflict report has been eliminated or mitigated, a coordination anomaly warning will be generated. Based on the level of collaborative anomaly warnings and logical anomalies, the strategy for handling differences is adjusted.

[0014] Optional features include methods for pre-filling missing information and risk guidance: The standard value range should be determined based on the standard information filling format and the specific missing locations. Pre-filled suggested values ​​are generated based on invoice element information and standard value range; Based on the pre-filled suggested values ​​and standard value ranges, the content is displayed in the preset content supplementation interface, and user modification operations are collected; Based on the user's modification actions, obtain the user's input content; If the user input exceeds the standard value range, combine the standard value range with the pre-filled suggested value to generate compliant input. Based on compliant input content, guide users to correct their input content to the standard value range and generate information for users to supplement missing information; The missing information provided by the user is integrated with the invoice element information to form a complete invoice dataset.

[0015] Optional, also includes: A multi-dimensional risk signal set is generated based on format differences, logical conflict reports, and collaborative anomaly alerts. The risk level of related bills is determined based on the bill code information; The fusion weight configuration is determined based on a multi-dimensional set of risk signals and the risk level of associated bills. Based on the fusion weight configuration, a multi-dimensional risk signal set is weighted and fused to generate a comprehensive risk assessment result and a response confidence level. If the confidence level of the handling is higher than the preset automatic handling threshold, the corresponding differential handling method, graded early warning prompt or collaborative abnormal warning response process will be automatically triggered and executed based on the comprehensive risk assessment result.

[0016] Secondly, this application provides an intelligent extraction and verification system for invoice elements, which adopts the following technical solution: A smart system for extracting and verifying document elements, comprising: The acquisition module is used to acquire image information of the invoice; The memory is used to store the program that implements a method for intelligent extraction and verification of document elements; A processor is used to load and execute programs stored in memory.

[0017] In summary, this application includes at least one of the following beneficial technical effects: 1. By linking the invoice image, extracted element information, and coding information, the system can not only identify the current filling format but also accurately match the standard format that should be used for this type of invoice. When inconsistencies are detected, the system can further analyze the specific manifestations of the differences and automatically determine the type of difference, triggering corresponding processing strategies. This transforms the originally vague error messages into specific and actionable correction guidelines, thereby improving the efficiency of invoice information processing. 2. When a logical conflict is detected, the system not only locates the core problematic field but also proactively traces related format differences and performs compliance simulation corrections on these differences. Through secondary logical verification after simulation, the system can intelligently determine whether the root cause of the conflict is an isolated data error or a systemic risk caused by the combined effect of multiple format differences. If the simulation correction can eliminate or significantly mitigate the conflict, it means that the current format mismatch problem faced by the user may be hiding a deeper level of logical contradiction risk. Based on this, the system generates a collaborative anomaly warning and dynamically adjusts subsequent processing strategies. This allows for the early identification and warning of complex and systemic errors in user operations, guiding users to correct problems at their root and avoiding new conflicts caused by partial corrections. This significantly improves the overall reliability, compliance, and effectiveness of decision support in document processing. 3. Based on the standard format and missing locations, the system dynamically determines the compliant value range and automatically generates intelligent pre-filled values ​​as a reference benchmark, combining these with identified elements. The system simultaneously displays pre-filled suggestions and the standard value range to the user in the interactive interface, guiding them to quickly confirm or correct their input. When user input exceeds the standard range, the system does not simply report an error; instead, based on the compliant range and pre-filled suggestions, it proactively generates guiding examples of compliant content or boundary prompts, guiding the user to adjust their input within the valid range. This process transforms passive verification into proactive and user-friendly compliance guidance, freeing users from memorizing complex format rules. It significantly reduces the operational barriers and interruption risks caused by missing or non-standard information, improving the completeness and quality of the invoice dataset while optimizing human-computer interaction efficiency and user experience. Attached Figure Description

[0018] Figure 1 This is a flowchart of a method for intelligent extraction and verification of document elements; Figure 2 This is a flowchart of the method for pre-filling missing information and risk guidance. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0020] Reference Figure 1 This application discloses a method for intelligent extraction and verification of invoice elements, including the following steps: S10: Collect ticket image information.

[0021] Ticket image information refers to the digital image of a ticket that requires element extraction and verification. Ticket image information is obtained by having a user take a picture of the physical ticket using a mobile device's camera, or by scanning the ticket using a high-precision scanner.

[0022] S11: Extracting document element information based on document image information.

[0023] Bill element information refers to the structured business data identified and parsed from the bill image. This information is obtained by parsing the bill image using optical character recognition (OCR) technology and natural language processing (NLP) algorithms, and includes, but is not limited to, the bill number, issue date, maturity date, bill amount, full name of the drawer, full name of the payee, full name of the acceptor, and bank information. OCR technology and NLP algorithms are common knowledge in this field and will not be elaborated upon here.

[0024] S12: Obtain the current information filling format based on the invoice image information and invoice element information.

[0025] The current information filling format refers to the actual layout and writing specifications of each element on the ticket identified from the ticket image.

[0026] The current information filling format is obtained by analyzing the visual position distribution, font type, font size, and text alignment of various elements in the ticket image. The specific analysis method is based on the principles of image processing and pattern recognition, which are well-known technologies in this field and will not be elaborated here.

[0027] S13: Obtain the bill code information based on the bill element information.

[0028] Ticket coding information refers to the feature code used to identify the type of ticket.

[0029] The bill coding information is obtained by parsing the bill number prefix rule and the bill category identifier field in the bill element information. The specific parsing rules are preset by those skilled in the art according to the bill type specification, and will not be elaborated here.

[0030] S14: Match the standard information filling format using the invoice code information.

[0031] Standard information filling format refers to the standardized filling format template for various types of invoices.

[0032] The standard information filling format is obtained by matching the invoice code information with a pre-set invoice template database. The database stores standard format specifications corresponding to different invoice types, including standard parameters such as element positions, font requirements, and seal positions. The invoice template database is set in advance by those skilled in the art and will not be described in detail here.

[0033] S15: If the current information filling format is inconsistent with the standard information filling format, generate a format difference report by combining the current information filling format and the standard information filling format.

[0034] Format differences refer to the specific differences between the current format for filling out invoice elements and the standard format.

[0035] The format difference analysis is performed by comparing the current information entry format with the standard information entry format element by element to identify the difference types such as position deviation, font mismatch, missing or redundant elements, and generate a structured difference report.

[0036] S16: Determine the method for handling differences based on the format differences, and provide prompts according to the method for handling differences.

[0037] Difference handling methods refer to the processing strategies adopted for different format difference types.

[0038] The method for handling differences is determined by querying a preset difference handling strategy table. The strategy table defines the handling methods corresponding to different difference types. The specific content of the difference handling strategy table is configured by those skilled in the art according to business specifications, and will not be elaborated here.

[0039] Once you receive the method for handling differences, you should follow the instructions provided.

[0040] The specific methods for handling differences will be explained in detail in subsequent sections S20 to S25 and S30 to S34, and will not be repeated here.

[0041] Formatting differences include incorrect information and incomplete information. The methods for handling differences arising from incorrect information include: S20: When the format difference situation is an information filling error situation, obtain the specific filling error content based on the format difference situation.

[0042] Incorrect information entry refers to situations where the location, format, or content of the invoice elements does not conform to the standard information entry format requirements.

[0043] The specific errors refer to the differences in the elements identified by comparing the current information filling format with the standard information filling format.

[0044] By comparing the current information filling format with the standard information filling format element by element, the elements that differ from each other are identified. Then, the actual content filled in for these differing elements in the current ticket image is extracted, i.e., the specific incorrect content.

[0045] S21: Determine the location of the error based on the ticket code information and the specific incorrect content.

[0046] The location of the error content refers to the position of the specific incorrectly filled content in the ticket image.

[0047] The location of the error is determined by matching the specific incorrect content with the standard coordinate template corresponding to the invoice code information. The standard coordinate template is pre-stored in the invoice template database.

[0048] S22: Fill in the format and location of errors according to the standard information to know the type of content to be filled in correctly.

[0049] The correct content type refers to the standard content type that should be filled in when there is an error.

[0050] The correct content type can be obtained by querying the content type definition table corresponding to the standard information filling format. The content type definition table records the standardized content types that should be filled in each standard position on the invoice.

[0051] The specific method is as follows: Based on the coordinate information of the location of the erroneous content, match the corresponding standard location identifier in the standard information filling format, and then query the content type definition table through this identifier to obtain the correct content type specified for that location, including but not limited to numeric, text, date, amount, and coded types. The content type definition table is pre-set by those skilled in the art according to the invoice specifications and will not be elaborated here.

[0052] S23: Obtain a single error type based on the specific error content.

[0053] A single incorrect content type refers to the error classification of a specific incorrect content in terms of format, type, or semantics.

[0054] Individual error types are determined by analyzing the characteristic attributes of the specific incorrect content and matching it with preset error type classification criteria. Error types include, but are not limited to, format errors, type mismatches, semantic errors, and positional deviations. Specific classification criteria are preset by those skilled in the art according to the requirements of invoice specifications and will not be elaborated here.

[0055] S24: When a single incorrectly filled content falls into the correct content type, combine the single incorrectly filled content with the standard information filling format to generate a content position correction plan.

[0056] The content location correction scheme refers to a scheme that automatically adjusts the specific incorrect content in the incorrect content location.

[0057] The content position correction scheme is obtained by calculating the coordinate offset between the specific incorrect content and the standard position, and generating the corresponding position adjustment parameters.

[0058] When a single incorrectly filled-in item falls into a correctly filled-in item type, it indicates an error in the placement of the content on the form. The system automatically generates a position exchange instruction to move the incorrectly filled-in item to the standard position. The calculation method for the coordinate offset and position adjustment parameters is based on the principle of geometric transformation in image processing, which is a well-known technique in the field and will not be elaborated here.

[0059] S25: Based on the content location correction scheme, the position of the specific incorrectly filled content at the incorrect content location is corrected.

[0060] Once the content position correction plan is obtained, the specific incorrect content in the incorrect position should be corrected according to the content position correction plan, so that the specific incorrect content is moved to the correct position specified in the standard information filling format.

[0061] Methods for handling discrepancies due to incomplete information include: S30: When the format difference situation is an incomplete information situation, the specific missing location is obtained based on the format difference situation.

[0062] Incomplete information refers to situations where the required elements of the standard information filling format are missing from the invoice.

[0063] The specific missing location refers to the standard position of the missing element in the invoice, identified by comparing the current information filling format with the standard information filling format.

[0064] By comparing the current information entry format with the standard information entry format element by element, the essential elements missing in the current document image are identified. Then, based on the positional definition of these essential elements in the standard information entry format, their standard position coordinates in the document template are determined. Finally, the standard position coordinates of all identified missing elements are recorded to form a specific set of missing positions.

[0065] S31: Combine the specific missing location with the standard information filling format to determine the specific missing content.

[0066] The specific missing content refers to the standardized content that should be filled in at the specific missing location.

[0067] By understanding the coordinates of the specific missing location, the corresponding standard location identifier is matched in the standard information filling format. Then, the identifier is used to query the pre-set content specification database to obtain the essential element content requirements for that location, including but not limited to element name, content format, value range, filling rules, and other specification descriptions. The content specification database records the complete content specifications corresponding to each standard location. The content specification database is pre-established by those skilled in the art based on the invoice type specifications, and will not be elaborated here.

[0068] S32: Generate a content supplementation request based on the specific missing content.

[0069] A content supplementation request is a structured request sent by a user to supplement missing information.

[0070] The content supplementation request encapsulates the standardized description information such as the element name, content format, and filling rules of the missing content, along with the location coordinates of the missing location, and generates a structured data request according to a preset request template. The request template defines necessary fields such as request identifier, missing element type, format requirements, and location hints. The specific template structure is set by those skilled in the art according to the system interaction specifications and will not be elaborated here.

[0071] S33: In response to a content supplementation request, generate a content supplementation interface and collect information from the user to supplement the missing information.

[0072] The content supplementation interface is a user interface specifically designed for collecting missing information. When a content supplementation request is received, this interface needs to be generated for the user to input additional information. The method for generating the content supplementation interface is common knowledge in this field and will not be elaborated upon here.

[0073] User-supplemented missing information refers to compliant content data entered by the user through the content supplementation interface.

[0074] Users supplement missing information by collecting user input data through interactive controls provided on the content supplementation interface. This data is then processed through front-end input validation and back-end format verification. The validation rules are consistent with the format constraints defined in the content specification database to ensure that the user input conforms to the specifications of the document elements. Specific validation methods include format verification, range checking, and logical validation, which are well-known technologies in the field and will not be elaborated upon here.

[0075] S34: Integrate the missing information provided by the user with the information on the invoice elements to form a complete invoice dataset.

[0076] A complete bill dataset is obtained by merging the missing information provided by the user into the original bill element information according to the specific location of the missing information.

[0077] It also includes methods for verifying the authenticity of invoices: S40: Extract digital watermark features based on ticket image information.

[0078] Digital watermark features refer to the digital characteristics of anti-counterfeiting marks embedded in the image of a ticket.

[0079] Digital watermark features are obtained by parsing the ticket image using a digital watermark extraction algorithm. The algorithm identifies and extracts hidden watermark information by analyzing the image's frequency domain features. The specific extraction algorithm is well-known in the field and will not be elaborated here.

[0080] S41: Determine the baseline watermark features based on the invoice code information.

[0081] The benchmark watermark feature refers to the standard anti-counterfeiting watermark feature template corresponding to this type of document.

[0082] The baseline watermark feature is obtained by matching the corresponding standard watermark feature data in a pre-set watermark feature database using the invoice coding information as the query condition. The watermark feature database is provided by the invoice issuing institution and stores standard watermark feature data for various types of invoices.

[0083] S42: Combine digital watermark features and baseline watermark features to calculate watermark matching degree.

[0084] Watermark matching degree refers to the quantitative value of the similarity between the extracted digital watermark features and the baseline watermark features.

[0085] The watermark matching degree is obtained by calculating the cosine similarity between the digital watermark feature vector and the reference watermark feature vector. The specific calculation method is based on the basic principle of feature matching and is a well-known technology in this field, so it will not be elaborated here.

[0086] S43: When the watermark matching degree is lower than the preset standard watermark matching degree, report an abnormal ticket prompt.

[0087] Standard watermark matching degree refers to the minimum similarity threshold used to determine the authenticity of digital watermarks on invoices.

[0088] The standard watermark matching degree is set in advance by those skilled in the art, and will not be elaborated here.

[0089] When the watermark matching degree is lower than the standard watermark matching degree, it indicates that the authenticity of the invoice is low. An invoice anomaly alert should be reported, and relevant risk control personnel should be notified to conduct manual review.

[0090] Methods for verifying the authenticity of invoices also include: S50: When the watermark matching degree is not lower than the preset standard watermark matching degree, the historical circulation status record is determined based on the ticket code information.

[0091] Historical circulation status records refer to the serialized data of the time, subject, and status changes of key business operations such as endorsement, discounting, and pledging of the bill within the financial system from the date of issuance.

[0092] Historical transaction status records are retrieved by querying the pre-defined bill registration and settlement system using the bill code information as an index. The bill registration and settlement system is accessed by those skilled in the art according to business requirements, and will not be elaborated upon here.

[0093] When the watermark matching degree is not lower than the preset standard watermark matching degree, in order to further verify the authenticity of the document, it is necessary to first determine the historical circulation status record for subsequent steps.

[0094] S51: Determine the current date and current holder information of the negotiable instrument based on the instrument's element information.

[0095] The current date of a bill refers to the current date recorded in the bill's information elements and used for business processing.

[0096] The current holder information refers to the identity information of the user who submitted the ticket for verification, including the user name and account identifier.

[0097] The current date and current holder information of the negotiable instrument are obtained by directly extracting the corresponding fields from the instrument's element information.

[0098] S52: Determine the current status statement by combining the current date of the negotiable instrument with the current holder's information.

[0099] A current status statement is a description of the ownership and operability of a negotiable instrument at the current moment, based on the current date and current holder information.

[0100] The current status statement is obtained by formatting the information that the current holder holds the instrument and has applied for processing on the current date of the instrument into a standard status description string.

[0101] S53: Perform a logical continuity comparison between the current state declaration and the historical transition state records.

[0102] Logical continuity comparison refers to checking whether the state described in the current state declaration can be deduced from the historical state records and whether there are any time paradoxes or ownership contradictions.

[0103] Logical continuity comparison involves parsing historical transaction records to construct a time-based state transition chain, and verifying whether the holder in the current state declaration is the legitimate recipient of the previous state, and whether the current date of the bill is later than the previous state change date. The specific comparison logic is pre-set by those skilled in the art based on the bill business rules, and will not be elaborated here.

[0104] S54: If the logical continuity comparison fails, the bill status is determined to be abnormal, and an abnormal bill prompt containing specific continuity conflicts is reported.

[0105] An abnormal status of a negotiable instrument refers to a situation where the current claimed status of the instrument cannot form a legitimate and continuous business logic chain with its historical circulation records.

[0106] If the logical continuity comparison fails, the bill status is determined to be abnormal. The system generates a bill abnormality prompt containing specific conflict points (such as "the holder is not the previous endorser" or "the application date is earlier than the most recent pledge date") and reports it.

[0107] It also includes a method for verifying the logical consistency of invoices: S60: Obtain key logical fields based on document element information.

[0108] Key logical fields refer to the fields in the bill element information that participate in the calculation of business rules and whose values ​​have inherent constraints.

[0109] The key logical fields are obtained by filtering out fields such as bill amount, issuance date, maturity date, acceptor type, and bill type from the bill element information. The specific filtering rules are preset by those skilled in the art based on the bill business rules, and will not be elaborated here.

[0110] S61: Match the relevant business rule base based on the invoice code information.

[0111] The associated business rule base refers to the set of rules associated with the current invoice type, used to verify the rationality of its business logic.

[0112] The associated business rule base is obtained by matching and querying invoice coding information with a pre-set business rule knowledge graph. This knowledge graph stores the laws, regulations, industry practices, and risk control rules that various types of invoices should comply with. The business rule knowledge graph is constructed in advance by those skilled in the art and will not be elaborated here.

[0113] S62: Verify the constraint relationships of key logical fields by associating them with the business rule base.

[0114] Constraint validation refers to checking whether the values ​​of key logical fields meet the established calculation relationships, range restrictions, or logical conditions based on business rules.

[0115] Constraint validation is performed by calling the rule engine defined in the associated business rule library. The rule engine substitutes the values ​​of key logical fields into preset rule conditions for logical judgment. The implementation of the rule engine is a well-known technology in this field and will not be described in detail here.

[0116] S63: If a key logical field violates a preset constraint relationship, a logical conflict report will be generated by combining the key logical field with the associated business rule base.

[0117] Constraint relationships refer to specific validation conditions defined in the business rule base, such as "the amount of the bill shall not exceed the credit limit" and "the due date must be later than the issuance date".

[0118] A logic conflict report is a structured document that records in detail which key logical fields violate which specific constraint.

[0119] If the verification violates the constraint relationship, the system extracts the field value of the violation, the specific rule clause that was triggered, and generates a logical conflict report in a formatted manner.

[0120] S64: Determine the level of logical anomaly based on the logical conflict report, and provide graded early warning prompts according to the level of logical anomaly.

[0121] Logical anomaly level refers to the classification based on the importance and risk level of the violated constraints.

[0122] The level of logical anomalies is determined by querying a pre-defined rule severity lookup table. This table defines the anomaly levels (such as "Warning," "General Anomaly," and "Severe Anomaly") corresponding to violations of different business rules. The system matches the corresponding level to the rule clauses in the logical conflict report and triggers the appropriate warning prompt (such as an interface pop-up, SMS notification, etc.). The rule severity lookup table is set by those skilled in the art based on risk control strategies and will not be elaborated here.

[0123] It also includes user behavior risk assessment methods: S70: Use logical conflict reports to locate the core conflict fields.

[0124] The core conflict field refers to the core document element field that is marked in the logical conflict report as directly causing a rule violation.

[0125] The core conflict fields are obtained by semantically parsing the logical conflict report and identifying the field names that are the subjects in the violation conditional sentences.

[0126] S71: Filter out related format difference items that are associated with the conflicting core field based on the format difference.

[0127] Related format difference items refer to format difference items that are adjacent to or belong to the same filling area on the document page as the conflicting core field in the case of format difference.

[0128] The correlation format difference items are obtained by calculating the position coordinates of the conflicting core field in the standard information entry format and performing spatial proximity analysis with the position coordinates of all format difference items, filtering out the difference items whose distance is less than a preset threshold. The specific threshold is set by those skilled in the art and will not be elaborated here.

[0129] S72: By comparing the information on bill elements with the standard information filling format, conduct compliance simulation corrections for differences in related formats.

[0130] Compliance simulation correction refers to the process of virtually correcting the related format differences within the system according to the standard information filling format, and updating the temporary copy of the invoice element information.

[0131] The compliance simulation correction calls processing logic similar to S24 or S31 to S34 to virtually correct the position or fill in the content of the identified related format differences, generating a corrected simulated version of the bill element information that assumes compliance.

[0132] S73: Based on the simulated and corrected bill element information, perform the constraint relationship verification in the logical consistency verification method again.

[0133] Using the simulated and corrected bill element information as input, the constraint relationship verification process of steps S60 to S62 is re-executed.

[0134] S74: If the conflict in the logical conflict report is eliminated or mitigated after verification, a coordination anomaly warning is generated.

[0135] Collaborative anomaly alerts are risk warning signals that indicate a potential correlation between "formatting errors" and "logical errors," possibly pointing to abnormal human operation or fraudulent attempts.

[0136] If the original logical conflict is eliminated or its severity is reduced after simulation correction, the system generates a collaborative anomaly warning, which includes the associated format differences and details of the logical conflict.

[0137] S75: Adjust the strategy for handling differences based on the level of collaborative anomaly warnings and logical anomalies.

[0138] Strategy adjustment refers to increasing the processing priority of the invoice, strengthening manual review requirements, or triggering a more stringent risk control review process based on the existence of collaborative anomaly alerts and the level of logical anomaly.

[0139] Strategy adjustments are achieved by querying a pre-defined risk control strategy matrix. This matrix defines the escalation plans for handling different levels of logical anomalies and the presence or absence of coordinated anomaly alerts. The risk control strategy matrix is ​​pre-configured by those skilled in the art and will not be elaborated upon here.

[0140] Reference Figure 2 It also includes methods for pre-filling missing information and risk guidance: S80: Determine the standard value range based on the standard information filling format and the specific missing location.

[0141] The standard value range refers to the range and format requirements of numerical or character values ​​that are allowed to be filled in at specific missing positions, in accordance with business specifications.

[0142] The standard value range is obtained by querying the same content specification database as in step S31, and the content format, value range and other constraints corresponding to the missing position are directly extracted.

[0143] S81: Generate pre-filled suggested values ​​based on bill element information and standard value range.

[0144] Pre-filled suggested values ​​refer to the values ​​that the system recommends to users based on the already filled in document elements and business rules, which are the most likely to be correct for filling in missing information.

[0145] The pre-filled suggested values ​​are generated by analyzing other related fields in the bill's information (such as inferring the issuer's frequently used banks based on historical data) and filtering for reasonableness within a standard value range. The specific generation algorithm can be designed by those skilled in the art based on business logic, and will not be elaborated here.

[0146] S82: Display the pre-filled suggested values ​​and standard value ranges in the preset content supplementation interface, and collect user modification operations.

[0147] The content supplementation interface refers to the interactive interface generated in step S33, which guides users to supplement information.

[0148] User modification operations refer to actions taken by users on this interface to edit, delete, or confirm pre-filled suggested values.

[0149] The system highlights the pre-filled suggested values ​​and their basis in the content supplementation interface, and simultaneously displays the standard value range as a prompt. All user keyboard and mouse operations are collected through an interface event listening mechanism.

[0150] S83: Obtain user input based on user modification operations.

[0151] User input refers to the text or numerical value that the user ultimately submits in the input box on the content completion interface.

[0152] User input is obtained by capturing the input field value when the user completes input and triggers a confirmation event.

[0153] S84: If the user input exceeds the standard value range, combine the standard value range with the pre-filled suggested value to generate compliant input.

[0154] Compliant input refers to a suggested correction value that conforms to the standard value range and is used to replace illegal user input.

[0155] The compliant input method compares the user's input with a standard value range. If the input exceeds the range, the system automatically selects the closest suggested value from pre-filled suggestions or range boundary values. The specific selection strategy is determined by those skilled in the art and will not be elaborated here.

[0156] S85: Based on compliant input content, guide users to correct the input content to the standard value range and generate user-supplemented missing information.

[0157] The user-supplemented missing information is defined in S33, and here it refers to the valid supplementary information that is ultimately accepted and recorded by the system.

[0158] The system displays compliant input to users through pop-up windows or inline prompts, guiding users to adopt the input or make secondary corrections based on it. When the value entered or ultimately adopted by the user falls within the standard value range, that value is recorded as the user's supplementary missing information.

[0159] S86: Integrate the missing information provided by the user with the information on the invoice elements to form a complete invoice dataset.

[0160] Similar to step S34, the finalized user-supplemented missing information is merged into the invoice element information.

[0161] Also includes: S90: Generate a multi-dimensional set of risk signals based on format differences, logical conflict reports, and collaborative anomaly alerts.

[0162] A multi-dimensional risk signal set refers to a collection of risk indicators extracted from different dimensions such as format compliance, logical rationality, and operational behavior relevance.

[0163] The multidimensional risk signal set is obtained by quantifying the key conclusions (such as the number of differences, the level of conflict, and the warning signs) in the above three types of reports / alerts into numerical risk signals and packaging them into a feature vector.

[0164] S91: Determine the risk level of related bills based on bill code information.

[0165] The risk level of related bills refers to the basic risk level of the bills pre-assessed based on static information such as bill type, acceptor qualifications, and industry attributes.

[0166] The risk level of the associated bill is obtained by matching the bill code information with a pre-set bill risk profile database. This database stores information such as historical default rates and risk scores for different types of bills, and is constructed by those skilled in the art based on historical data, which will not be elaborated here.

[0167] S92: Determine the fusion weight configuration based on a multi-dimensional set of risk signals and the risk level of associated bills.

[0168] The fusion weight configuration refers to the importance coefficients assigned to risk signals and basic risk levels in various dimensions in the comprehensive risk assessment model.

[0169] The fusion weight configuration is determined by querying a preset dynamic weight table. This table assigns different weights to various risk signals based on the risk level of the associated bills. For example, for bills with high basic risk, logical conflicts have a higher weight. The dynamic weight table is configured by those skilled in risk control based on their experience and will not be elaborated upon here.

[0170] S93: Based on the fusion weight configuration, perform weighted fusion calculation on the multi-dimensional risk signal set to generate a comprehensive risk judgment result and disposal confidence level.

[0171] The comprehensive risk assessment result refers to the quantitative scoring or classification of the overall risk level of the bill.

[0172] The confidence level of the treatment refers to the system's self-assessment score of the reliability of the above-mentioned automatic judgment results.

[0173] A weighted summation fusion model is employed to calculate multi-dimensional risk signals according to fusion weights, outputting a comprehensive risk score and the confidence probability of the model's prediction based on that score. The specific fusion model is a well-known technique in the field and will not be elaborated upon here.

[0174] S94: If the confidence level of the handling is higher than the preset automatic handling threshold, the corresponding differential handling method, graded early warning prompt or collaborative abnormal warning response process will be automatically triggered and executed based on the comprehensive risk assessment result.

[0175] The automatic handling threshold refers to the lowest confidence level set by the system to determine whether a handling process can be executed automatically without human intervention.

[0176] The automatic processing thresholds are set in advance by those skilled in the art and will not be elaborated here.

[0177] If the confidence level of the action is higher than the threshold, the system will map the comprehensive risk assessment result to the preset action (such as automatic rejection, increased verification strength, marking as high risk, etc.) and automatically execute the processing, warning or response process defined in steps S16, S54, S75, etc.

[0178] It also includes a method for applying the above-mentioned intelligent extraction and verification method of bill elements to the intelligent inquiry scenario of bill discounting, which includes the following steps: S100: Collects users' historical transaction data, partner institutions' historical service data, and users' expected conditions.

[0179] User historical transaction data refers to the collection of behavioral and result data generated when users conduct historical discounting operations on this platform, including discounting frequency, historical transaction interest rates, and the period from inquiry to transaction.

[0180] Historical service data of partner institutions refers to the collection of performance data accumulated by various discounting institutions that cooperate with the platform during the historical service process, including authentication efficiency, business processing time, and customer repeat transaction rate.

[0181] User expectations refer to the target parameters proposed by the user for this discounting transaction, mainly the expected interest rate range.

[0182] User historical transaction data and partner institution historical service data are obtained by extracting and aggregating them from a pre-set platform business database. User expectations are obtained through user input collected via an interactive interface.

[0183] S101: Generate customer profiles based on complete bill datasets and user historical transaction data.

[0184] Customer profiling refers to a data model used to quantitatively describe a user's discounting behavior preferences and risk characteristics.

[0185] Customer profiles are constructed by analyzing the bill type, amount, and acceptor information in a complete bill dataset, and combining this with features such as discount frequency, price sensitivity, and transaction cycle from the user's historical transaction data, using a feature fusion algorithm. The specific algorithm used for construction was designed by those skilled in the art and will not be elaborated here.

[0186] S102: Generate an institutional profile based on the complete bill dataset and historical service data of partner institutions.

[0187] Institutional profiling refers to a data model used to quantitatively assess the service capabilities and suitability characteristics of discounting institutions.

[0188] Institutional profiles are constructed by analyzing the business needs (such as amount and term) of a complete bill dataset and combining them with features such as processing efficiency, risk preference, and customer matching degree from the historical service data of partner institutions, using a feature fusion algorithm.

[0189] S103: Combine customer profiles, organizational profiles, and user expectations to generate a target execution plan.

[0190] The target execution plan refers to the plan that the system calculates and recommends to the customer the best discounting institution and its quotation.

[0191] The target execution plan is generated through the following sub-steps: S1031: Calculate the matching score based on customer profiles and organization profiles. The matching score is obtained by calculating the similarity between the vectors of the two entities in a multi-dimensional feature space.

[0192] S1032: Select institutions with a matching score higher than the threshold and whose quotes meet the user's expectations as candidate institutions.

[0193] S1033: Obtain real-time quotes from candidate organizations and rank them based on a comprehensive score, taking into account their service ratings in the organization profile.

[0194] S1034: Identify the candidate organization with the highest score and its offer as the target implementation plan and push it to the user.

[0195] In this application scenario, the intelligent extraction and verification method of bill elements recorded in steps S10 to S94 provides accurate, reliable and risk-controllable core bill data for subsequent profile construction and intelligent matching, which is the foundation for ensuring the intelligence and security of the entire discount inquiry process.

[0196] Based on the same inventive concept, embodiments of the present invention provide an intelligent extraction and verification system for invoice elements, including: The data acquisition module is used to collect ticket image information, allow users to supplement missing information, and handle user modification operations. The memory is used to store the program that implements a method for intelligent extraction and verification of document elements; A processor is used to load and execute programs stored in memory.

[0197] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0198] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for intelligent extraction and verification of document elements, characterized in that, include: Collect ticket image information; Extracting essential information from invoice images; The current information filling format is obtained based on the invoice image information and invoice element information; Based on the information of the elements of the bill, the bill code information is obtained; The standard information filling format is matched using the invoice code information; If the current information filling format is inconsistent with the standard information filling format, the format difference will be generated by combining the current information filling format and the standard information filling format. Based on the format differences, determine the method for handling the differences, and then process the prompts according to the method for handling the differences.

2. The method for intelligent extraction and verification of document elements according to claim 1, characterized in that, Formatting differences include incorrect information and incomplete information. The methods for handling differences arising from incorrect information include: When the format difference situation indicates an information filling error, the specific filling error content is obtained based on the format difference situation; Determine the location of the error based on the invoice code information and the specific incorrect content; Based on the standard information filling format and the location of incorrect content, you can know the type of content that should be filled in correctly; The individual error content type is obtained based on the specific incorrect content; When a single incorrect entry falls into the correct entry type, a content position correction plan is generated by combining the single incorrect entry with the standard information entry format. The content location correction scheme corrects the specific errors in the content.

3. The intelligent extraction and verification method for document elements according to claim 2, characterized in that, Methods for handling discrepancies due to incomplete information include: When the format difference situation results in incomplete information, the specific missing location is determined based on the format difference situation. Determine the specific missing content by combining the specific location of the missing information with the standard information filling format; generate a content supplementation request based on the specific missing content; In response to a content supplementation request, generate a content supplementation interface and collect the user's supplemented missing information; The missing information provided by the user is integrated with the invoice element information to form a complete invoice dataset.

4. The intelligent extraction and verification method for document elements according to claim 3, characterized in that, It also includes methods for verifying the authenticity of invoices: Extracting digital watermark features based on ticket image information; The basic watermark features are determined based on the invoice coding information; The watermark matching degree is calculated by combining digital watermark features and baseline watermark features. When the watermark matching degree is lower than the preset standard watermark matching degree, an abnormal ticket notification will be reported.

5. The intelligent extraction and verification method for document elements according to claim 4, characterized in that, The method for verifying the authenticity of invoices also includes: When the watermark matching degree is not lower than the preset standard watermark matching degree, the historical circulation status record is determined based on the ticket code information. The current date and current holder information of the negotiable instrument are determined based on its key information. The current status statement is determined by combining the current date of the negotiable instrument with the current holder's information; Perform a logical continuity comparison between the current state declaration and the historical transition state records; If the logical continuity comparison fails, the bill status is determined to be abnormal, and an abnormal bill notification containing specific continuity conflicts is reported.

6. The intelligent extraction and verification method for document elements according to claim 1, characterized in that, It also includes a method for verifying the logical consistency of invoices: Key logical fields are obtained based on the information of the invoice elements; Based on the invoice code information, the relevant business rule base is matched; Constraint relationships of key logical fields are validated by associating them with the business rule base. If a key logical field violates a preset constraint, a logical conflict report will be generated by combining the key logical field with the associated business rule base. Logical conflict reports are used to determine the level of logical anomalies, and graded early warning prompts are issued according to the level of logical anomalies.

7. The intelligent extraction and verification method for document elements according to claim 6, characterized in that, It also includes user behavior risk assessment methods: Use logical conflict reports to pinpoint the core conflict fields; Based on the format differences, filter out the related format difference items that are associated with the core conflicting fields; By comparing the format of invoice element information with the standard information, compliance simulation and correction are performed on the differences in related formats. Based on the simulated and corrected bill element information, the constraint relationship verification in the logical consistency verification method is executed again. If the verification shows that the conflict in the logical conflict report has been eliminated or mitigated, a coordination anomaly warning will be generated. Based on the level of collaborative anomaly warnings and logical anomalies, the strategy for handling differences is adjusted.

8. The intelligent extraction and verification method for document elements according to claim 3, characterized in that, It also includes methods for pre-filling missing information and risk guidance: The standard value range should be determined based on the standard information filling format and the specific missing locations. Pre-filled suggested values ​​are generated based on invoice element information and standard value range; Based on the pre-filled suggested values ​​and standard value ranges, the content is displayed in the preset content supplementation interface, and user modification operations are collected; Based on the user's modification actions, obtain the user's input content; If the user input exceeds the standard value range, combine the standard value range with the pre-filled suggested value to generate compliant input. Based on compliant input content, guide users to correct their input content to the standard value range and generate information for users to supplement missing information; The missing information provided by the user is integrated with the invoice element information to form a complete invoice dataset.

9. The intelligent extraction and verification method for document elements according to claim 7, characterized in that, Also includes: A multi-dimensional risk signal set is generated based on format differences, logical conflict reports, and collaborative anomaly alerts. The risk level of related bills is determined based on the bill code information; The fusion weight configuration is determined based on a multi-dimensional set of risk signals and the risk level of associated bills. Based on the fusion weight configuration, a multi-dimensional risk signal set is weighted and fused to generate a comprehensive risk assessment result and a response confidence level. If the confidence level of the handling is higher than the preset automatic handling threshold, the corresponding differential handling method, graded early warning prompt or collaborative abnormal warning response process will be automatically triggered and executed based on the comprehensive risk assessment result.

10. A smart system for extracting and verifying document elements, characterized in that, include: The acquisition module is used to acquire image information of the invoice; A memory for storing a program that implements the intelligent extraction and verification method for document elements as described in any one of claims 1 to 9; A processor is used to load and execute programs stored in memory.