A method for improving the association rate of bank transaction details and electronic receipts
By automatically acquiring bank transaction details and electronic receipts through the direct bank-enterprise connection interface, and optimizing the matching using multi-factor association rules, the problems of cumbersome data collection and low association rate are solved. This achieves efficient and accurate association of transaction details and electronic receipts, generating auditable and traceable electronic vouchers to support enterprise fund management and compliance checks.
Patent Information
- Application Number
- CN202511472456.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-10-15
AI Technical Summary
When enterprises process the association between bank transaction details and electronic receipts, they face problems such as cumbersome data collection, long processing time, and low association rate, which increases the workload of financial staff and leads to untimely preparation of audit materials.
The system automatically acquires transaction details and electronic receipts from multiple bank accounts through a direct bank-enterprise connection interface. It then uses identification and analysis to generate structured receipt data, constructs multi-factor association rules, and matches these rules with transaction serial numbers, dates, amounts, and counterparty information. The system also optimizes the matching rules through a feedback mechanism to achieve automatic association.
It reduces manpower and time consumption for data collection, improves correlation accuracy, ensures data consistency, detects abnormal transactions in real time, reduces risks, generates auditable and traceable electronic vouchers, and supports fund allocation and compliance management.
Smart Images

Figure CN120931387B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method for improving the correlation rate between bank transaction details and electronic receipts. Background Technology
[0002] When enterprises use their treasury systems to link bank transaction details with electronic receipts, they often find it difficult to systematically obtain electronic receipts from multiple banks in bulk. Financial personnel must log into each bank's corporate online banking system to download the receipts or go to the bank counter to print paper receipts. For example, when a manufacturing company processes monthly receipts from the Industrial and Commercial Bank of China (ICBC) and a local city commercial bank, logging into ICBC's corporate online banking requires verification via a USB key, followed by a process of account management, receipt query, and downloading each receipt individually. Downloading 50 receipts per account at a time takes an average of about 20 minutes. The local city commercial bank's online banking only supports downloading 10 receipts at a time, and requires manual selection of the date range. Processing monthly receipts for two accounts at this bank takes an additional 15 minutes. Furthermore, manual operation can easily lead to missing receipts due to missed account selections or omissions of certain date ranges.
[0003] Furthermore, although the treasury system can obtain transaction details, the association with electronic receipts largely relies on manual matching or can only achieve simple association of basic fields, resulting in a low overall association rate. When the treasury system obtains transaction details from the bank, if it only matches electronic receipts by transaction amount, it is easy to cause confusion where multiple transactions of the same amount correspond to multiple receipts. Financial personnel need to manually verify information such as counterparty name and transaction date one by one, which not only increases the workload of financial personnel, but may also affect the preparation of audit data due to untimely receipt completion. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide a method to improve the correlation rate between bank transaction details and electronic receipts, so as to enable enterprises to efficiently complete electronic receipt reconciliation and archiving, and improve the electronic and digital process of treasury enterprises.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, a method for improving the correlation rate between bank transaction details and electronic receipts, the method comprising:
[0007] Step 1: Automatically retrieve transaction details and electronic receipt data from multiple bank accounts opened by the enterprise through the bank-enterprise direct connection interface;
[0008] Step 2: Identify and analyze the electronic receipt data, extract key accounting elements including transaction serial number, date, amount and counterparty information, and generate structured receipt data;
[0009] Step 3: Construct a transaction data feature set based on structured return receipt data, determine the benchmark data point from the feature set, and generate two sets of data association features based on the benchmark data point;
[0010] Step 4: Based on the correlation features of the two sets of data, calculate the time, amount and the difference between the counterparty to construct a multidimensional feature dataset, and determine the effective range to select positive and negative samples; map the samples into three-dimensional spatial points according to the time series to form polygon boundaries, determine the position of sample points by the odd-inside-even-outside rule, count the number of inside and outside points to calculate the density of inside points, construct the correlation matching rule system and generate matching calibration parameters;
[0011] Step 5: Using the matching calibration parameters, based on the transaction serial number matching, and combined with the multi-factor association rules of transaction date, account, amount and counterparty name, the transaction details and electronic receipts are matched to obtain the matching results;
[0012] Step 6: Using the matching results and existing order data resources, the matching rules are continuously trained and dynamically adjusted through a feedback mechanism to obtain the adjusted matching rules.
[0013] Step 7: Based on the matching rules, automatically associate the successfully matched transaction details with the electronic receipts and generate auditable and traceable electronic vouchers.
[0014] In a second aspect, a computing device includes:
[0015] One or more processors;
[0016] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0017] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0018] The above-described solution of the present invention has at least the following beneficial effects:
[0019] By automatically acquiring transaction details and electronic receipts from multiple bank accounts through direct bank-enterprise connections, the system reduces manpower and time consumption in data collection. Real-time data acquisition and intelligent matching algorithms help corporate treasurers monitor fund transaction status in real time, providing timely data support for fund allocation and accounting. Data cleaning removes irrelevant characters and formatting interference from electronic receipts, and combined with preset rules and a keyword library, accurately extracts key accounting elements such as transaction serial numbers, amounts, and counterparties, avoiding errors in element extraction due to format differences and ambiguous information during manual receipt identification. Furthermore, standardized date formats and amount verification further ensure data consistency. Using the transaction serial number as the core, a multi-factor weighted association rule is constructed based on transaction date, account number, amount, and counterparty name. Compared to traditional single-field matching, this effectively avoids mismatches in scenarios such as different transactions of the same amount or multiple transactions with the same counterparty. Simultaneously, dynamic adjustment of rules through matching calibration parameters, combined with a feedback mechanism, continuously optimizes the model, ensuring stable association accuracy.
[0020] Through automatic matching and trajectory analysis, abnormal transactions with details but no receipts or receipts but no details can be detected in real time, triggering timely risk warnings and avoiding problems such as unverifiable fund flows and outstanding accounts due to missing receipts. This reduces the risks of fund misappropriation and transaction omissions. Through automated data collection and intelligent correlation matching, real-time integration and unified management of transaction data are achieved, helping enterprises build a data-driven treasury management model. This reduces process bottlenecks caused by manual intervention and automatically generates electronic vouchers containing key transaction information, receipt image indexes, associated timestamps, and audit trace identifiers. This fully records the correlation process and data source, meeting the compliance requirements of audits for traceable transactions and verifiable data, and avoiding problems such as lost vouchers and difficulties in tracing due to manual receipt processing. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating a method for improving the correlation rate between bank transaction details and electronic receipts, provided by an embodiment of the present invention.
[0022] Figure 2 This is a schematic diagram provided by an embodiment of the present invention, which describes the construction of a transaction data feature set based on structured return data, the determination of a benchmark data point from the feature set, and the generation of two sets of data association features based on the benchmark data point. Detailed Implementation
[0023] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0024] like Figure 1 As shown, an embodiment of the present invention proposes a method for improving the correlation rate between bank transaction details and electronic receipts, the method comprising the following steps:
[0025] Step 1: Automatically retrieve transaction details and electronic receipt data from multiple bank accounts opened by the enterprise through the bank-enterprise direct connection interface;
[0026] Step 2: Identify and analyze the electronic receipt data, extract key accounting elements including transaction serial number, date, amount and counterparty information, and generate structured receipt data;
[0027] Step 3: Construct a transaction data feature set based on structured return receipt data, determine the benchmark data point from the feature set, and generate two sets of data association features based on the benchmark data point;
[0028] Step 4: Based on the correlation features of the two sets of data, calculate the time, amount and the difference between the counterparty to construct a multidimensional feature dataset, and determine the effective range to select positive and negative samples; map the samples into three-dimensional spatial points according to the time series to form polygon boundaries, determine the position of sample points through the odd-inside-even-outside rule, count the number of inside and outside points to calculate the density of inside points, construct the correlation matching rule system and generate calibration parameters;
[0029] Step 5: Using the matching calibration parameters, based on the transaction serial number matching, and combined with the multi-factor association rules of transaction date, account, amount and counterparty name, the transaction details and electronic receipts are matched to obtain the matching results;
[0030] Step 6: Using the matching results and existing order data resources, the matching rules are continuously trained and dynamically adjusted through a feedback mechanism to obtain the adjusted matching rules.
[0031] Step 7: Based on the matching rules, automatically associate the successfully matched transaction details with the electronic receipts and generate auditable and traceable electronic vouchers.
[0032] In this embodiment of the invention, transaction details and electronic receipt data are automatically obtained from multiple bank accounts through a direct bank-enterprise connection interface. This eliminates the need for manual downloading and aggregation of data from different bank systems, avoiding the tediousness and time-consuming nature of manual operations and reducing manpower investment in the data collection process. Furthermore, the identification, analysis, extraction of key accounting elements, and structuring of electronic receipt data require no manual intervention throughout the entire process, shortening the data processing cycle and allowing corporate treasury personnel to focus more on core cash management tasks, thus improving overall work efficiency. During the data processing stage, the identification and analysis of electronic receipt data accurately extracts key accounting elements such as transaction serial numbers, dates, amounts, and counterparty information, generating structured receipt data. This avoids errors in element extraction caused by format differences and ambiguous information during manual receipt identification. By constructing a feature set based on structured data and determining benchmark data points, and combining multi-factor association rules for matching, and by determining the effective range and generating calibration parameters through difference measurement analysis, the mismatch problem that may occur in a single matching dimension is further reduced, and the accuracy of the association between transaction details and electronic receipts is improved. By utilizing the matching results and existing receipt data resources, the matching rules are continuously trained and dynamically adjusted through a feedback mechanism, which allows the rules to adapt to the diverse transaction scenarios of enterprises. For example, when new transaction types, changes in counterparty information formats, or adjustments to bank data fields occur, the feedback mechanism can promptly capture matching deviations. By iteratively optimizing rule parameters and association logic, the matching rules are ensured to maintain high applicability, avoiding a decrease in association rate due to changes in transaction scenarios, and ensuring the long-term stability of the association work.
[0033] After the transaction details and electronic receipts are automatically linked, auditable and traceable electronic vouchers can be generated, fully recording key transaction information, receipt linking process, and data sources. This reduces the manpower and storage costs of voucher management and better meets the requirements of corporate audit compliance for traceable and verifiable transaction data. When facing internal audits or external regulatory inspections, electronic vouchers and linked records can be quickly retrieved, improving the efficiency of compliance inspections and reducing compliance risks caused by missing vouchers or difficulties in tracing. Through efficient and accurate linking of transaction details and electronic receipts, corporate treasuries can have a real-time and clear understanding of the fund transactions of each bank account, including key information such as fund flow, counterparties, and transaction amounts. This avoids blind spots in fund management caused by untimely or inaccurate data linking. At the same time, the structured linked data can further support fund analysis work. For example, by statistically analyzing the frequency and amount of fund transactions between different counterparties, reliable data can be provided for corporate fund allocation, credit assessment, and cooperation decisions, helping to improve the precision of corporate treasury fund management.
[0034] In a preferred embodiment of the present invention, step 1 above, which automatically obtains transaction details and electronic receipt data from multiple bank accounts opened by the enterprise through a direct bank-enterprise connection interface, may include:
[0035] In this embodiment of the invention, information on all settlement accounts opened by the enterprise at various banks is collected, including account name, account number, bank code, and enterprise identification code assigned by the bank. The completeness and validity of this information are verified one by one to confirm that each account has completed the bank-enterprise direct connection permission activation and filing with the corresponding bank, and that the account status is normal. This eliminates subsequent data acquisition failures caused by incorrect account information or permission issues. For each bank's bank-enterprise direct connection interface specifications, dedicated interface access parameters are configured, including the bank-provided interface address (URL), the enterprise's dedicated API key, signature algorithm type (e.g., MD5, SHA256), request timeout (typically set to 30-60 seconds to avoid request anomalies due to network latency), and data transmission format (e.g., JSON, XML). Simultaneously, based on the enterprise's... The system allows businesses to set rules for the time range of data acquisition. By default, it acquires all transaction data from the previous day at midnight every day. It also supports custom time periods for businesses, such as incremental data acquisition rules by hour or by specific date ranges. According to preset time cycles, such as automatically triggering data acquisition tasks at fixed times every day, it generates independent transaction detail query requests for each enterprise account of each bank. The request content includes the configured interface parameters, account identifier, account number + bank code, data query time range, such as 2024-05-2000:00:00 to 2024-05-2023:59:59, and the required transaction detail fields, such as transaction serial number, transaction date, transaction time, transaction type, transaction amount, counterparty account name, counterparty account number, summary, handling fee amount, account balance, etc.
[0036] After the request is sent to the corresponding bank's direct bank-enterprise connection server, the bank's system first verifies the validity of the API key and signature information in the request. Once the enterprise's identity and account permissions are confirmed, the system retrieves the corresponding transaction records for that account in the bank's core transaction system based on the account identifier and time range in the request. During the retrieval process, the bank system filters invalid transactions, such as test transactions and reversed transactions, and organizes the valid transaction details data in chronological order to ensure data accuracy and completeness. Subsequently, the transaction details are packaged and returned to the enterprise's treasury system according to the agreed data format. The enterprise treasury system receives the transaction details data returned by the bank. Next, the system first checks whether the data transmission is complete. This includes checking if the number of fields and the length of the returned data match the agreement made in the request, and whether there are any missing fields or garbled characters. If the data is incomplete, the system automatically re-initiates the request, with the default number of retries set to 3, with a 5-second interval between each retrieval. If the system still cannot obtain complete data after retries, an exception log is recorded, including account information, request time, error message, and an alarm notification is triggered. If the data is complete, the system further verifies the validity of key fields, such as whether the transaction amount format is numeric, whether the transaction serial number is unique, and whether the account balance is positive, filtering out abnormal data that does not meet the format requirements.
[0037] Based on the acquired transaction details, a unique identifier and transaction serial number (generated by the bank) are extracted for each transaction, with each transaction corresponding to a unique serial number. This is then linked to basic fields such as account information and transaction date to form a receipt query index table. This ensures that each transaction can be matched with a unique query condition. Following the principle of grouping by bank and batching by account, the query conditions in the receipt query index table are converted into electronic receipt query requests in batches. Each request includes bank interface parameters, account identifier, transaction serial number, and receipt format requirements, such as PDF format and electronic signature. To avoid congestion of the bank interface due to a large number of requests in a single batch, the number of requests in each batch is typically controlled within 100-200. The system dynamically adjusts based on the interface capacity of different banks. The bank system retrieves the corresponding electronic receipt file based on the transaction serial number in the receipt query request. This file contains transaction details, the bank's electronic signature, and the receipt number. The receipt file is then returned to the enterprise treasury system as a binary stream or file link. Upon receiving the receipt data, the system associates the electronic receipt with the acquired transaction details one-to-one using the transaction serial number. It verifies key information in the receipt, such as transaction amount, transaction date, and account information, to ensure they match the transaction details. If they match, the association is completed; otherwise, if they do not match, such as a discrepancy between the receipt amount and the transaction details, it is marked as an abnormal receipt, the discrepancy is recorded, and a manual review process is triggered.
[0038] After completing the first round of data acquisition and matching, the total number of transaction details and electronic receipts for each bank account are counted, and the two are compared to see if they match. If there are cases where there are transaction details but no electronic receipts, the transaction serial number corresponding to the missing receipt is automatically identified, a request to retrieve the missing receipt is generated, and a second query is initiated for the missing receipt. The number of retries for retrieving the missing receipt is set to 2, with an interval of 10 seconds. If the second query still fails to retrieve the receipt, the account, transaction serial number, query time, and other information of the missing receipt are recorded, a receipt missing report is generated, and relevant operators are notified so that the receipt can be retrieved manually through the bank's online banking or counter channels. For successfully matched transaction details and electronic receipts, the number of transaction details is recorded. The data is structured and stored in a relational database, such as MySQL or Oracle, with fields including account identifier, transaction serial number, transaction time, amount, and counterparty information. Electronic receipt files are stored in a directory structure of bank-account-year-month on a file server, such as NAS or cloud storage, and metadata such as storage path, file size, and creation time of the receipt files are recorded in the database. This ensures that the corresponding details and receipts can be quickly retrieved by transaction serial number. At the same time, the complete data acquired daily is automatically backed up, and the backup data retention period is set according to the enterprise's financial audit requirements, usually 5-10 years, to ensure data security and traceability.
[0039] In a preferred embodiment of the present invention, step 2 above, which involves identifying and analyzing the electronic receipt data to extract key accounting elements including transaction serial number, date, amount, and counterparty information, and generating structured receipt data, may include:
[0040] In this embodiment of the invention, step 220 involves cleaning the original electronic receipt data obtained from bank accounts to remove irrelevant characters and formatting information, resulting in cleaned standardized text data. Specifically, this includes: reading the original electronic receipt data obtained from various bank accounts of the enterprise; identifying the specific data type using a format detection tool; if it is a PDF, further distinguishing between text-based and scanned PDFs; text-based PDFs can have their text directly extracted, while scanned PDFs require OCR technology to convert the image into text; if it is an XML or HTML format, marking the tag levels enclosing the text content; subsequently, performing preliminary splitting for different formats; for text-based PDFs, extracting the text content of each page and concatenating them in page number order; for scanned PDFs, converting the text in the image into editable text using OCR recognition, while filtering out blurry characters generated during OCR recognition; for XML / HTML formats, removing all tag symbols and retaining only the plain text content within the tags to form the original text dataset; and constructing... The bank receipt irrelevant information database contains three main categories: first, format-irrelevant characters such as spaces, tabs, line breaks, page breaks, and special symbols; second, redundant bank identification information such as the text corresponding to the bank logo and receipt generation auxiliary information; and third, invalid blank paragraphs, i.e., empty text segments containing only spaces or line breaks. Next, the original text dataset is scanned line by line, comparing each character in each line with the irrelevant information database. If a character belongs to the format-irrelevant category, it is deleted directly; if a text segment matches redundant bank identification information, the entire segment is removed; if a blank paragraph is detected, that paragraph is removed. The encoding format of the cleaned text is checked, and all text is uniformly converted to UTF-8 encoding to avoid garbled characters caused by encoding incompatibility. Simultaneously, full-width characters in the text are converted to half-width characters to ensure consistent formatting of numbers, letters, and punctuation marks; and English words in the text are uniformly converted to lowercase to avoid affecting subsequent keyword matching due to case differences. Finally, standardized text data with uniform formatting and no redundancy is generated.
[0041] Step 221: Based on the preset rules and keyword library, analyze and identify the cleaned standardized text data, extract key fields, including transaction serial number, transaction date, transaction amount, and counterparty information, and generate a key field set. Specifically, this includes: loading a key field identification rule library built for the characteristics of multi-bank receipts. This library contains two core parts: first, a keyword-field type mapping table, which records all possible expressions corresponding to each key accounting element; second, field extraction rules, which clarify the content positioning method of each field, split the standardized text data into independent sentences according to line breaks, scan each sentence character by character, and compare the keyword-field type mapping. If a keyword is found in the table, the sentence is determined to contain the corresponding field type. The field content is then located according to the field extraction rules. For the format "keyword + colon + content", the content is extracted from the first character after the colon until the end of the sentence or the next keyword appears. For the format "keyword + space + content", the content is extracted from the space until the end of the sentence. For the transaction amount field, if the sentence contains both uppercase and lowercase indicators, the uppercase amount and lowercase amount are extracted separately. For counterparty information, if the sentence contains the payer's name, such as "XX Company Payer Account: 6220XXXX1234", the counterparty's name and account number are extracted separately.
[0042] After scanning all sentences, the extracted field content is organized to form a preliminary key field set. It is then verified whether all necessary key accounting elements are included, such as transaction serial number, transaction date, transaction amount, and counterparty information. If any fields are missing, such as the transaction serial number not being extracted, sentences in the standardized text that do not match the keywords are rescanned to check for expressions not covered by the keyword library. For example, in a bank statement, the business number corresponds to the transaction serial number. If an uncovered expression is found, it is temporarily included in the keyword scope, and the corresponding content is extracted. If fields are still missing after rescanning, the statement is marked as having missing fields, and the type of missing field and the bank from which the statement originated are recorded. A manual review process is then triggered. Finally, a key field set containing all key accounting elements is generated, with each field in the set labeled with its field type, the source sentence from which it was extracted, and its content.
[0043] Step 222: Combine and transform the key field set according to a preset format to generate a preliminary structured receipt data unit. Specifically, this includes: loading a preset electronic receipt structured data template. This template is designed based on enterprise accounting needs and treasury system data storage specifications, clearly defining the field's unique identifier, field name, data type, and field order. For example, flow_id - transaction serial number - string - 1; trans_date - transaction date - date type - 2; trans_amt_upper - transaction amount (uppercase) - string - 3; trans_amt_lower - transaction amount (lowercase) - numeric type - 4; counterparty_name - counterparty name - string - 5; counterparty_account - counterparty account - string - 6. It also defines the length limit for each field, such as a maximum length of 32 characters for the transaction serial number and a maximum length for the counterparty account. 20 characters; Read each field in the key field set one by one, and perform precise matching between the field name and the field name in the structured template. Map the transaction serial number to the template flow_id field, the transaction date to the trans_date field, the transaction amount (uppercase) to the trans_amt_upper field, the transaction amount (lowercase) to the trans_amt_lower field, the counterparty name to the counterparty_name field, and the counterparty account to the counterparty_account field. If the length of a field in the key field set exceeds the limit specified by the template, such as the counterparty account being 25 characters long and exceeding the 20-character limit, the first 20 characters are truncated as temporary content and marked as excessive, to be processed in subsequent verification stages. If the field content is empty, such as the counterparty account not being extracted, NULL is filled into the corresponding field in the template and the content is marked as empty. Following the field order specified in the template, the mapped and matched field content is sequentially filled into the corresponding positions in the template to form a preliminary structured receipt data unit. For example, based on the above template, the assembled unit format is: flow_id: 78901234; trans_date: 2024-05-20; trans_amt_upper: One Thousand Yuan; trans_amt_lower: 1000.00; counterparty_name: XX Company; counterparty_account: 6220XXXX1234. At the same time, metadata information is added to this data unit, including the name of the receipt source bank, the receipt acquisition time, and the data unit generation time. The preliminary structured data unit is then temporarily stored in a temporary database with the naming rule of bank name-receipt acquisition date-transaction serial number, awaiting subsequent verification.
[0044] Step 223 involves validating and standardizing the initially structured receipt data units. This includes uniformly converting the transaction date format, standardizing the counterparty name and transaction amount numerical formats, and finally generating structured receipt data. Specifically, this includes reading the content of the trans_date field in the initial structured data unit, determining its original format using date format recognition rules. If it contains " / ", it is in YYYY / MM / DD or MM / DD / YYYY format; if it contains "-", it is in YYYY-MM-DD or MM-DD-YYYY format; and if it contains "year / month / day", it is in YYYY-MM-DD or MM-DD-YYYY format. The YY year MM month DD day format is used to perform conversion according to the preset standard date format. For the YYYY / MM / DD format, " / " is replaced with "-"; for the MM / DD / YYYY format, it is adjusted to YYYY-MM-DD; for the YYYY year MM month DD day format, the year / month / day is deleted and replaced with "-". After conversion, the validity of the date is verified. First, the range of year / month / day values is verified. Second, the date is verified to be within a reasonable transaction range. If the date is invalid, it is marked as an anomaly, the reason for the anomaly is recorded, and manual review is triggered. If the date is valid, the trans_date field is updated to the standard format.
[0045] Scan the `counterparty_name` field, removing irrelevant and redundant characters. This includes deleting parentheses and their corresponding comments, removing special symbols, and standardizing the abbreviation and full name. The processed counterparty name is then standardized to the format of administrative division + trade name + industry + organizational form. If the administrative division is missing, it is added based on the location of the counterparty's bank. If the industry description is not standardized, it is standardized to "XX Technology Co., Ltd.", ensuring consistent counterparty name format and complete information. Read the `trans_amt_lower` field, removing non-numeric characters. The decimal places are standardized to two. If the amount is negative, a minus sign is retained before the amount, and a refund is marked in the transaction type supplementary field. The Chinese uppercase amount to lowercase conversion rule library is loaded, converting the uppercase amount in the `trans_amt_upper` field to lowercase. The converted lowercase amount is compared with the standardized `trans_amt_lower` field. If they match, the validation passes; otherwise, the amount is marked as abnormal, the difference between the uppercase and lowercase conversion results is recorded, and manual review is triggered.
[0046] After completing all validation and standardization processes, the standardized content of the fields flow_id, trans_date, trans_amt_upper, trans_amt_lower, counterparty_name, and counterparty_account is integrated, and fields such as transaction currency, transaction type, and receipt status are added to generate the final structured receipt data. This data is stored in the structured database of the enterprise treasury system, with the transaction serial number flow_id used as the primary key to ensure data uniqueness. At the same time, the structured receipt data is associated with the corresponding original electronic receipt file (PDF / XML), and the storage path of the original file is recorded in the database.
[0047] By removing irrelevant characters, formatting information, and redundant prompts, the system avoids identification errors caused by invalid information during subsequent key field extraction, ensuring that only core transaction information is extracted and improving the accuracy of key field identification. Based on preset rules and keyword libraries, it can quickly match the feature identifiers of core accounting elements such as transaction serial numbers, dates, amounts, and counterparties, avoiding omissions or misjudgments during manual extraction and ensuring that no key information is missing. This provides complete data support for structured processing. By using preset templates, scattered key fields are integrated into data units with fixed structures, transforming originally disordered text information into structured data with clearly defined fields. This facilitates system storage, querying, and data interaction, providing data format support for automated processing in bank-enterprise direct connection scenarios. Structured data units do not require repeated text parsing and can directly enter the verification stage, reducing repetitive workload in data processing. At the same time, the fixed field structure facilitates alignment with transaction detail data fields, laying the foundation for rapid matching and shortening the overall data processing cycle.
[0048] like Figure 2 As shown, in a preferred embodiment of the present invention, step 3 above, which involves constructing a transaction data feature set based on structured return data, determining a benchmark data point from the feature set, and generating two sets of data association features based on the benchmark data point, may include:
[0049] In this embodiment of the invention, step 330 involves extracting features from the structured receipt data, including transaction amount, transaction timestamp, payment account information, receiving account information, and counterparty identification information, to construct an initial transaction data feature set. Specifically, this includes: firstly, sorting through all fields in the structured receipt data, including transaction serial number, transaction date, transaction amount (lowercase and uppercase), counterparty name, counterparty account, payment account name, payment account number, receiving account name, receiving account number, receipt generation time, etc. Based on the principle of direct correlation with bank transaction details, core fields are selected. The transaction amount corresponds to the transaction amount (lowercase) field, which has been standardized to two decimal places. The transaction timestamp needs to be generated based on the combination of the transaction date and receipt generation time. The payment account information includes the payment account name and payment account number, which must fully reflect the unique identifier of the account. The receiving account information includes the receiving account name and receiving account number. Similarly, the counterparty identification information needs to be determined in conjunction with the transaction direction to ensure that it points to a unique counterparty. Fields such as transaction serial number and receipt number, which are only used for internal identification and have no correlation analysis value, are excluded.
[0050] The system directly reads the standardized value from the lowercase field of the transaction amount, verifying the format digit by digit to ensure there are no letters or special symbols, containing only numbers and a decimal point, with two fixed decimal places. If any abnormal format is found, it automatically adds 0 to two decimal places at the end to determine the final transaction amount characteristic value, ensuring that the amount data has no format differences in subsequent comparisons. It first reads the standard format of the transaction date field, then reads the receipt generation time from the cell data, concatenating the two into a complete timestamp in YYYY-MM-DDHH:MM:SS format. If the receipt generation time is only accurate to the minute, it is automatically supplemented. The second count is 00; if the receipt generation time is missing, it will be filled with 12:00:00 of the transaction date as a temporary time, and the timestamp will be marked as to be completed to ensure that each transaction has a complete timestamp feature; the payment account name and payment account number are concatenated in the fixed separator format of name|account. Before concatenation, the account name is checked to see if it contains special characters. If so, the special characters are removed. The account number needs to be checked for digits, usually 16-20 digits. If the number of digits is abnormal, the account number is marked as to be verified. Finally, a payment account information feature value with a uniform format is formed to ensure that the payment account can be uniquely identified.
[0051] The processing logic is consistent with that of the payment account information. The receiving account name and receiving account number are concatenated in the format of name|account number. Similarly, special characters in the name are removed and the number of digits in the account number are verified. This generates a characteristic value for the receiving account information, ensuring the integrity and standardization of the receiving account identifier. First, the transaction direction is determined by the positive / negative transaction amount or the transaction type field. For example, a positive amount indicates a receipt, and a negative amount indicates a payment. Alternatively, the transaction type field can explicitly indicate receipt or payment. If it is a payment transaction, the counterparty is the recipient, and the receiving account name|receiving account number is used as the counterparty identifier. If it is a receipt transaction, the counterparty is the payer, and the payment type is used as the payer identifier. The payment account name and payment account number serve as the counterparty identification information, ensuring that the counterparty identification completely matches the actual transaction relationship. The five types of features extracted and formatted above—transaction amount, transaction timestamp, payment account information, receiving account information, and counterparty identification information—are organized one by one according to the correspondence between feature name and feature value to form an initial transaction data feature set. After assembly, each feature is checked for null values or abnormal markers. If any exist, the feature set is marked as needing improvement and temporarily stored in a temporary folder. It will be re-included into the set after manual supplementation of information, ensuring that all features in the initial feature set are valid and complete.
[0052] Step 331 involves statistically analyzing the initial transaction data feature set, calculating the probability distribution and frequency of each feature value, and selecting data points with prominent transaction amount frequencies and representative transaction time distributions as benchmark data points. Specifically, this includes: collecting all verified initial transaction data feature sets within a fixed period of the enterprise to form a full feature dataset; classifying and splitting this dataset according to feature type to construct separate subsets: transaction amount subset (containing the amount feature values of all transactions), transaction timestamp subset (containing the timestamp feature values of all transactions), and payment account information subset (containing the payment account feature values of all transactions), etc. Each subset needs to be labeled with a unique identifier for the corresponding transaction, such as a transaction serial number, for subsequent tracking. The process involves: First, deduplicate all transaction amounts in the transaction amount subset to generate a unique list of amounts, ensuring no duplicates. Then, iterate through the entire dataset, counting the frequency of each unique amount. During this process, verify the transaction serial number for each amount to avoid duplicate or missed counts. Next, calculate the total number of transactions in the full feature dataset (e.g., 1500 transactions). Divide the frequency of each unique amount by the total number of transactions to obtain its probability distribution, keeping two decimal places. Finally, sort the unique amounts by frequency from highest to lowest, selecting the top 20% and marking them as high-frequency transactions to ensure their prevalence in transactions.
[0053] All timestamps in the transaction timestamp subset are grouped by hourly intervals, resulting in 24 time periods, such as 00:00-01:00, 01:00-02:00...23:00-24:00. Each time period contains the timestamps of all transactions within that hour; for example, the 09:00-10:00 time period includes timestamps such as 09:05:30 and 09:42:18. The number of transactions within each time period is counted to generate an hourly transaction frequency distribution table. For example, the 09:00-10:00 time period has 280 transactions, and the 14:00- There were 220 transactions at 15:00, 190 transactions between 10:00 and 11:00, and 5 transactions between 00:00 and 01:00. During the statistics, it was ensured that each timestamp belonged to only one time period, with no duplicate counts across time periods. The hourly transaction frequency distribution table was analyzed to identify the top three transaction frequency periods, such as 09:00-10:00, 14:00-15:00, and 10:00-11:00, which were marked as high-frequency trading periods. The transaction timestamps within these periods reflect the time patterns of the company's regular transactions and are representative of the distribution.
[0054] Data points that simultaneously meet the following two conditions are selected from the full feature dataset: the transaction amount of the data point belongs to the high-frequency transaction amount; the transaction timestamp of the data point falls within the high-frequency transaction period. For the selected data points, their payment account information, receiving account information, and counterparty identification information are further verified to ensure they are complete (no empty values, no abnormal markers). If they are complete, they are determined as the benchmark data points; if there is missing information, the data points are excluded and the selection is repeated. The final benchmark data points need to record the corresponding transaction serial number and bank name to facilitate the association with bank transaction details and electronic receipt data, ensuring that the benchmark data points have both the advantage of transaction frequency and conform to the time distribution pattern, and can represent the characteristics of most transactions.
[0055] Step 332: Based on the benchmark data points, extract corresponding feature information from the bank transaction details data and electronic receipt data respectively, and generate a first feature sequence representing the characteristics of the bank transaction details and a second feature sequence representing the characteristics of the electronic receipts. Specifically, this includes: for each benchmark data point, locating the corresponding bank transaction details data and electronic receipt data in the enterprise treasury system database using its two unique identifiers, the transaction serial number and the payment account number; entering the bank transaction details database of the treasury system; and performing a precise query by inputting the payment account number and transaction serial number of the benchmark data point into the query conditions. The system matches the single bank transaction detail corresponding to the benchmark data point, including fields such as transaction amount, transaction time, payment account, receiving account, and counterparty information recorded by the bank. If the query result is empty or multiple details appear, the detail association is marked as abnormal, triggering manual verification. If the query result is unique, the detail is locked as the target data. The system then enters the electronic receipt structured database of the treasury system, inputs the transaction serial number of the benchmark data point to execute the query, matches the corresponding structured receipt data, and verifies the uniqueness of the query result to ensure that only the single electronic receipt corresponding to the benchmark data point is associated, avoiding confusion from multiple receipts.
[0056] From the locked bank transaction details data, the following core features are extracted according to the principle of one-to-one correspondence with the features of the baseline data points:
[0057] Bank transaction amount: Read the transaction amount field from the bank statement, verify the amount unit, and if the bank statement amount is 1000, automatically fill it with 1000.00 as the transaction amount item in the first feature sequence;
[0058] Bank transaction timestamp: Read the transaction time field from the bank details. If the bank details only record the transaction date, combine it with the receipt generation time of the benchmark data point to supplement it into a complete timestamp, which is used as the transaction timestamp item.
[0059] Bank payment account information: Read the payment account name and payment account number from the bank statement, and concatenate them in the format of name|account number to form the payment account information item;
[0060] Bank account information: Similarly, combine the account name and number from the bank statement to form the account information field.
[0061] Bank counterparty identification information: Based on the direction of the transaction (receiving / paying), the counterparty is determined to be either the payee or the payer. The corresponding account name and account number are concatenated to form the counterparty identification information.
[0062] The extracted five features are arranged in a fixed order: transaction amount - transaction timestamp - payment account information - receiving account information - counterparty identification information, separated by semicolons, to generate the first feature sequence representing the characteristics of bank transaction details. Each feature in the sequence must have the same meaning as the corresponding feature field of the benchmark data point to ensure the effectiveness of subsequent comparisons.
[0063] From the matched structured receipt data, extract 5 features that completely correspond to the first feature sequence. The extraction logic is the same as that for the first feature sequence, but the data source is electronic receipts.
[0064] Transaction amount at the receipt end: Read the lowercase field of the transaction amount from the structured receipt and use it directly as the transaction amount item in the second feature sequence;
[0065] Transaction timestamp on the return end: Read the transaction timestamp field generated in step 330 and use it as the transaction timestamp item;
[0066] Payment account information on the receipt end: Read the payment account information field from the structured receipt and use it directly as the payment account information item;
[0067] Receipt end payment account information: Read the payment account information field from the structured receipt and use it as the payment account information item;
[0068] Counterparty identification information on the receipt end: Read the counterparty identification information field from the structured receipt and use it as the counterparty identification information item.
[0069] Similarly, arrange the transaction amount, transaction timestamp, payment account information, receiving account information, and counterparty identification information in the order of transaction amount, separated by semicolons, to generate a second feature sequence that represents the characteristics of the electronic receipt. After generation, the number of fields in the second feature sequence needs to be verified to ensure that there are 5 items to avoid missing fields and to ensure that it is completely aligned with the structure of the first feature sequence, so that it can be directly used for item-by-item comparison.
[0070] Step 333: Calculate the feature matching degree between the first feature sequence and the second feature sequence, including transaction timestamp difference, transaction amount matching degree, and counterparty information similarity. Based on the matching degree calculation results, generate two sets of data association features for association matching. Specifically, this includes: extracting the transaction timestamp items from the first and second feature sequences respectively, converting the two timestamps into second-level timestamp values, i.e., the total number of seconds from January 1, 1970, 00:00:00 (UTC time) to that point in time. During the conversion, ensure that the time zone is consistent (both are Beijing time) to avoid errors caused by time zone differences. For example, 2024-06-15 09:35:20 is converted to 1718434520 seconds, and 2024-06-15 09:35:22 is converted to 1718434522 seconds; calculate the absolute difference between the two second-level timestamp values, and preset the time difference scoring rules according to the precise matching requirements in the document:
[0071] If the time difference is ≤3 seconds: it is determined to be a time high match, the difference is assigned a value of 0, and the difference is minimal;
[0072] If 3 seconds < time difference ≤ 5 seconds: it is judged as a basic time match, and the difference is assigned a value of 0.2;
[0073] If 5 seconds < time difference ≤ 10 seconds: it is judged as a slight time mismatch, and the difference value is assigned 0.5;
[0074] If the time difference is greater than 10 seconds: it is judged as a serious time mismatch, and the difference is assigned a value of 1.0, which is the largest difference.
[0075] According to the above rules, the difference degree corresponding to 2 seconds is 0. Record the difference degree result of this set of timestamps. The smaller the difference degree value, the stronger the time consistency between the bank details and the receipt.
[0076] Extract the bank transaction amount from the first feature sequence and the receipt transaction amount from the second feature sequence. Convert both amounts to numeric values and verify the unit. If there is a unit difference, convert them to yuan before proceeding with calculations. Assign a matching degree and determine if the amounts are completely equal. If they are completely equal (e.g., both are 1000.00), the amounts are considered a perfect match, and a matching degree of 1.0 is assigned (highest matching degree). If there is a difference (e.g., bank statement amount 1000.00, receipt amount 999.99), first calculate the absolute value of the difference |1|. 000.00-999.99|=0.01, then divide the absolute value of the difference by the bank statement amount 0.01 / 1000.00=0.001%. Assign a value according to the preset rules: if the ratio is ≤0.01%, the difference is extremely small and may be due to bank rounding error, assign a matching degree of 0.9; if 0.01% < ratio ≤0.1%, assign a matching degree of 0.7; if 0.1% < ratio ≤1%, assign a matching degree of 0.5; if the ratio is >1%, assign a matching degree of 0, indicating a serious mismatch in amount. Record the amount matching degree result. The closer the matching degree is to 1.0, the stronger the consistency of the amount.
[0077] The counterparty identification information items of the first and second feature sequences are split into counterparty name and counterparty account by "|". The similarity between the counterparty account and counterparty name is compared separately. For counterparty account similarity, since the account is a unique identifier, if the two sets of accounts are completely identical, the account similarity is assigned a value of 1.0; if there is any difference in any single digit, the account similarity is assigned a value of 0, with no intermediate values, ensuring absolute accuracy in account matching. For counterparty name similarity, a character matching rate is used for calculation. First, special characters in both sets of names are removed to obtain simplified names. Then, the number of identical characters in the simplified names is counted. Finally, the number of identical characters is divided by the total number of characters in the longer simplified name to obtain the character matching rate, i.e., the name similarity. If the name is... If the simplified character matching rate is ≥90%, the name similarity is also assigned a value of 0.8. Considering that the uniqueness of the account has a higher priority than the name, the comprehensive similarity is calculated with a weight of 60% for account similarity and 40% for name similarity. The formula logic is: Comprehensive Similarity = (Account Similarity × 60%) + (Name Similarity × 40%). For example, when the account similarity is 1.0 and the name similarity is 0.8, the comprehensive similarity = (1.0 × 0.6) + (0.8 × 0.4) = 0.6 + 0.32 = 0.92. If the account similarity is 0 and the name similarity is 1.0, the comprehensive similarity = (0 × 0.6) + (1.0 × 0.4) = 0.4. The comprehensive similarity result is recorded. The higher the value, the more accurate the matching of the opponent's information.
[0078] The three indicators calculated above—transaction timestamp difference, transaction amount matching degree, and comprehensive similarity of counterparty information—are linked and integrated with the core information of the corresponding benchmark data points (transaction serial number, bank name), the first feature sequence (bank detail features), and the second feature sequence (receipt features) to generate two sets of data association features:
[0079] The first set of associated features (bank transaction details associated features): assembled in the order of baseline identifier - bank transaction details feature - matching degree index, in the format of transaction serial number: XXX; bank name: XXX; first feature sequence: XXX; transaction timestamp difference: XXX; transaction amount matching degree: XXX; comprehensive similarity of counterparty information: XXX; example: transaction serial number: 2024061500123; bank name: XX Bank; first feature sequence: 1000.00; 2024-06-15 09:35:20; XX Trading Co., Ltd. in Region A | XXXXXXXX; YY Technology Co., Ltd. in Region B | XXXXXXXX; YY Technology Co., Ltd. in Region B | XXXXXXXX; transaction timestamp difference: 0; transaction amount matching degree: 1.0; comprehensive similarity of counterparty information: 0.92.
[0080] The second set of associated features (electronic receipt associated features): assembled in the order of baseline identifier - receipt feature - matching degree index, in the format of transaction serial number: XXX; bank name: XXX; second feature sequence: XXX; transaction timestamp difference: XXX; transaction amount matching degree: XXX; comprehensive similarity of counterparty information: XXX; example: transaction serial number: 2024061500123; bank name: XX Bank; second feature sequence: 1000.00; 2024-06-15 09:35:22; XX Trading Co., Ltd. in Region A | XXXXXXXX; YY Technology Co., Ltd. in Region B | XXXXXXXX; YY Technology Co., Ltd. in Region B | XXXXXXXX; transaction timestamp difference: 0; transaction amount matching degree: 1.0; comprehensive similarity of counterparty information: 0.92.
[0081] After generation, the integrity and consistency of the two sets of associated features need to be verified, and then they are stored in the data association feature library of the treasury system for easy direct access to accurately match bank transaction details and electronic receipts in the future.
[0082] This algorithm filters core features such as transaction amount, timestamp, and account information from structured receipt data, excluding fields like serial numbers that are only used for identification and have no value for matching analysis. This allows feature analysis and matching to focus on key dimensions, avoiding interference from invalid features and providing a clear data foundation for accurately linking bank details and receipts. By statistically analyzing high-frequency transaction amounts and high-frequency transaction periods, data points that conform to the company's regular transaction patterns are selected as benchmarks to avoid feature bias caused by random selection. This ensures that the benchmark data points can reflect the characteristics of most transactions, improving the overall efficiency of data association processing. The first and second feature sequences are arranged in a fixed order, and the fields correspond one-to-one, ensuring that the features of bank transaction details and electronic receipts are completely aligned in structure. This avoids comparison errors caused by disordered field order or missing fields. The matching degree calculation rules can be flexibly adjusted according to the characteristics of different bank details and receipt data. The generated association features can adapt to the differences in data from multiple banks, enhancing the algorithm's adaptability to multi-bank scenarios.
[0083] In a preferred embodiment of the present invention, step 4 above, which involves calculating time, amount, and opponent difference based on the correlation features of two sets of data to construct a multidimensional feature dataset, and determining the effective range by selecting positive and negative samples; mapping the samples to three-dimensional spatial points according to the time series to form polygonal boundaries, determining the position of sample points through the odd-inside-even-outside rule, counting the number of inside and outside points to calculate the density of inside points, constructing a correlation matching rule system, and generating calibration parameters, may include:
[0084] In this embodiment of the invention, step 440 involves calculating the transaction timestamp difference, transaction amount difference, and counterparty information similarity based on the first feature sequence and the second feature sequence, forming a multi-dimensional feature difference dataset including three dimensions: time difference, amount difference, and information similarity. Specifically, the first feature sequence comes from bank transaction details, containing the specific timestamp of each transaction accurate to the second, the transaction amount accurate to the minute, and the complete name of the counterparty; the second feature sequence comes from electronic receipts, containing the transaction timestamp of the receipt record accurate to the second, the transaction amount accurate to the minute, and the complete name of the counterparty. When calculating the transaction timestamp difference, the timestamp of a certain transaction in the first feature sequence and the timestamp of the corresponding receipt to be matched in the second feature sequence are first found. The timestamp value of the former is subtracted from the timestamp value of the latter, and the absolute value of this difference is taken to obtain the timestamp difference of the matching objects, in seconds. The transaction amount is then calculated. When calculating the difference, the amount of the transaction details in the first feature sequence is subtracted from the amount of the electronic receipt in the second feature sequence, and the absolute value of this difference is taken to obtain the amount difference of the matching objects, in yuan. When calculating the counterparty information similarity, the counterparty names in the first and second feature sequences are first split into Chinese characters one by one, and common suffixes such as "Limited Company" and "Stock Company" are removed, leaving only the core name part. Then, the number of the same Chinese characters in the two sets of core names is counted, and the number of the same Chinese characters is divided by the sum of the total number of Chinese characters in the two sets of core names to obtain the counterparty information similarity of the matching objects. The result is represented by a value between 0 and 1. The timestamp difference, amount difference, and information similarity of all the transaction details and electronic receipts to be matched are recorded in sequence to form a multi-dimensional feature difference dataset containing these three dimensions. One record in each dataset corresponds to a set of difference features between transaction details and electronic receipts.
[0085] Step 441 involves statistically analyzing the multidimensional feature difference dataset to determine the confidence intervals for the difference values of each dimension and generating the effective range for data association. Specifically, this includes: for the time difference dimension in the multidimensional feature difference dataset, collecting the timestamp difference values of all records under this dimension; first, calculating the sum of these values, then dividing the sum by the total number of records to obtain the average value of the time difference dimension; next, calculating the difference between each timestamp difference value and this average value, squaring each difference, then summing these squared values, dividing this sum by the total number of records to obtain the squared average; finally, taking the square root of the squared average to obtain the standard deviation of the time difference dimension. At a preset 95% confidence level, the confidence interval for the time difference dimension is the range between the average value minus twice the standard deviation and the average value plus twice the standard deviation. For the monetary difference dimension, the same method is used: first, calculating the sum of all monetary difference values divided by the total number of records to obtain the average value, then calculating the square root of each monetary difference value... The standard deviation of the monetary difference dimension is obtained by taking the square root of the sum of the squares of the differences between the numerical values and the average value, divided by the total number of records. This determines the 95% confidence interval for the monetary difference dimension as the range between the average value minus twice the standard deviation and the average value plus twice the standard deviation. Similarly, for the information similarity dimension, the average value is obtained by dividing the sum of all information similarity values by the total number of records. The standard deviation is then obtained by taking the square root of the sum of the squares of the differences between each information similarity value and the average value, divided by the total number of records. The 95% confidence interval for the information similarity dimension is also the range between the average value minus twice the standard deviation and the average value plus twice the standard deviation. Integrating these three confidence intervals, when the timestamp difference between a set of transaction details and electronic receipts falls within the confidence intervals for the time difference dimension, the monetary difference dimension, and the information similarity dimension, the set of objects is preliminarily determined to be potentially related. These three intervals together constitute the effective range for data association.
[0086] Step 442: Select positive sample points within the effective range and negative sample points in the outer boundary region to form a verification sample set. Specifically, this includes: traversing all records in the multidimensional feature difference dataset and checking whether the values of each record's three dimensions are all within the corresponding dimension confidence intervals determined in step 441. If a record's timestamp difference is within the time difference dimension confidence interval, its amount difference is within the amount difference dimension confidence interval, and its information similarity is within the information similarity dimension confidence interval, then the sample point corresponding to that record is marked as a positive sample point. This indicates that the transaction details and electronic receipts corresponding to that sample point are highly correlated matching pairs. For those pairs with at least one... For records where the dimension values exceed the corresponding confidence interval, sample points close to the effective range boundary are selected as negative sample points. The specific selection criteria are that the portion of the dimension value exceeding the confidence interval does not exceed 50% of the standard deviation of that dimension. For example, the value of the timestamp difference exceeding the upper limit of the confidence interval for the time difference dimension does not exceed 50% of the standard deviation of the time difference dimension, or the value of the amount difference below the lower limit of the confidence interval for the amount difference dimension does not exceed 50% of the standard deviation of the amount difference dimension, etc. These negative sample points represent matching pairs where the transaction details and electronic receipts have a low probability of correlation. All selected positive and negative sample points are summarized to form a validation sample set for correlation trajectory analysis.
[0087] Step 443: Based on the time sequence of transactions, map the sample points in the verification sample set to three-dimensional spatial coordinate points, and connect the baseline data points to construct a polygonal trajectory boundary. Specifically, this includes: constructing a three-dimensional space using the three dimensions determined in Step 441 as coordinate axes, where the X-axis represents timestamp difference (in seconds); the Y-axis represents amount difference (in yuan); and the Z-axis represents information similarity (ranging from 0 to 1). Map each sample point in the verification sample set to a coordinate point in the three-dimensional space according to its timestamp difference value corresponding to the X-axis coordinate, its amount difference value corresponding to the Y-axis coordinate, and its information similarity value corresponding to the Z-axis coordinate; the baseline data points are from... Representative points were selected from the initial transaction data feature set. The selection criteria were that the transaction amount ranked in the top 20% of the frequency of similar transactions and the transaction time was distributed across different dates and time periods to ensure coverage of the common situation of daily transactions. At least 30 benchmark data points were selected for each bank account. According to the actual time sequence of the transactions, these benchmark data points were connected sequentially in three-dimensional space. The endpoint of the previous benchmark data point was connected to the starting point of the next benchmark data point, which ultimately formed a closed polygon. This polygon is the trajectory boundary for judging the rationality of the association of sample points. Its shape can reflect the typical distribution range of details and receipts in three dimensions in normal transactions.
[0088] Step 444: Draw a ray from the sample point and calculate the number of intersections between the ray and the polygon boundary. If the number of intersections is odd, the sample point is determined to be inside the polygon; otherwise, it is outside. Specifically, for each sample point mapped to 3D space in the verification sample set, draw a ray from that point along the positive X-axis. The direction of this ray remains unchanged, always parallel to the X-axis and extending in the direction of increasing value. Check one by one whether this ray intersects with each edge of the polygon trajectory boundary constructed in step 443. During the check, it is necessary to determine whether the ray passes through the part between the two endpoints of the edge. If the ray has exactly one intersection with an edge and the intersection is not on the endpoint of the edge, it is recorded as a valid intersection. Count the total number of valid intersections between the ray and all edges of the polygon trajectory boundary. If the total number is odd, the sample point is determined to be inside the polygon, indicating that the correlation between the transaction details and electronic receipts corresponding to the sample point conforms to the distribution characteristics of normal transactions. If the total number is even, the sample point is determined to be outside the polygon, indicating that its correlation deviates from the distribution characteristics of normal transactions.
[0089] Step 445: Count the number of verification sample points inside and outside the polygon, calculate the interior point distribution density, and construct an association matching rule system based on the interior point distribution density analysis results, generating matching calibration parameters. Specifically, this includes: counting the total number of sample points determined to be inside the polygon in step 444, then counting the total number of sample points in the verification sample set, dividing the total number of interior sample points by the total number of sample points in the verification sample set to obtain the interior point distribution density, and analyzing the value of the interior point distribution density. If the interior point distribution density is high, it indicates that the current effective range and the polygon trajectory boundary can better cover sample points with high association probability. At this point, based on the distribution characteristics of positive sample points in three dimensions, further clarify the weight of each dimension in association matching. For example, timestamp difference has a high weight in the matching of ICBC. Because its time record is accurate, and information similarity has a higher weight in the matching of local city commercial banks, and because its counterparty name on the receipt is more complete, if the inner point distribution density is low, the confidence level of the effective range is adjusted, such as adjusting the 95% confidence level to 90%, or reselecting the benchmark data points, increasing the number of benchmark points for high-frequency transaction types, and then reconstructing the polygon trajectory boundary. Steps 443 to 444 are repeated until the inner point distribution density reaches the expected level. Based on the final inner point distribution density analysis results, the specific weights and judgment thresholds of timestamp difference, amount difference, and information similarity in different banks and different transaction types are determined, forming a complete association matching rule system. The specific values of these weights and thresholds are the matching calibration parameters, which are used for accurate matching calculations between transaction details and electronic receipts.
[0090] By analyzing the differences in three key dimensions—transaction time, amount, and counterparty information—we can comprehensively capture the correlation characteristics between transaction details and electronic receipts, avoiding confusion caused by matching only the amount. Statistical analysis determines the effective range of each dimension, making the correlation judgment more consistent with the actual transaction patterns of enterprises and reducing misjudgments caused by differences in bank receipt formats. The accurate selection of positive and negative samples provides a reliable basis for the verification of correlation rules and improves the applicability of the rules.
[0091] In a preferred embodiment of the present invention, step 5 above, using matching calibration parameters, based on transaction serial number matching and combined with multi-factor association rules of transaction date, account, amount, and counterparty name, matches transaction details and electronic receipts to obtain matching results, may include:
[0092] In this embodiment of the invention, step 550 involves constructing a multi-factor weighted association rule, including transaction serial number, transaction date, account information, transaction amount, and counterparty name, using matching calibration parameters. Specifically, this includes extracting matching calibration parameters, including sample weight parameters, time adjustment parameters, amount tolerance parameters, and information matching parameters. For example, a sample weight parameter of 1.1 indicates a high proportion of positive samples and strong data reliability; a time adjustment parameter of 1.05 indicates a slight delay in receipt generation, requiring a wider time range; an amount tolerance parameter of 1.02 indicates acceptable minor amount errors; and an information matching parameter of 0.99 indicates that the lower limit of information similarity can be slightly lowered. These calibration parameters are then mapped to each association factor. The sample weight parameter adjusts the confidence level of the overall rule; the time adjustment parameter corrects the matching range of the transaction date; the amount tolerance parameter adjusts the error tolerance of the transaction amount; and the information matching parameter optimizes the counterparty name and account information. The similarity judgment criteria were established. Based on the importance of the enterprise's business and historical matching data, basic weights were set for each factor: the transaction serial number, as the core unique identifier, had a basic weight of 40%; the transaction date, strongly correlated with the time of fund flow, had a basic weight of 20%; account information, as the identifier of the transaction entity, had a basic weight of 15%; the transaction amount, directly related to the fund size, had a basic weight of 15%; and the counterparty name, assisting in confirming the transaction object, had a basic weight of 10%. The final weights of each factor were adjusted based on the sample weight parameter. If the sample weight parameter was 1.1, the weights of each factor were scaled proportionally. After adjustment, the transaction serial number's weight was 40% × 1.1 = 44%, the transaction date's was 20% × 1.1 = 22%, the account information's was 15% × 1.1 = 16.5%, the transaction amount's was 15% × 1.1 = 16.5%, and the counterparty name's was 10% × 1.1 = 11%, ensuring that the overall total weight remained 100%.
[0093] The transaction serial number rule requires that the transaction serial number in the bank transaction details and the serial number in the electronic receipt be completely identical. If they match, this factor scores 44 points (weight 44% × 100 points); otherwise, it scores 0 points. The transaction date rule, based on a time adjustment parameter of 1.05, adjusts the previously allowed date difference range (e.g., ±1 day of the transaction date) to ±1.05 days, or approximately ±25.2 hours. If the transaction date in the details and the receipt date are within the adjusted range, 22 points are awarded; otherwise, points are deducted proportionally, with 1 point deducted for each hour exceeding the limit, up to a maximum deduction of 22 points. The account information rule compares the payment account and receiving account in the details and the receipt. If they match completely, 16.5 points are awarded; if only the payment account or only the receiving account matches, 8.25 points are awarded; if neither matches, 0 points are awarded. The information matching parameter of 0.99 is also considered; if the account characters are similar... For transactions with a similarity score ≥ 99%, if individual character differences are due to blurry receipt printing, they will be treated as consistent. For transaction amount rules, based on the amount tolerance parameter of 1.02, the original allowed amount error range, such as ±0.5%, is adjusted to ±0.5% × 1.02 = ±0.51%. If the difference between the detailed amount and the receipt amount is within the adjusted range, 16.5 points are awarded; otherwise, points are deducted proportionally based on the error, such as deducting 5 points for a 0.6% error, up to a maximum deduction of 16.5 points. For counterparty name rules, name character similarity is calculated, with a reference information matching parameter of 0.99. A similarity score ≥ 99% earns 11 points, 95%-99% earns 5.5 points, and below 95% earns 0 points. All factor scoring rules are integrated to form a multi-factor weighted association rule. A total score ≥ 80 points is considered a preliminary match; 50-79 points require further verification; and below 50 points is considered a mismatch.
[0094] Step 551: Based on multi-factor weighted association rules, match and calculate the bank transaction details data and standardized receipt data to obtain preliminary matching results. Specifically, this includes: determining the core feature dimensions of the details; extracting five core features from the bank transaction details: transaction serial number, transaction date, account information, transaction amount, and counterparty name, as the five dimensions for constructing multi-dimensional coordinate points; standardizing the feature values; and converting non-numerical features into calculable numerical forms. Specifically, the transaction serial number is converted into a unique numerical code using coding rules, serving as the first dimension value of the coordinate points. The transaction date is converted to a timestamp format, which is the total number of seconds from a specific start time to the transaction date, and used as the second dimension value. Account information, such as the account string, is used to generate a fixed-length hash value through a hash algorithm, which serves as the third dimension value. The transaction amount is retained in its original value and used as the fourth dimension value. The opponent's name is also hashed using a hash algorithm, serving as the fifth dimension value. This ultimately forms a multi-dimensional coordinate point representing the details of the transaction. .
[0095] Determine the center coordinates of the electronic receipt, extract five core features corresponding to the details from the electronic receipt: transaction serial number, transaction date, account information, transaction amount, and counterparty name. Using the same rules as for detail feature value conversion, convert these features into numerical form to obtain... Using these coordinates as the center of the circle, the radii of each dimension of the circle are set. Based on the error tolerance range of each factor in the multi-factor weighted association rule, a corresponding radius is set for each dimension. The transaction serial number serves as a unique identifier and must be completely consistent with the transaction details; therefore, the radius of this dimension... Set to 0; the transaction date dimension reference time adjustment parameter, based on the allowed date difference range, such as the allowed number of hours of fluctuation before and after the transaction date, is converted into the corresponding timestamp difference, which serves as the radius of this dimension. The account information dimension is calculated based on an allowed similarity standard, such as 99% similarity, to determine the allowable range of differences in hash values, which serves as the radius. The transaction amount dimension references the amount tolerance parameter. Based on the allowed amount error ratio, such as ±0.51%, the allowable fluctuation range of the amount is calculated and used as the radius. The opponent's name dimension is used to calculate the allowable difference range of hash values based on the allowed similarity standard, such as ≥95% similarity, and this range is used as the radius. By using the center coordinates and radii of each dimension, a multi-dimensional circular structure representing this transaction order is constructed.
[0096] To calculate the weighted distance from the detail point to the center of the return receipt, first calculate the basic distance difference for each dimension, then calculate the difference between the coordinates of the detail point and the center of the receipt for each dimension. The process involves squaring each difference to obtain the squared difference for each dimension. A weighted sum of squares is then calculated based on the factor weights. According to the weights set in the multi-factor weighted association rules (e.g., transaction serial number 44%, transaction date 22%, account information 16.5%, transaction amount 16.5%, counterparty name 11%), the squared difference for each dimension is multiplied by its corresponding weight, and all results are summed to obtain the weighted sum of squares. The weighted distance is then calculated, and the square root of this weighted sum of squares is taken to obtain the weighted distance from the detail point to the center of the return receipt circle. To calculate the weighted radius of a circle, first calculate the weighted square of the radii in each dimension, then sum the squares of the radii in each dimension. After squaring, multiply by the corresponding factor weights to obtain the weighted squared value of the radius for each dimension; calculate the weighted radius by summing the weighted squared values of the radii for each dimension, and then taking the square root of the sum to obtain the weighted radius of the circle. ; Determine the positional relationship and matching results, if the weighted distance of the details points Less than or equal to the weighted radius of the circle If the detail point is within the return slip circle, and the corresponding multi-factor weighted total score reaches 80 points, the initial match is successful. If the weighted distance... Greater than the weighted radius Then calculate the ratio of the difference between the two. If the difference ratio is ≤20%, the corresponding multi-factor weighted total score is between 50 and 79 points, and it is marked as needing further verification; if the difference ratio is >20%, the corresponding multi-factor weighted total score is below 50 points, and it is judged as a preliminary mismatch. Iterate through all bank transaction details to be matched, convert each detail into a multi-dimensional coordinate point according to the above steps, and then determine the positional relationship between the point and the circle in the multi-dimensional circular structure corresponding to all electronic receipts. Determine the matching status of each detail with each receipt, record the matching status for each detail, and simultaneously record the corresponding weighted distance. Weighted radius Key data such as the difference ratio and the multi-factor weighted total score are used to organize all detailed matching information into a preliminary matching result list according to a preset format.
[0097] Step 552: Perform conflict detection on the preliminary matching results, identify data items with matching conflicts, and resolve these conflicts to generate the final matching results. Specifically, this includes: duplicate matching conflict detection, checking for one-to-many or many-to-one situations in the preliminary matching results, i.e., one detail matching multiple receipts, or multiple details matching the same receipt. For example, detail A matches receipts B1 and B2 simultaneously, and details C1 and C2 match receipt E simultaneously. Such situations are marked as duplicate matching conflicts; score contradiction conflict detection, targeting the preliminary... For matched details, if there are obvious contradictions in the scores of each factor, such as the transaction serial number getting 0 points but the total score being ≥80 points, or the amount difference exceeding the tolerance range but getting full marks, it is marked as a score contradiction conflict. For example, a detail's serial number is completely inconsistent with the receipt, getting 0 points, but other factors have extremely high scores, resulting in a total score of 82 points, which needs to be marked as a conflict. Data anomaly conflict detection involves checking for anomalies in the detail or receipt data itself, such as the detail lacking core fields, the receipt information being vague and unrecognizable, or the transaction direction of the detail and the receipt being opposite. Such cases are marked as data anomaly conflict.
[0098] For duplicate matching conflict resolution, for one-to-many matching, calculate the weighted distance between the detail and each receipt, select the receipt with the smallest distance as the unique match, and unmatch the remaining receipts. For many-to-one matching, similarly calculate the weighted distance between each detail and the receipt, retain the detail with the smallest distance, and mark the remaining details as pending rematch. For example, if the weighted distance between detail A and receipt B1 is 5, and the distance between detail A and receipt B2 is 8, retain the match between A and B1, and unmatch A and B2; if the distance between detail C1 and receipt E is 6, and the distance between C2 and E is 10, retain the match between C1 and E, and mark C2 as pending rematch. For score conflict resolution, re-examine the score calculation process for each factor of the conflicting detail. If the score is abnormal due to incorrect parameter application, such as time adjustment parameters not taking effect, correct the parameters and recalculate the score. If the score is abnormal due to incorrect serial number recognition, such as... For instances of fuzzy or misread serial numbers printed on receipts, manual verification is performed to check the original receipt serial numbers against the details to confirm a match. The score is then corrected, and the match status is reassessed. For example, if a detail's date score is low due to a missing time adjustment parameter, the parameter is corrected, and the score is recalculated. If the score is still ≥80, the match is retained; otherwise, it is adjusted to pending verification or mismatch. For resolving data anomalies and conflicts, for details or receipts lacking core fields, attempts are made to supplement them from other data sources, such as retrieving receipt information from bank online banking. For cases of reversed transaction directions, the actual transaction situation is manually verified, such as whether a refund transaction caused the reversed direction. If it is confirmed to belong to the same transaction, the match is retained; otherwise, it is judged as mismatch. For example, if the detail shows a payment of 1000 yuan and the receipt shows a receipt of 1000 yuan, and it is verified to be a supplier refund, it is confirmed to belong to the same transaction, and the match is retained.
[0099] After integrating the results of conflict resolution, the matching status of all details is finally confirmed, and the matching receipt number, the mark of needing manual intervention or the reason for non-matching are clearly defined for each detail. A final matching result report is generated, which includes the number of successfully matched records, the success rate, the number of records needing manual intervention, the number of non-matching records and the classification of reasons. For example, in a certain batch, 920 records were successfully matched, with a success rate of 92%, 30 records needing manual intervention, and 50 records were non-matching (of which 20 records had no corresponding receipts and 30 records had abnormal data).
[0100] It covers five core factors: transaction serial number, date, account, amount, and counterparty name, avoiding the limitations of matching based on a single factor and improving the matching coverage of receipts with different data quality. It sets weights based on business importance and incorporates calibration parameters generated in the early stage, enabling the rules to dynamically adapt to the characteristics of receipts from different periods and banks, improving the rules' adaptability to changes in business scenarios. Through multi-dimensional detection of duplicate matching, score contradictions, and data anomalies, it comprehensively identifies potential problems in the initial matching results and avoids including erroneous matching results in the final statistics.
[0101] In a preferred embodiment of the present invention, step 6 above, which utilizes the matching results and existing return receipt data resources to continuously train and dynamically adjust the matching rules through a feedback mechanism to obtain the adjusted matching rules, may include:
[0102] In this embodiment of the invention, step 660 compares the matching results with the actual receipt data, identifies matching errors and analyzes rule defects, and generates error analysis results. Specifically, this includes: exporting all matching result data generated during this matching process from the system, including the matching status corresponding to each bank transaction detail, the matched electronic receipt number, the weighted score during matching, and the difference values for each dimension, forming a matching result list; and collecting the actual receipt data corresponding to these transaction details, including scanned copies of paper receipts obtained from bank counters and original electronic receipt files downloaded from corporate online banking, to ensure that the actual receipt data matches the matching results. The transaction details in the list are matched one by one. The matching results are manually reviewed and compared with the actual receipt data for each transaction. For details marked as matching, it is verified whether the matched receipt is the actual corresponding receipt for that detail. For details to be verified, it is confirmed whether the actual corresponding receipt exists and whether it is consistent with a certain unmatched receipt. For unmatched details, it is checked whether there is a corresponding actual receipt. If it exists, the reason for the unmatch is determined, the comparison results are statistically analyzed, and the matching error of each detail is recorded. For example, matching is successful but there is no actual corresponding receipt, there is an actual corresponding receipt but the matching result is unmatched, matching error, etc., forming an error record ledger.
[0103] To calculate the matching error rate, first determine the total number of transaction details within the statistical period as N, and then determine the number of details with matching errors as M. The matching error rate is then calculated as (M ÷ N) × 100%. For example, if the total number of details in a certain period is 1000, and 25 details have errors, then the matching error rate is (25 ÷ 1000) × 100% = 2.5%. Simultaneously, further calculate the error rate for each type of error, such as the false match rate (number of matches that passed but were actually incorrect ÷ total number of details × 100%), the missed match rate (number of matches that were actually incorrect ÷ total number of details × 100%), and the missing match rate (number of matches that were actually incorrect ÷ total number of details × 100%). (Number of returned but unmatched transactions ÷ Total number of details × 100%); Error contribution analysis for each dimension: For details with errors, extract the difference data of each dimension during matching, and calculate the proportion of error caused by each dimension. For example, in 25 error details, 8 errors are caused by unreasonable time difference judgment, 5 errors are caused by inappropriate tolerance range of amount difference, 7 errors are caused by deviation in account information similarity calculation, 3 errors are caused by incorrect identification of counterparty name, and 2 errors are caused by loopholes in serial number matching rules. Therefore, the error contribution of each dimension is... The contribution rates of the errors are as follows: time dimension 32% (8÷25×100%), amount dimension 20% (5÷25×100%), account dimension 28% (7÷25×100%), counterparty name dimension 12% (3÷25×100%), and transaction number dimension 8% (2÷25×100%). Rule defect identification involves analyzing the defects of the matching rules based on the contribution rates of errors in each dimension and specific error cases. For example, the high proportion of errors in the time dimension may be due to the time adjustment parameters not considering the special scenario of delays in the generation of bank month-end statements, causing reasonable time differences to be judged as mismatches. Errors in the amount dimension may be due to the amount tolerance parameters not being set according to the transaction amount, applying the same error ratio to small and large transactions, causing small transactions to be misjudged due to tiny absolute errors exceeding the tolerance range. Errors in the account information dimension may be due to the hash value calculation algorithm being overly sensitive to tiny differences in account characters, causing actually identical accounts to be judged as different. Errors in the counterparty name dimension may be due to the similarity calculation not considering the correspondence between abbreviations and full names, causing truly matching statements to be excluded.
[0104] Compile the calculated data such as matching error rate, error rate of each type, and error contribution of each dimension, and form a quantitative statistical table to clearly show the overall error situation and the degree of influence of each dimension. Combine the result of rule defect location, describe in detail the specific manifestation, scope of influence and typical error cases of each defect, and form a text analysis report. Integrate the quantitative table and text report to generate a complete error analysis result document.
[0105] Step 661: Based on the error analysis results, a training sample set is constructed using existing receipt data resources. This set is used to iteratively train and adjust the parameters of the similarity calculation model in the matching rules, resulting in an updated matching model. Specifically, this includes: selecting data from existing receipt data resources, prioritizing transaction details and corresponding receipts involving error cases in the error analysis results, and supplementing with details and receipt data from different banks, transaction types, and receipt formats to ensure comprehensive sample coverage and avoid overfitting of the matching model due to a single scenario. For example, 5000 data entries are selected, including 1000 error case data entries and 4000 data entries from other banks, transaction types, and receipt formats. For each selected data entry, labels are manually added based on the actual business scenario to clarify its true matching relationship. Labels are divided into positive samples (details and receipts have a true correspondence) and negative samples (details and receipts do not have a true correspondence). During the labeling process, auxiliary materials such as transaction vouchers provided by banks and corporate financial accounting records are referenced to ensure the accuracy of the labels. For example, in 5000 data points, 3500 are labeled as positive samples and 1500 as negative samples.
[0106] For each sample, features related to the matching rules are extracted, including transaction serial number, transaction date, account information, transaction amount, counterparty name, and receipt format. The extracted features are preprocessed, converting non-numerical features into numerical codes and normalizing numerical features, such as converting the amount to a value in the range [0, 1] using the formula: Normalized Amount = (Original Amount - Minimum Amount) ÷ (Maximum Amount - Minimum Amount). This eliminates the impact of differences in feature magnitudes on the matching model training. The preprocessed sample data is divided into training, validation, and test sets in a 7:2:1 ratio. The training set is used for learning the matching model parameters, the validation set is used to adjust the hyperparameters of the matching model and prevent overfitting, and the test set is used to evaluate the final performance of the matching model. For example, in 5000 samples, the training set has 3500 samples (2450 positive, 1050 negative), the validation set has 1000 samples (700 positive, 300 negative), and the test set has 500 samples (350 positive, 150 negative).
[0107] A suitable model for multi-dimensional feature similarity calculation is selected. Here, the Weighted K-Nearest Neighbors (WKNN) model is used as the basic framework. This model can assign different weights according to the importance of features, which is logically compatible with the multi-factor weighted association rule of this invention. During model initialization, an initial K value is set, such as K=5, which means selecting 5 nearest neighbor samples for judgment. The initial feature weights refer to the weights in the original matching rule: serial number 44%, date 22%, account 16.5%, amount 16.5%, and counterparty name 11%. The similarity calculation uses the Euclidean distance formula; the first round... Training involves inputting the training set data into the initialized WKNN model. For each training set sample, the weighted distance between it and other samples in the training set is calculated. The K nearest samples are selected, and their labels (positive / negative) are used to vote on the predicted label for that sample. The predicted label is then compared with the true label, and the following percentages are calculated: accuracy (number of correctly predicted samples ÷ total number of training set samples × 100%), precision (number of correctly predicted positive samples and actual positive samples ÷ total number of correctly predicted positive samples × 100%), and recall (number of correctly predicted positive samples and actual positive samples ÷ total number of actual positive samples × 100%). For example, after the first round of training, the training set accuracy is 85%, precision is 82%, and recall is 80%.
[0108] The WKNN model performance is evaluated using validation set data. If the validation set accuracy is more than 10% lower than the training set accuracy (e.g., 85% for training and 70% for validation), the WKNN model is considered overfitting. In this case, the K value is reduced (e.g., from 5 to 3), and a regularization term is added. If both the validation and training set accuracies are low (e.g., both below 80%), the feature weights are adjusted. Referring to the error contribution of each dimension in the error analysis results, the weights of dimensions with high error contribution are reduced (e.g., the time dimension, with a 32% error contribution, is reduced from 22% to 18%), while the weights of dimensions with low error contribution (e.g., serial number) are increased. The dimensional error contribution was 8%, so its weight was increased from 44% to 48%. At the same time, the feature preprocessing method was optimized, such as adding abbreviation and full name mapping encoding to the opponent's name feature. The above training and parameter adjustment process was repeated. After each round of training, the performance indicators of the training set and the validation set were calculated until the WKNN model performance was stable. In three consecutive rounds of training, the accuracy of the validation set fluctuated by no more than 2%. For example, after 5 rounds of iteration, the accuracy of the training set reached 4%, precision 92%, and recall 91%, and the accuracy of the validation set reached 92%, precision 90%, and recall 89%. The performance of the WKNN model met the requirements.
[0109] Input the test set data into the trained and stabilized WKNN model, and calculate the accuracy, precision, recall, and F1 score (a combined metric of precision and recall, F1 = 2 × (precision × recall) ÷ (precision + recall)) to evaluate the generalization ability of the WKNN model. For example, if the test set accuracy is 91%, precision is 89%, recall is 88%, and F1 score is 88.5%, and all metrics are higher than preset thresholds (e.g., accuracy ≥ 90%, F1 ≥ 85%), then the WKNN model passes the evaluation and is selected. The evaluated WKNN model parameters, including the optimal K value (e.g., 3), the final weights of each feature (serial number 48%, date 18%, account 17%, amount 15%, counterparty name 2%), and feature preprocessing rules such as counterparty name abbreviation mapping encoding, amount normalization range, and similarity calculation thresholds (e.g., weighted distance ≤ a certain value is considered similar), form an updated matching model parameter document. The determined parameters are then integrated into the original similarity calculation model framework, replacing the initial parameters of the original model, and generating the updated matching model.
[0110] Step 662: Based on the updated matching model, dynamically adjust the threshold and weight strategies used in the matching process to form adjusted matching rules. Specifically, this includes: extracting matching thresholds from the original matching rules, including total score thresholds (e.g., ≥80 points for successful matching, 50-79 points for pending verification, <50 points for no match), and difference thresholds for each dimension (e.g., time difference ≤60 seconds, amount difference ≤1%, information similarity ≥90%); statistically analyzing the performance data of matching results under the original thresholds (e.g., matching success rate 78%, false match rate 3%, missed match rate 4%); and analyzing whether the original thresholds are suitable for the new model based on the similarity score distribution output by the updated matching model. For example, if the original total score threshold is 80 points, in the new model, 95% of samples below 80 points are negative samples, and 90% of samples above 80 points are positive samples, indicating that the original thresholds are basically suitable, but can be fine-tuned to further reduce misjudgments; if in the new model, positive and negative samples are mixed in the 78-82 point range (each dimension...). If the threshold is 50%, then the threshold needs to be adjusted to avoid the fuzzy range. Combining the error analysis results and the performance indicators of the matching model, the direction of the threshold adjustment can be determined. For example, the original time difference threshold was 60 seconds. Error analysis shows that the delay in the end-of-month reconciliation caused some reasonable transactions (time difference of 70 seconds) to be misjudged. Moreover, the new model reduces the weight of the time dimension (from 22% to 18%). Therefore, the time difference threshold can be widened to 75 seconds. The original amount difference threshold was 1%. In the new model, the weight of the amount dimension has been reduced to 15%. Small transactions (<100 yuan) are misjudged because the absolute error is small (e.g., 1 yuan) but the proportional error is high (1%). Therefore, the amount difference threshold can be set according to the transaction amount (<100 yuan threshold 2%, 100-10000 yuan threshold 1%, >10000 yuan threshold 0.5%). The original total score threshold was 80 points. Referring to the score distribution of the new model, it can be adjusted to ≥82 points for passing the match, 60-81 points for pending verification, and <60 points for not matching, to reduce misjudgments in the fuzzy range.
[0111] The final weights of each feature in the updated matching model (serial number 48%, date 18%, account 17%, amount 15%, counterparty name 2%) are synchronously applied to the weight strategy of the matching rules, replacing the original weights (serial number 44%, date 22%, account 16.5%, amount 16.5%, counterparty name 11%). This ensures that the rule weights are consistent with the feature importance learned by the matching model. For example, the counterparty name dimension has a high error contribution (12%) and the matching model has learned that its impact on matching is relatively small, so its weight is significantly reduced from 11% to 2% to reduce the interference of this dimension on the total score. Based on the differences in transaction record characteristics among different banks, The design incorporates a dynamic weighting adaptation mechanism. For example, error analysis reveals that a city commercial bank's account information accuracy is extremely low (frequent printing errors) but its transaction number accuracy is 100%. In this case, the matching rules for that bank would reduce the weight of account information from 17% to 10% and increase the weight of transaction number from 48% to 55%. Similarly, if a joint-stock commercial bank's counterparty name is properly labeled (without abbreviation), the weight of the counterparty name could be increased from 2% to 5% to enhance its auxiliary matching role. Furthermore, a weighting adjustment cycle is set, and the weighting strategies for each bank are updated periodically based on the latest error analysis results to ensure that the weights are adapted to the characteristics of different banks' transaction records.
[0112] The adjusted matching thresholds (tiered amount difference threshold, relaxed time difference threshold, new total score threshold), dynamic weighting strategies (basic weights, bank-specific weights), and updated similarity calculation models (feature preprocessing rules, similarity calculation methods) are integrated to form a complete adjusted matching rule document. The document clearly defines the applicable scenarios for each rule type, such as general scenarios, specific bank scenarios, and specific transaction type scenarios, as well as the calculation logic, such as the determination process for tiered amount thresholds, the method for invoking dynamic weights, and parameter values, such as the weight values for each dimension and the threshold values. A small number of new transaction details are also selected. Using the returned order data, perform matching tests with the adjusted matching rules, calculate the matching accuracy, false match rate, and missed match rate, and compare them with the performance data before the adjustment. For example, before the adjustment, the accuracy was 85%, the false match rate was 3%, and the missed match rate was 4%, while after the adjustment, the accuracy was 92%, the false match rate was 1.5%, and the missed match rate was 2%, showing a significant performance improvement, indicating that the rule adjustment is effective. If the performance does not meet the standard, return to step 661 to re-optimize the matching model and parameters until the rule is verified. The verified adjusted matching rule is then added to the system rule base, replacing the original matching rule, and the rule description document is updated.
[0113] By comparing the matching results with the actual receipt data one by one, and combining the quantitative error rate and the error contribution of each dimension, the specific error type and rule defects in the matching process can be accurately located, rather than relying solely on experience. The training sample set covers error cases, data from different banks, different transaction types, and different receipt formats, and has undergone precise manual annotation and feature preprocessing to avoid model overfitting caused by a single sample. The amount difference threshold is set according to the transaction amount, and the time difference threshold is adjusted according to the characteristics of bank receipts to avoid misjudgment of a single threshold in special scenarios.
[0114] In a preferred embodiment of the present invention, step 7 above, which involves automatically associating successfully matched transaction details with electronic receipts according to matching rules and generating auditable and traceable electronic vouchers, may include:
[0115] In this embodiment of the invention, step 770, based on the adjusted matching rules, performs matching operations on unrelated transaction detail data and electronic receipt data to identify associatable transaction receipt pairs and generate a successfully matched list. Specifically, this includes: extracting all bank transaction detail data and electronic receipt data in an unrelated state from the treasury system database. The transaction detail data must include core fields such as transaction serial number, transaction date, payer account, payee account, transaction amount, counterparty name, and transaction summary; the electronic receipt data must include receipt number, receipt generation date, receipt serial number, payer information, payee information, amount information, and receipt text. Key information such as file storage path is collected; unlinked data is preprocessed, and the data format is standardized by converting transaction date and receipt generation date to the YYYY-MM-DDHH:MM:SS standard format, and retaining two decimal places for transaction amount and receipt amount in yuan; invalid data, such as details lacking transaction serial numbers or receipt data that is corrupted and cannot be read, is removed to ensure the integrity of data participating in the matching operation; the multi-factor weighted association logic in the adjusted matching rules is called to select each unlinked valid detail one by one and perform matching operation with all unlinked valid receipts, matching transaction serial numbers and comparing detailed transaction flows. For transaction serial number matching, if the serial number and receipt serial number are completely identical, the score for this factor is the serial number weight score set by the rules; otherwise, 0 points are awarded. For transaction date matching, the time difference between the detailed transaction date and the receipt generation date is calculated. If the time difference is within the valid range allowed by the rules (e.g., ≤75 seconds), a date weight score is awarded; if it exceeds the range, points are deducted proportionally. For account information matching, the detailed payment account is compared with the receipt payment account, and the detailed payment account is compared with the receipt payment account. If both sets of accounts are completely identical, an account weight score is awarded; if only one set of accounts is identical, 8.5 points are awarded; if neither is identical, 0 points are awarded, while also referring to the account information similarity standards in the rules. For transaction amount matching... For matching, calculate the error ratio between the detailed transaction amount and the receipt amount (|detail amount - receipt amount| ÷ detailed amount × 100%). If the error ratio is within the rule-based threshold, a weighted score is awarded; if it exceeds the threshold, points are deducted according to the excess ratio. For counterparty name matching, calculate the character similarity between the counterparty name in the details and the corresponding counterparty name in the receipt (count of identical characters ÷ number of characters in the longer name × 100%). If the similarity is ≥ the rule-set standard, a name weighted score is awarded; if it is below the standard, 0 points are awarded. Calculate the total score for each detail and its corresponding receipt. If the total score is ≥ the matching threshold set by the rule, the detail and receipt are determined to be a matching transaction receipt pair.
[0116] Collect all transaction receipt pairs that are determined to be correlateable, record the detailed core information, receipt core information, matching score, and matching time for each pair, and sort the records in ascending order by transaction date and transaction serial number to form a structured list of successfully matched transactions. The list format includes a header, detailed transaction number, detailed date, detailed amount, receipt number, receipt serial number, total matching score, matching time, and corresponding data row, ensuring that the list can be directly used for subsequent association storage and voucher generation.
[0117] Step 771: Based on the successfully matched list, extract the corresponding transaction details and electronic receipt original information, bind and store the two in a related relationship to form a linked data record with a clear correspondence. Specifically, this includes: based on the transaction serial number in the successfully matched list, extracting the complete data of the detail from the treasury system's transaction details database, including basic information (transaction serial number, transaction initiation time, transaction status, full name of payer, full name of payee), account information (payer account number, payer's bank, payee account number, payee's bank), amount information (transaction amount, handling fee amount, actual amount received), and business information (transaction summary, business type, related contract number), ensuring that the extracted information covers the core detailed dimensions required for financial auditing; based on the receipt number and receipt file path in the successfully matched list, extracting the original receipt information from the treasury system's receipt storage module, including... This includes electronic receipt files (PDF format), receipt unit data (receipt generation time, receipt version number, bank identifier, receipt verification code), and receipt content parsing data (consistency verification results of the transaction serial number, amount, payment / receipt information marked on the receipt and the details). Simultaneously, the electronic receipt files undergo integrity verification to ensure the original receipt text is not damaged or tampered with. A unique association ID is generated for each successfully matched detail and receipt pair. This association ID serves as the core identifier for subsequent retrieval and tracing of the association, ensuring uniqueness and rapid location within the system. The extracted complete detail data is bound to the original receipt information according to preset field mapping rules, such as binding the detail transaction serial number to the receipt's marked serial number, binding the detail transaction amount to the receipt amount, and binding the full name of the detail payer to the receipt payer information. The consistency results of the bound fields are recorded, and any discrepancies must be noted along with their reasons.
[0118] A structure for constructing associated data records is established, comprising four core areas: an associated identifier area (association ID, successfully matched list number, associated operator, and associated timestamp), a detailed data area (extracted complete detailed data fields and values), a return receipt data area (return receipt number, return unit data, return receipt file storage path, and return receipt parsing result), and an associated verification area (field binding consistency result, difference notes, and return receipt file integrity verification result). This ensures a clear record structure and complete information. The constructed associated data records are then partitioned and stored in the treasury system's associated database by association ID. An index is created using association ID, detailed transaction serial number, and return receipt number as index fields to improve the efficiency of subsequent queries and retrievals of associated records. After storage, a final verification is performed on the associated data records to confirm that all field values are correct, the return receipt file is accessible, and the index has been successfully created. Records that pass the verification are considered associated data records with a clear correspondence and can be used for subsequent electronic voucher generation.
[0119] Step 772: Based on the associated data records, automatically generate a structured electronic voucher including key transaction information, receipt image index, associated timestamp, and audit trail identifier. Specifically, this includes: reading the core information from the associated identifier area, detailed data area, and receipt data area of the associated data records; parsing and extracting the key transaction information required for the electronic voucher, including basic transaction information (transaction serial number, transaction date, business type, transaction status), transaction entity information (full name and account number of payer, full name and account number of payee, and banks of both parties), transaction amount information (transaction amount, handling fee, actual amount received, and currency), and business association information (transaction summary, associated contract number, and remarks), ensuring that the extracted key information meets the requirements of financial audit for voucher elements; obtaining the storage path, receipt number, and receipt verification code of the electronic receipt file from the receipt data area of the associated data records, and generating a receipt image index. This index can be directly used to quickly retrieve the electronic receipt image in the system without re-searching the associated records.
[0120] Generate a basic audit trail identifier, including the associated ID, voucher generation timestamp, voucher generation operator ID, and data source identifier, ensuring that the identifier can trace the voucher's generation source and time. Calculate the hash value of extracted key transaction information using the SHA-256 algorithm to obtain the key information hash value; calculate the hash value of the return receipt image file to obtain the return receipt image hash value; combine the two hash values to generate a data verification identifier, which can be used to verify whether the key voucher information and return receipt image have been tampered with during subsequent audits. Integrate the basic identifier and verification identifier to form a complete audit trail identifier, ensuring that the traceability identifier of each electronic voucher is unique and contains full traceability information; construct structured electronic vouchers using standardized XML or JSON format, with the voucher including a voucher header (voucher number, voucher type, generation time, audit trail identifier), transaction details, and more. The system comprises four core sections: Key Information Area (parsed and extracted transaction basis, subject, amount, and business-related information), Receipt Image Index Area (receipt image index information), and Relationship Description Area (consistency results of binding related IDs, details, and receipt fields). This ensures the voucher structure conforms to industry standards for financial electronic vouchers. The extracted key information, receipt image index, and audit trail identifier are populated into the voucher structure. The populated voucher undergoes format validation, such as checking for missing transaction amount fields, standard transaction date formats, and consistency between the receipt image hash value and the calculation result in the receipt file. If validation fails, the system returns to re-extract information until validation passes. The validated structured electronic vouchers are then stored in the treasury system's electronic voucher archive in a year-month-voucher type directory structure. Simultaneously, a voucher retrieval index is established to complete the automatic generation and archiving of vouchers.
[0121] Matching operations are performed based on the adjusted matching rules, which can more accurately distinguish between reasonable and abnormal associations compared to the original rules. This reduces the chances of misjudging non-corresponding details and receipts as related, while also lowering the probability of missing genuine related pairs. The complete data of the details and the original text and metadata of the receipts are extracted, rather than just the core fields. This ensures that the associated data records contain all the information required for auditing. Subsequent audits do not require additional retrieval of data from other systems; verification can be completed directly through the associated records, improving audit convenience. Electronic vouchers are constructed using a standardized format and include key transaction information, receipt image indexes, and audit tracking identifiers required for auditing. No manual supplementation or adjustment of voucher content is required; they can be directly used as audit basis, avoiding audit rework due to non-compliant voucher formats and improving the reliability and security of audits.
[0122] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0123] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0124] The bank account information (including account name, account number, bank code, etc.), transaction details (including transaction serial number, transaction amount, transaction date, counterparty information, etc.) and electronic receipt data (including receipt number, receipt generation time, receipt verification code, etc.) involved in the embodiments of the present invention all fall within the scope of enterprise or individual privacy data.
[0125] In the actual application of the method of the present invention, before collecting the above-mentioned privacy data, it is necessary to clearly inform the data provider (such as enterprise users, account holders, etc.) of the purpose, scope, use and storage period of data collection, and ensure that the data provider is fully aware and clearly agrees to the collection and subsequent processing of data (including but not limited to data cleaning, feature extraction, association matching and electronic certificate generation, etc.) through legal forms such as written authorization and electronic confirmation.
[0126] If private data is collected or used without the data provider's valid consent, or if data leakage or misuse occurs due to violations of laws and regulations during data processing, the data collector and user shall bear the relevant responsibilities, which are unrelated to the technical solution of this invention. Simultaneously, the data processor must establish a robust data security mechanism, adopting measures such as encrypted storage, access control, and operation log retention to ensure the security and integrity of private data throughout the entire processing flow, preventing unauthorized access, tampering, or leakage of data.
[0127] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for improving the correlation rate between bank transaction details and electronic receipts, characterized in that, The method includes: Step 1: Automatically retrieve transaction details and electronic receipt data from multiple bank accounts opened by the enterprise through the bank-enterprise direct connection interface; Step 2: Identify and analyze the electronic receipt data, extract key accounting elements including transaction serial number, date, amount and counterparty information, and generate structured receipt data; Step 3: Construct a transaction data feature set based on structured return receipt data, determine the benchmark data point from the feature set, and generate two sets of data association features based on the benchmark data point; Step 4: Based on the correlation features of the two sets of data, calculate the time, amount, and opponent's difference to construct a multidimensional feature dataset, and determine the effective range by selecting positive and negative samples; map the samples to three-dimensional spatial points according to the time series to form polygon boundaries, determine the position of sample points through the odd-inside-even-outside rule, count the number of inside and outside points to calculate the density of inside points, construct the correlation matching rule system, and generate matching calibration parameters, including: Based on the first feature sequence and the second feature sequence, the transaction timestamp difference, transaction amount difference, and counterparty information similarity are calculated to form a multidimensional feature difference dataset including time difference, amount difference, and information similarity. By statistically analyzing the multidimensional feature difference dataset, the confidence intervals of the difference values in each dimension are determined, and the effective range of data association is generated. Positive sample points are selected within the effective range, and negative sample points are selected in the outer boundary region to form a validation sample set; Based on the time sequence of transactions, the sample points in the verification sample set are mapped to three-dimensional spatial coordinate points, where the X-axis represents the timestamp difference in seconds; the Y-axis represents the amount difference in yuan; and the Z-axis represents the information similarity. The baseline data points are then connected to construct the boundary of a polygonal trajectory. Draw a ray along the positive X-axis starting from the sample point, and calculate the number of intersections between the ray and the polygon boundary. If the number of intersections is odd, the sample point is determined to be inside the polygon; otherwise, it is outside. The number of validation sample points inside and outside the polygon is counted, the distribution density of the interior points is calculated, and based on the analysis results of the interior point distribution density, an association matching rule system is constructed and matching calibration parameters are generated. Step 5: Using the matching calibration parameters, based on the transaction serial number matching, and combined with the multi-factor association rules of transaction date, account, amount and counterparty name, the transaction details and electronic receipts are matched to obtain the matching results; Step 6: Using the matching results and existing order data resources, the matching rules are continuously trained and dynamically adjusted through a feedback mechanism to obtain the adjusted matching rules. Step 7: Based on the matching rules, automatically associate the successfully matched transaction details with the electronic receipts and generate auditable and traceable electronic vouchers.
2. The method for improving the correlation rate between bank transaction details and electronic receipts according to claim 1, characterized in that, The electronic receipt data is identified and analyzed to extract key accounting elements, including transaction serial number, date, amount, and counterparty information, to generate structured receipt data, including: The raw electronic receipt data obtained from the bank account is cleaned to remove irrelevant characters and formatting information, resulting in cleaned standardized text data. Based on preset rules and keyword library, the cleaned standardized text data is analyzed and identified to extract key fields, including transaction serial number, transaction date, transaction amount and counterparty information, and generate a set of key fields; The key field set is combined and transformed according to a preset format to generate a preliminary structured receipt data unit; The preliminary structured receipt data units are validated and standardized, including uniformly converting the transaction date format, standardizing the counterparty name and the transaction amount into numerical formats, and finally generating structured receipt data.
3. The method for improving the correlation rate between bank transaction details and electronic receipts according to claim 2, characterized in that, A transaction data feature set is constructed based on structured transaction receipt data. A baseline data point is determined from this feature set, and two sets of data correlation features are generated based on this baseline data point, including: Based on structured receipt data, features including transaction amount, transaction timestamp, payment account information, receiving account information, and counterparty identification information are extracted to construct an initial transaction data feature set; Statistical analysis was performed on the initial set of transaction data features to calculate the probability distribution and frequency of each feature value. Data points whose transaction amounts ranked in the top 20% of the frequency of similar transactions and whose transaction times were distributed across different dates and time periods were selected as benchmark data points. Based on the benchmark data points, corresponding feature information is extracted from bank transaction details data and electronic receipt data respectively to generate a first feature sequence representing the features of bank transaction details and a second feature sequence representing the features of electronic receipts. The feature matching degree is calculated for the first feature sequence and the second feature sequence, including the difference in transaction timestamps, the matching degree of transaction amounts, and the similarity of counterparty information. Based on the matching degree calculation results, two sets of data association features are generated for association matching.
4. The method for improving the correlation rate between bank transaction details and electronic receipts according to claim 3, characterized in that, By statistically analyzing the multidimensional feature difference dataset, confidence intervals for the difference values of each dimension are determined, and the effective range of data association is generated, including: Perform statistical distribution analysis on the multidimensional feature difference dataset and calculate the mean and standard deviation of the difference values for each feature dimension; Based on the mean and standard deviation, the confidence intervals of the difference values of each dimension are calculated according to the preset confidence level, and the confidence intervals are determined as the effective range boundary values of each feature dimension; By integrating the effective range boundary values of each feature dimension, an effective range space for multidimensional data association is constructed, forming a complete effective range boundary for data association.
5. The method for improving the correlation rate between bank transaction details and electronic receipts according to claim 4, characterized in that, Using matching calibration parameters, based on transaction serial number matching, and combined with multi-factor association rules for transaction date, account, amount, and counterparty name, transaction details and electronic receipts are matched to obtain matching results, including: Using matching calibration parameters, a multi-factor weighted association rule is constructed, which includes transaction serial number, transaction date, account information, transaction amount, and counterparty name. Based on multi-factor weighted association rules, the bank transaction details data and standardized receipt data are matched and calculated to obtain preliminary matching results; The initial matching results are subjected to conflict detection to identify data items with matching conflicts, and conflict resolution is performed on the data items with matching conflicts to generate the final matching results.
6. The method for improving the correlation rate between bank transaction details and electronic receipts according to claim 5, characterized in that, Using the matching results and existing order data resources, the matching rules are continuously trained and dynamically adjusted through a feedback mechanism to obtain the adjusted matching rules, including: The matching results are compared with the actual return receipt data to identify matching errors and analyze rule defects, and generate error analysis results. Based on the error analysis results, a training sample set was constructed using existing return receipt data resources. This set was used to iteratively train and adjust the parameters of the similarity calculation model in the matching rules, resulting in an updated matching model. Based on the updated matching model, the threshold and weight strategies used in the matching process are dynamically adjusted to form the adjusted matching rules.
7. The method for improving the correlation rate between bank transaction details and electronic receipts according to claim 6, characterized in that, According to the matching rules, the successfully matched transaction details are automatically associated with the electronic receipts, and auditable and traceable electronic vouchers are generated, including: Based on the adjusted matching rules, the system performs matching operations on unrelated transaction details and electronic receipt data to identify associatable transaction receipt pairs and generate a list of successfully matched transactions. Based on the successfully matched list, extract the corresponding transaction details and original electronic receipt information, bind and store the two together to form a related data record with a clear correspondence. Based on associated data records, structured electronic certificates are automatically generated, including key transaction information, receipt image indexes, associated timestamps, and audit trail identifiers.
8. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Bank receipt electronic management system and method
CN115221854A
Electronic receipt matching method and device based on interconnection and intercommunication
CN120563221A