Intelligent audit analysis method and system for multi-source heterogeneous bank flow data
Patent Information
- Application Number
- CN202610872025.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]第一,标准化困难,数据质量一致性差
[0020] 1. By building a rule engine based on feature-based row matching, the system can automatically identify the standard data templates to which different bank transaction files belong. It automatically matches appropriate parsing strategies for various formats, such as single-line headers, multi-line headers, and formats with preceding descriptions or trailing statistics. This completes header location, field extraction, and structured parsing, transforming the inefficient process of manually reviewing, organizing, and repeatedly verifying each file into an automated workflow. Simultaneously, through the rule configuration and management center, business personnel can flexibly define field mapping, cleaning, and transformation rules via a visual interface. The data processing module strictly and consistently applies these rules to all data, accumulating data processing knowledge into reusable and version-manageable rule assets. This ensures that output data conforms to audit standard templates from the source, avoiding omissions and inconsistencies caused by manual operations, and significantly shortening the data preparation cycle for audit projects.
Smart Images

Figure CN122656789A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and in particular to an intelligent audit analysis method and system for multi-source heterogeneous bank transaction data. Background Technology
[0002] In the fields of financial auditing and related regulation, the processing and analysis of bank transaction data is a fundamental yet demanding core task. Different banks have different field names, formats, and business definitions, requiring manual comparison, cleaning, and mapping of each data point to unify them into the standard templates needed for audit analysis.
[0003] Existing technologies have the following main drawbacks when dealing with bank transaction data processing:
[0004] First, standardization is difficult, and data quality consistency is poor. Transaction fields vary greatly between different banks and in different formats. Existing technologies mostly rely on manual mapping and cleaning by auditors, lacking a unified, configurable rule engine for centralized management. This makes it difficult to unify data processing standards, and the same rule may produce differences when executed by different people or at different times, and the cost of changes and adjustments is high.
[0005] Second, the process is opaque and lacks auditability. The transformation, cleaning, and deduplication operations that the data undergoes lack systematic records. When questions arise about the final data, it is impossible to quickly trace the specific source of the original data and the processing procedures. This "black box" process cannot meet the core audit requirements of full-chain traceability and verifiability.
[0006] Third, the rules are rigid, resulting in poor flexibility and reusability. Existing semi-automated scripts are usually written once for specific banks and specific formats. When new templates or rule adjustments are made, technical personnel need to modify the code again. Business personnel cannot directly configure and manage them, leading to slow technical response and difficulty in knowledge accumulation.
[0007] Fourth, quality checks rely on manual post-processing and lack closed-loop control. For critical quality checks such as data integrity and uniqueness, existing technologies often rely on manual spot checks or post-processing verification using independent tools after processing. This is inefficient, lacks comprehensive coverage, and cannot achieve real-time closed-loop quality control in the processing flow, which may lead to erroneous data flowing into subsequent analysis stages.
[0008] Therefore, there is a need to provide intelligent auditing and analysis methods and systems for multi-source heterogeneous bank transaction data to improve the automation level and quality of bank transaction data auditing. Summary of the Invention
[0009] This invention provides an intelligent audit analysis system for multi-source heterogeneous bank transaction data, comprising: a rule configuration and management module for constructing a standard template library and a rule library, wherein the standard template library includes standard data templates required for audit analysis; a data access and parsing module for accessing bank transaction files, identifying standard data templates for the bank transaction files based on the rule library and the standard template library, and parsing the bank transaction files according to the standard data templates to generate structured intermediate data; a data processing module for generating standardized data results that meet the requirements of the standard templates based on the mapping rules corresponding to the standard data templates of the bank transaction files and the structured intermediate data; a quality inspection and traceability module for performing quality inspection on the standardized data results and recording the parsing process data; and a result output and reporting module for visualizing the standardized data results.
[0010] Furthermore, the data access and parsing module includes: a file access submodule, used to access bank transaction files through a unified file access interface; a template recognition submodule, used to identify the standard data template of the bank transaction file according to the multi-type template recognition rules in the rule base and the standard template library, and to retrieve the corresponding parsing strategy; and a classification parsing submodule, used to parse the bank transaction file based on the parsing strategy corresponding to the standard data template to which the bank transaction file belongs, and to generate structured intermediate data.
[0011] Furthermore, the template recognition submodule identifies the standard data template of bank statement documents based on multiple template recognition rules in the rule base and the standard template library, including: identifying the key identifier lines of the bank statement documents according to the feature line recognition rules; determining the location of the data header of the bank statement documents according to the header positioning rules; identifying the start position, end position, and end summary information area of the transaction details of the bank statement documents according to the data area recognition rules; and identifying the standard data template of the bank statement documents based on the key identifier lines, the location of the data header, the start position, end position, and end summary information area of the transaction details.
[0012] Furthermore, the classification and parsing submodule parses the bank statement file based on the parsing strategy corresponding to the standard data template to which the bank statement file belongs, generating structured intermediate data. This includes: if the standard data template to which the bank statement file belongs is a single-row header template, extracting data row by row after locating the header row and data area of the bank statement file to generate structured intermediate data; if the standard data template to which the bank statement file belongs is a multi-row header template, merging, standardizing, and merging fields of the multi-row headers, establishing column index relationships, extracting data, and generating structured intermediate data.
[0013] Furthermore, the mapping rules include field mapping rules for mapping field names to platform standard fields, data conversion rules for converting original cell content into data formats that conform to standard data template requirements, and field combination and splitting rules.
[0014] Furthermore, the data processing module generates standardized data results that meet the requirements of the standard template based on the mapping rules corresponding to the standard data template of the bank transaction file and the structured intermediate data. This includes: transforming the fields of the structured intermediate data column by column according to the field mapping rules corresponding to the standard data template; formatting and normalizing the transformed field values according to the data transformation rules corresponding to the standard data template; and performing field merging or splitting processing according to the field combination and splitting rules corresponding to the standard data template.
[0015] Furthermore, the quality inspection and traceability module includes: a task and file information traceability submodule, used to store basic task information and file information corresponding to bank transaction files; a parsing result traceability submodule, used to store structured intermediate data; a field mapping result traceability submodule, used to store information on field relationships before and after field mapping, field conversion results, and mapping rules; and a result association storage module, used to establish associations between bank transaction files, structured intermediate data, field mapping results, and standardized data results through task identifiers, file identifiers, and record association relationships, forming a processing traceability data link.
[0016] Furthermore, the rule configuration and management module is also used to maintain multiple template recognition rules and mapping rules for different standard data templates through a graphical interface.
[0017] Furthermore, the quality inspection and traceability module is also used to: perform key field non-empty verification, date format validity verification, amount field value validity verification, and duplicate transaction record identification on standardized data results, and generate quality inspection results; and generate anomaly markers based on the quality inspection results.
[0018] This invention provides an intelligent audit analysis method for multi-source heterogeneous bank transaction data, comprising: constructing a standard template library and a rule library, wherein the standard template library includes standard data templates required for audit analysis; accessing bank transaction files and, based on the rule library and the standard template library, identifying the standard data templates of the bank transaction files; parsing the bank transaction files according to the standard data templates to generate structured intermediate data; generating standardized data results that meet the requirements of the standard templates based on the mapping rules corresponding to the standard data templates of the bank transaction files and the structured intermediate data; performing quality checks on the standardized data results and recording the parsing process data; and visualizing the standardized data results.
[0019] Compared with existing technologies, the intelligent auditing and analysis method and system for multi-source heterogeneous bank transaction data provided by this invention has at least the following beneficial effects:
[0020] 1. By building a rule engine based on feature-based row matching, the system can automatically identify the standard data templates to which different bank transaction files belong. It automatically matches appropriate parsing strategies for various formats, such as single-line headers, multi-line headers, and formats with preceding descriptions or trailing statistics. This completes header location, field extraction, and structured parsing, transforming the inefficient process of manually reviewing, organizing, and repeatedly verifying each file into an automated workflow. Simultaneously, through the rule configuration and management center, business personnel can flexibly define field mapping, cleaning, and transformation rules via a visual interface. The data processing module strictly and consistently applies these rules to all data, accumulating data processing knowledge into reusable and version-manageable rule assets. This ensures that output data conforms to audit standard templates from the source, avoiding omissions and inconsistencies caused by manual operations, and significantly shortening the data preparation cycle for audit projects.
[0021] 2. By assigning a unique identifier to each piece of data and transmitting it throughout the entire process, combined with operation records at key processing nodes, a complete data lineage map is constructed. Auditors can use a visual traceability interface to trace any result record back to its original source and view all intermediate transformation steps, meeting the core requirements of auditing for process credibility and result verifiability. Simultaneously, the quality inspection and traceability modules are deeply integrated into the core processing flow, automatically performing checks such as non-empty verification of key fields, date format verification, monetary value verification, and duplicate record identification. Abnormal data is identified and marked in real time rather than discarded, forming a proactive quality management closed loop of "processing-verification-reporting-correction," effectively preventing problematic data from flowing into downstream analysis stages.
[0022] 3. When new bank templates or business rules change, business personnel can quickly adjust the rules without technical development intervention. Project experience is accumulated and shared, reducing maintenance costs and enabling rapid adaptation to different audit scenarios. The final output structured data fully conforms to the preset audit standard template. Dates, amounts, enumeration values, etc., are all uniformly formatted. At the same time, uploaded file information, parsed content, field mapping content, and final processing results are classified, stored, and linked. This facilitates quick location of the problem in template recognition, field extraction, or field mapping when results are abnormal, comprehensively improving the transparency, standardization, and verifiability of data processing. Attached Figure Description
[0023] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0024] Figure 1 This is a schematic diagram of a module for an intelligent audit and analysis system for multi-source heterogeneous bank transaction data, as shown in some embodiments of this specification.
[0025] Figure 2 This is a schematic diagram of the architecture of an intelligent audit and analysis system for multi-source heterogeneous bank transaction data, as shown in some embodiments of this specification.
[0026] Figure 3 This is a flowchart illustrating an intelligent audit analysis method for multi-source heterogeneous bank transaction data, based on some embodiments of this specification.
[0027] Figure 4 This is a schematic diagram of an electronic device according to some embodiments of this specification. Detailed Implementation
[0028] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0029] Figure 1 This is a schematic diagram of a module for an intelligent audit and analysis system for multi-source heterogeneous bank transaction data, as shown in some embodiments of this specification. Figure 1 As shown, the intelligent audit analysis system for multi-source heterogeneous bank transaction data can include a rule configuration and management module, a data access and parsing module, a data processing module, a quality inspection and traceability module, and a result output and reporting module. Considering that bank transaction data is primarily provided in structured table format in actual business operations, but the templates for structured tables corresponding to different banks, different export channels, and different account types vary significantly in terms of header positions, field naming, description row content, merged cell formats, loan amount layout, and data start and end ranges, this system adopts a unified access, feature recognition, rule matching, and classification parsing technical approach to achieve automated parsing of various bank transaction files. This changes the inefficient traditional processing method that relies on manual review, manual organization, and repeated verification. By transforming the originally scattered, repetitive, and experience-dependent operations into an automated system process, the efficiency of bank transaction data processing is significantly improved, manual input is reduced, and the data preparation cycle for audit projects is shortened.
[0030] The rule configuration and management module is used to build standard template libraries and rule libraries.
[0031] The standard template library includes standard data templates required for audit analysis, which clearly define the standard field names, field types, and format requirements, such as "transaction time", "transaction amount", "account balance", "counterparty name", and "transaction summary".
[0032] The rule base includes multiple types of template recognition rules, which at least include the following:
[0033] Feature line identification rules: used to identify key identifier lines in bank statement files, such as header lines containing fields such as "transaction date", "summary", "counterparty account name", "debit amount", "credit amount", "balance", etc., or description lines containing bank name, account information, start and end dates, etc.
[0034] Header positioning rules: used to determine the location of the header in bank statement documents, and whether there are multiple header rows or cross-column merging, etc.
[0035] Data area identification rules: used to identify the start and end positions of the transaction details in bank statement documents, as well as the summary information area at the end.
[0036] The rule base can also include parsing strategies corresponding to different standard data templates. If the standard data template to which the bank statement file belongs is a single-row header template, the parsing strategy is: after locating the header row and data area of the bank statement file, extract the data row by row to generate structured intermediate data; if the standard data template to which the bank statement file belongs is a multi-row header template, the parsing strategy is: merge, standardize, and consolidate the fields of the multi-row headers, establish column index relationships, extract the data, and generate structured intermediate data; for standard data templates with a pre-description area or a tail statistics area, the parsing strategy is: filter out non-transaction details to avoid erroneous extraction.
[0037] Mapping rules include field mapping rules for mapping field names to platform standard fields, data conversion rules for converting raw cell content into data formats that conform to standard data template requirements, and field combination and splitting rules:
[0038] Field mapping rules: Map field names in the source file to platform standard fields, such as mapping "transaction date" and "accounting date" to the standard field "transaction time";
[0039] Data conversion rules: Convert the original cell content into a data format that conforms to the requirements of the standard template, such as standardizing the date format, converting the amount field to a numeric type, and standardizing the income and expenditure direction, etc.
[0040] Field combination and splitting rules: When the same information is stored in a scattered or combined manner, the fields are merged or split according to preset rules.
[0041] In some embodiments, the rule configuration and management module is also used for:
[0042] The system maintains various template recognition rules and mapping rules for different standard data templates through a graphical interface.
[0043] By managing rules in a unified, visual, and versionable manner, a shift from rigid program logic to the accumulation of business rules has been achieved. The adaptation experience of standard data templates for various banks no longer relies on the memory of individual developers, but rather forms reusable, accumulative, and maintainable rule assets, which is conducive to improving the organization's data processing capabilities and continuous adaptation capabilities. Business personnel can directly configure and maintain rules based on a visual interface, without relying on developers to repeatedly modify program code, thereby improving the efficiency of rule adjustment and template adaptation capabilities.
[0044] The data access and parsing module is used to access bank transaction files, identify standard data templates for bank transaction files based on the rule base and standard template base, parse the bank transaction files according to the standard data templates, and generate structured intermediate data.
[0045] Specifically, the data access and parsing module provides a web interface and API interface, allowing users to upload bank transaction files to be processed for specific audit tasks, such as supporting structured table formats like .xlsx, .xls, and .csv.
[0046] In some embodiments, the data access and parsing module includes:
[0047] The file access submodule is used to access bank transaction files through a unified file access interface;
[0048] The template recognition submodule is used to identify standard data templates for bank statement documents based on multiple template recognition rules in the rule base and the standard template library, and to retrieve the corresponding parsing strategy.
[0049] The classification and parsing submodule is used to parse bank transaction files based on the parsing strategy corresponding to the standard data template to which the bank transaction file belongs, and generate structured intermediate data.
[0050] In some embodiments, the template recognition submodule identifies standard data templates for bank statement documents based on multiple template recognition rules in the rule base and a standard template library, including:
[0051] Based on the feature line recognition rules, identify the key identifier lines in the bank statement documents;
[0052] Based on the header positioning rules, determine the location of the data header in the bank statement file;
[0053] Based on the data area identification rules, identify the start position, end position, and end summary information area of the bank statement details;
[0054] Identify the standard data template of bank statement documents based on key identifier lines, data header locations, start and end positions of transaction details, and summary information areas at the end.
[0055] Specifically, the key identifier lines, header locations, start and end positions of transaction details, and tail summary information areas of bank statement documents are compared with standard data templates in the rule base to automatically identify the standard data template to which the bank statement document belongs. For example, the key identifier lines, header locations, start and end positions of transaction details, and tail summary information areas of bank statement documents are converted into structured feature vectors. The rule base pre-sets standard data templates for various banks and formats. Each standard data template defines corresponding key feature descriptions, including: the set of keywords that the identifier line should contain, the names and order of header fields, the format characteristics of the detail data (such as date format, whether the amount contains thousands separators, the way payment and receipt indicators are expressed, etc.), and the names and relative positions of the tail summary fields, etc. The structured feature vector of the bank statement document is matched against each standard template in the rule base using multi-dimensional features: First, keyword matching is performed to determine whether the identifier row and header field match the keyword set defined in the standard data template; second, positional feature matching is performed to determine whether the header row number, detail start and end row numbers, and summary area row number fall within the range defined in the standard data template; third, format feature matching is performed to determine whether the regular expressions of each column of detail data conform to the format specifications defined in the standard data template; finally, the matching results of the above dimensions are combined, and a weighted scoring mechanism is used to calculate the overall similarity score between the bank statement document and each standard data template. The standard data template with the highest score that exceeds a preset threshold is determined to be the standard data template to which the document belongs.
[0056] In some embodiments, the classification and parsing submodule parses the bank statement file based on the parsing strategy corresponding to the standard data template to which the bank statement file belongs, generating structured intermediate data, including:
[0057] If the standard data template of the bank statement file is a single-row header template, the data is extracted row by row after locating the header row and data area of the bank statement file to generate structured intermediate data. Specifically, according to the standard data template of the bank statement file, the row number of the header row and the starting and ending row numbers of the data area are located. A one-to-one column index relationship is established between the column field names of the header row (such as transaction date, summary, debit amount, credit amount, balance, etc.) and the columns of the data area. Then, starting from the starting row of the data area, the data is read row by row, and the field values of each row are extracted by column. The data is then formatted and converted according to the data type specifications defined in the template (such as converting the date to YYYY-MM-DD format, removing the thousands separator of the amount and converting it to a floating-point number, mapping the payment and receipt flags to standard codes, etc.). Null values or abnormal values are marked and processed according to the template rules. Finally, structured intermediate data with unified field names as keys and column data as values is generated.
[0058] If the standard data template of the bank statement file is a multi-row header template, the multi-row headers are merged, standardized, and their fields are consolidated. After establishing column index relationships, the data is extracted to generate structured intermediate data. Specifically, since the header information is scattered across multiple rows (for example, the first row may contain major categories such as "income" and "expenses," while the second row may contain subfields such as "amount" and "summary," or there may be cases of cross-row merged cells), the multi-row headers are first logically concatenated and consolidated according to column affiliation based on the predefined header merging rules of the template. This eliminates interference from cross-row and merged cells, merging the multi-row headers into a complete single-row field name list and establishing accurate column index relationships. Then, the values of each field are extracted row by row from the detailed data area and formatted to generate structured intermediate data consistent with the output format of the single-row header template.
[0059] For standard data templates with a pre-explanation area or a tail statistics area, the parsing strategy is as follows: filter out non-transaction details to avoid erroneous extraction. Specifically, based on the predefined area boundary rules of the template, the pre-explanation area (such as bank statements, printed information, page numbers, and other non-transaction content) and the tail statistics area (such as totals, total number of transactions, total amount, and other summary information) are first filtered out. Only the transaction details data area is retained for parsing. When extracting data line by line, it is simultaneously verified whether each line of data conforms to the detail line format characteristics defined by the template (such as whether the date column is a valid date and whether the amount column is a valid value). Lines that do not conform to the characteristics are automatically marked as abnormal lines and skipped to avoid erroneously extracting explanatory text or statistical data as transaction records, thereby ensuring that the generated structured intermediate data only contains valid transaction details.
[0060] This allows the system to adapt to the differences in structure and field representation of standard data templates from different banks, stably converting raw heterogeneous transaction files into unified structured intermediate data, thus providing a foundation for subsequent field mapping and standardization processing.
[0061] The data processing module is used to generate standardized data results that meet the requirements of the standard template based on the mapping rules corresponding to the standard data template of the bank transaction file and the structured intermediate data.
[0062] Specifically, it includes:
[0063] Based on the field mapping rules corresponding to the standard data template, the fields of the structured intermediate data are transformed column by column;
[0064] Based on the data transformation rules corresponding to the standard data template, the transformed field values are formatted and standardized;
[0065] Based on the field combination and splitting rules corresponding to the standard data template, perform field merging or splitting processing.
[0066] Specifically, the process first invokes the field mapping rules corresponding to the standard data template of the bank transaction files to map the field names from different bank original files (such as each bank's custom "Transaction Date", "Occurrence Date", "Posting Date", "Counterparty Account Number", "Counterparty Account Name", etc.) to the unified field names defined in the standard template (such as "Transaction Date", "Counterparty Account Number", "Counterparty Account Name", etc.), achieving automatic alignment of different bank field definitions to the audit standard definition. After completing the field mapping, the process further formats and standardizes the field values according to the data conversion rules corresponding to the standard data template. For example, it converts different date formats from various banks (such as YYYY / MM / DD, YYYY.MM.DD, DD / MM / YYYY, etc.) to the YYYY-MM-DD standard format and converts amount strings containing thousands separators to standard floating-point numbers. The different payment and receipt symbols used by various banks (such as "debit / credit", "income / expenditure", "D / C", "inflow / outflow", etc.) are uniformly mapped to standard codes (such as "income" being recorded as a positive value and "expenditure" as a negative value), ensuring that the same field has a consistent data type and expression standard across different data sources. In addition, flexible processing is performed according to the field combination and splitting rules corresponding to the standard data template: for fields that need to be combined, such as "transaction date" and "transaction time" belonging to two separate columns in the original data, they are merged into a unified "transaction date and time" field according to the rules; for fields that need to be split, such as the "summary" field in the original data containing the combination information of "transaction type + counterparty name", it is split into two independent fields, "transaction type" and "counterparty name", so that the final output standardized data results fully comply with the requirements of the standard template in terms of field names, data format, and field granularity.
[0067] The quality inspection and traceability module is used to perform quality checks on standardized data results and record data during the parsing process.
[0068] In some embodiments, the quality inspection and traceability module includes:
[0069] The Task and File Information Tracking submodule is used to store basic task information and file information corresponding to bank transaction files;
[0070] The parsing result logging submodule is used to store structured intermediate data;
[0071] The Field Mapping Result Tracking Submodule is used to store information about the field relationships before and after field mapping, the field conversion results, and the mapping rules.
[0072] The result association storage module is used to establish associations between bank transaction files, structured intermediate data, field mapping results, and standardized data results through task identifiers, file identifiers, and record association relationships, forming a data chain for processing and leaving traces.
[0073] Specifically, the task and file information logging submodule independently stores basic task information and uploaded file information in the database when each processing task is created, including task identifier, file name, file source, upload time, task to which it belongs, and processing status, providing a unified index entry for all subsequent logging data. The parsing result logging submodule independently stores the intermediate structured content obtained after file parsing in the database, including the original worksheet structure, identified header information, parsed row data, and matched template rule information. This data can be used to check whether the template identification is accurate and whether the original transaction record has been completely extracted. The field mapping result logging submodule independently stores the field relationship before and after mapping, field conversion results, and rule version information used in the database during field mapping and standardization. This method can clearly identify which column in the original file a certain standard field comes from and what mapping logic and format conversion it has undergone. The result association storage module establishes a complete association between uploaded files, parsed content, field mapping results, and final standardization results through task identifiers, file identifiers, and record association relationships, forming a queryable processing logging data chain. Unlike traditional methods that only record simple operation logs, this system stores key data objects from the processing flow in separate databases, creating a clear link between uploaded files, parsing results, field mapping results, and the final output. When problems occur, users can directly locate the specific parsed and mapped content based on the task or file, quickly determining whether the problem occurred in the template recognition, field extraction, or field mapping stage. This improves troubleshooting efficiency and processing transparency, meeting the high requirements of auditing for verifiability and process reliability.
[0074] In some embodiments, the quality inspection and traceability module is also used for:
[0075] The standardized data results are validated for key fields (non-empty), date format validity, amount field value validity, and duplicate transaction records, generating quality inspection results.
[0076] Anomaly markers are generated based on the quality inspection results.
[0077] Specifically, each standardized data record is scanned to check for null values or missing values in key fields (such as transaction date, counterparty account, and amount). The date field is verified to conform to the preset standard format (e.g., YYYY-MM-DD), and the amount field is checked for valid values and absence of abnormal symbols. Suspected duplicate transaction records are identified by comparing combinations of transaction date, amount, and counterparty account fields. After completing these checks, based on the quality inspection results, records that do not conform to preset rules are marked with anomaly flags in the database. These flags include the anomaly type (e.g., "abnormal date format," "illegal amount," "suspected duplicate," etc.) and the specific rule number that triggered the anomaly. Instead of isolating or discarding the anomaly records, this process preserves the complete dataset while providing a clear basis for subsequent manual review and problem correction.
[0078] The results output and reporting module is used to visualize standardized data results.
[0079] The workflow of the intelligent audit and analysis system for multi-source heterogeneous bank transaction data includes the following steps:
[0080] Step 1: Create a task: The user first creates a specific processing task in the system, determines the task name, and the system generates a corresponding task identifier;
[0081] Specific handling:
[0082] 1. The system receives a task creation request submitted by the user;
[0083] 2. Generate a unique task identifier and create a task master record in the database;
[0084] 3. Record basic information such as task creation time, creator, and task status, which will serve as a unified associated object for subsequent file uploads, parsing, mapping, and exporting;
[0085] Output: Generates a unique task record and returns the task identifier for subsequent file processing.
[0086] Step 2, File Upload and Parsing: The user uploads bank statement files related to the created task. After receiving the file, the system performs preprocessing, feature line matching, template recognition, and classification parsing to extract structured intermediate data. The uploaded file information and the parsed content are stored independently in the database.
[0087] Specific handling:
[0088] 1. File Upload and Storage: Users upload one or more bank transaction files to the system, save the files to the file storage area, and record information such as file name, associated task, upload time, and file status in the database.
[0089] 2. Template Recognition: Reads multiple template recognition rules and identifies the template type of the current file by matching feature lines in the file. Feature lines include, but are not limited to, header lines containing fields such as "Transaction Date", "Summary", "Counterparty Name", "Debit Amount", "Credit Amount", and "Balance".
[0090] 3. Classification and Analysis: Based on the identified standard data template, the corresponding analysis method is called to extract the data.
[0091] For a standard data template with a single-row header, directly locate the header and data area and then extract the data row by row.
[0092] For standard data templates with multiple header rows, first concatenate and organize the header rows, then establish column index relationships;
[0093] For standard data templates that include a preface description area or a tail statistics area, filter out non-transaction details according to rules.
[0094] 4. Parsing Traceability: The intermediate parsed data is stored independently in the database, and the corresponding file, worksheet, header recognition result, template type and parsing result content are saved to facilitate subsequent investigation of template recognition and field extraction issues.
[0095] Output: Generates a structured parsed result dataset associated with tasks and files, and saves it to the database.
[0096] Step 3, Field Mapping: Based on the identified standard data template, load the corresponding field mapping rules and transformation rules, map the parsed structured intermediate data into standard fields, and store the mapping relationship and mapping result in the database.
[0097] Specific handling:
[0098] 1. Load the pre-set standard audit template and corresponding field mapping rules;
[0099] 2. Select the appropriate field mapping rule set based on the template type of the current file;
[0100] 3. Convert the original fields to standard fields according to the mapping rules. For example: the source fields "Transaction Date" and "Posting Date" are mapped to the standard field "Transaction Time".
[0101] 4. For fields that require rule transformation, execute preset data transformation logic, including field splitting, field merging, type conversion, content replacement, etc.
[0102] 5. Store the correspondence between fields before and after mapping, the rule version used, and the mapping results independently in the database.
[0103] Output: Generates an intermediate dataset with field names and business meanings unified to the standard data template, and saves the field mapping content.
[0104] Step 4: Standardization: Based on the identified standard data template, the field format, field meaning, and field values are uniformly processed to form standardized results that can be directly used for audit analysis. Simultaneously, quality rules are invoked to check the processing results, and abnormal records are marked in the database.
[0105] Specific handling:
[0106] 1. Date Standardization: Converts date content in different formats into a preset standard format, such as "YYYY-MM-DD" or "YYYY-MM-DD HH:MM:SS".
[0107] 2. Amount Standardization: Remove non-numeric content such as thousands separators and currency symbols from the amount field and convert them into a standard numerical format.
[0108] 3. Text cleaning: Remove leading and trailing spaces, unify full-width and half-width characters, and standardize special characters in text fields such as summary, recipient's name, and purpose.
[0109] 4. Enumeration value normalization: unify field values with the same meaning but different representations into a standard expression, such as converting loan identifiers in templates from different sources into a unified enumeration value.
[0110] 5. Quality Rule Check and Anomaly Marking: Pre-defined quality check rules are invoked to validate standardized results, such as checking for null values in key fields, validity of amounts, validity of dates, continuity of balances, and identification of suspected duplicate records. Records that trigger quality rules are marked as abnormal in the database.
[0111] 6. Standardized Traceability: The system saves standardized processing results and anomaly marking results to the database and establishes associations with tasks, files, parsing results, and field mapping results.
[0112] Output: Generate a standardized result dataset that meets the requirements of the standard audit template, and save the anomaly labeling information in the database.
[0113] Step 5: Export Results: Export the standardized data results that meet the standard template requirements to the specified format for audit analysis, subsequent calculations, or use by external systems.
[0114] Specific handling:
[0115] 1. Based on the export method selected by the user, organize the standardized results into the target output format (CSV or Excel);
[0116] Output: Generates standardized result files that can be used for audit analysis, data verification, and subsequent processing.
[0117] Figure 3 This is a flowchart illustrating an intelligent audit analysis method for multi-source heterogeneous bank transaction data, as shown in some embodiments of this specification. Figure 3 As shown, the intelligent audit analysis method for multi-source heterogeneous bank transaction data may include the following steps.
[0118] Build a standard template library and a rule library, whereby the standard template library includes standard data templates required for audit analysis;
[0119] It accesses bank transaction records and identifies standard data templates for bank transaction records based on a rule base and a standard template base. It then parses the bank transaction records based on these standard data templates to generate structured intermediate data.
[0120] Based on the mapping rules corresponding to the standard data template of bank transaction files, standardized data results that meet the requirements of the standard template are generated according to the structured intermediate data.
[0121] Perform quality checks on the standardized data results and record the data from the parsing process;
[0122] Visualize the results of standardized data.
[0123] The intelligent audit analysis method for multi-source heterogeneous bank transaction data can be applied to intelligent audit analysis systems for multi-source heterogeneous bank transaction data, which will not be elaborated here.
[0124] Figure 4 These are schematic diagrams of electronic devices illustrated according to some embodiments of this specification, such as... Figure 4As shown, the electronic device includes a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on a program stored in Read-Only Memory (ROM) or a program loaded from storage into Random Access Memory (RAM), such as executing the methods described in the above embodiments. Various programs and data required for system operation are also stored in the RAM. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0125] The following components are connected to the I / O interface: input sections including a keyboard, mouse, etc.; output sections 507 including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers, etc.; storage sections including hard disks, etc.; and communication sections including network interface cards such as LAN (Local Area Network) cards and modems, etc. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as needed.
[0126] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs various functions defined in the system of this application.
[0127] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0129] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0130] Another aspect of this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a computer's processor, causes the computer to perform the method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.
[0131] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.
[0132] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. An intelligent auditing and analysis system for multi-source heterogeneous bank transaction data, characterized in that, include: The rule configuration and management module is used to build a standard template library and a rule library, wherein the standard template library includes standard data templates required for audit analysis; The data access and parsing module is used to access bank transaction files, identify standard data templates for bank transaction files based on the rule base and standard template base, parse the bank transaction files according to the standard data templates, and generate structured intermediate data. The data processing module is used to generate standardized data results that meet the requirements of the standard template based on the mapping rules corresponding to the standard data template of the bank transaction file and the structured intermediate data. The quality inspection and traceability module is used to perform quality checks on standardized data results and record data during the parsing process. The Results Output and Reporting module is used to visualize standardized data results.
2. The intelligent auditing and analysis system for multi-source heterogeneous bank transaction data according to claim 1, characterized in that, The data access and parsing module includes: The file access submodule is used to access bank transaction files through a unified file access interface; The template recognition submodule is used to identify standard data templates for bank statement documents based on multiple template recognition rules in the rule base and the standard template library, and to retrieve the corresponding parsing strategy. The classification and parsing submodule is used to parse bank transaction files based on the parsing strategy corresponding to the standard data template to which the bank transaction file belongs, and generate structured intermediate data.
3. The intelligent auditing and analysis system for multi-source heterogeneous bank transaction data according to claim 2, characterized in that, The template recognition submodule identifies standard data templates for bank statement documents based on multiple template recognition rules in the rule base and a standard template library, including: Based on the feature line recognition rules, identify the key identifier lines in the bank statement documents; Based on the header positioning rules, determine the location of the data header in the bank statement file; Based on the data area identification rules, identify the start position, end position, and end summary information area of the bank statement details; Identify the standard data template of bank statement documents based on key identifier lines, data header locations, start and end positions of transaction details, and summary information areas at the end.
4. The intelligent auditing and analysis system for multi-source heterogeneous bank transaction data according to claim 2, characterized in that, The classification and parsing submodule parses the bank statement files based on the parsing strategy corresponding to the standard data template to which the bank statement files belong, generating structured intermediate data, including: If the standard data template of the bank statement file is a single-line header template, extract the data line by line after locating the header line and data area of the bank statement file to generate structured intermediate data; If the standard data template of the bank statement file is a multi-row header template, the multi-row headers are merged, standardized, and fields are consolidated. After establishing column index relationships, the data is extracted and structured intermediate data is generated.
5. The intelligent auditing and analysis system for multi-source heterogeneous bank transaction data according to any one of claims 2-4, characterized in that, The mapping rules include field mapping rules for mapping field names to platform standard fields, data conversion rules for converting original cell content into data formats that conform to standard data template requirements, and field combination and splitting rules.
6. The intelligent auditing and analysis system for multi-source heterogeneous bank transaction data according to claim 5, characterized in that, The data processing module, based on the mapping rules corresponding to the standard data template of bank transaction files, generates standardized data results that conform to the requirements of the standard template according to the structured intermediate data, including: Based on the field mapping rules corresponding to the standard data template, the fields of the structured intermediate data are transformed column by column; Based on the data transformation rules corresponding to the standard data template, the transformed field values are formatted and standardized; Based on the field combination and splitting rules corresponding to the standard data template, perform field merging or splitting processing.
7. The intelligent auditing and analysis system for multi-source heterogeneous bank transaction data according to claim 6, characterized in that, The quality inspection and traceability module includes: The Task and File Information Tracking submodule is used to store basic task information and file information corresponding to bank transaction files; The parsing result logging submodule is used to store structured intermediate data; The Field Mapping Result Tracking Submodule is used to store information about the field relationships before and after field mapping, the field conversion results, and the mapping rules. The result association storage module is used to establish associations between bank transaction files, structured intermediate data, field mapping results, and standardized data results through task identifiers, file identifiers, and record association relationships, forming a data chain for processing and leaving traces.
8. The intelligent auditing and analysis system for multi-source heterogeneous bank transaction data according to any one of claims 1-4, characterized in that, The rule configuration and management module is also used for: The system maintains various template recognition rules and mapping rules for different standard data templates through a graphical interface.
9. The intelligent auditing and analysis system for multi-source heterogeneous bank transaction data according to any one of claims 1-4, characterized in that, The quality inspection and traceability module is also used for: The standardized data results are validated for key fields (non-empty), date format validity, amount field value validity, and duplicate transaction records, generating quality inspection results. Anomaly markers are generated based on the quality inspection results.
10. An intelligent auditing and analysis method for multi-source heterogeneous bank transaction data, characterized in that, The intelligent audit and analysis system for multi-source heterogeneous bank transaction data as described in any one of claims 1-9 includes: Construct a standard template library and a rule library, wherein the standard template library includes standard data templates required for audit analysis; It accesses bank transaction records and identifies standard data templates for bank transaction records based on a rule base and a standard template base. It then parses the bank transaction records based on these standard data templates to generate structured intermediate data. Based on the mapping rules corresponding to the standard data template of bank transaction files, standardized data results that meet the requirements of the standard template are generated according to the structured intermediate data. Perform quality checks on the standardized data results and record the data from the parsing process; Visualize the standardized data results.