Metadata generation method for account book file based on custom rule definition

Through the metadata method of book archive generation defined by custom rules, the traditional financial system's lack of flexibility, inefficiency and poor accuracy caused by rule solidification is solved, and the flexibility and efficiency of book archive generation is improved, ensuring the accuracy and consistency of data.

CN120197595APending Publication Date: 2025-06-24INSPUR GENERSOFT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510270609.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Due to the solidification of rules, traditional financial systems have problems such as insufficient flexibility, inefficiency and poor accuracy in book file generation.

Method used

The metadata method of generating books and archives defined based on custom rules is adopted, and the original data is extracted through the extraction rules, the metadata model is constructed and the field differences are processed, and a structured metadata file is generated.

Benefits of technology

It realizes the flexibility and efficiency of book archive generation, ensures data accuracy and consistency, supports differentiated generation logic for multiple book types, and meets the needs of complex business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197595A_ABST
    Figure CN120197595A_ABST
Patent Text Reader

Abstract

The invention provides an account book file metadata generation method based on custom rule definition, and relates to the technical field of financial management, and the method comprises the steps: carrying out the field extraction of original data of an account book file based on an extraction rule, so as to obtain a key field; constructing a metadata model, associating the key fields with corresponding attributes of the metadata model through a mapping rule, and processing field differences to obtain a mapping result; and generating a structured metadata file based on the metadata model and the mapping result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of financial management, and specifically relates to a method for generating metadata of ledger files based on custom rule definitions. Background Art

[0002] With the complication of enterprise financial management, traditional paper ledgers have gradually been replaced by digital systems. Early computerized financial software achieved automated bookkeeping through preset rules, significantly improving efficiency and accuracy. However, such systems are usually based on fixed templates and are difficult to adapt to the personalized needs of enterprises. In recent years, the rise of metadata management technology has provided new ideas for financial data standardization. By defining data attributes and relationships, it enhances the interpretability and reusability of data. Nevertheless, there are still significant bottlenecks in the dynamic definition of metadata, flexible configuration of rules, and cross-system compatibility in the existing technologies.

[0003] Most financial software relies on predefined accounting set templates, and users need to enter data in fixed fields and formats. It is impossible to adjust the data structure according to specific enterprise operations (such as multi-currency accounting, custom account classification). When adding new business scenarios, it is necessary to modify the underlying code, which has a long development cycle and high costs. Some systems adopt standardized metadata models, but the model structures are rigid and only support limited field mappings. When there are large differences between the original data and the metadata model, manual intervention is required to handle field conflicts (such as unit conversion, coding standardization). When updating the metadata model, it is necessary to re-develop, and it is difficult to adapt to changes in accounting standards or regulatory policies. Some advanced systems introduce rule engines, but rule definitions rely on technical teams, and business personnel cannot directly participate in configuration. Business requirements need to be converted into technical rules by developers, resulting in high communication costs and easy to generate deviations. Rule optimization requires downtime maintenance, affecting system continuity. Summary of the Invention

[0004] This application provides a method for generating metadata of ledger files based on custom rule definitions to solve the technical problems of insufficient flexibility, low efficiency, and poor accuracy in generating ledger files caused by rigid rules in traditional financial systems.

[0005] The technical solution adopted by this application is as follows:

[0006] An embodiment of this application provides a method for generating metadata of ledger files based on custom rule definitions, including:

[0007] Based on extraction rules, extract fields from the original data of the ledger file to obtain key fields;

[0008] Construct a metadata model, associate the key fields with the corresponding attributes of the metadata model through mapping rules, and handle field differences to obtain a mapping result;

[0009] Generate a structured metadata file based on the metadata model and the mapping result.

[0010] According to an embodiment of the present application, the acquisition of the original data of the account book file is specifically as follows: differentiating the data by the account book category;

[0011] Performing integrity verification and confidentiality verification on the differentiated data;

[0012] The data after passing the integrity verification and the confidentiality verification is used as the original data of the account book file;

[0013] The original data includes: account set information, subject information, currency information, and transaction information.

[0014] According to an embodiment of the present application, the field extraction of the original data of the account book file based on the extraction rule to obtain key fields is specifically as follows:

[0015] Parsing the original data, and identifying and extracting the key fields related to the account book file from the parsed original data;

[0016] The parsing method includes: matching through regular expressions, positioning fields, and identifying keywords.

[0017] According to an embodiment of the present application, after identifying and extracting the key fields related to the account book file from the parsed original data, it further includes verifying and optimizing the key fields, specifically as follows:

[0018] Adjusting the extraction field range, and dividing the key fields into required fields and optional fields;

[0019] Checking whether there are any missing required fields;

[0020] Verifying the formats of the required fields and the optional fields;

[0021] If the extraction of the required fields and / or the optional fields fails, adjust the parsing method.

[0022] According to an embodiment of the present application, the construction of the metadata model is specifically as follows:

[0023] Presetting a basic metadata model, which includes general attributes;

[0024] Adding custom attributes to the preset basic metadata model according to user requirements to construct the metadata model.

[0025] According to an embodiment of the present application, the mapping rule is specifically as follows:

[0026] Direct mapping, directly associating when the field name is the same as the model attribute name;

[0027] Conversion mapping, including: unit conversion and encoding standardization;

[0028] Composite mapping, merging and / or splitting fields.

[0029] According to an embodiment of the present application, the difference processing is specifically:

[0030] If a field is missing, fill in the default value according to the rule;

[0031] If the same field is mapped to multiple attributes, select according to the priority.

[0032] According to an embodiment of the present application, after generating the structured metadata file based on the metadata model and the mapping result, it further includes the verification of the metadata, specifically:

[0033] Verify whether the structured metadata file completely reflects the original data structure;

[0034] Automatically mark abnormal data and trigger an alarm.

[0035] According to an embodiment of the present application, after the verification of the metadata, it further includes the generation of archives, specifically:

[0036] Automatically generate a standardized ledger archive based on the structured metadata file;

[0037] When the structured metadata file or the original data changes, trigger the regeneration process.

[0038] According to an embodiment of the present application, after the verification of the metadata, it further includes maintenance management, specifically:

[0039] Record the version history of the structured metadata file, supporting backtracking and difference comparison;

[0040] Regularly back up the structured metadata file and the ledger archive.

[0041] Due to the adoption of the above technical solution, the beneficial effects obtained by the present application are:

[0042] This application accurately identifies required fields (such as transaction date, amount) from heterogeneous data sources through custom extraction rules (such as regular expressions, field positioning). It supports structured and unstructured data (such as log files, CSV / XML), covering the requirements of various ledger types. It converts keyword fields into a unified format according to mapping rules (such as currency exchange rate conversion, account code standardization) to eliminate data ambiguity. By dynamically processing field differences (such as default value filling, conflict priority), it ensures seamless connection between the original data and the target model. It generates standardized metadata files (such as JSON, XML) based on a template engine to reduce manual splicing errors. It supports different generation logics for multiple ledger types (general ledger, subsidiary ledger) to meet the needs of complex business scenarios. It realizes full-process automation (extraction → mapping → generation) through a rule engine, and the generation efficiency is increased by more than 80% compared with traditional manual operations. Dynamic rule adjustment does not require system downtime to ensure continuous operation of the system. Users can customize rules and metadata models through a visual interface to quickly respond to business changes (such as adding accounts, adjusting report formats). It supports multi-standard adaptation (such as IFRS, GAAP) to meet the compliance requirements of multinational enterprises. A multi-level verification mechanism (integrity verification, logical verification) ensures that the metadata complies with business rules by 100%. Permission control and audit log functions (such as recording rule modifications, data access history) ensure the security and traceability of financial data. Modular design supports function expansion (such as adding verification rules, docking external systems), reducing the difficulty of secondary development. Metadata version management simplifies the backtracking of historical data and the comparison of differences, supporting quick repair and compliance auditing. Brief Description of the Drawings

[0043] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0044] Figure 1 It is a schematic flowchart of a method for generating metadata of ledger files based on custom rule definition provided by an embodiment of the present application. Detailed Embodiments

[0045] In order to more clearly illustrate the overall concept of the present application, the following will be described in detail by way of examples in conjunction with the drawings of the specification.

[0046] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application. However, the present application may be implemented in other ways different from those described herein. Therefore, the scope of protection of the present application is not limited by the specific embodiments disclosed below. It should be noted that, without conflict, the embodiments of the present application and the features in each embodiment may be combined with each other.

[0047] In this application, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may mean that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0048] Embodiment 1

[0049] As Figure 1 shown, a method for generating metadata of ledger files based on custom rules includes:

[0050] Extract fields from the original data of the ledger file based on the extraction rules to obtain key fields.

[0051] Specifically, first, it is necessary to distinguish data by ledger category. This means that the system can identify different types of ledger files (for example, cash journal, bank deposit journal, etc.) and process the corresponding data according to their types. Then, integrity verification and confidentiality verification will be performed on the distinguished data to ensure that the data is both complete and secure. Only after these verifications pass will the data be used as the original data of the ledger file. For the confirmed original data, parsing will be carried out next. The methods mentioned here include using regular expressions for matching to locate fields and identify keywords. The goal of this step is to identify and extract key fields related to the ledger file from the original data. For example, the transaction information may include keywords such as date, amount, and counterparty. Once the key fields are extracted, further verification and optimization are required. This step includes: classifying the key fields into required fields and optional fields. Check whether there are any missing required fields. If any are found missing, measures need to be taken to supplement or adjust. Verify the formats of the required fields and optional fields to ensure that they meet the expected standards. If any problems are found during the above process, such as missing required fields or incorrect formats, the parsing method needs to be adjusted and the correct information extraction needs to be tried again.

[0052] For example, first, distinguish data according to the ledger category. In this example, focus on the "Bank Deposit Journal". Obtain relevant data from the company's financial system, which may include but is not limited to information such as transaction date, summary, debit amount, credit amount, balance, etc. Then, perform integrity verification on this data (for example, ensure that each transaction has a corresponding date and amount) and confidentiality verification (ensure that sensitive information such as account numbers is not leaked).

[0053] Next, apply extraction rules to identify and extract key fields related to the ledger file. For example:

[0054] Regular expression matching: Use regular expressions to locate the date format (such as 2025-02-26), so that the transaction date can be accurately extracted from each record.

[0055] Keyword identification: Look for specific keywords in the summary column, such as "salary", "rent", "purchase", which helps in subsequent analysis of the use of funds.

[0056] Once the key fields are extracted, further verify and optimize them:

[0057] Required field check: Confirm that all transactions have necessary information such as date and amount. If it is found that a certain record is missing the transaction amount, mark this record as abnormal and try to re-extract the correct information by adjusting the parsing method.

[0058] Format verification: Ensure that all dates follow a unified format (such as YYYY-MM-DD), and the amount has no negative numbers (unless specifically indicated as a refund or other special circumstances).

[0059] Adjust the parsing method: If certain key fields cannot be successfully extracted during the initial parsing (such as errors caused by format problems), the parsing logic needs to be modified. For example, change the pattern of the regular expression to adapt to different date writing methods (such as DD / MM / YYYY or MM / DD / YYYY).

[0060] Furthermore, according to actual needs, divide the extracted key fields into required fields and optional fields. This helps to make the subsequent data verification process more efficient. Ensure that all required fields exist and are complete. If a missing situation is found, corresponding measures need to be taken for supplementation or marked as abnormal. Perform format verification on all extracted key fields to ensure that they meet the expected standards. For example, dates must follow a unified format, and amounts cannot contain illegal characters, etc. If problems are encountered during the initial extraction process (such as certain key fields not being correctly extracted), the parsing method needs to be adjusted. This may involve modifying the pattern of the regular expression or adopting other more effective parsing methods.

[0061] Construct a metadata model, associate the keyword fields with the corresponding attributes of the metadata model through mapping rules, and process field differences to obtain a mapping result.

[0062] Specifically, first, a basic metadata model containing common attributes needs to be set. This model usually covers the most common attributes and structures in the account book files. According to user requirements, custom attributes are added on top of the basic metadata model to adapt to specific business needs or special application scenarios. For example, for some enterprises, additional financial information may need to be recorded or specific industry standards need to be followed. When the field names in the original data are the same as the attribute names in the metadata model, the association can be established directly. In this case, the mapping process is relatively simple and straightforward. If there are inconsistencies in units or different coding formats, corresponding conversions are required. For example, converting the amount unit from "cents" to "yuan", or unifying different coding standards (such as unifying multiple date formats to YYYY-MM-DD). In some scenarios, multiple fields may need to be merged into one attribute, or one field may need to be split into multiple attributes. For example, splitting the name field into two independent attributes: surname and given name. If a field is found to be missing during the mapping process, default values can be filled according to preset rules. This helps to ensure the integrity and consistency of the generated metadata file. If a field maps to multiple attributes, the most appropriate mapping method needs to be selected according to the preset priority. This can avoid duplicate or conflicting information in the final metadata. Through the above steps, the mapping from keyword fields to metadata model attributes is completed, and the possible field differences are resolved, thus obtaining the final mapping result. These mapping results are the basis for constructing a structured metadata file and the key to ensuring that the account book files accurately reflect the original data. The entire process not only improves the efficiency of data processing but also enhances the consistency and reliability of the data.

[0063] For example, assume that a bank deposit journal of a company is being processed, and some keyword fields need to be extracted from the original data and mapped to a preset metadata model.

[0064] Example of original data:

[0065] Date: 2025-02-26, Description: Salary income, Income amount: 3000.00 CNY, Expenditure amount, Account balance: 15000.00 CNY

[0066] Preset basic metadata model (including common attributes):

[0067] transactionDate (transaction date)

[0068] description (abstract)

[0069] incomeAmount (Income amount)

[0070] expenseAmount (Expense amount)

[0071] accountBalance (Account balance)

[0072] Considering certain specific business requirements, additional information may also need to be recorded, such as the transaction type (income / expense), which requires adding a custom attribute transactionType to the metadata model.

[0073] Application of mapping rules

[0074] Direct mapping:

[0075] The date field can be directly mapped to the transactionDate attribute.

[0076] The summary field can be directly mapped to the description attribute.

[0077] The income amount field can be directly mapped to the incomeAmount attribute.

[0078] The expense amount field can be directly mapped to the expenseAmount attribute.

[0079] The account balance field can be directly mapped to the accountBalance attribute.

[0080] Conversion mapping:

[0081] If the currency unit in the original data is in "cents" while the metadata model requires "yuan", unit conversion is needed. However, in this example, all amounts are already given in the form of "CNY" and meet the requirements, so no unit conversion is required.

[0082] For the expense amount field that is not filled, a default value of 0 can be set to ensure that all numerical attributes have values.

[0083] Composite mapping: Determine the transactionType based on the existence of income amount or expense amount. In this example, since there is an income amount but no expense amount, it can be inferred that the transactionType is "income".

[0084] Handling Field Differences: In this case, since the transaction type is not explicitly specified in the original data, the value of transactionType needs to be determined based on the presence or absence of the income amount and expenditure amount fields. If both fields are empty, this record can be marked as an anomaly or filled with a special default value to indicate an unknown transaction type.

[0085] Obtain the mapping result

[0086] After the above steps, the mapping result finally obtained may be such structured metadata:

[0087]

[0088] This process demonstrates how to utilize the extracted key fields, associate them with the corresponding attributes of the metadata model through mapping rules, and handle field differences to generate the final mapping result. This not only ensures data consistency and integrity but also provides a solid foundation for subsequent data analysis and management.

[0089] Furthermore, first define a basic metadata model that includes common attributes. These attributes are usually designed according to the general requirements of the accounting records, such as basic financial information like transaction date, summary information, amount, etc. On this basis, users are allowed to add custom attributes according to specific business requirements. This step is very crucial because it enables the metadata model to flexibly adapt to different types of accounting records and enterprise-specific requirements. For example, certain industries may need to record additional information such as project numbers or contract numbers.

[0090] When the field names in the original data are exactly the same as the attribute names in the metadata model, a mapping relationship can be directly established. In this case, the mapping process is relatively straightforward and simple, without the need for additional conversion.

[0091] When the units used for some values in the original data (such as amounts) are inconsistent with the metadata model, unit conversion is required. For example, the amount in the original data may be in "cents", while the metadata model requires "yuan", and in this case, the numerical value needs to be divided by 100 for conversion.

[0092] For some data with specific formats (such as dates), if the original data format is different from the requirements of the metadata model, format conversion is required. For example, if the date format in the original data is MM / DD / YYYY, while the metadata model requires the YYYY-MM-DD format, corresponding conversion logic is needed to ensure consistency.

[0093] Sometimes, it is necessary to combine multiple raw data fields into an attribute in a metadata model. For example, the name field may be split into two separate fields, the surname and the given name, but only a complete name field may be required in the metadata model.

[0094] Conversely, the opposite situation may also occur, where a raw data field needs to be split into multiple attributes. For example, address information may contain multiple parts such as street, city, and postal code, which need to be split and mapped to the corresponding attributes respectively during mapping.

[0095] If a certain field is found to be missing during the mapping process, default values can be filled according to preset rules. For example, for required but missing fields, they can be set to "unknown" or "not provided" to ensure the integrity of the data structure.

[0096] When the same field is mapped to multiple attributes, the final mapped target attribute should be selected according to the preset priority. This helps to solve the problems of data duplication or conflict and ensures that each attribute has a uniquely determined value.

[0097] Generate a structured metadata file based on the metadata model and the mapping results.

[0098] Specifically, first, there is a preset basic metadata model, which contains common attributes (such as transaction date, summary information, amount, etc.). According to specific business requirements, users can add custom attributes to the basic metadata model to form a metadata model suitable for specific scenarios. Next, the key fields obtained from the previous steps have been associated with the corresponding attributes in the metadata model according to the set mapping rules, and possible field difference problems have been processed. These mapping results include but are not limited to:

[0099] Direct mapping: The field name in the raw data directly matches the attribute in the metadata model.

[0100] Conversion mapping: The mapping result after converting the unit or encoding format.

[0101] Composite mapping: The mapping result involving the combination or splitting of multiple fields.

[0102] Using the above metadata model and mapping results, organize all relevant information into a data format with a clear hierarchical structure. This usually means using standardized data formats such as XML, JSON, etc. to represent metadata for subsequent processing and parsing. During the verification phase, it is necessary to ensure that the generated structured metadata file completely reflects the structure of the original data. This means that all required fields have been correctly filled and no important information points have been omitted. If any data that does not meet expectations (such as format errors or logical contradictions) is found during the verification process, the system should automatically mark these abnormal data and trigger an alarm mechanism to prompt relevant personnel to check or correct them. The last step is to output the verified structured metadata in the form of a file. This file not only contains all the key information extracted from the original ledger archives but also follows the predefined metadata model specifications to ensure that it can be accurately read and understood by other systems or applications.

[0103] For example, assume that you are processing the bank deposit journal archives of a company and have completed the extraction of key fields and the mapping to the metadata model. Now, generate a structured metadata file based on this information.

[0104] Example of original data:

[0105] Date: 2025-02-26, Description: Salary income, Income amount: 3000.00 CNY, Account balance: 15000.00 CNY

[0106] Metadata model (including general attributes):

[0107] transactionDate (transaction date)

[0108] description (abstract)

[0109] incomeAmount (income amount)

[0110] accountBalance (account balance)

[0111] Mapping result: According to the previous steps, we have successfully mapped the key fields in the original data to the corresponding attributes in the above metadata model.

[0112] Generate a structured metadata file

[0113] Step 1: Determine the output format

[0114] Select a structured format suitable for storing financial data, such as JSON or XML. In this example, JSON format will be used.

[0115] Step 2: Integrate the mapping results

[0116] Integrate the mapping results into the selected format. For the above example, the JSON structure might be as follows:

[0117]

[0118] Step 3: Integrity Check

[0119] During the verification stage, ensure that all required fields are correctly filled and no important information points are missed. For example, in this case, we need to check if all necessary fields such as transactionDate, description, incomeAmount, and accountBalance are correctly filled.

[0120] Step 4: Anomaly Marking and Alarming

[0121] If any data that does not meet expectations (such as format errors or logical contradictions) is found during the verification process, the system should automatically mark these abnormal data and trigger the alarm mechanism. For example, if the value of the incomeAmount field is negative, it is regarded as an abnormal situation and requires manual review or adjustment of the data source.

[0122] Step 5: Output File

[0123] Finally, output the verified structured metadata in the form of a file. It can be saved in a form similar to bank_deposit_journal_20250226.json. This file not only contains all the key information extracted from the original ledger archive but also follows the predefined metadata model specification to ensure that it can be accurately read and understood by other systems or applications.

[0124] Furthermore, select a suitable structured format for storing metadata according to actual needs, such as JSON, XML, etc. These formats provide a clear data hierarchy and are easy to parse by machines. JSON: A lightweight data interchange format that is easy to read and write, and also easy to parse and generate. XML: A standard for defining structured information in a flexible way, suitable for complex data structures.

[0125] Organize data according to the metadata model: Ensure that all key fields extracted from the original data and mapped to the metadata model are correctly organized into the selected structured format. For example, in the JSON format, each attribute name corresponds to an attribute in the metadata model, and its value is the corresponding mapping result.

[0126] Before generating the structured metadata file, it is necessary to perform an integrity check on the data to confirm that all required fields are filled and no important information is missing. This includes, but is not limited to, checking for null values or missing necessary attributes. Ensure that the converted data accurately reflects the content of the original data. For example, whether the amount field has been correctly converted in terms of units and whether the date format is unified, etc.

[0127] If any data that does not meet the expectations (such as format errors or logical contradictions) is found during the verification process, the system should be able to automatically identify and mark these abnormal data points. For the detected abnormal situations, the system should be able to trigger an alarm to notify the relevant personnel for review or correction. This is crucial for maintaining data quality and reliability.

[0128] Once the above steps are completed, the sorted structured data can be output in the form of a file. The file naming usually follows certain rules for easy management and retrieval, such as including a date-time stamp or other identifiers. Ensure that the generated file uses a consistent character encoding (such as UTF-8) to avoid problems caused by encoding mismatches.

[0129] According to an embodiment of the present application, the acquisition of the original data of the account book file is specifically: differentiating the data by the account book category;

[0130] Performing an integrity check and a confidentiality check on the differentiated data;

[0131] The data after passing the integrity check and the confidentiality check is used as the original data of the account book file;

[0132] The original data includes: account set information, subject information, currency information, and transaction information.

[0133] Specifically, differentiating the data by the account book category

[0134] Account book category: First of all, the system needs to identify and classify different types of account book files. The account book category may include, but is not limited to, cash journal, bank deposit journal, accounts receivable journal, etc. This step ensures that different types of data can be processed separately, thereby improving the efficiency and accuracy of subsequent processing.

[0135] Performing an integrity check and a confidentiality check on the differentiated data

[0136] Integrity check

[0137] Purpose: Ensure that each record contains the necessary information and no important fields are missing.

[0138] Operation: For example, in a bank deposit journal, each record should contain information such as the transaction date, description, income amount or expenditure amount (at least one), account balance, etc. If any of the above required fields are missing from a record, the record will be considered incomplete and may require further manual review or automatic correction.

[0139] Confidentiality check

[0140] Purpose: To protect sensitive information from unauthorized access or disclosure.

[0141] Operation: For data containing sensitive information (such as personal identity information, account numbers, etc.), measures must be taken to ensure the security of this information. This may involve measures such as encrypted storage, restricted access rights, etc. In addition, it is necessary to check whether a secure protocol is used during data transmission to prevent information from being intercepted during transmission.

[0142] The data that has passed the integrity and confidentiality checks is used as the original data for the account book files

[0143] Verify data quality: Only when the data has passed both the integrity and confidentiality checks can it be considered reliable original data for further processing.

[0144] Ensure data reliability: This means that all data used is accurate and secure, providing a solid foundation for generating metadata files later.

[0145] Content of the original data

[0146] Account set information: Refers to information related to the accounting period, such as the fiscal year, quarter, etc.

[0147] Subject information: Covers detailed information on various financial subjects, such as the specific subject names and their numbers under classifications such as assets, liabilities, owner's equity, income, expenses, etc.

[0148] Currency information: Records the currency types in which transactions occur, which is particularly important for multinational companies or multi-currency operations.

[0149] Transaction information: Includes specific transaction details, such as the transaction date, description, amount (income or expenditure), balance changes, etc.

[0150] According to an embodiment of the present application, based on the extraction rules, the original data of the account book files is subjected to field extraction to obtain key fields, specifically:

[0151] Parse the original data, and identify and extract the key fields related to the account book files from the parsed original data;

[0152] The parsing method includes: matching through regular expressions, locating fields, and identifying keywords.

[0153] Specifically, parsing of raw data

[0154] Parsing objective: The objective is to identify and extract all keyword fields related to the ledger file from the raw data of the ledger file. These keyword fields are the basis for generating metadata and are crucial for subsequent data processing and analysis.

[0155] Parsing method: Matching through regular expressions: Use regular expressions (regex) to locate data or keywords in a specific format. For example, regular expressions can be used to identify date formats (such as YYYY-MM-DD), currency amount formats, etc.

[0156] Field location and keyword identification: During the parsing process, not only specific format data needs to be identified, but also the position of each field needs to be accurately located, and those keywords crucial for understanding the ledger content need to be identified.

[0157] Detailed explanation of the field extraction process

[0158] Data parsing

[0159] First, the raw data in the ledger file needs to be converted into a processable form. This may involve reading text files, database records, or other forms of data sources. Use regular expressions to perform pattern matching on the raw data. For example, if you want to extract the transaction date, you can use a regular expression to find all strings that match the date format.

[0160] Identifying and extracting keyword fields

[0161] After successfully parsing the raw data, the next step is to identify and extract all keyword fields directly related to the ledger file from this data. Keyword fields usually include but are not limited to:

[0162] Transaction date: Used to record the time when each transaction occurs.

[0163] Summary information: Briefly describes the content or reason of the transaction.

[0164] Income / expense amount: The amount of funds involved, and it is necessary to clearly indicate whether it is income or expense.

[0165] Account balance: The remaining amount in the account after each transaction.

[0166] The application of regular expressions is particularly important here because it can help to accurately locate these keyword fields. For example, for "abstract information", relevant records can be identified by searching for specific keywords (such as "salary", "rent", etc.).

[0167] Verification and optimization

[0168] After completing the preliminary field extraction, it is also necessary to verify the extraction results to ensure their accuracy and integrity. This includes checking whether all required fields have been correctly filled and confirming whether the data format meets the expectations. If it is found that some fields are not correctly extracted (such as errors caused by format problems), the parsing logic or regular expressions need to be adjusted and the correct information needs to be extracted again.

[0169] According to an embodiment of the present application, after identifying and extracting the keyword fields related to the account book file from the parsed original data, it further includes the verification and optimization of the keyword fields, specifically:

[0170] Adjust the range of the extracted fields, and divide the keyword fields into required fields and optional fields;

[0171] Check whether there are any missing required fields;

[0172] Verify the formats of the required fields and the optional fields;

[0173] If the extraction of the required fields and / or the optional fields fails, adjust the parsing method.

[0174] Specifically, adjust the range of the extracted fields

[0175] Distinguish required fields and optional fields: First, it is necessary to classify the extracted keyword fields to clarify which are required fields (i.e., data that must be included) and which are optional fields (i.e., information that is not necessary but helps to more comprehensively understand the data). For example, in the bank deposit journal, "transaction date", "abstract information", and "account balance" may be required fields, while "remark" may be an optional field.

[0176] Check whether there are any missing required fields

[0177] Integrity check: Conduct a comprehensive check on all fields marked as required to confirm whether they have all been correctly filled. If it is found that any required field is missing, the record is considered incomplete and may require further manual review or an automatic supplementation mechanism to solve.

[0178] Processing Strategy: For missing required fields, various measures can be taken, such as filling in default values through preset rules, prompting the user to complete them manually, or triggering an alarm to notify relevant personnel for handling.

[0179] Verify the formats of required fields and optional fields

[0180] Format Consistency Check: Ensure that the content of each field conforms to the expected format requirements. For example, the "Transaction Date" should follow a unified date format (such as YYYY - MM - DD), and the "Amount" field should be of numeric type and contain no illegal characters, etc.

[0181] Data Validity Verification: In addition to format, the validity of the data also needs to be verified. For example, the amount cannot be negative (unless specifically stated as a refund or other special circumstances), and the date should be within a reasonable range, etc.

[0182] If the extraction of required fields and / or optional fields fails, adjust the parsing method

[0183] Problem Diagnosis and Correction: When encountering problems with incorrect extraction of required fields or optional fields, analyze the reasons and adjust the parsing strategy. This may involve modifying the matching pattern of regular expressions to adapt to different data formats, or improving the field positioning logic to enhance accuracy.

[0184] Dynamic Adjustment of Parsing Method: In practical applications, due to the diversity and complexity of data sources, the parsing method may need to be continuously adjusted. For example, using different parsing templates for data from different sources, or trying alternative parsing solutions after the initial parsing fails.

[0185] According to an embodiment of the present application, the construction of the metadata model is specifically as follows:

[0186] Preset a basic metadata model containing general attributes;

[0187] Based on user requirements, add custom attributes to the preset basic metadata model to construct the metadata model.

[0188] Specifically, the preset basic metadata model refers to a set of standard frameworks designed for account books and files, which contains general attributes applicable to most situations. These general attributes are set based on common financial record requirements and can meet basic data management and analysis needs.

[0189] Examples of General Attributes:

[0190] Transaction Date: Records the date when each transaction occurs.

[0191] Abstract Information (Description): Briefly describe the content or reason of the transaction.

[0192] Income Amount: If the transaction involves income, record its amount.

[0193] Expense Amount: If the transaction involves expenses, record its amount.

[0194] Account Balance: Update the current balance of the account after each transaction.

[0195] These general attributes form the basis of the metadata model, ensuring that different types of ledger files can have a common starting point for data structure, facilitating system processing and user understanding.

[0196] Add custom attributes to the preset basic metadata model according to user requirements to construct the metadata model.

[0197] Although the preset basic metadata model provides a widely applicable framework, in actual applications, different enterprises or organizations may have specific requirements, which requires adding custom attributes on this basis to expand the functions of the model.

[0198] Add custom attributes: Based on the specific business requirements of users, add additional fields or attributes to the basic metadata model. For example:

[0199] Project Number: For enterprises that manage finances by project, it is very important to record the project number corresponding to each transaction.

[0200] Contract Number: When it comes to contract management, associating each transaction with the corresponding contract number helps with tracking and auditing.

[0201] Approval Status: Used to mark whether a certain transaction has gone through the necessary approval process.

[0202] In this way, the metadata model can be flexibly adjusted to better meet the needs of actual application scenarios. This not only enhances the adaptability of the system but also improves the relevance and practicality of the data.

[0203] According to an embodiment of the present application, the mapping rule is specifically as follows:

[0204] Direct mapping: When the field name is the same as the model attribute name, directly associate them;

[0205] Transformation mapping, including: unit conversion and encoding standardization;

[0206] Composite mapping, merging and / or splitting fields.

[0207] Specifically, direct mapping: When the field name in the original data is exactly the same as the attribute name in the metadata model, a mapping relationship can be directly established.

[0208] Application scenario: This mapping method is applicable to cases where there are no differences in field names between the original data and the metadata model. For example, if there is a field named "transaction date" in the original data and there is also an attribute with the same name in the metadata model, then these two fields can be directly associated.

[0209] Advantages: Simple and direct, easy to implement, reducing complexity.

[0210] Transformation mapping: When there are differences in format or unit between the original data and the metadata model, certain conversion rules are needed to match the two.

[0211] Including: Unit conversion: For example, the amount in the original data may be in "cents", while the metadata model requires "yuan", in which case the value needs to be divided by 100 for conversion.

[0212] Encoding standardization: For some data with specific formats (such as dates), if the original data format is MM / DD / YYYY and the metadata model requires the YYYY-MM-DD format, then corresponding conversion logic is needed to ensure consistency.

[0213] Application scenario: When processing data from different systems or sources, since each system may adopt different standards or units, appropriate conversions are required to correctly map to the metadata model.

[0214] Advantages: Improves data consistency and comparability, enabling data from different sources to be analyzed and managed under a unified standard.

[0215] Composite mapping: Involves merging or splitting operations between multiple fields to better meet the requirements of the metadata model.

[0216] Including: Merging fields: Sometimes it is necessary to merge multiple original data fields into one attribute in the metadata model. For example, the name field may be split into two independent fields, surname and given name, and when mapping, they need to be recombined into a complete name field.

[0217] Split fields: Conversely, in some cases, it is also necessary to split an original data field into multiple attributes. For example, address information may contain multiple parts such as street, city, and zip code, which need to be split and mapped to the corresponding attributes respectively during mapping.

[0218] Application scenario: When the design of the metadata model does not exactly correspond to the original data structure, composite mapping can be used to adjust the data structure to meet the expected model design.

[0219] Advantages: It enhances flexibility, can handle complex mapping requirements, and makes the data structure more reasonable and convenient to use.

[0220] According to an embodiment of the present application, the difference processing is specifically as follows:

[0221] If a field is missing, fill in the default value according to the rules;

[0222] If the same field is mapped to multiple attributes, select according to the priority.

[0223] Specifically, if a field is missing, fill in the default value according to the rules

[0224] Purpose: Ensure that all necessary data fields exist, even if some fields may be missing in the original data.

[0225] Identify missing fields: First, it is necessary to check whether each key field exists in the original data. For those fields marked as required but actually non-existent (or empty), the system will identify them as "missing fields".

[0226] Define default value rules: For each field that may be missing, set a set of default value rules in advance. These rules can be customized according to the actual situation. For example:

[0227] For amount fields, the default value can be set to 0.

[0228] For date fields, if the specific date cannot be obtained, the day when the transaction occurs can be used as the default value.

[0229] For text description fields, the default value can be set to "not provided" or "unknown".

[0230] Apply default values: Once a missing field is detected, the system will automatically fill in the default value of the field according to the preset rules. This not only ensures the integrity of the data structure but also facilitates subsequent data processing.

[0231] If the same field is mapped to multiple attributes, select according to the priority

[0232] Objective: To solve the conflict problem that occurs when a single raw data field needs to be mapped to multiple attributes in the metadata model.

[0233] Identifying multi-attribute mapping conflicts: Sometimes, due to business requirements or data structure design reasons, a single raw data field may need to be mapped to multiple different attributes simultaneously. For example, a certain amount field may need to record both total income and total expenditure (even though logically they should be separate).

[0234] Setting priorities: To avoid data inconsistencies or duplications caused by such conflicts, priorities must be set for each possible target attribute. The priority determines which attribute should receive the data of this field first in case of a conflict.

[0235] For example, in the above example of income and expenditure, if a record is clearly marked as "income", the corresponding amount should be preferentially mapped to the "income amount" attribute; vice versa.

[0236] Implementing priority selection: During the mapping operation, the system will determine how to allocate data according to the set priority order. This means that only the attribute with the highest priority will be assigned the value of this field, and other attributes with lower priorities will not be affected unless there are additional instructions on how to handle the remaining data.

[0237] According to an embodiment of the present application, after generating the structured metadata file based on the metadata model and the mapping result, it further includes validating the metadata, specifically:

[0238] Checking whether the structured metadata file completely reflects the original data structure;

[0239] Automatically marking abnormal data and triggering an alarm.

[0240] Specifically, checking whether the structured metadata file completely reflects the original data structure

[0241] Objective: To ensure that the generated structured metadata file contains all necessary information and that this information is completely consistent with the original data.

[0242] Integrity check:

[0243] Field coverage: Checking whether the generated metadata file contains all the key fields extracted from the original data. For example, if the original data contains key fields such as transaction date, summary information, income amount, etc., these fields should all be reflected in the metadata file.

[0244] Required field check: Pay special attention to whether all fields defined as required are correctly filled. Missing any required field will cause the file to be regarded as incomplete.

[0245] Data format consistency: Confirm that the data format of each field conforms to the expected standard. For example, dates should follow a unified format (such as YYYY - MM - DD), and amounts should be of numeric type, etc.

[0246] Accuracy verification:

[0247] Data value verification: Ensure that each field value in the metadata file is exactly the same as the corresponding value in the original data. This includes the results after direct mapping, conversion mapping (such as unit conversion), and composite mapping.

[0248] Logical consistency: Check whether the logical relationships between data are reasonable. For example, changes in account balances should match the amounts of income and expenditure.

[0249] Automatically mark abnormal data and trigger alerts

[0250] Purpose: To promptly detect and handle any data that does not meet expectations, preventing errors from spreading to subsequent data processing flows.

[0251] Automatically mark abnormal data:

[0252] Abnormal detection mechanism: The system should have the ability to automatically identify abnormal data. For example, if a negative value appears in a certain amount field (while such a situation is not allowed in business logic), then this record will be marked as abnormal.

[0253] Format or logic errors: When the data format is incorrect (such as an incorrect date format) or logically unreasonable (such as the account balance not matching the actual transactions), it should also be marked as abnormal.

[0254] Trigger alerts:

[0255] Notification mechanism: Once abnormal data is detected, the system should immediately trigger the alert mechanism. This may include sending emails, text messages to notify relevant personnel, or displaying warning messages on the system interface.

[0256] Detailed report: In addition to simple alerts, a detailed exception report should also be provided, listing the locations of all abnormal data, specific problems, and recommended solutions. This helps to quickly locate problems and take appropriate corrective measures.

[0257] According to an embodiment of the present application, after the verification of the metadata, it further includes the generation of archives, specifically:

[0258] Automatically generate a standardized ledger archive based on the structured metadata file;

[0259] When the structured metadata file or the original data changes, a regeneration process is triggered.

[0260] According to an embodiment of the present application, after the verification of the metadata, it further includes maintenance management, specifically:

[0261] Record the version history of the structured metadata file to support backtracking and difference comparison;

[0262] Regularly back up the structured metadata file and the ledger files.

[0263] Specifically, the generation of files

[0264] Automatically generate a standardized ledger file based on the structured metadata file

[0265] Purpose: Use the verified structured metadata file to create a ledger file that conforms to the standard format, facilitating storage, query, and auditing.

[0266] Implementation method:

[0267] Standardized format: According to preset standards or specifications (such as financial statement formats, industry-specific requirements, etc.), convert the metadata into a structured ledger file. For example, PDF, Excel, or other formats that are easy to read and share can be generated.

[0268] Automated process: Automatically execute this process through programming scripts or dedicated software tools to reduce human intervention, improve efficiency, and reduce error rates.

[0269] When the structured metadata file or the original data changes, a regeneration process is triggered

[0270] Dynamic update mechanism: To maintain the timeliness and accuracy of the ledger file, the system needs to have a monitoring function. Once it detects changes in the structured metadata file or the original data, it automatically starts the regeneration process.

[0271] Change detection: The data can be identified whether it has been modified through timestamps, version numbers, etc.

[0272] Automated response: Once a change is detected, the system immediately triggers the regeneration process, updates the ledger file with the latest data, and ensures that the information accessed by all users is the latest and accurate.

[0273] Maintenance management

[0274] Record the version history of the structured metadata file to support backtracking and difference comparison

[0275] Version control: For each generated structured metadata file, the system records its version information, including the creation time, modifier, etc.

[0276] Backtracking function: Allows users to view metadata files of any historical version, which is very useful for audit trails.

[0277] Difference comparison: Provides tools to help users compare differences between different versions, quickly understand which parts have changed, and helps analyze the cause of problems or perform compliance checks.

[0278] Regularly back up the structured metadata file and the ledger file

[0279] Data protection measures: Regularly back up the structured metadata file and the finally generated ledger file to prevent data loss caused by hardware failures, software errors, or human errors.

[0280] Backup strategy: Develop a reasonable backup plan, such as backing up once a day, once a week, or once a month, and ensure that the backup data is stored in a safe place, preferably off-site storage to increase security.

[0281] Recovery test: Regularly test the recoverability of the backup data to ensure that normal operations can be quickly restored in case of an emergency.

[0282] What is not described in this application can be implemented by adopting or referring to existing technologies.

[0283] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between each embodiment, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments.

[0284] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A method for generating metadata for account book archives based on user-defined rules, characterized in that: include: Based on the extraction rules, the original data of the account book archive is extracted to obtain the key fields; Constructing a metadata model, associating the key fields with corresponding attributes of the metadata model through mapping rules, and processing field differences to obtain a mapping result; A structured metadata file is generated based on the metadata model and the mapping result.

2. The method according to claim 1, characterized in that The acquisition of the original data of the account book archives is specifically: distinguishing the data by account book category; Performing integrity check and confidentiality check on the differentiated data; The data after passing the integrity check and the confidentiality check is used as the original data of the account book file; The original data includes: account set information, subject information, currency information and transaction information.

3. The method according to claim 1, characterized in that Based on the extraction rules, the original data of the account book archive is subjected to field extraction to obtain key fields, specifically: Parsing the original data, identifying and extracting the key fields related to the account book archive from the parsed original data; The parsing method includes: matching through regular expressions, locating fields and identifying keywords.

4. The method according to claim 1, characterized in that After identifying and extracting the key fields related to the account book archive from the parsed original data, the key fields are verified and optimized, specifically: Adjust the range of the extracted fields and divide the key fields into required fields and optional fields; Check whether the required fields are missing; Verifying the formats of the required fields and the optional fields; If the required fields and / or the optional fields fail to be extracted, the parsing method is adjusted.

5. The method according to claim 4, characterized in that The metadata model is constructed as follows: Preset basic metadata model, including common attributes; The preset basic metadata model is added with custom attributes according to user needs to construct the metadata model.

6. The method according to claim 4, characterized in that The mapping rules are specifically: Direct mapping, direct association when the field name is consistent with the model attribute name; Conversion mapping, including: unit conversion and coding standardization; Composite mappings, merging and / or splitting fields.

7. The method according to claim 1, characterized in that The difference processing is specifically as follows: If the field is missing, fill in the default value according to the rules; If the same field is mapped to multiple attributes, the one that is selected is based on priority.

8. The method according to claim 1, characterized in that After generating the structured metadata file based on the metadata model and the mapping result, the metadata verification is further included, specifically: Verifying whether the structured metadata file completely reflects the original data structure; Automatically mark abnormal data and trigger alerts.

9. The method according to claim 1, characterized in that: After the metadata is verified, the archive is also generated, specifically: Automatically generate a standardized account book archive based on the structured metadata file; When the structured metadata file or the original data is changed, a regeneration process is triggered.

10. The method according to claim 1, characterized in that After the verification of the metadata, maintenance management is also included, specifically: Record the version history of the structured metadata file and support backtracking and difference comparison; The structured metadata files and the account book archives are backed up regularly.