Intelligent accounting data management and compliance system and method
Through intelligent accounting data management and compliance systems, combined with blockchain technology, the problems of low automation and insufficient data accuracy of existing accounting information systems are solved, efficient and accurate accounting data processing and intangible asset value assessment are achieved, and the scientificity and rationality of enterprise asset management are improved.
Patent Information
- Application Number
- CN202510003994.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
The existing accounting information system has low degree of automation and insufficient data accuracy and timeliness, and it is impossible to comprehensively evaluate the value of intangible assets of enterprises, and it is difficult to meet the needs of modern enterprises for efficient and accurate data processing.
The intelligent accounting data management and compliance system is adopted to ensure the immutability and transparency of data through automated data collection, data processing and cleaning, data storage and management, data analysis and machine learning, data visualization and reporting, etc. through block chain technology.
It improves the automation and efficiency of accounting data processing, enhances the accuracy and timeliness of data, comprehensively evaluates the value of intangible assets, improves the scientificity and rationality of enterprise asset management, and provides reliable audit support and real-time decision-making support.
Smart Images

Figure CN119940715A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of accounting information technology, and specifically, to an intelligent accounting information enhancement system that combines automated data collection, big data analysis, artificial intelligence algorithms and blockchain technology, and in particular to an intelligent accounting data management and compliance system and method. Background Art
[0002] With the expansion of enterprise scale and the complexity of business, traditional accounting information systems can no longer meet the needs of modern enterprises for data processing. In existing technical solutions, problems such as low automation, insufficient data accuracy and timeliness, and incomplete evaluation of intangible assets are becoming increasingly prominent. Existing accounting information systems mainly rely on manual data input and processing, which is prone to human errors and inefficient. In addition, existing systems have obvious deficiencies in processing large-scale data and real-time data analysis, and cannot meet the needs of enterprises for efficient and accurate data processing.
[0003] In the process of development, in addition to tangible assets, many companies will also be accompanied by the increase of intangible assets. Therefore, when auditing a company, it is necessary to evaluate and audit the company's intangible assets.
[0004] The existing invention patent with publication number CN115375417A discloses a comprehensive financial audit system based on big data, including: a data acquisition module, a data analysis module, a learning prediction module and a database. The data acquisition module collects structured and semi-structured enterprise intangible asset evaluation data and data information of related industries based on big data technology, and collects enterprise intangible asset audit data at the same time. The data analysis module analyzes and compares the data collected by the data acquisition module with the historical data in the database to comprehensively analyze the audit data of the enterprise's intangible assets.
[0005] As enterprises expand in size and their businesses become more complex, the above traditional accounting information systems can no longer meet the needs of modern enterprises. Problems include low automation, insufficient data accuracy and timeliness, and incomplete evaluation of intangible assets. Summary of the invention
[0006] In view of the deficiencies in the prior art, the present invention provides an intelligent accounting data management and compliance system and method.
[0007] According to an intelligent accounting data management and compliance system and method provided by the present invention, the scheme is as follows:
[0008] In a first aspect, an intelligent accounting data management and compliance system is provided, the system comprising:
[0009] Automated data collection module: collects data and uploads it to the blockchain, encrypting the data during transmission and storage;
[0010] Data processing and cleaning module: standardize and clean the collected data, and perform data correction and anomaly detection;
[0011] Data storage and management module: adopts multi-database storage to classify and store the cleaned data;
[0012] Data Analysis and Machine Learning Module: Use machine learning algorithms for financial forecasting and anomaly detection, and perform data modeling;
[0013] Data visualization and reporting module: Generate reports and dynamic charts to display data changes and prediction results.
[0014] Preferably, the automated data collection module includes various types of sensors and API interfaces for collecting financial data, operational data, market data and accounting policy data;
[0015] Wherein, the data sources of the financial data and operational data are automatically collected through the data interface of the enterprise resource planning ERP system;
[0016] The data source of the market data is collected through relevant APP interfaces including Enterprise Early Warning to collect market trends and competition analysis data;
[0017] The data source of the accounting policy data collects the latest accounting policy update and change information through the data interface of the target department.
[0018] Preferably, the automated data collection module further includes:
[0019] Data encryption: All data are encrypted using the AES algorithm during transmission and storage to ensure data confidentiality;
[0020] Hash generation: Each encrypted data is generated with a unique hash value through the SHA-256 hash algorithm. The hash value of each piece of data will be used as a block in the blockchain to ensure the integrity and immutability of the data.
[0021] Smart contract records: Each operation on each data can trigger a smart contract to record the detailed information of the operation. The smart contract writes the operation record into the blockchain and generates a new block to ensure the transparency and traceability of the operation.
[0022] Preferably, in the data processing and cleaning module, processing financial data, operational data, market data and accounting policy data includes:
[0023] Data correction:
[0024] 1) Data collection:
[0025] Acquiring raw data from the automated data acquisition module, including financial data, operational data, market data, and accounting policy data;
[0026] 2) Data preprocessing:
[0027] Data deduplication: Use algorithms to remove duplicate data to ensure data uniqueness;
[0028] Missing value processing: fill in missing data through interpolation or machine learning models;
[0029] 3) Data correction:
[0030] Linear regression: used to correct systematic errors in data. Linear regression models are used to correct biases in sensor data.
[0031] 4) Data standardization: Convert the bias-corrected data into a standard normal distribution to ensure that data from different data sources are comparable;
[0032] Anomaly Detection:
[0033] 1) Data preprocessing:
[0034] Data denoising: using filters or smoothing algorithms to remove noise from data;
[0035] 2) Anomaly Detection:
[0036] K-means clustering: used to detect outliers in data. Through cluster analysis, data points that are far away from the cluster center are identified as outliers.
[0037] Preferably, the system also includes real-time analysis of accounting policy changes of the target department, specifically including:
[0038] 1) Obtain the latest accounting policy data in real time through the API interface;
[0039] 2) Use NLP models to parse policy texts and extract entity information in policy documents through named entity recognition;
[0040] 3) Generate data processing rules based on the extracted entity information;
[0041] 4) Use the rule engine to execute the generated data rules and perform data-related processing and cleaning;
[0042] Among them, data processing rules include: data cleaning rules, data conversion rules, outlier processing rules, integrity rules, and security and compliance rules;
[0043] 5) storing the processed and cleaned data in the data storage and management module;
[0044] By analyzing policy changes in real time, continuously updating data processing rules and cleaning data again, integrating old and new data, accounting data processing is kept consistent with the latest policies, and data compliance and accuracy are ensured in a timely manner.
[0045] Preferably, the data storage and management module adopts a multi-database storage form, supporting structured and unstructured data storage, including: relational database, NoSQL database and data lake;
[0046] Wherein, the relational database stores relevant structured data including financial data and operational data;
[0047] The NoSQL database stores relevant unstructured data including market data and accounting policy data;
[0048] The data lake stores large-scale unstructured data and supports flexible query and analysis;
[0049] Data synchronization between different databases is achieved through data synchronization tools, and data integration between relational databases and NoSQL databases is achieved through data integration tools;
[0050] After the automated data acquisition module collects data, it stores the data in corresponding databases according to the data type; the data processing and cleaning module extracts data from different databases for preprocessing and cleaning, and puts it back into the corresponding database after cleaning; the data analysis and machine learning module extracts data from different databases for analysis and modeling, and stores the analysis results and models in the corresponding database for subsequent use; the data visualization and reporting module: extracts data from different databases for visualization and report generation, and stores the visualization results and reports in the corresponding database for viewing;
[0051] Data mobilization between multiple databases enables data synchronization, data integration and data access, and performs real-time data processing, batch data processing and comprehensive data analysis.
[0052] Preferably, the data analysis and machine learning module uses regression models, time series models and classification models for financial forecasting and anomaly detection, including linear regression, ARIMA, LSTM and support vector machine algorithms;
[0053] Among them, financial forecasting: predicting changes in financial indicators through regression models and time series models, and providing future financial trend forecasts;
[0054] Anomaly detection: Detect anomalies in data through classification models and clustering algorithms.
[0055] Preferably, the data visualization and reporting module uses Tableau and Power BI tools to create interactive dashboards and reports, and uses Python's Matplotlib and Seaborn libraries to create dynamic charts to display data changes and prediction results;
[0056] Among them, Tableau: connect to data sources, create charts, add interactive functions, publish and share dashboards;
[0057] Power BI: connect to data sources, create charts, add interactive features, publish and share reports;
[0058] Matplotlib: Create static and dynamic charts with fine-grained control over chart appearance and behavior;
[0059] Seaborn: Quickly create beautiful statistical charts suitable for data analysis and exploratory data analysis (EDA).
[0060] In a second aspect, a smart accounting data management and compliance method is provided, the method comprising:
[0061] Automated data collection steps: Collect data and upload it to the blockchain, encrypting the data during data transmission and storage;
[0062] Data processing and cleaning steps: standardize and clean the collected data;
[0063] Data storage and management steps: Use multi-database storage to classify and store the cleaned data;
[0064] Data Analysis and Machine Learning Steps: Use machine learning algorithms for financial forecasting and anomaly detection, and perform data modeling;
[0065] Data visualization and reporting steps: Generate reports and dynamic charts to display data changes and prediction results.
[0066] Preferably, the automated data collection step includes various types of sensors and API interfaces for collecting financial data, operational data, market data and accounting policy data;
[0067] Wherein, the data sources of the financial data and operational data are automatically collected through the data interface of the enterprise resource planning ERP system;
[0068] The data source of the market data is collected through relevant APP interfaces including Enterprise Early Warning to collect market trends and competition analysis data;
[0069] The data source of the accounting policy data is to collect the latest accounting policy update and change information through the data interface of the target department;
[0070] The automated data collection step also includes:
[0071] Data encryption: All data are encrypted using the AES algorithm during transmission and storage to ensure data confidentiality;
[0072] Hash generation: Each encrypted data is generated with a unique hash value through the SHA-256 hash algorithm. The hash value of each piece of data will be used as a block in the blockchain to ensure the integrity and immutability of the data.
[0073] Smart contract records: Each operation on each data can trigger a smart contract to record the detailed information of the operation. The smart contract writes the operation record into the blockchain and generates a new block to ensure the transparency and traceability of the operation.
[0074] In the data processing and cleaning step, the processing of financial data, operational data, market data and accounting policy data includes:
[0075] Data correction:
[0076] 1) Data collection:
[0077] Acquiring raw data from the automated data collection step, including financial data, operational data, market data, and accounting policy data;
[0078] 2) Data preprocessing:
[0079] Data deduplication: Use algorithms to remove duplicate data to ensure data uniqueness;
[0080] Missing value processing: fill in missing data through interpolation or machine learning models;
[0081] 3) Data correction:
[0082] Linear regression: used to correct systematic errors in data. Linear regression models are used to correct biases in sensor data.
[0083] 4) Data standardization: Convert the bias-corrected data into a standard normal distribution to ensure that data from different data sources are comparable;
[0084] Anomaly Detection:
[0085] 1) Data preprocessing:
[0086] Data denoising: using filters or smoothing algorithms to remove noise from data;
[0087] 2) Anomaly Detection:
[0088] K-means clustering: used to detect outliers in data. Through cluster analysis, data points that are far away from the cluster center are identified as outliers.
[0089] The system also includes real-time analysis of accounting policy changes for target departments, including:
[0090] 1) Obtain the latest accounting policy data in real time through the API interface;
[0091] 2) Use NLP models to parse policy texts and extract entity information in policy documents through named entity recognition;
[0092] 3) Generate data processing rules based on the extracted entity information;
[0093] 4) Use the rule engine to execute the generated data rules and perform data-related processing and cleaning;
[0094] Among them, data processing rules include: data cleaning rules, data conversion rules, outlier processing rules, integrity rules, and security and compliance rules;
[0095] 5) storing the processed and cleaned data in the data storage and management step;
[0096] By analyzing policy changes in real time, constantly updating data processing rules and re-cleaning data, integrating new and old data, accounting data processing is kept consistent with the latest policies, and data compliance and accuracy are ensured in a timely manner;
[0097] The data storage and management steps adopt a multi-database storage form, supporting structured and unstructured data storage, including: relational databases, NoSQL databases and data lakes;
[0098] Wherein, the relational database stores relevant structured data including financial data and operational data;
[0099] The NoSQL database stores relevant unstructured data including market data and accounting policy data;
[0100] The data lake stores large-scale unstructured data and supports flexible query and analysis;
[0101] Data synchronization between different databases is achieved through data synchronization tools, and data integration between relational databases and NoSQL databases is achieved through data integration tools;
[0102] After the automated data collection step collects data, the data is stored in corresponding databases according to the data type; the data processing and cleaning step extracts data from different databases for preprocessing and cleaning, and puts the data back into the corresponding database after cleaning; the data analysis and machine learning step extracts data from different databases for analysis and modeling, and the analysis results and models are stored in corresponding databases for subsequent use; the data visualization and reporting step extracts data from different databases for visualization and report generation, and the visualization results and reports are stored in corresponding databases for viewing;
[0103] Data mobilization between multiple databases enables data synchronization, data integration and data access, and performs real-time data processing, batch data processing and comprehensive data analysis.
[0104] The data analysis and machine learning steps use regression models, time series models, and classification models for financial forecasting and anomaly detection, including linear regression, ARIMA, LSTM, and support vector machine algorithms;
[0105] Among them, financial forecasting: predicting changes in financial indicators through regression models and time series models, and providing future financial trend forecasts;
[0106] Anomaly detection: Detect anomalies in data through classification models and clustering algorithms;
[0107] The data visualization and reporting step uses Tableau and Power BI tools to create interactive dashboards and reports, and uses Python's Matplotlib and Seaborn libraries to create dynamic charts to display data changes and prediction results;
[0108] Among them, Tableau: connect to data sources, create charts, add interactive functions, publish and share dashboards;
[0109] Power BI: connect to data sources, create charts, add interactive features, publish and share reports;
[0110] Matplotlib: Create static and dynamic charts with fine-grained control over chart appearance and behavior;
[0111] Seaborn: Quickly create beautiful statistical charts suitable for data analysis and exploratory data analysis (EDA).
[0112] Compared with the prior art, the present invention has the following beneficial effects:
[0113] 1. Through the application of blockchain technology, the present invention enables the system to ensure the immutability and transparency of financial data, operational data, market data and accounting policy data, provide reliable audit support, and enhance the security and credibility of corporate financial management;
[0114] 2. The present invention ensures the accuracy and integrity of data by effectively performing data correction and anomaly detection, providing a reliable data basis for subsequent data analysis and prediction;
[0115] 3. Through the application of multiple database storage forms, the intelligent accounting data management system of the present invention can efficiently manage and process different types of data, provide flexible storage and analysis capabilities, and meet the needs of enterprises for efficient and accurate data processing;
[0116] 4. The present invention can create interactive dashboards and reports to intuitively display data changes and forecast results, helping enterprise management to quickly understand and analyze data;
[0117] 5. The present invention can efficiently perform financial forecasting and anomaly detection through the application of machine learning models, and provide accurate financial data and trend forecasts;
[0118] 6. The present invention significantly improves the automation and efficiency of accounting data processing, enhances the accuracy and timeliness of data, comprehensively evaluates the value of intangible assets, improves the scientificity and rationality of enterprise asset management, and enhances the comprehensibility and ease of use of information through data visualization.
[0119] Other beneficial effects of the present invention will be explained in the specific implementation manner through the introduction of specific technical features and technical solutions. Through the introduction of these technical features and technical solutions, those skilled in the art should be able to understand the beneficial technical effects brought about by the technical features and technical solutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0120] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0121] Figure 1 This is the overall architecture diagram of the system;
[0122] Figure 2 Workflow diagram for automated data acquisition and processing. DETAILED DESCRIPTION
[0123] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0124] Embodiment 1:
[0125] The embodiment of the present invention provides an intelligent accounting data management and compliance system and method, which improves the accuracy and timeliness of accounting information by integrating advanced technologies such as big data, artificial intelligence (AI), and blockchain, and fully reflects the true value of an enterprise's intangible assets and long-term assets. The system automatically collects, processes, stores, analyzes, and visualizes accounting data from multiple data sources to improve the efficiency of enterprise operations and the quality of decision-making. Figure 1 As shown, the system specifically includes: an automated data acquisition module, a data processing and cleaning module, a data analysis and machine learning module, and a data visualization and reporting module.
[0126] Among them, the automated data collection module collects data and uploads it to the blockchain, encrypting the data during transmission and storage.
[0127] This module is responsible for automatically collecting data from various data sources (such as sensors, databases, APIs, etc.) and encrypting the data to ensure the security and integrity of the data. Data encryption uses the Advanced Encryption Standard (AES) to ensure the security of data during transmission and storage.
[0128] Specifically, the automated data collection module includes various types of sensors and API interfaces for collecting financial data, operational data, market data and accounting policy data.
[0129] The data sources of financial data and operational data are automatically collected through the data interface of the enterprise resource planning ERP system; the data source of market data is collected through relevant APP interfaces including the enterprise early warning system to collect market trend and competition analysis data; the data source of accounting policy data is collected through the department's data interface to collect the latest accounting policy updates and changes.
[0130] The data interface for market data and accounting policy data uses RESTful API for data transmission; the data interface for financial data and operational data uses RESTful API and SOAP protocol for data transmission; all data supports JSON and XML formats to ensure data compatibility and readability.
[0131] The data security of financial data and operational data uses the AES-256 encryption standard to encrypt data during transmission to ensure data security and privacy protection. Access control uses the OAuth 2.0 protocol to ensure that only authorized users can access the data. The data security of market data uses the TLS protocol to ensure the security of data during transmission. The data security of accounting policy data uses the RSA encryption algorithm to encrypt data to ensure data security and privacy protection.
[0132] Data processing and cleaning module: standardize and clean the collected data. Specifically, data standardization: convert the deviation-corrected data into a standard normal distribution to ensure that data from different data sources are comparable. Data cleaning: includes missing value filling, outlier detection and processing, and deduplication processing.
[0133] The data processing and cleaning module also includes the use of machine learning algorithms for data correction and anomaly detection, including linear regression, ridge regression, K-means clustering, and Isolation Forest algorithms.
[0134] Data standardization is to ensure the consistency and comparability of data from different sources, using data type conversion, unit unification, and format standardization; data cleaning uses the tool Talend to remove duplicate data, fill in missing values, and correct data errors.
[0135] The system also includes real-time analysis of accounting policy changes for target departments, including:
[0136] 1) Obtain the latest accounting policy data in real time through the API interface;
[0137] 2) Use NLP models to parse policy texts and extract entity information in policy documents through named entity recognition;
[0138] 3) Generate data processing rules based on the extracted entity information;
[0139] 4) Use the rule engine to execute the generated data rules and perform data-related processing and cleaning;
[0140] Among them, data processing rules include: data cleaning rules, data conversion rules, outlier processing rules, integrity rules, and security and compliance rules;
[0141] 5) storing the processed and cleaned data in the data storage and management module;
[0142] By analyzing policy changes in real time, continuously updating data processing rules and cleaning data again, integrating old and new data, accounting data processing is kept consistent with the latest policies, and data compliance and accuracy are ensured in a timely manner.
[0143] Data storage and management module: Data is stored in the processed and cleaned data classification format using multiple database storage methods, including relational databases (such as MySQL), NoSQL databases (such as MongoDB), and data lakes (such as Apache Hadoop) to ensure data flexibility and scalability.
[0144] Among them, relational databases store relevant structured data including financial data and operational data; NoSQL databases store relevant unstructured data including market data and accounting policy data; data lakes store large-scale unstructured data and support flexible query and analysis. The coordination relationship between multiple databases and other modules: Data synchronization and integration between different databases are achieved through data synchronization and integration tools to ensure data consistency and availability.
[0145] Data Analysis and Machine Learning Module: Use machine learning algorithms for financial forecasting and anomaly detection, and perform data modeling.
[0146] In the data analysis and machine learning module, machine learning algorithms include regression models (such as linear regression and ridge regression), time series models (such as ARIMA and LSTM), and classification models (such as logistic regression and support vector machine).
[0147] Among them, financial forecasting: predicting changes in financial indicators through regression models and time series models;
[0148] Anomaly detection: Detect anomalies in data through classification models and clustering algorithms.
[0149] Specific details of this Data Analysis and Machine Learning module:
[0150] 1. Financial forecast:
[0151] Regression Models: Use linear regression and ridge regression models to predict changes in financial indicators.
[0152] Time Series Models: Forecasting time series data using ARIMA and LSTM models.
[0153] 2. Anomaly Detection:
[0154] Classification model: Use support vector machine (SVM) for anomaly detection.
[0155] Data visualization and reporting module: Generate reports and dynamic charts to display data changes and prediction results.
[0156] The data visualization and reporting module uses Tableau and Power BI tools to create interactive dashboards and reports, and uses Python's Matplotlib and Seaborn libraries to create dynamic charts to display data changes and prediction results;
[0157] Among them, Tableau: connect to data sources, create charts, add interactive functions, publish and share dashboards;
[0158] Power BI: connect to data sources, create charts, add interactive features, publish and share reports;
[0159] Matplotlib: Create static and dynamic charts with fine-grained control over chart appearance and behavior;
[0160] Seaborn: Quickly create beautiful statistical charts suitable for data analysis and exploratory data analysis (EDA).
[0161] The present invention also provides an intelligent accounting data management method that integrates big data and blockchain technology, using the ERP system and market monitoring tools to automatically collect data, big data technology to analyze financial and non-financial data, and predictive analysis and machine learning models to identify trends and anomalies. Figure 2 As shown, the method includes: automated data collection steps, data processing and cleaning steps, data analysis and machine learning steps, and data visualization and reporting steps, which can be implemented by executing the module of the intelligent accounting data management system of the integrated big data and blockchain technology. That is, those skilled in the art can understand the intelligent accounting data management system of the integrated big data and blockchain technology as a preferred implementation of the intelligent accounting data management method of the integrated big data and blockchain technology.
[0162] Embodiment 2:
[0163] This embodiment is described in more detail based on the previous embodiment.
[0164] This embodiment provides an intelligent accounting data management and compliance system, which is specifically as follows:
[0165] Automatic data acquisition module:
[0166] 1) Financial data and operating data:
[0167] Data source: Financial and operational data is automatically collected through the data interface of the enterprise resource planning (ERP) system.
[0168] Data interface: Use RESTful API and SOAP protocol for data transmission. In this embodiment, RESTful API stands for Representational State Transfer (REST). REST is an architectural style, not a protocol. It defines a set of principles for API design, aiming to transmit data through standard HTTP methods (such as GET, POST, PUT, DELETE).
[0169] In this embodiment, the full name of the SOAP protocol is Simple Object Access Protocol (SOAP), which stands for Simple Object Access Protocol. SOAP is a standardized protocol used for communication between different software systems.
[0170] Data format: Supports JSON and XML formats to ensure data compatibility and readability.
[0171] Data processing: During the data collection process, the data cleaning tool Talend is used to remove duplicate data, fill in missing values, and correct data errors to ensure data accuracy and consistency.
[0172] Data security: AES-256 encryption standard is used to encrypt data transmission to ensure data security and privacy protection. Access control adopts OAuth 2.0 protocol to ensure that only authorized users can access data.
[0173] 2) Market data:
[0174] Data source: Market trend and competition analysis data are collected through the enterprise early warning interface.
[0175] Data interface: Use RESTful API for data transmission.
[0176] Data format: Supports JSON and XML formats to ensure data compatibility and readability.
[0177] Data processing: During the data collection process, the data cleaning tool Talend is used to remove duplicate data, fill in missing values, and correct data errors to ensure data accuracy and consistency.
[0178] Data security: Use the TLS protocol to ensure the security of data during transmission.
[0179] 3) Accounting policy data:
[0180] Data source: Collect the latest accounting policy updates and changes through the target department’s data interface.
[0181] Obtain the latest accounting policy updates and changes in real time through the API interfaces provided by the target departments (such as the Ministry of Finance, the Tax Bureau, etc.). These API interfaces usually provide structured data formats (such as JSON, XML) to facilitate system automation processing.
[0182] Specific steps for parsing policies:
[0183] Data acquisition: Obtain the latest accounting policy data regularly or in real time through the API interface.
[0184] Data parsing: Use natural language processing (NLP) technology to parse policy documents and extract key information. This includes loading pre-trained NLP models, parsing policy texts (official documents issued by target departments or relevant regulatory agencies, which contain regulations, guidelines or legal requirements in specific areas (such as finance, taxation, accounting, etc.). These policy texts are official guidance documents that companies and organizations must follow in their daily operations and decision-making), and extracting key information (the so-called key information refers to important data or instructions in policy texts that have a direct impact on corporate operations, compliance or strategic decisions. For example, a new tax policy may specify in detail new measures that companies need to take when filing taxes, or change the method of calculating taxes. These specific operating guidelines or changes constitute key information. By parsing these policy texts, NLP technology can automatically identify and extract this key information, helping companies to adjust their strategies in a timely manner to meet the latest regulatory requirements).
[0185] Among them, the application of NLP models in policy analysis
[0186] 1. Application of NLP Model:
[0187] 1.1: Model selection: Use a pre-trained NLP model (BERT) for text parsing.
[0188] Using the BERT (Bidirectional Encoder Representations from Transformers) model for text parsing is an advanced natural language processing technology. The core advantage of the BERT model lies in its bidirectional training mechanism, which can more comprehensively understand the context of the language. The following are the basic steps and methods for using BERT to parse policy text:
[0189] Loading of pre-trained model:
[0190] Before performing specific policy text analysis, you first need to load a pre-trained BERT model. This model is usually pre-trained on large-scale text data and has rich language understanding capabilities.
[0191] Text preprocessing:
[0192] The policy text first needs to go through preprocessing steps, including word segmentation, stop word removal, and normalization, to adapt to the input requirements of the BERT model.
[0193] Feature extraction:
[0194] The BERT model, through its multi-layer Transformer network structure, can capture the complex relationship between words and extract high-quality features from the text. These features represent the deep semantic information of each word in the text.
[0195] Contextual understanding:
[0196] Due to BERT’s bidirectional training characteristics, it is able to take into account the context of words, thereby more accurately understanding the global semantics of the text. This is especially important for understanding the complex sentences and terminology that often appear in policy documents.
[0197] Key information extraction:
[0198] Based on feature extraction and context understanding, specific strategies (such as classification, entity recognition or keyword extraction techniques) can be used to extract key information from policy texts. This may involve identifying key elements such as regulatory clauses, obligations, rights, etc. in the text.
[0199] Output parsing results:
[0200] The results of the analysis can be directly specific clauses or summaries in the text, or they can be structured representations of key information, such as a list of key clauses, a clear division of responsibilities, etc., to facilitate further analysis or decision support.
[0201] In this way, BERT can effectively parse and extract key information that has a decisive impact on the enterprise from complex policy documents, supporting compliance monitoring and decision-making processes.
[0202] 1.2: Information Extraction: Extract key entities (such as date, amount, policy name, etc.) in policy documents through named entity recognition (NER).
[0203] 1.3: Rule generation: Generate corresponding rules based on the extracted information for subsequent data processing and cleaning.
[0204] In the intelligent accounting information enhancement system, the description of "generating corresponding rules" usually refers to the creation of specific criteria or algorithms for data processing and cleaning. These rules are key components of the automated system to process data to ensure its accuracy, consistency and availability. These rules can specifically include:
[0205] a. Data cleaning rules:
[0206] Deduplication: Define rules to identify and remove duplicate records, ensuring data uniqueness in the database.
[0207] Missing value handling: Setting rules to fill in or ignore missing data, which may include using the mean, median, mode, or other statistical methods to estimate missing values.
[0208] Data format normalization: Ensure that all data follows the same format and standards, such as dates in a unified format of YYYYMMDD and phone numbers in a unified format.
[0209] b. Data conversion rules:
[0210] Data type conversion: Automatically convert numbers in text format to numeric data, or convert dates in string format to date type for subsequent processing.
[0211] Unit unification: Unify different units in the data (such as kilometers and miles, or monetary units) into a standard unit.
[0212] c. Outlier processing rules:
[0213] Anomaly detection: Define rules to identify anomalies or outliers in the data, for example, using statistical thresholds (such as the standard deviation method) or boundaries based on business logic.
[0214] Anomaly handling: Processing detected outliers, which may include issuing warnings, correcting data, or excluding outlier data.
[0215] d. Integrity rules:
[0216] Dependency checking: Make sure that the data of certain fields depends on other fields, for example, the order date should not be earlier than the customer's registration date.
[0217] Data validity: Checks whether the data is within a reasonable range, for example, the age field should be between 0 and 150.
[0218] e. Security and Compliance Rules:
[0219] Data encryption: Encrypt sensitive information such as personal identity information, financial data, etc.
[0220] Access control: Set data access rules based on user roles and permissions to ensure data security.
[0221] These rules not only help improve data quality and processing efficiency, but also support ensuring data reliability and compliance with business and regulatory requirements. In the intelligent accounting information enhancement system, these rules may be implemented through programming or integrated into the system using specialized data processing software (such as Trifacta, Talend), so that the data processing process is automated and meets the preset business logic and compliance requirements.
[0222] 2. Specific steps to form rules:
[0223] Named Entity Recognition (NER): Identify key entities in policy texts. This includes loading pre-trained NLP models, parsing policy texts, and extracting key information.
[0224] Rule generation: Generate rules based on the extracted entities, and then print the generated rules (i.e., rules for data processing and cleaning).
[0225] Figure 2 The connection between the policy change analysis step and the data cleaning step includes:
[0226] 1. Application of rule engine:
[0227] Rule Engine: Use a rule engine (such as Drools, Nected) to manage and execute generated rules. The rule engine can automatically process and cleanse data according to predefined rules.
[0228] Rule execution: The rule engine automatically performs corresponding data processing and cleansing operations based on the rules generated by the parsed policy information.
[0229] 2. Connect to the specific implementation of the data cleaning step:
[0230] Data cleaning: In the data processing and cleaning module, the rule engine is used to execute the generated rules to perform data standardization and cleaning. This includes initializing the rule engine, loading the generated rules, and executing the rule engine to perform data cleaning.
[0231] Among them, in the intelligent accounting information enhancement system, the steps of analyzing policy changes and cleaning data again are of great significance, mainly because accounting data processing needs to be consistent with the latest regulations and policies to ensure data compliance and accuracy. The following explains this process and its purpose in detail:
[0232] Analyzing the purpose of policy changes
[0233] A. Ensure compliance:
[0234] Policy changes often involve key factors such as tax laws, financial reporting standards, and compliance requirements, which directly affect the way accounting data is processed and reported.
[0235] By parsing the latest policies, the system can adjust data processing rules and logic in a timely manner to ensure that the company complies with the latest regulations in accounting and financial reporting.
[0236] B. Update data processing rules:
[0237] Policy updates may cause existing data processing rules to become obsolete or no longer applicable, so these rules need to be updated to adapt to new policy requirements.
[0238] For example, if the tax rate changes, the relevant data processing rules also need to be adjusted accordingly to ensure the accuracy of the calculation.
[0239] C. Dynamically respond to market and legal environment:
[0240] Accounting data processing is not a static process; it needs to be able to adapt to the ever-changing external environment.
[0241] Analyzing policy changes enables the system to flexibly respond to environmental changes and improve the adaptability and competitiveness of enterprises.
[0242] Purpose and methods of data cleaning after analyzing policy changes:
[0243] Data cleaning is particularly important after analyzing policy changes because:
[0244] A. Integrate new and old data:
[0245] Policy changes may require new data fields or changes in data formats, and the original data needs to be adjusted and cleaned according to the new requirements.
[0246] The cleaning process includes correcting data errors, filling missing values, formatting data, etc. to adapt to new reporting requirements.
[0247] B. Cleaning specific content:
[0248] Remove redundancy: Delete data fields that are no longer needed or discard outdated information based on new policies.
[0249] Correction of errors: Correction of erroneous data entry based on updated logic of new policies.
[0250] Standardize data: Ensure that all data formats and structures meet the standards set by the new policy.
[0251] C. Tools and techniques used:
[0252] Rule Engine: Automates data cleaning tasks using predefined rules or dynamically generated rules. The rule engine can quickly adapt to policy changes and automatically apply new data processing logic.
[0253] Data cleaning software (e.g. Trifacta, Talend): These tools provide advanced data cleaning capabilities such as data fusion, quality checking, transformation, and integration to support complex data cleaning needs.
[0254] Relevance:
[0255] The acquired data, such as policy data and market data, needs to be consistent with the latest accounting policies, so analyzing policy changes is to update and verify the current validity and compliance of this data.
[0256] The data cleansing process ensures the accuracy and applicability of this data for financial reporting and decision analysis, especially when there are major policy changes.
[0257] In short, analyzing policy changes and performing data cleansing are crucial steps to ensure that the intelligent accounting information enhancement system provides accurate, timely and compliant financial information. This not only helps companies respond to policy adjustments in a timely manner, but also ensures that the company's data processing process is continuously optimized and meets the latest regulatory requirements.
[0258] 3. Flowchart diagram:
[0259] Step 1: Get the latest accounting policy data through the API interface.
[0260] Step 2: Use the NLP model to parse the policy text and extract key information.
[0261] Step 3: Generate rules based on the extracted information.
[0262] Step 4: Use the rule engine to execute the generated rules and perform data cleansing.
[0263] Step 5: Store the cleaned data in the data storage and management module.
[0264] Through the above steps, the intelligent accounting information enhancement system can efficiently parse the latest accounting policy change information and automatically perform corresponding data processing and cleaning operations to ensure the accuracy and consistency of the data.
[0265] Data interface: Use RESTful API for data transmission.
[0266] Data format: Supports JSON and XML formats to ensure data compatibility and readability.
[0267] Data processing: During the data collection process, the data cleaning tool Talend is used to remove duplicate data, fill in missing values, and correct data errors to ensure data accuracy and consistency.
[0268] Data security: Use the RSA encryption algorithm to encrypt data to ensure data security and privacy protection.
[0269] Furthermore, the application of blockchain technology in this system:
[0270] Blockchain technology is mainly used to ensure the immutability and transparency of data. The specific implementation methods are as follows:
[0271] 1. Data encryption and storage:
[0272] Data encryption: During data transmission and storage, the Advanced Encryption Standard (AES) is used to encrypt data. Before each piece of data is written into the blockchain, it is encrypted using the AES algorithm to ensure the confidentiality of the data.
[0273] Hash algorithm: Use the SHA-256 hash algorithm to generate a unique identifier (hash value) for the data. The hash value of each piece of data will be used as a block in the blockchain to ensure the integrity and immutability of the data.
[0274] 2. Blockchain architecture:
[0275] Combination of public and private chains: The system adopts a hybrid blockchain architecture, with the public chain used for open and transparent data records and the private chain used for the management of sensitive data within the enterprise. The public chain is Ethereum and the private chain is Hyperledger Fabric.
[0276] Node deployment: Deploy multiple nodes within the enterprise to ensure distributed storage and verification of data. Each node stores complete blockchain data to ensure high availability and security of data.
[0277] 3. Smart Contracts:
[0278] Definition and deployment: Smart contracts are automated programs that run on the blockchain to execute and verify transactions. Smart contracts are written in Solidity and deployed on the Ethereum network.
[0279] Functional implementation: Smart contracts are used to automatically record detailed information about each data operation, including operation time, operation type, operator, etc. Each data operation triggers the smart contract and generates a new blockchain record to ensure transparency and traceability of the operation.
[0280] The blockchain application in the automated data collection module of the present invention is as follows:
[0281] 1. Financial data, operational data, and market data:
[0282] Data encryption: Financial data, operational data, and market data are encrypted using the AES algorithm during transmission and storage to ensure data confidentiality. For example, a company’s financial statements, transaction records, and operating data of production equipment are encrypted before being transmitted to the blockchain.
[0283] Hash generation: Encrypted financial data, operational data, and market data are generated with a unique hash value using the SHA-256 hash algorithm. The hash value of each piece of data will be used as a block in the blockchain to ensure data integrity and immutability.
[0284] Smart contract records: Every operation of financial data, operational data, and market data (such as adding, modifying, and deleting) will trigger a smart contract to record the detailed information of the operation. The smart contract writes the operation record to the blockchain and generates a new block to ensure the transparency and traceability of the operation.
[0285] 2. Accounting policy data:
[0286] Data encryption: Accounting policy data is encrypted using the AES algorithm during transmission and storage to ensure data confidentiality. For example, updated information on IFRS and US GAAP is encrypted before being transmitted to the blockchain.
[0287] Hash generation: The encrypted accounting policy data generates a unique hash value through the SHA-256 hash algorithm. The hash value of each piece of data will be used as a block in the blockchain to ensure the integrity and immutability of the data.
[0288] Smart contract records: Each operation of accounting policy data (such as adding, modifying, and deleting) will trigger a smart contract to record the detailed information of the operation. The smart contract writes the operation record into the blockchain and generates a new block to ensure the transparency and traceability of the operation.
[0289] Through the application of the above-mentioned blockchain technology, the intelligent accounting information enhancement system can ensure the immutability and transparency of financial data, operational data, market data and accounting policy data, provide reliable audit support, and enhance the security and credibility of corporate financial management.
[0290] Data processing and cleaning module:
[0291] Use the data cleaning tool Talend to remove duplicate data, fill missing values, and correct data errors.
[0292] Data standardization ensures the consistency and comparability of data from different sources by using data type conversion, unit unification, and format standardization.
[0293] Furthermore, the data processing and cleaning module specifically includes data correction and anomaly detection:
[0294] Among them, the specific implementation steps of data correction are:
[0295] 1. Data collection: Obtain raw data from the automated data collection module, including financial data, operational data, market data, and accounting policy data.
[0296] 2. Data preprocessing: 1) Data deduplication: Use algorithms to remove duplicate data to ensure data uniqueness. 2) Missing value processing: Fill in missing data through interpolation or machine learning models (such as K-nearest neighbor algorithm).
[0297] 3. Data Correction: Linear Regression: Used to correct systematic errors in the data. Use a linear regression model to correct for biases in sensor data.
[0298] 4. Data standardization: Convert the bias-corrected data into a standard normal distribution to ensure that data from different data sources are comparable.
[0299] Specific implementation steps of anomaly detection:
[0300] 1. Data preprocessing: Data denoising, using filters or smoothing algorithms to remove noise from the data.
[0301] 2. Anomaly detection: K-means clustering is used to detect anomalies in the data. Through cluster analysis, data points that are far away from the cluster center are identified as anomalies.
[0302] Through the above detailed implementation steps, the data processing and cleaning module can effectively perform data correction and anomaly detection, ensure the accuracy and completeness of the data, and provide a reliable data foundation for subsequent data analysis and prediction.
[0303] Data storage and management module:
[0304] Relational database: stores structured data (such as financial data, operational data). Examples: MySQL, PostgreSQL.
[0305] NoSQL database: stores unstructured data (such as market data, accounting policy data). Examples: MongoDB, Cassandra.
[0306] Data Lake: Stores large amounts of unstructured data and supports flexible query and analysis. Examples: Apache Hadoop (open source distributed storage and processing framework), Amazon S3 (object storage service provided by Amazon Web Services (AWS)).
[0307] Application of Data Lake in Intelligent Accounting Information Enhancement System
[0308] A data lake is a storage architecture for storing large-scale unstructured data that supports flexible query and analysis. The main feature of a data lake is that it can store various types of data, including structured, semi-structured, and unstructured data, and provide efficient data processing and analysis capabilities.
[0309] Large-scale unstructured data:
[0310] In this system, large-scale unstructured data includes log data, multimedia data and text data.
[0311] Supports the implementation of flexible query and analysis:
[0312] Data lakes integrate multiple data processing and analysis tools to enable flexible query and analysis of large-scale unstructured data. The following are specific implementation methods and examples:
[0313] 1. Data storage:
[0314] Apache Hadoop: Hadoop is an open source distributed storage and processing framework that can efficiently store and process large amounts of data. Hadoop's HDFS (Hadoop Distributed File System) is used to store data, and MapReduce is used for data processing.
[0315] Amazon S3 (Simple Storage Service): S3 is an object storage service provided by Amazon that can store any amount of data and provides high availability and high scalability. S3 supports multiple data formats and is suitable for storing large-scale unstructured data.
[0316] 2. Data processing and analysis:
[0317] Apache Spark: Spark is a fast, general-purpose distributed data processing engine that can run on Hadoop. Spark supports a variety of data processing and analysis tasks, including batch processing, stream processing, machine learning, etc.
[0318] Presto: Presto is a distributed SQL query engine that can interactively query data sources such as Hadoop and S3. Presto supports standard SQL syntax and can efficiently query and analyze large-scale data.
[0319] 3. Examples of flexible query and analysis:
[0320] Example 1: Log data analysis
[0321] Data storage: The log files generated by the system are stored in Hadoop's HDFS.
[0322] Data processing: Use Apache Spark to batch process log data and extract useful information, such as error logs, access logs, etc.
[0323] Data query: Use Presto to interactively query the processed log data to analyze the system's operating status and performance bottlenecks.
[0324] Through the above implementation methods and examples, the data lake can efficiently store and process large-scale unstructured data, support flexible query and analysis, and meet the needs of enterprises for data processing and analysis.
[0325] Specifically, the advantages of multi-database storage in the data storage and management module are:
[0326] 1. Diversity of data types:
[0327] Relational databases (such as MySQL): Suitable for storing structured data, such as financial statements, transaction records, etc. This type of data has fixed patterns and relationships, and relational databases can query and operate efficiently.
[0328] NoSQL databases (such as MongoDB): Suitable for storing unstructured or semi-structured data, such as market trends, customer feedback, sensor data, etc. The schema of this type of data is not fixed, and NoSQL databases provide greater flexibility and scalability.
[0329] Data lake (such as Apache Hadoop): Suitable for storing large-scale unstructured data, such as log files, images, videos, etc. Data lake can efficiently store and process big data and support batch processing and real-time analysis.
[0330] 2. Performance optimization:
[0331] Query performance: Relational databases have advantages in processing complex queries and transactions, and can provide efficient query performance and data consistency.
[0332] Scalability: NoSQL databases and data lakes have advantages in processing large-scale data and high concurrent access, and can provide good scalability and high availability.
[0333] 3. Flexibility and scalability:
[0334] Flexible data model: NoSQL databases support flexible data models that can adapt to rapidly changing business needs.
[0335] Big data processing: Data lakes can store and process large amounts of data and support a variety of data processing and analysis tools, such as MapReduce and Spark.
[0336] The coordination relationship between multiple databases and other modules:
[0337] 1. Architecture of data storage and management module:
[0338] Relational database (MySQL): used to store structured data, such as financial data, transaction records, etc.
[0339] NoSQL database (MongoDB): used to store unstructured or semi-structured data, such as market data, sensor data, etc.
[0340] Data Lake (Apache Hadoop): Used to store large-scale unstructured data, such as log files, images, videos, etc.
[0341] 2. Data storage and management module and other modules cooperation:
[0342] Automated data collection module: collects data from sensors and other data sources and stores it in the corresponding database according to the data type. For example, structured data is stored in MySQL, unstructured data is stored in MongoDB, and large-scale data is stored in Hadoop.
[0343] Data processing and cleaning module: extract data from different databases for preprocessing and cleaning. The cleaned data is stored back in the corresponding database as needed.
[0344] Data analysis and machine learning module: extract data from different databases for analysis and modeling. The analysis results and models can be stored in the corresponding database for subsequent use.
[0345] Data visualization and reporting module: extract data from different databases for visualization and report generation. The visualization results and reports can be stored in the corresponding database for management to view.
[0346] 3. Data transfer between multiple databases:
[0347] Data synchronization: Data synchronization between different databases can be achieved through data synchronization tools (such as Apache Kafka and Apache NiFi). For example, unstructured data in NoSQL databases can be synchronized to the data lake for big data analysis.
[0348] Data integration: Data integration between relational databases and NoSQL databases can be achieved through data integration tools (such as Apache Sqoop and Apache Flume). For example, structured data in MySQL can be imported into MongoDB for flexible query.
[0349] Data access: Extract data from different databases for processing and analysis based on business needs. For example, the data analysis and machine learning module can extract data from MySQL and MongoDB at the same time for comprehensive analysis.
[0350] Data transfer in specific situations:
[0351] 1. Real-time data processing:
[0352] Sensor data: The automated data acquisition module collects real-time data from sensors and stores it in MongoDB. The data processing and cleaning module extracts data from MongoDB for cleaning and correction, and the cleaned data is stored back in MongoDB.
[0353] Financial data: The automated data collection module collects real-time data from the financial system and stores it in MySQL. The data processing and cleaning module extracts data from MySQL for cleaning and correction, and the cleaned data is stored back in MySQL.
[0354] 2. Batch data processing:
[0355] Market data: The automated data collection module collects batch data from market data sources and stores it in Hadoop. The data processing and cleaning module extracts data from Hadoop for cleaning and correction, and the cleaned data is stored back in Hadoop.
[0356] Log data: The automated data collection module collects batch data from system logs and stores it in Hadoop. The data processing and cleaning module extracts data from Hadoop for cleaning and correction, and the cleaned data is stored back in Hadoop.
[0357] 3. Comprehensive data analysis:
[0358] Multi-source data fusion: The data analysis and machine learning module extracts data from MySQL, MongoDB, and Hadoop for comprehensive analysis. For example, financial data can be extracted from MySQL, market data from MongoDB, and log data from Hadoop for comprehensive analysis and modeling.
[0359] The specific contents of multi-data fusion are as follows:
[0360] The purpose of data fusion is to combine data from different sources into a unified view for deeper analysis and decision making.
[0361] Steps of data fusion:
[0362] Step 1: Data extraction (Extract), extract data from multiple data sources, and use ETL (Extract, Transform, Load) tools to extract data from data sources such as MySQL, MongoDB, and Hadoop.
[0363] Step 2: Data transformation (Transform), data cleaning and standardization, clean and standardize the extracted data to ensure the consistency and comparability of the data.
[0364] Step 3: Data loading (Load), data storage, store the cleaned and standardized data in a data warehouse or data lake for subsequent analysis.
[0365] Step 4: Data fusion, merge the data from different data sources to create a unified data view.
[0366] Implementation of comprehensive analysis:
[0367] Specific implementation of data analysis and machine learning modules:
[0368] Step 1: Data preprocessing, data cleaning and standardization to ensure data consistency and comparability.
[0369] Step 2: Data modeling, using machine learning algorithms for financial forecasting and anomaly detection:
[0370] Regression Models: Used for financial forecasting.
[0371] Classification models: used for anomaly detection.
[0372] Step 3: Data visualization, creating interactive dashboards and reports using Tableau and Power BI:
[0373] Tableau: Connect to data sources, create charts, add interactivity, publish and share dashboards.
[0374] Power BI: Connect to data sources, create charts, add interactive features, publish and share reports.
[0375] Through the above steps, the intelligent accounting information enhancement system can effectively realize multi-source data fusion and conduct comprehensive analysis to provide enterprises with accurate financial forecasts and anomaly detection, helping enterprises make wise decisions.
[0376] Predictive Analysis: The Data Analysis and Machine Learning module extracts data from MySQL and MongoDB for predictive analysis. For example, it extracts historical financial data from MySQL and market trend data from MongoDB to perform financial and market forecasts.
[0377] Through the application of the above-mentioned multi-database storage format, the intelligent accounting data management system can efficiently manage and process different types of data, provide flexible storage and analysis capabilities, and meet the needs of enterprises for efficient and accurate data processing.
[0378] Data Analysis and Machine Learning Modules:
[0379] Use machine learning algorithms such as regression models, time series models, classification models for financial forecasting and anomaly detection.
[0380] Use machine learning algorithms for financial forecasting and anomaly detection. In the intelligent accounting information enhancement system, multiple machine learning algorithms (regression model, time series model, classification model) are used for financial forecasting and anomaly detection.
[0381] 1. The following example shows how a model can achieve specific steps for financial forecasting;
[0382] Example: Classification Model
[0383] Logistic Regression Model:
[0384] Introduction: The logistic regression model is used for binary classification problems, and predicts binary classification results by fitting a logistic function.
[0385] Input: Historical financial data and related characteristics.
[0386] Output: classification results (such as whether the company will go bankrupt).
[0387] Specific implementation steps:
[0388] 1. Data preparation: Collect financial data and select relevant features and labels.
[0389] 2. Model training: Use the logistic regression model to train the data and fit the prediction formula.
[0390] 3. Forecasting: Use the trained model to predict future financial data.
[0391] The fitting process of the prediction formula:
[0392] Step 1: Assume we have a set of financial data X and label Y (0 for not bankrupt, 1 for bankrupt).
[0393] Step 2: Use the logistic regression model to fit the data and obtain the prediction formula P(Y=1|X)=\frac{1}{1+e^{(\beta_0+\beta_1X)}}, where \beta_0 and \beta_1 are model parameters obtained by maximizing the likelihood function.
[0394] Step 3: Use the fitted model to predict future data and obtain classification results.
[0395] Suppose we have financial data X and label Y (0 means not bankrupt, 1 means bankrupt). We split the data into training set and test set. Use logistic regression model to train the data and fit the prediction formula. Then, use the trained model to predict future financial data.
[0396] 2. The implementation of anomaly detection in this module specifically includes:
[0397] Anomaly detection models include clustering models, support vector machines (SVMs), and isolation forests;
[0398] (1) Clustering model
[0399] Kmeans clustering: It is used to divide data points into K clusters and identify outliers by minimizing the square error of data points within the cluster.
[0400] Input: Financial data and related characteristics.
[0401] Output: Identification results of outliers.
[0402] Specific implementation steps:
[0403] 1. Data preparation: Collect financial data and select relevant features.
[0404] 2. Model training: Use Kmeans clustering model to train data and identify cluster centers.
[0405] 3. Anomaly detection: Calculate the distance from the data point to the cluster center and identify outliers.
[0406] Anomaly detection process:
[0407] Step 1: Assume we have a set of financial data X.
[0408] Step 2: Use the Kmeans clustering model to divide the data points into K clusters and get the center \mu_k of each cluster.
[0409] Step 3: Calculate the distance from each data point to the cluster center d(x,\mu_k)=\sqrt{\sum_{i=1}^{n}(x_i\mu_{k,i})^2}.
[0410] Step 4: Set a threshold and identify outliers based on the distance. If the distance exceeds the threshold, the data point is considered an outlier.
[0411] Suppose we have financial data X. We use the Kmeans clustering model to divide the data points into K clusters and get the center of each cluster. Then, we calculate the distance from each data point to the cluster center. We set a threshold and identify outliers based on the distance. If the distance exceeds the threshold, the data point is considered an outlier.
[0412] (2) Support Vector Machine (SVM)
[0413] OneClass SVM: used for anomaly detection, identifying outliers by finding the best segmentation hyperplane.
[0414] Input: Financial data and related characteristics.
[0415] Output: Identification results of outliers.
[0416] Specific implementation steps:
[0417] 1. Data preparation: Collect financial data and select relevant features. Standardize the data to ensure consistency and comparability.
[0418] 2. Model training: Use the OneClass SVM model to train the data and fit the segmentation hyperplane. OneClass SVM is an unsupervised model specifically used for anomaly detection. Its goal is to learn a boundary that surrounds normal data points and excludes abnormal points.
[0419] 3. Find the best splitting hyperplane:
[0420] The best splitting hyperplane is the one that maximizes the distance between normal data points and the hyperplane.
[0421] SVM finds the best hyperplane by optimizing the following objective function:
[0422] \min_{\mathbf{w},b}\frac{1}{2}\|\mathbf{w}\|^2;
[0423] Among them, \mathbf{w} is the weight vector and b is the bias term.
[0424] The constraints are:
[0425] y_i(\mathbf{w}^T\mathbf{x}_i+b)\geq 1\xi_i
[0426] Among them, yi is the label of the data point (+1 for a normal point and 1 for an abnormal point), and \xi is a slack variable used to allow some data points to violate the constraints.
[0427] Calculate the distance from the data point to the splitting hyperplane
[0428] 1. Definition of Hyperplane:
[0429] The equation of the hyperplane is:
[0430] \mathbf{w}^T\mathbf{x}+b=0
[0431] Among them, \mathbf{w} is the weight vector and b is the bias term.
[0432] 2. Calculate the distance:
[0433] The distance from the data point \mathbf{x} to the hyperplane can be calculated by the following formula:
[0434] d(\mathbf{x})=\frac{|\mathbf{w}^T\mathbf{x}+b|}{\|\mathbf{w}\|}
[0435] where \|\mathbf{w}\| is the norm of the weight vector.
[0436] Identify outliers based on distance
[0437] 1. Set the threshold:
[0438] Set a threshold \delta to determine whether a data point is an outlier.
[0439] If the distance d(\mathbf{x}) from a data point to the hyperplane exceeds the threshold \delta, the data point is considered an outlier.
[0440] 2. Anomaly detection process:
[0441] For each data point \mathbf{x}, calculate its distance d(\mathbf{x}) to the hyperplane.
[0442] If d(\mathbf{x})>\delta, mark the data point as an outlier.
[0443] Through the above specific implementation process, the intelligent accounting information enhancement system can efficiently perform anomaly detection, provide accurate anomaly point identification, and help enterprises discover and handle abnormal situations in a timely manner.
[0444] (3) Isolation Forest
[0445] Isolation Forest: Used for anomaly detection, it builds a tree structure by randomly selecting features and split values to identify isolated data points as outliers.
[0446] Input: Financial data and related characteristics.
[0447] Output: Identification results of outliers.
[0448] Specific implementation steps:
[0449] 1. Data preparation: Collect financial data and select relevant features.
[0450] 2. Model training: Use the Isolation Forest model to train data and build a tree structure.
[0451] 3. Anomaly detection: Calculate the anomaly score of the data point and identify the anomaly points.
[0452] Anomaly detection process:
[0453] Step 1: Assume we have a set of financial data X.
[0454] Step 2: Use the Isolation Forest model to train the data and build a tree structure.
[0455] Step 3: Calculate the anomaly score s(x) = 2^{\frac{E(h(x))}{c(n)}} for each data point, where E(h(x)) is the expected value of the path length and c(n) is the adjustment factor.
[0456] Step 4: Identify outliers based on the anomaly score. If the anomaly score exceeds the threshold, the data point is considered an outlier.
[0457] Suppose we have financial data X. We use the Isolation Forest model to train the data and build a tree structure. Then, we calculate the anomaly score for each data point. We identify outliers based on the anomaly score. If the anomaly score exceeds the threshold, the data point is considered an outlier.
[0458] Combine with other data for prediction and detection:
[0459] In the process of financial forecasting and anomaly detection, combining other data (such as accounting policy data, operational data, market data, etc.) can improve the accuracy and robustness of the model. The following is a specific description of how to combine these data:
[0460] Data Fusion:
[0461] Data Extraction: Extract data from different data sources like MySQL, MongoDB, Hadoop.
[0462] Data cleaning and standardization: The extracted data are cleaned and standardized to ensure data consistency and comparability.
[0463] Data fusion: Merge data from different data sources to create a unified data view.
[0464] Extract data from different data sources (such as MySQL, MongoDB, Hadoop). Clean and standardize the extracted data to ensure data consistency and comparability. Merge data from different data sources to create a unified data view.
[0465] Combine with other data for prediction and detection:
[0466] Financial Forecast:
[0467] Incorporate accounting policy data: Input accounting policy data into the model as features to improve the accuracy of predictions.
[0468] Assume we have accounting policy data policy_data. We fuse accounting policy data with financial data. We use a linear regression model to make financial forecasts. We input accounting policy data into the model as a feature to improve the accuracy of the forecast.
[0469] Anomaly Detection:
[0470] Combine operational data and market data: Input operational data and market data into the model as features to improve the accuracy of anomaly detection.
[0471] Assume we have operational data and market data. We fuse operational data and market data with financial data. Use SVM model for anomaly detection. Input operational data and market data into the model as features to improve the accuracy of anomaly detection.
[0472] Through the application of the above-mentioned machine learning models, the intelligent accounting information enhancement system can efficiently perform financial forecasting and anomaly detection, provide accurate financial data and trend forecasts, and help enterprises make forward-looking decisions.
[0473] Use Python's Scikit-learn (a Python-based machine learning library that provides simple and efficient tools for data mining and data analysis), TensorFlow (an open source deep learning framework developed by Google, widely used in machine learning and deep learning tasks), and Keras (a high-level neural network API that simplifies the construction and training process of deep learning models and provides a user-friendly interface) for data modeling.
[0474] Here is an example of the specific steps of data modeling:
[0475] Data Modeling with Scikit-learn
[0476] Example: Using Linear Regression for Financial Forecasting;
[0477] 1. Data preparation:
[0478] Collect historical financial data: Collect the company's historical financial data, including revenue, expenditure, market trends and other relevant characteristics.
[0479] Data Splitting: Split the data into training set and test set. The training set is used to train the model, and the test set is used to evaluate the performance of the model.
[0480] 2. Model training:
[0481] Create a linear regression model: Use the linear regression model from the Scikitlearn library for training.
[0482] Fitting prediction formula: By minimizing the mean square error (MSE), the prediction formula \hat{Y}=\beta_0+\beta_1X is fitted, where \beta_0 and \beta_1 are model parameters.
[0483] 3. Prediction:
[0484] Use the model for prediction: Use the trained model to predict the test set and future financial data to obtain the predicted value \hat{Y}.
[0485] 4. Evaluate the model:
[0486] Calculate mean squared error: To evaluate the performance of the model, calculate the mean squared error (MSE) between the predicted values and the actual values.
[0487] Through the above specific implementation process, the intelligent accounting information enhancement system can efficiently perform financial forecasting and anomaly detection, provide accurate financial data and trend forecasts, and help enterprises make forward-looking decisions.
[0488] In the intelligent accounting information enhancement system, the data processing and cleaning module is responsible for preliminary data correction and anomaly detection, while the data analysis and machine learning module further conducts in-depth data analysis and prediction. Therefore, the data processing and cleaning module mainly handles the preliminary cleaning and correction of data.
[0489] The Data Analysis and Machine Learning module is responsible for in-depth data analysis and prediction, including further anomaly detection, trend analysis, and predictive modeling. This module uses more complex machine learning and deep learning algorithms for data analysis.
[0490] Data visualization and reporting module:
[0491] Use tools such as Tableau (a data visualization tool that helps users create interactive and shareable dashboards) and Power BI (a set of business analysis tools developed by Microsoft that provides interactive visualization and business intelligence capabilities) to create interactive dashboards and reports, and use Python's Matplotl ib (a Python 2D drawing library that can generate a variety of static, dynamic, and interactive charts) and Seaborn (an advanced data visualization library based on Matplotlib that provides a simpler API and more beautiful default styles) to create dynamic charts to intuitively display data changes and prediction results.
[0492] Specifically, create interactive dashboards and reports:
[0493] 1. Create interactive dashboards and reports using Tableau
[0494] 1. Connect to the data source: Open Tableau, select the "Connect" option, and select the data source (such as Excel, SQL database, CSV file, etc.).
[0495] 2. Data preparation: In the Data Source tab, preview and clean the data to ensure that the data format is correct.
[0496] 3. Create a chart: In the "Worksheet" tab, select the chart type (such as line chart, bar chart, pie chart, etc.), drag and drop data fields into rows and columns to generate a chart.
[0497] 4. Create dashboards: In the Dashboard tab, drag and drop multiple worksheets, adjust the layout and style, and create interactive dashboards.
[0498] 5. Add interactivity: Use the Filters, Parameters, and Actions options to add interactive features to make the dashboard more dynamic and user-friendly.
[0499] 6. Publish and share: Publish dashboards to Tableau Server or Tableau Public to share with your team or the public.
[0500] 2. Create interactive dashboards and reports using Power BI
[0501] 1. Connect to the data source: Open Power BIDesktop, select the "Get Data" option, and select the data source (such as Excel, SQL database, CSV file, etc.).
[0502] 2. Data preparation: In the Query Editor, preview and clean the data to ensure that the data format is correct.
[0503] 3. Create a chart: In the Report view, select a chart type (such as a line chart, bar chart, pie chart, etc.), drag and drop data fields into the visualization panel, and generate a chart.
[0504] 4. Create dashboards: In the Report view, drag and drop multiple charts, adjust the layout and style, and create interactive dashboards.
[0505] 5. Add interactivity: Use Slicers, Filters, and Bookmarks options to add interactive features to make the dashboard more dynamic and user-friendly.
[0506] 6. Publish and share: Publish reports to Power BI Service and share them with your team or the public.
[0507] 3. Create dynamic charts using Matplotlib and Seaborn
[0508] 1. Install the library: Make sure you have installed the Matplotlib and Seaborn libraries.
[0509] 2. Import libraries: Import Matplotlib and Seaborn libraries in the Python script.
[0510] 3. Data preparation: Load and prepare data, and ensure that the data format is correct.
[0511] 4. Create charts: Create dynamic charts using Matplotlib and Seaborn.
[0512] Through the above tools and libraries, the intelligent accounting information enhancement system can create interactive dashboards and reports to intuitively display data changes and forecast results, helping corporate management to quickly understand and analyze data and make wise decisions.
[0513] The embodiment of the present invention provides an intelligent accounting data management and compliance system and method, which can automatically collect and process accounting data from ERP systems, market monitoring tools and target departments, improve the efficiency and accuracy of data processing, comprehensively evaluate the value of intangible assets, and provide real-time decision support for enterprises.
[0514] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.
[0515] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. An intelligent accounting data management and compliance system, characterized in that: include: Automated data collection module: collects data and uploads it to the blockchain, encrypting the data during transmission and storage; Data processing and cleaning module: standardize and clean the collected data, and perform data correction and anomaly detection; Data storage and management module: adopts multi-database storage to classify and store the cleaned data; Data Analysis and Machine Learning Module: Use machine learning algorithms for financial forecasting and anomaly detection, and perform data modeling; Data visualization and reporting module: Generate reports and dynamic charts to display data changes and prediction results.
2. The intelligent accounting data management and compliance system according to claim 1, characterized in that: The automated data collection module includes various types of sensors and API interfaces for collecting financial data, operational data, market data and accounting policy data; Wherein, the data sources of the financial data and operational data are automatically collected through the data interface of the enterprise resource planning ERP system; The data source of the market data is collected through relevant APP interfaces including Enterprise Early Warning to collect market trends and competition analysis data; The data source of the accounting policy data collects the latest accounting policy update and change information through the data interface of the target department.
3. The intelligent accounting data management and compliance system according to claim 2, characterized in that: The automated data acquisition module also includes: Data encryption: All data are encrypted using the AES algorithm during transmission and storage to ensure data confidentiality; Hash generation: Each encrypted data is generated with a unique hash value through the SHA-256 hash algorithm. The hash value of each piece of data will be used as a block in the blockchain to ensure the integrity and immutability of the data. Smart contract records: Each operation on each data can trigger a smart contract to record the detailed information of the operation. The smart contract writes the operation record into the blockchain and generates a new block to ensure the transparency and traceability of the operation.
4. The intelligent accounting data management and compliance system according to claim 1, characterized in that: In the data processing and cleaning module, the processing of financial data, operational data, market data and accounting policy data includes: Data correction: 1) Data collection: Acquiring raw data from the automated data acquisition module, including financial data, operational data, market data, and accounting policy data; 2) Data preprocessing: Data deduplication: Use algorithms to remove duplicate data to ensure data uniqueness; Missing value processing: fill in missing data through interpolation or machine learning models; 3) Data correction: Linear regression: used to correct systematic errors in data. Linear regression models are used to correct biases in sensor data. 4) Data standardization: Convert the bias-corrected data into a standard normal distribution to ensure that data from different data sources are comparable; Anomaly Detection: 1) Data preprocessing: Data denoising: using filters or smoothing algorithms to remove noise from data; 2) Anomaly Detection: K-means clustering: used to detect outliers in data. Through cluster analysis, data points that are far away from the cluster center are identified as outliers.
5. The intelligent accounting data management and compliance system according to claim 1, characterized in that: The system includes real-time analysis of accounting policy changes for target departments, including: 1) Obtain the latest accounting policy data in real time through the API interface; 2) Use NLP models to parse policy texts and extract entity information in policy documents through named entity recognition; 3) Generate data processing rules based on the extracted entity information; 4) Use the rule engine to execute the generated data rules and perform data-related processing and cleaning; Among them, data processing rules include: data cleaning rules, data conversion rules, outlier processing rules, integrity rules, and security and compliance rules; 5) storing the processed and cleaned data in the data storage and management module; By analyzing policy changes in real time, continuously updating data processing rules and cleaning data again, integrating old and new data, accounting data processing is kept consistent with the latest policies, and data compliance and accuracy are ensured in a timely manner.
6. The intelligent accounting data management and compliance system according to claim 1, characterized in that: The data storage and management module adopts a multi-database storage form, supporting structured and unstructured data storage, including: relational database, NoSQL database and data lake; Wherein, the relational database stores relevant structured data including financial data and operational data; The NoSQL database stores relevant unstructured data including market data and accounting policy data; The data lake stores large-scale unstructured data and supports flexible query and analysis; Data synchronization between different databases is achieved through data synchronization tools, and data integration between relational databases and NoSQL databases is achieved through data integration tools; After the automated data acquisition module collects data, it stores the data in corresponding databases according to the data type; the data processing and cleaning module extracts data from different databases for preprocessing and cleaning, and puts it back into the corresponding database after cleaning; the data analysis and machine learning module extracts data from different databases for analysis and modeling, and stores the analysis results and models in the corresponding database for subsequent use; the data visualization and reporting module: extracts data from different databases for visualization and report generation, and stores the visualization results and reports in the corresponding database for viewing; Data mobilization between multiple databases enables data synchronization, data integration and data access, and performs real-time data processing, batch data processing and comprehensive data analysis.
7. The intelligent accounting data management and compliance system according to claim 1, characterized in that: The data analysis and machine learning module uses regression models, time series models, and classification models for financial forecasting and anomaly detection, including linear regression, ARIMA, LSTM, and support vector machine algorithms; Among them, financial forecasting: predicting changes in financial indicators through regression models and time series models, and providing future financial trend forecasts; Anomaly detection: Detect anomalies in data through classification models and clustering algorithms.
8. The intelligent accounting data management and compliance system according to claim 1, characterized in that: The data visualization and reporting module uses Tableau and Power BI tools to create interactive dashboards and reports, and uses Python's Matplotlib and Seaborn libraries to create dynamic charts to display data changes and prediction results; Among them, Tableau: connect to data sources, create charts, add interactive functions, publish and share dashboards; Power BI: connect to data sources, create charts, add interactive features, publish and share reports; Matplotlib: Create static and dynamic charts with fine-grained control over chart appearance and behavior; Seaborn: Quickly create beautiful statistical charts suitable for data analysis and exploratory data analysis (EDA).
9. An intelligent accounting data management and compliance method, characterized in that: include: Automated data collection steps: Collect data and upload it to the blockchain, encrypting the data during data transmission and storage; Data processing and cleaning steps: standardize and clean the collected data; Data storage and management steps: Use multi-database storage to classify and store the cleaned data; Data Analysis and Machine Learning Steps: Use machine learning algorithms for financial forecasting and anomaly detection, and perform data modeling; Data visualization and reporting steps: Generate reports and dynamic charts to display data changes and prediction results.
10. The intelligent accounting data management and compliance method according to claim 9, characterized in that: The automated data collection step includes various types of sensors and API interfaces for collecting financial data, operational data, market data and accounting policy data; The automated data collection step also includes: Data encryption: All data are encrypted using the AES algorithm during transmission and storage to ensure data confidentiality; Hash generation: Each encrypted data is generated with a unique hash value through the SHA-256 hash algorithm. The hash value of each piece of data will be used as a block in the blockchain to ensure the integrity and immutability of the data. Smart contract records: Each operation on each data can trigger a smart contract to record the detailed information of the operation. The smart contract writes the operation record into the blockchain and generates a new block to ensure the transparency and traceability of the operation. In the data processing and cleaning step, the processing of financial data, operational data, market data and accounting policy data includes: data correction and anomaly detection; The system includes real-time analysis of accounting policy changes for target departments, including: 1) Obtain the latest accounting policy data in real time through the API interface; 2) Use NLP models to parse policy texts and extract entity information in policy documents through named entity recognition; 3) Generate data processing rules based on the extracted entity information; 4) Use the rule engine to execute the generated data rules and perform data-related processing and cleaning; 5) storing the processed and cleaned data in the data storage and management step; The data storage and management steps adopt a multi-database storage form, supporting structured and unstructured data storage, including: relational databases, NoSQL databases and data lakes; The data analysis and machine learning steps use regression models, time series models, and classification models for financial forecasting and anomaly detection, including linear regression, ARIMA, LSTM, and support vector machine algorithms; The data visualization and reporting steps use Tableau and Power BI tools to create interactive dashboards and reports, and use Python's Matplotlib and Seaborn libraries to create dynamic charts to display data changes and prediction results.
Citation Information
Patent Citations
Comprehensive financial auditing system based on big data
CN115375417A
Cited By
Water affair data analysis and management system and method based on data model
CN120087853A
Comprehensive data asset value evaluation device
CN120198168A
Automatic process processing method for accounting data
CN120256507A