Enterprise finance and taxation intelligent management system based on data analysis

By building an intelligent enterprise financial and tax management system, real-time fusion and dynamic perception of multi-source data have been achieved, solving the problems of data silos and response delays in existing systems, improving the real-time performance and intelligence level of financial and tax management, and enabling proactive risk identification and intelligent decision-making.

CN120912153APending Publication Date: 2025-11-07RANDIAN (NANTONG) TECH ENTREPRENEURSHIP SERVICE CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202511318242.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing financial and tax management systems cannot integrate multi-source heterogeneous data in real time and lack the ability to dynamically perceive business scenarios, resulting in lag in risk identification and compliance analysis. Relying on manual interpretation of regulatory updates is inefficient and prone to errors.

Method used

We will build an intelligent enterprise financial and tax management system based on data analysis. Through a multi-source heterogeneous data real-time access layer, a central data processing and knowledge graph construction module, a real-time analysis engine based on stream computing, and a strategy execution and feedback interface, we will achieve deep integration and real-time perception of all enterprise financial and tax data, dynamically perceive business scenarios, and provide intelligent decision-making.

Benefits of technology

It has enabled a shift from post-event review to in-event intervention, improving the real-time nature, accuracy, and intelligence of financial and tax management, enhancing enterprises' risk control and compliance capabilities, proactively identifying invoice risks and compliance deviations, dynamically matching tax incentive policies, and improving the efficiency of financial and tax planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912153A_ABST
    Figure CN120912153A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of enterprise finance and taxation management, and discloses an enterprise finance and taxation intelligent management system based on data analysis. According to the method, the problems of data islands, response delay and business disjunction in the prior art are solved by constructing a multi-source data real-time access and dynamic knowledge graph and stream calculation analysis engine collaborative architecture. Specifically, deep fusion and real-time perception of enterprise finance and taxation total data are realized, invoice risks and compliance deviations in a business flow can be actively identified, and the conversion from post-event recheck to in-event intervention is completed; through combination of a dynamic knowledge graph and an AI model, an analysis decision deeply accords with a business scene, a tax preferential policy is matched, and the finance and tax planning efficiency is intelligently improved; meanwhile, the system has self-learning and self-adaptive capabilities and can be continuously optimized along with policy and regulation and business changes, the risk control level and compliance capability of an enterprise are enhanced, and the real-time performance, accuracy and intelligent degree of finance and tax management are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of enterprise financial and tax management, in particular to an enterprise financial and tax intelligent management system based on data analysis. BACKGROUND

[0002] Currently, the financial and tax management systems widely used by enterprises, such as traditional financial software or the financial and tax modules in ERP, mainly realize the informatization and automation of business processes such as accounting voucher generation, report preparation, invoice storage and tax declaration. These systems have improved the processing efficiency of basic businesses, but are still essentially post-record tools based on rules. Their core limitation is that they store financial, tax and business data in different subsystems in isolation, forming "data islands" that are difficult to break through. The system lacks the ability to deeply integrate, correlate and semantically understand multi-source heterogeneous data (such as bill images, contract texts, bank statements, policies and regulations), so that the vast amount of data is only "sleeping assets", and the deep business insights and risk values behind them are not effectively mined.

[0003] Further, the defects of the prior art are concentrated in the "passive response" and "static management" modes. Risk control relies on periodic report review and manual experience judgment, and cannot identify and intervene in invoice abnormalities, compliance deviations or planning opportunities occurring in real time in transaction flow within milliseconds. At the same time, the analysis models embedded in the system are mostly general and static algorithms, which cannot be deeply coupled with dynamically changing business context and tax law provisions, resulting in insufficient practicality and accuracy of the analysis results. In addition, the update of regulations completely depends on manual interpretation and manual input, which is low in efficiency and prone to errors, exposing enterprises to compliance risks. Therefore, the technical problem to be solved in the field is how to build a management system that can real-time integrate multi-source data, dynamically perceive business context, and actively provide intelligent decision-making to completely reverse the lagging and passive situation of financial and tax management. Therefore, an enterprise financial and tax intelligent management system based on data analysis is proposed. SUMMARY

[0004] In view of the deficiencies of the prior art, the present application provides an enterprise financial and tax intelligent management system based on data analysis to solve the problems in the background art.

[0005] To achieve the above-mentioned purpose, the present application provides the following technical solution: an enterprise financial and tax intelligent management system based on data analysis, comprising: A multi-source heterogeneous data real-time access layer adopts a configurable adapter interface to real-time stream collect structured and unstructured financial and tax related data from the business operation system, the tax billing system, the bank payment system and the external database of the enterprise; A central data processing and knowledge graph construction module connected to the data real-time access layer, configured to clean, fuse and semantically label the incoming data, and dynamically construct and update the enterprise finance and tax business knowledge graph based on the processed data; the knowledge graph takes enterprise entities, invoices, tax items, contracts and bank accounts as nodes, and takes transactions, ownership and correlation as edges, wherein at least part of the edges are accompanied by time stamp and amount weight attributes; A real-time analysis engine based on stream computing, which has built-in pluggable model containers for loading and running multiple micro-service AI analysis models; the analysis engine listens to real-time update events of the knowledge graph and calls corresponding AI analysis models to perform instant calculation on the changed subgraph; A policy execution and feedback interface for receiving instructions output by the real-time analysis engine, including risk blocking, planning suggestion prompts or compliance verification results, and returning the execution effect of the instructions as feedback data to the corresponding AI analysis model for online self-learning optimization of the model.

[0006] Preferably, the central data processing and knowledge graph construction module further has: An unstructured data processing unit integrated with an OCR recognition model and a natural language understanding model, configured to extract key entities and relationship triples from scanned contract and invoice images and automatically map them to corresponding nodes and edges of the knowledge graph.

[0007] Preferably, the real-time analysis engine based on stream computing further has: A real-time invoice risk early warning model for traversing the associated upstream and downstream enterprise nodes and historical transaction paths of each new invoice when it is entered into the graph, calculating the risk probability of the invoice existing in the fake tax evasion chain through a graph neural network algorithm, and generating real-time interception instructions to the policy execution and feedback interface for risk invoices exceeding a preset threshold.

[0008] Preferably, the judgment basis of the real-time invoice risk early warning model is: The depth of equity association, the similarity of historical fund reflux path, and the tax registration status change event of the invoicing enterprise and the receiving enterprise in the graph, and the comprehensive risk probability assessment.

[0009] Preferably, the real-time analysis engine based on stream computing further includes: A dynamic tax preferential policy matching model that continuously collects and analyzes data from official policy and regulation websites, automatically matches new policies with business scenarios in the enterprise knowledge graph using a text semantic similarity model, and actively pushes policy applicability analysis and measurement reports.

[0010] Preferably, the system further includes: A blockchain anchoring evidence module is configured to upload the hash values of the key risk early warning events, policy execution instructions and corresponding knowledge graph data snapshots issued by the real-time analysis engine to the blockchain, so as to ensure the auditability and non-tamperability of the full-link analysis decision.

[0011] Preferably, the policy execution and feedback interface is specifically configured to: In the procurement approval process, if a "high risk" instruction is received for an invoice, the payment process is automatically interrupted and a work order is pushed to the risk control personnel; At the same time, the final processing result of the risk control personnel is recorded and fed back to the real-time invoice risk early warning model as labeled training data for online incremental learning of the model.

[0012] Preferably, the knowledge graph construction module, when constructing relationships, is configured to: Preferably, the knowledge graph construction module, when constructing relationships, is configured to:

[0013] Compared with the prior art, the present application has the following beneficial effects: The present application solves the problems of data silos, response delays and business disconnection in the prior art by constructing a multi-source data real-time access, dynamic knowledge graph and flow calculation analysis engine collaborative architecture. Specifically, it realizes deep fusion and real-time perception of enterprise financial tax full data, can actively identify invoice risks and compliance deviations in business flow, and complete the transition from post-review to intervention; through the combination of dynamic knowledge graph and AI model, the analysis and decision making are deeply matched with the business scenario, the tax preferential policy is matched, and the efficiency of financial tax planning is intelligently improved; at the same time, the system has self-learning and self-adaptive ability, can continuously optimize with the changes of policies and regulations and business, and enhance the enterprise risk control level and compliance ability, and improve the real-time, accuracy and intelligent degree of financial tax management.

[0014] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and achieved by the structures indicated in the specification, claims and drawings. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 The figure is a system overall architecture diagram of the present application; Figure 2 The figure is a multi-source data access and processing flowchart of the present application; Figure 3 The figure is a real-time analysis engine workflow diagram of the present application; Figure 4 Strategy execution and feedback loop chart for the present application; Figure 5 Blockchain storage process chart for the present application. DETAILED DESCRIPTION

[0016] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0017] Please refer to Figures 1-5 The enterprise finance and tax intelligent management system based on data analysis in the present application realizes deep fusion and real-time perception of enterprise finance and tax full data through the collaborative operation of multi-source heterogeneous data real-time access layer, central data processing and knowledge graph construction module, real-time analysis engine based on stream computing and strategy execution and feedback interface, and completes the transformation from passive response to active intelligent management.

[0018] 1. Operation process of multi-source heterogeneous data real-time access layer This module realizes real-time and seamless data collection of various types of finance and tax related systems in the enterprise and external data sources, and provides comprehensive and timely data basis for subsequent analysis.

[0019] In the system initialization stage, the multi-source heterogeneous data real-time access layer first loads the pre-defined adapter configuration library. This library contains special adapters for different types of data sources: for relational databases (such as MySQL database in the enterprise's ERP system), the adapter establishes a persistent connection by configuring JDBC connection parameters (including connection address, credentials, connection pool size, etc.); for API interface type data sources (such as RESTful API of bank payment system), the adapter configuration includes API endpoint address, authentication method (OAuth 2.0 client credential mode), request frequency limit and other parameters; for file type data sources (such as invoice image files generated by the tax invoicing system), the adapter monitors the file change events of the specified cloud storage directory.

[0020] In particular, each adapter establishes a heartbeat detection mechanism during initialization, periodically reporting status to the central monitoring service, including the last data collection time, the number of collection records, the number of errors, and other indicators. For database-type data sources, the adapter first reads the metadata information of the database, automatically identifies the table structure and field type, and maps it to an internal unified data model. For API-type data sources, the adapter implements an automatic retry mechanism that retries according to an exponential backoff strategy when network anomalies occur, with a maximum of 5 retries and a retry interval starting at 1 second and doubling each time.

[0021] During data collection, for structured data, the adapter implements incremental collection through the following methods: 1) Set the last_modified timestamp field in the database table, and the adapter records the last collection time point. Each time the query is only for newly added or modified records after the time point. Query statement example: SELECT * FROM accounting_vouchers WHERE update_time>{last_fetch_time} ORDER BY update_time ASC.

[0022] 2) For unstructured data such as scanned contract documents, the adapter monitors file system events in the specified directory and triggers the collection process when a new file is detected.

[0023] In actual collection process, the system can use double timestamp mechanism to ensure data integrity: in addition to the database's own timestamp field, the adapter also records the collection waterline of each data source locally. This waterline uses vector clock technology and can handle the time synchronization problem of multiple data sources in a distributed environment. For large data tables, the adapter supports paging query and breakpoint resume, avoiding memory overflow caused by large data volume in a single query.

[0024] In the data transmission link, Apache Kafka message queue is used as the data bus, and the topic partition strategy is set to hash by taxpayer identification number, ensuring that related data of the same taxpayer is always routed to the same partition, guaranteeing data order. The message format uses Avro serialization, and the message header contains data schema version number, data source identification, collection timestamp, and other metadata. The message body contains actual business data, such as invoice code, number, issue date, amount, tax rate, and other fields for invoice data.

[0025] To ensure the reliability of data transmission, the system implements an end-to-end acknowledgement mechanism: the producer waits for all synchronous replicas to confirm before considering the message sent successfully, and the consumer manually commits the offset after successfully processing the message. The message queue sets a retention policy, retaining all messages for 7 days to facilitate fault recovery and data replay. In addition, the system also implements data quality monitoring, real-time statistics of each data source collection success rate, data delay and data integrity and other indicators.

[0026] This module realizes the unified access of various financial and tax data sources inside and outside the enterprise, ensures the real-time and integrity of the data, and provides high-quality data input for subsequent processing.

[0027] 2. Technical details of the central data processing and knowledge graph construction module This module cleans, fuses and semantically enhances the original data, and constructs a dynamically updated enterprise financial and tax knowledge graph, providing a structured knowledge base for intelligent analysis.

[0028] Detailed implementation process: The data cleaning stage implements a multi-level verification mechanism: The first level is syntax checking, which verifies whether the data format conforms to the specification, such as whether the invoice code conforms to the coding rules, and whether the amount value is within a reasonable range.

[0029] The second level is logical checking, which verifies data consistency through predefined business rules, such as verifying the matching relationship between input tax invoices and output tax invoices.

[0030] The third level uses machine learning anomaly detection, using the Isolation Forest algorithm to identify abnormal data points.

[0031] In the syntax checking stage, the system has a rich built-in verification rule library. For example, for value-added tax invoices, the checking rules include: the invoice code must be 10 digits, the invoice number must be 8 digits, the invoice date must conform to the date format, the amount must be greater than 0 and less than 1000 million, etc. These rules are defined through DSL (Domain Specific Language), supporting dynamic loading and hot updating. For data that fails verification, the system records detailed error information and transfers to the repair queue.

[0032] For numerical anomaly detection, a statistical distribution-based method is used. First, calculate the mean and standard deviation of historical data, then calculate the Z-score for the current data point:

[0033] Where X is the current data value, μ is the historical mean, and σ is the standard deviation. For example, when |Z|>3, it is determined to be an outlier, triggering the manual review process.

[0034] The system maintains a dynamic statistical indicator library, calculating statistics for each numerical field within a rolling time window. The size of the time window can be configured according to business needs, typically set to 30 days. Statistical calculations use a streaming algorithm, supporting incremental updates to avoid the overhead of full recalculation. For data with obvious seasonality, the system also supports using seasonal decomposition algorithms (such as STL) to eliminate the impact of seasonal factors.

[0035] The knowledge graph construction uses Neo4j graph database as the storage engine. During node creation, each entity is assigned a unique identifier, such as using the unified social credit code as the primary key for enterprise entities and using the combination of invoice code and invoice number as the primary key for invoice entities. Node attributes use a dynamic schema design, allowing flexible expansion of attribute fields according to different entity types.

[0036] In the entity resolution process, the system uses a multi-stage matching strategy: first using exact matching (such as complete matching of the unified social credit code), then using fuzzy matching (such as the edit distance similarity of enterprise names), and finally using machine learning-based entity linking algorithms. For cases where the matching confidence is below the threshold, an artificial review process is triggered. The entity resolution algorithm also considers the time dimension, capable of handling changes such as enterprise name changes and mergers and acquisitions.

[0037] Edge relationships are established based on association rules between entities. For example, when establishing the relationship between an invoice and the invoicing enterprise, the sales identification number on the invoice is matched with the unified social credit code of the enterprise entity. Edge attributes include relationship strength weights, with weight calculation using a normalization algorithm based on transaction frequency and amount:

[0038] where represents the number of transactions between entities i and j, represents the total transaction amount, and represent the maximum number of transactions and the maximum transaction amount in the system, respectively, and α and β are adjustment coefficients (default α = 0.4, β = 0.6).

[0039] The knowledge graph update uses an event-driven mechanism, when the source data changes, the system will automatically trigger the recalculation of the affected subgraph. To ensure performance, large-scale update operations use incremental calculation, only recalculating the affected part. The system also implements version management functions, allowing queries of the historical state of the knowledge graph, supporting time series analysis and backtracking queries.

[0040] The unstructured data processing unit deploys multi-modal analysis models. OCR recognition uses an end-to-end recognition architecture based on deep learning, which first preprocesses the image (binarization, noise reduction, skew correction), and then uses a CRNN network for text recognition. Natural language understanding uses a BERT model fine-tuned for training to extract key clause information from contract text.

[0041] The OCR processing flow includes multiple quality control steps: first, detect image quality and reject images that are blurry, too dark, or too bright; then perform layout analysis to identify tables, text areas, etc.; finally, perform text recognition and structured output. For areas with low recognition confidence, manual review is triggered. The NLP model is optimized for the finance and tax field and continues to be trained on finance and tax data to improve the accuracy of professional terminology and expression recognition.

[0042] This module builds an enterprise finance and tax knowledge graph with rich semantic relationships, enabling deep integration and semantic interconnection of multi-source data and providing a structured knowledge base for subsequent intelligent analysis.

[0043] 3. Execution mechanism of real-time analysis engine based on stream computing This module listens to real-time update events of the knowledge graph and performs real-time analysis of the changing subgraphs of the knowledge graph based on stream computing technology and AI models, enabling intelligent analysis functions such as risk early warning, tax planning, and compliance checking.

[0044] Detailed implementation process: The stream computing engine is built on the Apache Flink framework and uses event time semantics to process data. The water level mechanism is used to handle out-of-order events. Checkpoints are configured to automatically create every 30 seconds to ensure precise once state consistency.

[0045] The stream processing job is deployed using a distributed architecture, with each job containing multiple parallel instances. The keyBy operation ensures that data for the same taxpayer is routed to the same instance for processing. The RocksDB state backend is used for state management, supporting large state storage and fast recovery. The job monitoring interface displays real-time processing delay, throughput, back pressure, and other key indicators, facilitating performance tuning and fault diagnosis.

[0046] Model containers are deployed using Docker containerization, with each AI model running in an independent container environment. Containers communicate with each other using the gRPC protocol and use Protocol Buffers as the interface definition language. Model version management uses a blue-green deployment strategy, with new models first running in shadow mode to verify their effectiveness before gradually switching traffic.

[0047] The model container implements automatic scaling function, dynamically adjusts the number of container instances according to the request volume. Each container is equipped with resource monitoring and flow protection to avoid single model overload affecting the overall system. The model service provides a standardized prediction interface, with uniform input and output formats, facilitating the combination and replacement of different models.

[0048] The workflow of real-time invoice risk early warning model is as follows: (1) Listen to the node creation event of the knowledge graph, and trigger the analysis process when a new invoice node is detected; (2) Extract the associated subgraph of the invoice, including the nodes and edges of the invoicing enterprise, the receiving enterprise, and the related historical transactions; (3) Construct a feature vector for model input, including: The stock ownership correlation degree of the invoicing and receiving enterprises: calculated by the holding ratio and investment path length; Historical fund reflux similarity: calculate the similarity of fund flow time series using dynamic time warping algorithm; Tax registration status change frequency: statistics the change frequency of tax registration information in the past year.

[0049] (4) Use graph attention network for risk prediction:

[0050] Where, is the output feature vector of node i at the l+1 layer, is a nonlinear activation function such as ReLU or ELU, is the trainable weight matrix of the lth layer, is the attention coefficient, calculated by the softmax function, represents the neighbor node set.

[0051] (5) The risk probability output uses the sigmoid function:

[0052] Where, is the predicted risk probability value, which is a value between 0 and 1, for example, set > 0.85, generate risk interception instructions; is a natural constant, approximately equal to 2.71828; is the weight vector obtained by training; is the bias term obtained by training; is the dot product of the weight vector and the feature vector.

[0053] The feature engineering step uses an automatic feature generation technique. The system automatically extracts hundreds of potential features from the knowledge graph, and then selects the most relevant feature subset through feature importance ranking. The model training uses a combination of offline batch training and online incremental training. Offline training updates the full volume once a week, and online training absorbs new labeled data in real time. Model evaluation uses multiple indicators, including accuracy, recall, F1 score, and AUC value, to ensure comprehensive evaluation of model effectiveness.

[0054] The dynamic tax preference policy matching model crawls the tax bureau and policy release website every week, and uses a deep learning-based content extraction algorithm to obtain the policy text. The text embedding uses the Sentence-BERT model to generate a 384-dimensional semantic vector, and the policy and business scenario matching degree calculation uses cosine similarity:

[0055] Where A and B are the semantic vectors of the policy text and business scenario, respectively. For example, when the similarity exceeds 0.78, a policy matching event is triggered.

[0056] The policy crawling system realizes intelligent update detection by comparing the hash value changes of web page content. The policy analysis module can automatically identify the structure of the policy document and extract structured information such as policy title, issuing authority, issuing date, effective date, applicable object, and policy content. The policy matching considers multiple factors, including the industry, scale, region, and business characteristics of the enterprise, and the matching result is accompanied by a confidence score and explanation.

[0057] Real-time intelligent analysis of tax and financial business is realized, which can timely identify risk opportunities and policy matching points, and provide proactive decision support for enterprises.

[0058] 4. Collaborative work of strategy execution and feedback interface This module converts the analysis results into specific business operations and collects execution feedback for model optimization, forming a closed-loop learning mechanism.

[0059] Detailed implementation process: The strategy execution interface is deeply integrated with the existing business systems of the enterprise. For risk blocking scenarios, when a high-risk warning is received, the interface executes the blocking through the following process: (1) Generate a standardized blocking request message containing risk type, risk level, evidence summary, and other information; (2) Call the approval process interface of the OA system to create a pending approval task; (3) Send a warning notification through instant messaging platforms such as WeChat and DingTalk; (4) Record the blocking operation log, including operation time, execution personnel, processing result and other information.

[0060] The policy execution adopts a flexible blocking strategy, which takes different processing methods according to the risk level: for extremely high risk, it is directly automatically blocked and alarmed, for medium risk, it triggers manual review, for low risk, it only records logs and reports regularly. During the execution process, a compensation transaction mechanism is adopted to ensure that it can be rolled back or retried in case of system exception, avoiding state inconsistency. All execution operations generate audit logs, which are convenient for subsequent tracing and analysis.

[0061] One of the core functions of the feedback interface is to build a feedback loop from business operations to AI models. It encapsulates the "execution effect" of policy instructions (such as risk blocking) into structured "feedback data", including the manual annotation results of risk control personnel and the real business outcomes of the invoice. These data are returned to the corresponding "model container" in the "real-time analysis engine based on stream computing" through a unified data bus. After receiving this feedback data, the specific service in the engine triggers the online incremental learning process of the corresponding AI analysis model (such as the real-time invoice risk early warning model), so as to realize the automatic optimization and adaptive update of model parameters.

[0062] During the implementation of the feedback learning mechanism, the manual processing results are collected in the following ways: (1) Embed a feedback collection interface in the approval system, and the approval personnel can mark the processing result (confirm risk / false alarm); (2) Automatically collect subsequent business data, such as the final processing result of the blocked invoice; (3) Construct the training sample format: 〈feature vector, model prediction result, actual result〉.

[0063] After multiple verifications and cleanings, feedback data removes low-quality or contradictory annotations. The system calculates the consistency score between annotators to identify potentially problematic annotations. For important or controversial cases, an expert review process is initiated. Feedback data is divided into training and test sets in chronological order to avoid data leakage and bias in model evaluation.

[0064] Online incremental learning uses the FTRL optimization algorithm to update model weights:

[0065] where, is the updated model weight vector after the t+1 iteration, and is the solution to the optimization objective function; represents the weight vector w that minimizes the objective function value; is the cumulative sum of gradients from the first to the tth iteration, that is, ​Gradient of the s-th iteration Learning rate parameter for the s-th iteration, usually defined as where is the learning rate; L2 regularization term, used to control the difference between the current weight ww and the historical weight wsws, preventing the update from changing too much and maintaining the stability of the model; L1 regularization coefficient, used to produce sparse weights, making the model easier to interpret and possibly preventing overfitting; denotes the L1 norm of the weight vector w (sum of the absolute values of each element). A rolling time window is used in the model update process, retaining only the training data from the last 90 days.

[0066] Various techniques are employed in the incremental learning process to ensure model stability: a dynamic learning rate adjustment strategy that automatically reduces the learning rate when the model performance decreases; a gradient clipping technique to prevent gradient explosion; a model snapshot feature that allows quick rollback to previous versions. The system also regularly performs concept drift detection, triggering model retraining when there is a significant change in data distribution.

[0067] This module implements closed-loop management of analysis results to business operations and continuously optimizes the accuracy and practicality of the analysis model through continuous learning.

[0068] 5. Technical implementation of the blockchain storage module This module provides an unalterable audit trail for critical decision-making operations, enhancing the credibility and reliability of the system.

[0069] Detailed implementation process: The hash calculation uses the SHA-256 algorithm to calculate the digest value after serializing the risk event data packet. The data packet contains the following fields: event timestamp (ISO 8601 format), enterprise identifier, risk type code, decision basis summary, analysis model version number, and operator identifier.

[0070] Data serialization uses the Protocol Buffers format, which defines a strict data schema to ensure forward and backward compatibility. Before hash calculation, the data is standardized, including field ordering and encoding uniformity, to avoid different hash values for the same content. The calculated hash value is stored together with the source data for subsequent verification.

[0071] The storage process calls the Hyperledger Fabric SDK to write the hash value into the distributed ledger through the Chaincode interface. The writing process includes: (1) Construct a transaction proposal, including the parameters for calling the chaincode; (2) The endorsing node simulates the transaction and generates an endorsement response; (3) The transaction is submitted to the ordering service; (4) The transaction is packaged into a block and distributed to each ledger node; (5) The node verifies the validity of the transaction and updates the ledger.

[0072] The blockchain network adopts a consortium chain mode jointly maintained by multiple organizations, and each organization runs multiple nodes to ensure high availability. The smart contract implements basic functions such as record storage, query, and verification, and sets access control policies to protect sensitive information. The transaction throughput optimization uses channel partitioning technology, and different types of data are stored through different channels to improve parallel processing capability.

[0073] In the query and verification phase, users can query the record storage record through the transaction ID, and the system returns information such as block height, transaction time, and record storage hash value. When verifying, the data hash value is recalculated and compared with the hash value stored on the chain to ensure data integrity.

[0074] The verification service provides multiple interface forms, including Web interface, API interface, and command line tool, to meet the use needs of different scenarios. The verification result contains detailed credibility evaluation, including block chain confirmation number, signature situation of participating nodes, timestamp authority, and other indicators. For important verification operations, a printable verification report can be generated, containing all technical details and verification results.

[0075] This module establishes a trusted audit tracking mechanism, providing unalterable evidence support for key decisions, enhancing the reliability and credibility of the system.

[0076] System cooperation work example: Taking the processing of a newly issued value-added tax invoice as an example, the cooperative work flow of each module is shown: After the enterprise invoicing system generates a new invoice, the data access layer captures the invoice data in real time through the API adapter, converts it to a standard format, and sends it to the Kafka message queue; Specific example: A certain enterprise issues a value-added tax invoice with an amount of 100,000 yuan. The invoicing system generates JSON format invoice data and writes it to the API gateway at the moment of invoicing. The invoice adapter of the data access layer detects new data through the polling mechanism and immediately obtains the basic information of the invoice and the image storage path. After the adapter verifies that the data format is correct, it serializes the invoice data into Avro format, adds metadata header information (including data source, collection time, serial number, etc.), and then sends it to the Kafka topic named "invoice-input". The entire collection process is completed within 50 milliseconds.

[0077] The central processing module cleans the data after consuming the message, verifies the legality of the invoice code, and checks the reasonableness of the amount value. Through OCR recognition of the text information in the invoice image, the buyer and seller information, commodity details, and other data are extracted; Continuation example: A consumer instance of the central processing module obtains the invoice message from Kafka, first performs syntax checking: verifies that the invoice code conforms to the GB / T 14658-2018 standard, the invoice number is 8 digits, and the tax rate is within the correct range. Then perform logical checking: verify that the total amount of price and tax is equal to the amount plus the tax, and the buyer identification number conforms to the unified social credit code format. At the same time, trigger OCR processing: download the invoice image file, identify the text area through the CNN network, recognize the text content through the CRNN model, and extract entity information through the NLP model. It is found that the buyer's name matches the identification number, the commodity name matches the tax classification code, and all checks pass, and the knowledge graph construction phase is entered.

[0078] The knowledge graph construction module creates an invoice node, matches the unified social credit code to the corresponding enterprise node, and establishes a "issuing" relationship edge. At the same time, update the transaction frequency and amount statistics of the related enterprises; Continuation example: The knowledge graph engine first finds the enterprise node through the buyer identification number in the graph, finds the corresponding "Enterprise A" node. Find the sales enterprise node "Enterprise B" through the seller identification number. Create a new invoice node with attributes including invoice code, number, invoice date, amount, tax rate, etc. Create a "issuing" relationship to connect Enterprise B and the invoice node, with the direction from Enterprise B to the invoice; create a "receiving" relationship to connect the invoice node and Enterprise A, with the direction from the invoice to Enterprise A. Update the transaction relationship edge between Enterprise A and Enterprise B, increase the transaction amount, and recalculate the relationship weight. The entire process is completed in memory, and then batch written to the graph database to ensure operation atomicity.

[0079] The real-time analysis engine listens to knowledge graph change events and triggers the risk warning model. The model extracts the associated subgraph of the invoice and calculates the risk probability. At the same time, the policy matching model checks whether the invoice conforms to the latest tax preferential policy; Continuation example: the update event of the knowledge graph is published to the event bus, and the risk warning model listens to the invoice addition event. According to the pre-defined risk analysis rules, a multi-level correlation subgraph centered on the invoice is immediately extracted, including the invoicing enterprise, the invoice receiving enterprise, the shareholders of both parties, the top management, and other transaction partners. The feature extractor calculates a 128-dimensional feature vector, including: the equity correlation degree of the invoicing and invoice receiving enterprises (it is found that there is no direct equity relationship between the two parties), the historical transaction mode similarity (the similarity is 0.92 compared with the normal transaction mode), the tax registration state (the state of both parties is normal), etc. The graph neural network model infers that the risk probability is 0.12, which is lower than the threshold value 0.85, and determines that it is a low risk. At the same time, the policy matching model detects that the goods purchased by the invoice belong to energy-saving and environmental protection equipment, and matches to the value-added tax refund policy, generating a policy prompt information.

[0080] When a high risk is detected, the strategy execution interface calls the OA system interface to suspend the payment process related to the invoice, and sends a warning notification to the risk control personnel; Suppose another high-risk example: a certain invoice has a risk probability of 0.91 calculated by the model, which exceeds the threshold value. The strategy execution interface immediately generates a blocking request, calls the OA system through the REST API, and adds a “risk control audit” state to the invoice in the procurement payment process. At the same time, the warning notification is sent to the mobile terminal and desktop terminal of the risk control personnel through the message service, and the notification content includes the invoice basic information, the risk score, the main risk factors (such as “the seller is a new enterprise with less than 3 months of registration time”), the recommended handling method, etc. After receiving the blocking request, the OA system automatically suspends the payment approval process of the invoice, and waits for the risk control personnel to handle it.

[0081] After the risk control personnel completes the processing, the feedback interface collects the processing results for updating the training data of the risk warning model; Continuation of the high-risk example: after receiving the notification, the risk control personnel logs in to the system to view the invoice details and risk analysis report. After verification, it is confirmed that it is a high-risk invoice, and the “confirmed risk” is marked in the system and the handling measure “reject payment, suspend cooperation with the supplier” is recorded. The feedback interface collects this marking result to build a training sample (feature vector, prediction probability 0.91, actual label 1). At the same time, the subsequent business data is collected: the invoice is finally returned and re-opened, and the supplier is listed on the monitoring list. After desensitization, these data are added to the training data set for model incremental learning.

[0082] The key decisions and operations in the entire processing process are recorded by the blockchain storage module for non-tamperable storage; In the following example, the entire process of a high-risk invoice is recorded as an audit event, including risk detection time, risk score, model version, manual review result, and key information such as executed measures. The blockchain evidence storage module serializes this information and calculates the SHA-256 hash value, then calls the Fabric SDK to initiate the evidence storage transaction. After endorsement, ordering, and verification, the transaction is written to the block, and the transaction ID "txn-123456" is returned. The system stores the transaction ID together with the original data in the database, completing the evidence storage process. During subsequent audits, auditors can query the evidence storage record on the blockchain using the transaction ID to verify the integrity and authenticity of the decision-making process at the time.

[0083] This complete process flow realizes full automation from data collection to decision execution, reflecting the efficient collaborative work capability of each module.

Claims

1. A data analysis-based enterprise finance and tax intelligent management system, characterized in that, Comprise: A multi-source heterogeneous data real-time access layer, which uses a configurable adapter interface to collect structured and unstructured financial and tax-related data in real time from enterprise business operation systems, tax invoicing systems, bank payment systems, and external databases; A central data processing and knowledge graph construction module connected to the data real-time access layer, which is used to clean, fuse, and semantically annotate the incoming data, and dynamically construct and update the enterprise financial and tax business knowledge graph based on the processed data; the knowledge graph takes enterprise entities, invoices, tax items, contracts, and bank accounts as nodes, and transactions, ownership, and correlation as edges, at least some of which are accompanied by timestamp and amount weight attributes; A real-time analysis engine based on stream computing, which has a pluggable model container for loading and running multiple microservice AI analysis models; the analysis engine listens to real-time update events of the knowledge graph and calls corresponding AI analysis models to perform instant calculations on the changed subgraph; A policy execution and feedback interface for receiving instructions output by the real-time analysis engine, including risk blocking, planning suggestion prompts, or compliance verification results, and returning the execution effect of the instructions as feedback data to the corresponding AI analysis model for online self-learning optimization of the model. 2.The enterprise financial and tax intelligent management system based on data analysis of claim 1, wherein, The central data processing and knowledge graph construction module also has: An unstructured data processing unit integrated with an OCR recognition model and a natural language understanding model, which is used to extract key entities and relationship triples from scanned contract and invoice images and automatically map them to the corresponding nodes and edges of the knowledge graph. 3.The enterprise financial and tax intelligent management system based on data analysis of claim 1, characterized in that, The real-time analysis engine based on stream computing also has: A real-time invoice risk early warning model for traversing the associated upstream and downstream enterprise nodes and historical transaction paths of each new invoice as it is entered into the graph, calculating the risk probability of the invoice existing in a fake tax evasion chain through a graph neural network algorithm, and generating real-time interception instructions for risk invoices exceeding a preset threshold to the policy execution and feedback interface.

4. The enterprise finance and tax intelligent management system based on data analysis according to claim 3, characterized in that, The judgment basis of the real-time invoice risk early warning model is: The depth of equity association between the invoicing enterprise and the recipient enterprise in the graph, the similarity of historical capital reflux paths, and tax registration status change events, and comprehensive risk probability assessment.

5. The enterprise finance and tax intelligent management system based on data analysis according to claim 1, characterized in that, The real-time analysis engine based on stream computing also includes: A dynamic tax preferential policy matching model that continuously collects and analyzes data from official policy and regulation websites, automatically matches new policies with business scenarios in the enterprise knowledge graph using a text semantic similarity model, and actively pushes policy applicability analysis and measurement reports.

6. The enterprise finance and tax intelligent management system based on data analysis according to claim 1, characterized in that, The system also includes: A blockchain anchoring evidence module for generating a hash value of key risk early warning events, policy execution instructions, and their corresponding knowledge graph data snapshots issued by the real-time analysis engine and uploading them to the blockchain to ensure the auditability and non-tamperability of the entire link analysis decision.

7. The enterprise finance and tax intelligent management system based on data analysis according to claim 1, characterized in that, The policy execution and feedback interface is specifically used for: In the procurement approval process, if a "high risk" instruction is received for a certain invoice, the payment process is automatically interrupted and a work order is pushed to the risk control personnel; Meanwhile, the final processing result of the risk control personnel is recorded, and the result is fed back to the real-time invoice risk early warning model as labeled training data for online incremental learning of the model. 8.The enterprise finance and tax intelligent management system based on data analysis of claim 1, wherein, When constructing the relationship, the knowledge graph construction module: Preferentially, the node fusion and relationship connection are based on the unique social credit code of the enterprise and the invoice number. For data that cannot be directly associated, a two-stage fuzzy matching algorithm based on rules and machine learning is used for entity alignment to ensure the accuracy and integrity of the graph data.

Citation Information

Cited By

  • Invoice automatic registration system and method based on multi-factor authentication

    CN121258718A

  • Invoice automatic registration system and method based on multi-factor authentication

    CN121258718B

  • Financial information management system and method based on big data analysis

    CN121304364A

  • Business whole-process generation system and method based on enterprise knowledge graph enhancement model

    CN121680791A

  • Intelligent data fusion system and method for construction project

    CN121935863A