Enterprise finance and tax management method and system based on big data

Through big data processing and dynamic graph structure analysis, corporate financial and tax management methods improve the efficiency of risk identification and resource utilization, provide clear risk assessment, solve the problem of insufficient risk identification in existing technologies, and realize the insight into the laws of risk propagation and intelligent allocation of resources.

CN121120287APending Publication Date: 2025-12-12ANHUI DINGXIN PROJECT MANAGEMENT LTD BY SHARE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511336298.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

In corporate financial and tax management, the efficiency of risk identification is low, and the pattern of risk propagation cannot be discerned. This leads to a passive approach after risks spread, with resources being scattered and invested in non-critical areas, making timely control impossible.

Method used

The big data-based enterprise financial and tax management method involves receiving multi-source data, standardizing it, using multi-pattern matching algorithms to identify risks, constructing a dynamic graph structure network, applying a risk propagation model for analysis, and allocating computing resources based on business load and risk outcomes.

Benefits of technology

It improves the efficiency and coverage of risk identification, ensures the timeliness and accuracy of key risk analysis, optimizes resource utilization efficiency, provides clear risk assessment reports, and helps enterprises develop scientific risk response strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120287A_ABST
    Figure CN121120287A_ABST
Patent Text Reader

Abstract

The invention provides an enterprise finance and taxation management method and system based on big data, relates to the field of finance and taxation management, and solves the technical problem of low risk identification efficiency in existing enterprise finance and taxation management. The method comprises the following steps: receiving multisource enterprise finance and taxation original data, and carrying out finance and taxation compliance verification and business logic calculation on the preprocessed original data to generate standardized finance and taxation data; performing risk identification on the standardized finance and tax data by using a multi-mode matching algorithm to generate risk identification data; constructing an enterprise finance and taxation relation network by adopting a dynamic graph structure, and performing financial risk analysis on the enterprise finance and taxation relation network based on the risk propagation model; and allocating computing resources according to the service load data and the risk analysis result, and outputting a risk assessment report.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of financial tax management, and particularly relates to an enterprise financial tax management method and system based on big data. BACKGROUND

[0002] In the field of enterprise financial tax management, with the promotion of digital business, the sources of enterprise financial tax data are increasingly diversified. Due to the characteristics of strong concealment and deep correlation with multiple business links of enterprise financial tax risks, in early screening, it is easy to cause inaccurate risk identification due to heterogeneous formats of multi-source data, lack of financial tax scenario compliance verification, and only focus on single risk data or independent enterprises, without integrating the business relationships of related subjects such as enterprise and transaction counterparties, shareholders, etc., unable to accurately construct the risk propagation path, and unable to predict the diffusion trend and potential impact range of risks in the associated network in advance, thus leading the enterprise to fall into a passive dilemma in financial tax risk prevention and control: either discovering the core risk point after the risk has spread to multiple cooperative enterprises through transaction, equity and other associated relationships, causing chain compliance problems, missing the opportunity to block early; or because the impact boundary of the risk cannot be clearly defined, the prevention and control resources are scattered into non-critical links, and the real high-risk associated transaction chain cannot be timely controlled due to lack of resources, and finally may face the consequences of punishment, financial loss, etc., increasing the compliance pressure and operation burden of enterprise financial tax management. SUMMARY

[0003] The application provides an enterprise financial tax management method and system based on big data, which solves the technical problems of low risk identification efficiency and inability to understand the risk propagation law in the prior art.

[0004] To achieve the above purpose, the application adopts the following technical solutions: In a first aspect, an enterprise financial tax management method based on big data is provided, comprising: receiving multi-source enterprise financial tax original data, and performing financial tax compliance verification and business logic calculation on the preprocessed original data to generate standardized financial tax data; using a multi-mode matching algorithm to identify risks from the standardized financial tax data, and generating risk identification data; constructing an enterprise financial tax relationship network using a dynamic graph structure, and performing financial risk analysis on the enterprise financial tax relationship network based on a risk propagation model; the input data of the risk propagation model is the risk identification data; allocating computing resources according to business load data and risk analysis results, and outputting a risk assessment report.

[0005] Based on the above technical scheme, in the enterprise financial tax management method based on big data provided in the application, the data quality and consistency are ensured through the standardized data processing process; the automatic multi-dimensional risk scanning of massive financial tax data is realized by using the multi-mode matching algorithm, which can improve the efficiency and coverage of risk identification; by constructing a dynamic graph relationship network and applying a risk propagation model, the transmission path and systematic influence of risk in the enterprise correlation network can be deeply mined, and finally the business load and risk analysis results are used as the core basis for computing resource allocation, so that the system resources can be intelligently and dynamically tilted to high-risk and high-value data processing tasks, not only improving the overall resource utilization efficiency of the system, but also ensuring the timeliness and accuracy of key risk analysis tasks, helping enterprises to timely develop risk response strategies.

[0006] Further, the generation process of the standardized financial tax data comprises: receiving multi-source original data from enterprise resource planning systems, tax declaration platforms, bank reconciliation systems, and electronic invoice systems; performing data cleaning on missing values, abnormal values, and duplicate data in the multi-source original data; performing financial compliance verification on the cleaned data through a rule engine; the financial compliance verification includes invoice authenticity, tax rate compliance, and accounting standards compliance; converting the verified data into a unified JSON-LD format, adding data bloodline metadata and timestamps, and obtaining standardized financial tax data; the data bloodline metadata represents complete traceability information of the data, including but not limited to data source system, data collection time, data processing process, and association with other data.

[0007] Further, the rule engine represents a dynamically configurable automated business rule execution system, comprising a rule base, an inference engine, and a rule management interface; wherein, the rule base stores various financial tax business rules, including tax law provisions, accounting standards, and enterprise control rules; the inference engine is responsible for matching and executing corresponding rules according to input data; the rule management interface is used to define, modify, and deploy specific verification rules in a visual manner.

[0008] Further, the risk identification of the standardized financial tax data using the multi-mode matching algorithm comprises: performing feature extraction on the standardized financial tax data to obtain a key feature vector; the feature extraction uses a principal component analysis-based method; using a local sensitive hashing algorithm to map the key feature vector into a hash bucket and compare it with a pre-set risk pattern hash bucket, and output a first risk set; inputting the first risk set into an abnormal rule engine based on a Drools framework, executing a predefined IF-THEN rule chain, and outputting a second risk set; transforming the number of triggered abnormal rules in the second risk set into a rule confidence score according to an established risk confidence fusion mechanism; performing abnormality detection on the standardized financial and tax data and outputting an abnormality score for each data; calculating the rule confidence score and the abnormality score by using a weighted summation method to obtain a comprehensive risk score; labeling the standardized financial and tax data with a comprehensive risk score greater than a preset risk threshold as risk identification data.

[0009] Further, the abnormal rules include invoice consecutive number detection rules, tax number blacklist matching rules, three-in-one verification rules, tax rate abnormality detection rules, and associated transaction cycle detection rules. The three-in-one verification rules represent rules for verifying the consistency of invoice flow, fund flow, and cargo flow, including verifying the consistency of invoice issuers and payees, the consistency of invoice content and actual goods or services, and the consistency of payment direction and amount, to ensure that the transaction background is true and reliable. The associated transaction cycle detection rules represent rules for identifying associated party circular transaction behaviors hidden by complex transaction structures, specifically including detecting patterns of fund closed-loop circulation formed by multiple associated enterprises in a short period of time, identifying abnormal patterns of artificially increasing revenue or cost through multi-layer transactions, and discovering irregular behaviors of tax planning through associated transactions.

[0010] Further, the risk confidence fusion mechanism includes: generating a rule confidence score based on the statistical results of rule triggering events. The generation process of the rule confidence score adopts a hierarchical accumulation strategy, increments the calculation by assigning a basic confidence score value and according to the number of triggered abnormal rules, and sets an upper threshold to control the score range to obtain the rule confidence score.

[0011] Further, the mathematical expression of the risk confidence fusion mechanism is: rule =min(1.0,s0+0.1×(N rules -1)); where S rule represents the rule confidence score, s0 represents the basic confidence score, and N rules represents the number of triggered abnormal rules.

[0012] Further, the abnormality detection on the standardized financial and tax data includes: Randomly select ψ sample subsets from the standardized financial and tax data as training data, and construct t isolated trees to form an isolated forest, wherein each isolated tree is divided by recursively randomly selecting the key feature vector and the split value of the sample space until the subspace contains only one sample or reaches the tree height limit; the split value is a division preset randomly generated within the current data range of the selected feature vector, and the tree height limit is a preset parameter representing the maximum growth depth allowed for a single isolated tree, used to control the complexity of the model; For the standardized financial and tax data point x to be detected, traverse each isolated tree to calculate the path length h(x) of x from the root node to the leaf node; According to the path length, calculate the average path length E(h(x)) and the anomaly score of the data point x; the calculation formula of the anomaly score s(x,n) is: s(x,n)=2 -E(h(x)) / c(n) ; wherein c(n) is the standardized path length under a given number of standardized financial and tax data samples n, used to correct the reference value of the path length.

[0013] Further, the enterprise financial and tax relationship network is constructed using a dynamic graph structure, comprising: Extracting enterprise entities, legal representatives, shareholder structures, transaction counterparties, and capital flow elements from standardized financial and tax data as graph nodes of the enterprise financial and tax relationship network; Based on equity relationships, transaction relationships, capital flow relationships, and personnel association relationships, the edge structure of the enterprise financial and tax relationship network is constructed; Using a graph database to store the dynamic changes of the enterprise financial and tax relationship network, and updating the edge weight in real time according to the transaction amount and transaction frequency.

[0014] Further, the financial risk analysis of the enterprise financial and tax relationship network based on the risk propagation model comprises: Mapping the risk identification data to the graph database nodes, initializing the node risk value R i (0) according to the comprehensive risk score; Using an improved SIR model of infectious diseases to simulate the risk propagation of the enterprise financial and tax relationship network, the risk value of node i at time t is calculated as: ; wherein β represents the propagation coefficient (default value is 0.65), γ represents the recovery coefficient, w ij is the edge weight, N(i) represents the neighbor set of node i, and i, j represent node indices; Based on the risk value, using a path search algorithm to calculate the shortest path of risk propagation, and outputting multiple risk propagation paths.

[0015] Further, the allocation of computing resources according to the business load data and the risk analysis result comprises: From the risk analysis result, nodes with a node risk value exceeding a high risk threshold in the enterprise financial tax relationship network are screened to obtain a high-risk task queue; The output risk propagation path is analyzed, and nodes with a common occurrence or an out-degree order in the path in the front m are extracted to obtain a key node queue; Using a time series analysis model, the total business load in a future preset time period is predicted according to historical business load data; the time series analysis model is constructed based on a deep learning algorithm; A preset proportion of resources is divided from the total load to form an elastic resource pool; Nodes that belong to both the high-risk task queue and the key node queue are marked as a first priority; Nodes that only belong to the high-risk task queue are marked as a second priority; Nodes that only belong to the key node queue are marked as a third priority; The remaining nodes are marked as a fourth priority; According to the priority of the nodes and the total load, the computing resources are hierarchically divided and allocated; When the backlog of nodes of the first priority or the second priority exceeds a preset number, resources are automatically allocated from the elastic resource pool.

[0016] In a second aspect, an enterprise financial tax management apparatus is provided, comprising a communication unit and a processing unit; wherein, The communication unit is configured to receive multi-source enterprise financial tax original data, and is further configured to output a finally generated risk assessment report, and realize data interaction with an external data provider or a financial tax management demander; The processing unit is configured to pre-process the multi-source enterprise financial tax original data received by the communication unit, and further configured to perform financial tax compliance verification and business logic calculation on the pre-processed original data to generate standardized financial tax data; The multi-mode matching algorithm is used to identify risks of the standardized financial tax data to generate risk identification data; A dynamic graph structure is used to construct an enterprise financial tax relationship network, and a financial risk analysis is performed on the enterprise financial tax relationship network based on a risk propagation model; the input of the risk propagation model is the risk identification data; The computing resources are allocated according to the business load data and the risk analysis result.

[0017] In a third aspect, an enterprise financial tax management apparatus is provided, comprising a processor and a storage medium; the storage medium comprises instructions, and the processor is configured to execute the instructions to implement the method described in the first aspect and any possible implementation manner of the first aspect. The enterprise financial tax management apparatus can be an electronic device or a chip in an electronic device.

[0018] In a fourth aspect, the present application provides an enterprise financial tax management system based on big data, comprising a data processing module, a risk identification module and a resource allocation module, wherein The data processing module is configured to receive multi-source enterprise financial tax original data, perform preprocessing on the original data, and then perform financial tax compliance verification and business logic calculation to generate standardized financial tax data. The risk identification module is configured to use a multi-mode matching algorithm to identify risks in the standardized financial tax data and generate risk identification data. The resource allocation module is configured to construct an enterprise financial tax relationship network using a dynamic graph structure, perform financial risk analysis on the enterprise financial tax relationship network based on a risk propagation model taking the risk identification data as input, allocate computing resources according to business load data and risk analysis results, and output a risk assessment report.

[0019] In a fifth aspect, the present application provides a computer-readable storage medium having instructions stored therein, which, when executed on an enterprise financial tax management device, causes the enterprise financial tax management device to perform the method described in the first aspect and any possible implementation manner of the first aspect.

[0020] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on an enterprise financial tax management device, causes the enterprise financial tax management device to perform the method described in the first aspect and any possible implementation manner of the first aspect.

[0021] The present application provides an enterprise financial tax management method and system based on big data. Firstly, through global data standardization management, a unified and reliable data foundation is constructed. Using a multi-dimensional risk penetration identification mechanism, rule matching and intelligent algorithms are integrated to improve the coverage and accuracy of financial tax risk identification. By constructing a dynamic enterprise relationship graph and applying a risk propagation model, systematic analysis from isolated risk points to global risk networks is achieved, which can deeply trace the risk transmission path and locate the key hub, providing a clear financial tax situation for enterprises. In addition, the present application introduces an intelligent resource dynamic scheduling mechanism, taking business load prediction and risk level as dual driving, ensuring that computing resources can be preferentially and efficiently served for high-risk disposal tasks, thereby optimizing the resource utilization efficiency and system response capability as a whole. Finally, a quantitative assessment report is output to provide a comprehensive risk view for enterprises and help managers make scientific risk response decisions.

[0022] It should be understood that the description of technical features, technical solutions, advantages or similar language in this application does not imply that all features and advantages can be achieved in any single embodiment. On the contrary, it can be understood that the description of features or advantages means that the specific technical features, technical solutions or advantages are included in at least one embodiment. Therefore, the description of technical features, technical solutions or advantages in this specification does not necessarily refer to the same embodiment. Further, the technical features, technical solutions and advantages described in this embodiment can also be combined in any appropriate manner. Those skilled in the art will understand that the embodiments can be implemented without one or more specific technical features, technical solutions or advantages of the specific embodiments. In other embodiments, additional technical features and advantages can be identified in specific embodiments without embodying all embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings required in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings from these drawings without creating labor.

[0024] Figure 1 A system architecture diagram of an enterprise financial and tax management system based on big data is provided for the embodiments of the present application; Figure 2 A flowchart of an enterprise financial and tax management method based on big data is provided for the embodiments of the present application; Figure 3 A flowchart of another enterprise financial and tax management method based on big data is provided for the embodiments of the present application; Figure 4 A flowchart of another enterprise financial and tax management method based on big data is provided for the embodiments of the present application; Figure 5 A structural diagram of an enterprise financial and tax management device is provided for the embodiments of the present application; Figure 6 A hardware structural diagram of an enterprise financial and tax management device is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0025] In the description of the present application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" herein is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can mean that A exists alone, A and B exist together, and B exists alone. In addition, "at least one" means one or more, and "multiple" means two or more. "First", "second", etc. do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.

[0026] It should be noted that in the present application, "exemplary" or "for example" means to serve as an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of "exemplary" or "for example" is intended to present the relevant concept in a specific manner.

[0027] The enterprise financial tax management method based on big data provided by the embodiments of the present application can be applied to an enterprise financial tax management system based on big data as shown in Figure 1 as shown in Figure 1 The communication system includes a data processing module, a risk identification module, and a resource allocation module. Wherein, The data processing module is configured to receive multi-source enterprise financial tax raw data, perform pre-processing on the raw data, and then perform financial tax compliance verification and business logic calculation to generate standardized financial tax data. The risk identification module is configured to use a multi-mode matching algorithm to identify risks in the standardized financial tax data and generate risk identification data. The resource allocation module is configured to construct an enterprise financial tax relationship network using a dynamic graph structure, perform financial risk analysis on the enterprise financial tax relationship network based on a risk propagation model taking the risk identification data as input, allocate computing resources according to the business load data and the risk analysis results, and output a risk assessment report.

[0028] To solve the technical problems of low risk identification efficiency and inability to understand the risk propagation law in the existing enterprise financial tax management, the embodiments of the present application provide an enterprise financial tax management method and system based on big data, which includes: Receiving multi-source enterprise financial tax raw data, and performing financial tax compliance verification and business logic calculation on the pre-processed raw data to generate standardized financial tax data; Using a multi-mode matching algorithm to identify risks in the standardized financial tax data and generating risk identification data; The enterprise financial tax relationship network is constructed by using a dynamic graph structure, and the enterprise financial tax relationship network is analyzed for financial risk based on a risk propagation model. The computing resources are allocated according to the business load data and the risk analysis result, and a risk assessment report is output.

[0029] Based on this, the method unifies the data base through standardization processing, improves the risk identification ability with multi-mode matching, digs the risk association based on the dynamic network and the propagation model, and optimizes the resource allocation according to the load and the risk, so that the standardization of enterprise financial tax management, the risk control efficiency and the rationality of resource utilization can be improved.

[0030] As shown in Figure 2 The enterprise financial tax management method based on big data provided by the embodiments of the application comprises: S1, receiving multi-source enterprise financial tax original data, and performing financial tax compliance verification and business logic calculation on the preprocessed original data to generate standardized financial tax data.

[0031] The multi-source enterprise financial tax original data includes financial voucher data generated in daily business operations, tax declaration related data, bank fund flow data, electronic invoice data, and financial tax related data from enterprise internal management systems and external cooperation platforms; the standardized financial tax data refers to financial tax data that has been processed to have uniform format, compliant content, and can be directly used for subsequent analysis, and the core features are data format consistency, data content compliance and data availability.

[0032] In some implementations, the preprocessing operation of the original data includes data cleaning, data conversion and data integration; the financial tax compliance verification refers to checking whether the original data meets the requirements of national financial tax laws and regulations, industry accounting standards and enterprise internal financial tax management system, which is specifically implemented by means of preset compliance verification rules, simple rule judgment module or manual sampling verification; the business logic calculation refers to calculating and processing the data according to the basic logic of financial tax business to obtain intermediate or final data that meets the needs of business scenarios, which is specifically implemented by means of automatic calculation script, calculation function of common office software such as Excel or simple business calculation module.

[0033] For example, voucher data from an enterprise internal financial system and declaration data from an external tax platform are received, duplicate records in the voucher data are deleted during preprocessing, and missing amounts in the bank flow data are filled with median values; during compliance verification, it is checked whether the invoice code format in the invoice data meets the requirements of the State Administration of Taxation; during business logic calculation, the amount of tax to be paid is calculated according to the invoice amount and the corresponding tax rate, and finally these data are converted into CSV (Comma-Separated Values) format to generate standardized financial tax data.

[0034] S2, risk identification is performed on the standardized financial and tax data by using a multi-mode matching algorithm to generate risk identification data.

[0035] The multi-mode matching algorithm is used to simultaneously match multiple risk modes for the standardized financial and tax data, find data records in the data that are consistent with or similar to preset risk modes (such as abnormal invoices, abnormal tax amounts, and abnormal transactions), and obtain risk identification data with risk labels. The core purpose is to cover the identification needs of multiple risk scenarios at one time and improve the risk identification efficiency. The preset risk modes include but are not limited to abnormal invoice modes, abnormal tax amount modes, and abnormal transaction modes.

[0036] In some implementations, the multi-mode matching algorithm can use a string matching algorithm (such as the AC automatic machine algorithm, the KMP algorithm (Knuth-Morris-Pratt Algorithm), a rule-based multi-mode comparison method, or a simple multi-condition matching module). The preset risk modes can be obtained by manually analyzing common financial and tax risk scenarios and defining risk characteristics to form a risk mode library, and then comparing the standardized financial and tax data with the modes in the risk mode library one by one to mark the data with a matching degree reaching a preset threshold as risk candidate data, and finally generating risk identification data after manual review or simple risk confirmation rule screening.

[0037] For example, the AC automatic machine algorithm is used as the multi-mode matching algorithm, and the preset risk modes include “invoice amount exceeding 1 million yuan without corresponding goods details” and “tax amount calculation error for the same transaction”. The standardized invoice data is input into the AC automatic machine, and the algorithm automatically matches the invoice records that meet the above risk modes. After manual review, these records are marked with “high risk” to generate risk identification data.

[0038] S3, a dynamic graph structure is used to construct an enterprise financial and tax relationship network, and a financial risk analysis is performed on the enterprise financial and tax relationship network based on a risk propagation model.

[0039] The enterprise financial and tax relationship network is a network structure that presents the financial and tax association relationship between an enterprise and associated subjects (such as other enterprises, legal persons, and shareholders) in a graphical manner. It usually uses nodes to represent enterprises and associated subjects and edges to represent transaction relationships, financial relationships, equity relationships, and other financial and tax associations between the two. It is used to represent the financial and tax interaction between enterprises and associated subjects and reflect the association status and association strength of an enterprise in the financial and tax network.

[0040] In some implementations, when constructing a corporate financial and tax relationship network using a dynamic graph structure, the information of nodes and edges can be stored using existing graph storage technologies (such as simple graph structure data tables and basic graph database functional modules). The node attributes and edge attributes are updated in real time according to the updates to corporate financial and tax data (such as new transaction records and equity change records). When conducting financial risk analysis on the corporate financial and tax relationship network based on a risk propagation model, common simple propagation models can be selected, such as risk diffusion models based on fixed propagation coefficients or simplified versions of infectious disease models. The risk information in the risk identification data is used as the initial risk source. After inputting into the risk propagation model, the spread range, spread speed, and impact of the risk among network nodes are calculated, thereby realizing financial risk analysis.

[0041] For example, using Company A, Company B, and Company A's legal representative C as network nodes, and the transaction relationship between Company A and Company B, and the equity relationship between Company A and C as edges, an initial corporate financial and tax relationship network is constructed. When Company B is marked as a risk node, a risk diffusion model is used to analyze the possibility and degree of impact of Company B's risk spreading to Company A, thus completing the financial risk analysis.

[0042] S4. Allocate computing resources based on business load data and risk analysis results, and output a risk assessment report.

[0043] The allocation of computing resources based on business load data and risk assessment results is intended to achieve intelligent optimization and precise scheduling of computing resources, ensuring that the system can dynamically respond to changes in business needs and prioritize the processing capacity of high-risk tasks.

[0044] In some implementations, computing resource allocation can employ priority-based resource scheduling strategies and load balancing techniques. Risk assessment reports can be generated using existing report generation tools or by manually compiling and documenting the analysis results. The report should include at least risk node information, risk diffusion analysis results, risk level assessment, and simple risk response recommendations. Priority-based resource scheduling strategies typically categorize risk levels in the risk analysis results into high, medium, and low levels, with high-risk tasks receiving higher resource allocation priority. Load balancing techniques typically allocate computing resources (such as server CPU, memory, and storage space) to different tasks based on business load data (e.g., the number of tasks currently being processed by the system, CPU utilization, and memory usage), preventing any single task from consuming excessive resources.

[0045] For example, the business load data shows that the current system CPU usage is 40%, and there is idle resource; the risk analysis result shows that enterprise A is a high-risk node and enterprise D is a low-risk node; according to the priority strategy, 60% of the idle CPU resources are allocated to the risk deep analysis task of enterprise A, and 20% are allocated to the regular risk check task of enterprise D; at the same time, a risk assessment report is generated through Excel, which lists the risk types, risk impact range and the suggestion of "prioritizing the recent large transaction of enterprise A for verification".

[0046] Based on the above technical solution, the enterprise financial tax management method based on big data provided in the application solves the problems of data format confusion and low availability in traditional financial tax management through standardized processing of multi-source original data, guarantees data quality and consistency, realizes synchronous identification of multiple risks by using a multi-mode matching algorithm, improves the efficiency and coverage of risk identification, analyzes the conduction of risks in the enterprise correlation network with the help of a dynamic graph structure and a risk propagation model, allocates computing resources in combination with business load and risk results, avoids waste of resources and lag in response to key tasks, optimizes the efficiency of system resource utilization, and finally outputs a risk assessment report to provide clear risk information and decision-making reference for enterprises, which can improve the scientific nature and risk resistance of enterprise financial tax management.

[0047] In a possible implementation manner of the embodiment of the application, the S1 can be implemented through the following S101, S102 and S103, which are specifically described as follows: S101, receiving enterprise financial tax original data transmitted from a multi-source data providing system.

[0048] The multi-source data providing system includes an enterprise resource planning system (ERP), a tax declaration platform, a bank reconciliation system and an electronic invoice system, and the received enterprise financial tax original data covers invoice data of enterprise procurement and sales links, bank fund income and expenditure flow data, monthly and quarterly tax declaration form data, cost accounting data and asset liability data recorded in the ERP system and other core financial tax information.

[0049] In some implementations, to ensure the integrity and timeliness of multi-source raw data reception, an asynchronous data receiving mechanism based on message queue (MQ) can be used, an independent message channel is configured for each data providing system, and data sharding verification rules are set at the data receiving end, that is, raw data files exceeding a preset size (such as 100 MB) are transmitted in slices, each slice of data carries a unique slice identifier and total slice number information, the receiving end splices all data pieces matching the slice identifier, and verifies the consistency of the total slice number and the actual received slice number, if there is a missing, a data retransmission request is triggered; at the same time, to avoid tampering during data transmission, a hash value based on SHA-256 algorithm can be calculated for the raw data before transmission, the receiving end recalculates the hash value after receiving the data and compares it with the hash value provided by the sending end, and if they are consistent, the reception is confirmed to be valid.

[0050] It should be noted that the receiving range of the multi-source raw data is not fixed and can be extended according to the actual financial and tax management needs of the enterprise, for example, when the enterprise adds a supply chain management system (SCM), the interface protocol (such as RESTful API) between the system and the receiving end can be configured to include the purchase order data and supplier settlement data in the SCM system into the receiving range, and the extension process does not need to modify the core receiving logic, only the data source configuration parameters of the corresponding system need to be added.

[0051] For example, the enterprise production cost detail data in March 2024 (file size 120 MB, divided into 3 data slices) from the ERP system, the value-added tax declaration form data of the first quarter of 2024 from the tax declaration platform, the basic deposit account transaction data of the enterprise in March 2024 from the bank reconciliation system, and all value-added tax special invoices and ordinary invoices data issued in March 2024 from the electronic invoice system; through 3 independent channels of the MQ message queue, the data of each system is received, the receiving end matches the 3 data slices of the ERP system, the slice identifiers are "ERP-202403-Cost-01", "ERP-202403-Cost-02" and "ERP-202403-Cost-03", the total slice number "3" is verified to be consistent with the actual received number, and then the data is spliced, and the SHA-256 hash value is compared to confirm that all multi-source raw data reception is valid.

[0052] S102, sequentially performing data preprocessing and financial and tax compliance verification on the received multi-source raw data, and executing business logic calculation.

[0053] Among them, the data preprocessing includes targeted processing of missing values, abnormal values and repeated data in the original data; the financial and tax compliance verification is performed through a rule engine, and the verification content covers invoice authenticity, tax rate compliance and accounting standards compliance; the business logic calculation is a numerical operation and logical derivation on the verified data based on financial and tax business rules to obtain intermediate data results that meet the business scenario.

[0054] In some implementations, the specific operations of data preprocessing are as follows: For missing values, if they are numerical data, they are filled based on the weighted calculation of the historical same period data mean and median of the business scenario to which the data belongs; if they are character data, they are completed by associating other data in the same business flow, and those that cannot be associated are marked as “to be manually verified”.

[0055] For abnormal values, a combination of business threshold and statistical method is used to determine the abnormal values—first, a business threshold is set, such as a single invoice amount upper limit of 5 million yuan, and data exceeding the threshold is preliminarily determined as abnormal; then, for data not exceeding the business threshold, the 3σ principle (σ is the standard deviation of the data) is used for further screening, and data falling outside the range of [mean-3σ, mean+3σ] is determined as abnormal. Abnormal data is uniformly marked and the reason for the abnormality is recorded, such as “exceeding the single invoice amount upper limit” and “deviating from the mean by more than 3σ”.

[0056] For repeated data, comparison is made by constructing a unique data identifier, and data with the same unique identifier is determined as repeated. The earliest received or the data with the highest integrity is retained, and the rest of the repeated data is archived and stored. The unique data identifier can be a combination of “invoice code + invoice number + date of issuance” of an invoice, a combination of “transaction date + transaction amount + transaction counterparty account number” of a bank transaction, etc.

[0057] In some implementations, the rule engine used for financial and tax compliance verification is an automated business rule execution system that can be dynamically configured, which specifically includes a rule base, an inference engine and a rule management interface. The rule base stores a variety of financial and tax business rules, specifically covering tax law provisions, accounting standards and enterprise control rules, and these rules need to be dynamically updated according to the latest financial and tax policies. The rule management interface supports defining, modifying and deploying specific verification rules through a visual way. For example, when the State Administration of Taxation adjusts the value-added tax rate, staff can import the verification rules corresponding to the new tax rate through the “batch rule update” function of the interface. The inference engine is responsible for automatically matching the corresponding rules in the rule base according to the business attributes of the input data (such as invoice issuance time and business type), and performing verification operations.

[0058] The business logic calculation specifically includes: calculating tax amount according to invoice amount and corresponding tax rate, summarizing monthly net inflow / outflow amount according to the type of bank flow, calculating gross profit rate according to cost data and sales data in the ERP system, and recording the basis for calculation, such as the source of tax rate, the scope of cost collection, etc.

[0059] S103, transforming the data passed by the financial and tax compliance verification and business logic calculation into a unified format, adding metadata information, and generating standardized financial and tax data.

[0060] Among them, the unified format adopts JSON-LD (JavaScript Object Notation for Linked Data) format, which can define the semantics of data terms through context and realize data semantic intercommunication between different systems; the added metadata includes data bloodline metadata and timestamp, the data bloodline metadata records the complete traceability information of the data, and the timestamp marks the completion time of the data standardization processing.

[0061] In some implementations, the transformation of JSON-LD format needs to define a unified context configuration file first, which needs to clearly define the term definition of financial and tax core data fields (such as “invoiceNumber” corresponding to “invoice number”, “taxRate” corresponding to “tax rate”, “taxAmount” corresponding to “tax amount”), data type (such as “taxAmount” defined as “number” type, “invoiceDate” defined as “date” type) and associated vocabulary (such as “invoiceRelatedOrder” associated with “purchase order number”), and automatically generate JSON-LD format data by parsing the mapping relationship between the fields of the verified data and the context configuration file in the transformation process.

[0062] Among them, the specific fields of data bloodline metadata need to be determined according to the enterprise financial and tax data traceability requirements, at least including data source system identifier, data collection time, data preprocessing operation record, compliance verification result, and business logic calculation associated data source, these fields are obtained by calling system log interface to obtain original operation record, and automatically filled according to the preset bloodline metadata template. The timestamp is recorded in UTC (Coordinated Universal Time) format.

[0063] In some implementations, the bloodline metadata template can be: “data source: {sourceSystem}, collection time: {collectTime}, processing record: [{processRecord1}, {processRecord2}]”.

[0064] It should be noted that the JSON-LD context configuration file needs to be dynamically maintained according to the addition or change of financial and tax data fields. When a company adds "environmental tax declaration data", the definitions of fields such as "environmentalTaxAmount" (environmental tax amount) and "taxablePollutantType" (taxable pollutant type) need to be added in the context configuration file, and the maintenance process needs to be version controlled. Each version of the context configuration file needs to be backed up to facilitate rollback to the historical version. The storage of data bloodline metadata needs to be associated with standardized financial and tax data. The association method of "data ID + bloodline metadata ID" can be used to ensure that the corresponding bloodline information can be quickly queried through any standardized financial and tax data.

[0065] For example, 1 piece of goods sales invoice data issued on March 15, 2024, contains: invoice number: 00123456, amount: 113,000 yuan, tax rate: 13%, tax: 13,000 yuan, converted to JSON-LD format, context configuration file defines "@context": { "@vocab": "http: / / example.com / finance-tax#", "invoiceNumber": "invoice number", "invoiceAmount": "invoice amount", "taxRate": "tax rate", "taxAmount": "tax amount", "invoiceDate": "issue date"}, the converted JSON-LD data is: { "@context": "http: / / example.com / finance-tax#", "invoiceNumber": "00123456", "invoiceAmount": 113000, "taxRate": 0.13, "taxAmount": 13000, "invoiceDate": "2024-03-15"}. The added data bloodline metadata is: "data source: electronic invoice system (ID: E_INVOICE-001), collection time: 2024-03-1509:23:45.123, processing record: [missing value processing-no, abnormal value processing-no, duplicate data processing-no], compliance verification result: [invoice authenticity verification-pass, tax rate compliance verification-pass, accounting standards compliance verification-pass], business logic calculation associated data source: [invoice amount field, tax rate field]". The timestamp is marked as "2024-03-15 10:15:30.456+08:00", and the complete standardized financial and tax data is finally generated.

[0066] It should be noted that "http: / / example.com / finance-tax#" is a semantic namespace reference set for semantic interoperation of JSON-LD format data, which provides a unified semantic reference for the core fields in the standardized tax data (such as "invoiceNumber", "invoiceAmount", "taxRate", etc.), and ensures that different systems can accurately identify and understand the corresponding tax business meaning of each field based on the reference when reading and analyzing the JSON-LD format tax data, thereby breaking down the semantic barriers between different systems and realizing data semantic interoperation between different systems.

[0067] Based on the above technical solutions, S1 realizes accurate reception, targeted preprocessing and compliance verification, standardized format conversion and traceability metadata addition of multi-source raw data, not only ensures the format uniformity, content compliance and semantic interoperation of standardized tax data, but also realizes the full-link traceability of tax data from the source to the processing result through data bloodline metadata, provides a high-quality and high-credibility data foundation for subsequent risk identification and risk analysis based on standardized data, and at the same time, the dynamically configurable rule engine and context configuration file also improve the adaptability of the method to changes in tax policies and business expansion of enterprises.

[0068] In a possible implementation manner of the embodiment of the application, in combination with Figure 2 As shown in Figure 3 S2 can be implemented by the following S201, S202 and S203, which will be described in detail below. S201, feature extraction is performed on the standardized tax data to obtain a key feature vector capable of reflecting data risk features.

[0069] The key feature vector is a core data dimension combination with high correlation degree to tax risk selected from the standardized tax data, such as "invoice amount, tax rate, taxpayer identification number of the issuer, credit rating of the transaction counterparty" in the invoice data, "transaction amount, transaction frequency, fund flow direction, transaction counterparty account attribution" in the fund flow data, etc.; the feature extraction can adopt a method based on principal component analysis, and the main features containing risk information in the data are retained by dimension reduction, and the interference of redundant dimensions on subsequent risk identification is reduced.

[0070] In some implementation manners, the specific operation process of feature extraction is as follows: Firstly, the numerical features (such as invoice amount, tax amount) in the standardized tax data are subjected to Z-score standardization processing to eliminate dimensional differences; Second, calculate the covariance matrix of the standardized data, and get the eigenvalues and corresponding eigenvectors of the covariance matrix through eigenvalue decomposition; Third, sort the eigenvalues from large to small, select the eigenvectors whose cumulative contribution rate reaches the preset threshold as the principal components, and form the key feature vectors.

[0071] The preset principal component cumulative contribution rate threshold needs to be determined in combination with the historical financial and tax risk identification effect of the enterprise: refer to the financial and tax risk case data of the enterprise disposed in the past three years, and calculate the coverage ratio of risk features under different contribution rate thresholds. Take the minimum cumulative contribution rate corresponding to the "risk feature coverage ratio ≥ 90%" as the threshold, usually set to 85%~92%, for example, in the risk cases of a certain enterprise in the past three years, when the cumulative contribution rate is 88%, it can cover 91% of the risk features, so the threshold is set to 88%.

[0072] It should be pointed out that feature extraction is not limited to principal component analysis method. In the scene of low data dimension (such as single type of invoice data, feature dimension ≤ 10), linear discriminant analysis method can also be used, but principal component analysis has better effect on risk feature preservation when dealing with high dimensional financial and tax data (such as comprehensive data integrating invoice, fund flow and equity data, dimension ≥ 30); In addition, the dimension of key feature vector needs to be dynamically adjusted according to business scenarios, for example, when the enterprise adds "environmental protection tax declaration data", it needs to add features such as "environmental protection tax payable pollutant discharge amount" and "applicable tax amount standard", and recalculate the covariance matrix and principal components to ensure that the feature vector can cover the risk dimension of the new business.

[0073] For example, taking the standardized electronic invoice data of the enterprise in March 2024 (1000 pieces, feature dimensions including invoice amount, tax rate, tax amount, issuing date, issuer tax number, recipient tax number, commodity classification code, invoice status, 8 dimensions) as an example, first standardize "invoice amount" and "tax amount" by Z-score, get the mean μ 金额 =15 million, the standard deviation σ 金额 =8 million; the mean μ 税额 =1.95 million, the standard deviation 税额= 1.04 million; after calculating the covariance matrix, 8 eigenvalues are obtained, sorted from large to small as [3.2, 2.5, 1.8, 1.1, 0.8, 0.5, 0.3, 0.2], and the cumulative contribution rate is calculated as: the cumulative contribution rate of the first 3 principal components is (3.2+2.5+1.8) / (3.2+2.5+1.8+1.1+0.8+0.5+0.3+0.2)=7.5 / 10.4≈72.1%, the cumulative contribution rate of the first 4 principal components is (7.5+1.1) / 10.4≈82.7%, the cumulative contribution rate of the first 5 principal components is (8.6+0.8) / 10.4≈89.4%, reaching the preset threshold of 88%, so the first 5 feature vectors are selected to form the key feature vector, and the feature extraction is completed.

[0074] In S202, the key feature vector is matched with risks and checked by rules through a multi-mode matching algorithm to generate a second risk set.

[0075] The multi-mode matching algorithm includes two core links: one is to map the key feature vector to a hash bucket based on a locality-sensitive hashing (LSH) algorithm, and compare it with a preset risk mode hash bucket to generate a first risk set; the other is to input the first risk set into an abnormal rule engine based on a Drools framework to execute a predefined IF-THEN rule chain to generate a second risk set, realizing double risk screening of "algorithm matching + rule checking".

[0076] In some implementations, the parameter configuration of the locality-sensitive hashing algorithm needs to be determined in combination with the historical risk data distribution: the number of hash buckets is set according to the risk mode data volume accumulated by the enterprise in the past year, if the historical risk mode data is 120,000, the number of hash buckets is set to 1,200 (i.e. 1 hash bucket for every 100 risk mode data), and it is ensured that the risk modes in each hash bucket have strong similarity; the selection of the hash function uses an LSH function based on cosine similarity, which maps the vectors with a similarity ≥ 0.75 to the same hash bucket by calculating the cosine similarity between the key feature vector and the risk mode vector.

[0077] The IF-THEN rule chain of the abnormal rule engine is designed according to the risk type, including invoice rules (such as "IF invoice consecutive number and issuer are the same enterprise AND transaction amount is more than 500,000 yuan, THEN mark as invoice abnormality"), tax number rules (such as "IF the receiver's tax number is in the blacklist AND transaction frequency is ≥ 3 times / month, THEN mark as tax number abnormality"), three-in-one rules (such as "IF the invoice issuer and the fund recipient are inconsistent AND there is no reasonable explanation file, THEN mark as three-flow inconsistency abnormality"), etc.

[0078] It should be noted that the preset risk mode is a basic data feature template refined based on enterprise historical financial and tax risk cases, tax authority reported irregularities cases and industry common risk scenarios, used for rapid matching of financial and tax data with potential risk inclination during the first risk identification, without complex logic judgment, stored in the preset risk mode hash bucket, and each risk mode corresponds to a group of core data feature dimensions and thresholds, as shown in Table 1: Table 1, risk mode example table

[0079] It should be noted that the preset risk mode hash bucket needs to establish a dynamic updating mechanism: at the end of each month, new risk mode feature vectors are extracted according to the newly added financial and tax risk cases in the month, and the similarity of the new risk mode feature vectors with the existing hash bucket modes is calculated. If the similarity is less than 0.75, a new hash bucket is added, and if the similarity is greater than or equal to 0.75, it is classified into the corresponding hash bucket, so as to ensure that the risk mode can cover the latest risk types; the Drools framework of the abnormal rule engine needs to support the hot deployment of rules, so that the rules can take effect without restarting the system after being modified, and the rule modification needs to go through “financial and tax expert review-technical personnel configuration-test environment verification-production environment deployment”, so as to prevent misjudgment caused by misconfiguration of rules.

[0080] For example, the key feature vector obtained in S201 is input into the local sensitive hash algorithm: the number of hash buckets is set to 1200, and the cosine similarity LSH function is used to calculate the similarity of the vector with the 120,000 risk modes in the risk mode hash bucket. The similarity of the three risk modes is 0.82, 0.78 and 0.76 respectively, all of which are greater than or equal to 0.75, and they are classified into three different hash buckets. The three corresponding invoice data are marked as the first risk set. The first risk set is input into the Drools abnormal rule engine: the first invoice data triggers the rule of “invoice consecutive number and same issuer AND amount exceeding 500,000 yuan”, the second triggers the rule of “tax number blacklist AND transaction frequency ≥ 3 times / month”, and the third triggers the rule of “three flow inconsistency and no reasonable explanation”. Finally, the three data all pass the rule check, and the second risk set containing three records is generated.

[0081] S203, calculate the comprehensive risk score, and mark the standardized financial and tax data with a score exceeding a preset threshold as risk identification data.

[0082] The comprehensive risk score is obtained by weighted summation of the rule confidence score and the abnormal score. The rule confidence score is calculated based on the number of abnormal rules triggered by the second risk set, and the abnormal score is obtained by abnormal detection of the standardized financial and tax data through the isolation forest algorithm. The combination of the two realizes the quantitative evaluation of the risk level.

[0083] In some implementations, the determination method of each parameter is as follows: (1) The rule confidence score calculation uses the formula S rule = min(1.0, s0+0.1 x (N rules -1)), where the base confidence score s0 is assigned according to the importance of the rule type: s0 = 0.6 for tax clause related rules, s0 = 0.5 for accounting standard related rules, and s0 = 0.4 for enterprise control rules. The assignment is based on the risk cases triggered by different rules in the past three years, in which the loss ratio caused by tax rules is the highest, reaching 65%, so the weight of tax clause related rules is the highest. (2) Isolation forest parameters for anomaly detection: the number of isolation trees t is determined according to the amount of data to be detected. When the data amount is ≤50,000, t = 80; when the data amount is 50,000-100,000, t = 100; and when the data amount is >100,000, t = 120. The sample subset number ψ is 6%-8% of the total amount of data to be detected. For example, when the data amount is 1,000, ψ = 70 (7%). The tree height limit is set to log2(ψ) rounded up. For example, when ψ = 70, log2(70) ≈ 6.13, rounded up to 7. (3) The weight distribution of weighted summation uses the analytic hierarchy process: five tax experts are invited to score the importance of rule confidence score and anomaly score. After consistency test, the weight of rule confidence score is determined to be 0.6, and the weight of anomaly score is determined to be 0.4. (4) The preset risk threshold refers to the scoring distribution of high-risk cases in the past two years. The minimum value of the scores of all high-risk cases is taken as the threshold. For example, the scores of high-risk cases in the past two years are 0.72, 0.75, 0.68, 0.71, and 0.69, and the minimum value is 0.68. Therefore, the threshold is set to 0.68.

[0084] It should be noted that the calculation of the anomaly score needs to be combined with the average path length of the isolation forest, and the formula is s(x, n) = 2 -E(h(x)) / c(n) c(n) is the normalized path length, and when n ≥ 2, c(n) = 2H(n-1) - (2(n-1)) / n, H(n-1) is the (n-1)th harmonic number.

[0085] For example, take the first invoice data of the second risk set in S202 as an example: this data triggers 2 tax-related abnormal rules N rules = 2, s0 = 0.6, then the rule confidence score S rule = min(1.0, 0.6 + 0.1 x (2-1)) = 0.7; calculate the anomaly score by isolation forest: the amount of data to be detected is 1,000, t = 80, ψ = 70, the tree height limit = 7, the average path length E(h(x)) of the data is calculated to be 5.2, n = 1,000, c(n) = 2H(999) - (2 x 999) / 1,000 ≈ 2 x 7.485 - 1.998 ≈ 13.972, then the anomaly score s(x, n) = 2-5.2 / 13.972 ≈0.77; the weighted sum of the comprehensive risk score = 0.7*0.6 + 0.77*0.4 = 0.42 + 0.308 = 0.728, which exceeds the preset threshold value 0.68, so the invoice data is marked as risk identification data; similarly, the second data comprehensive score is 0.69, and the third data comprehensive score is 0.71, both of which exceed the threshold value, and finally generate risk identification data containing 3 records.

[0086] Based on the above technical solutions, S2 realizes efficient dimensionality reduction and similarity matching of high-dimensional financial and tax data by means of principal component analysis and local sensitive hashing algorithm through feature extraction, risk matching and comprehensive scoring, reduces the interference of redundant data on risk identification; and realizes the dual risk screening of rule-based verification + intelligent anomaly detection through the combination of Drools abnormal rule engine and isolation forest algorithm, and ensures the accuracy and timeliness of the risk score through the explicit preset parameter determination method and dynamic calibration mechanism, so that the finally generated risk identification data can accurately cover multiple potential risks in enterprise financial and tax, and provide accurate initial risk source for subsequent risk analysis of enterprise financial and tax relationship network.

[0087] In a possible implementation manner of the embodiment of the application, the combination of Figure 2 As shown in Figure 4 The above S3 can be implemented through the following S301, S302 and S303, which will be described in detail as follows: S301, extract enterprise financial and tax correlation elements, and construct a dynamic graph structure of enterprise financial and tax relationship network.

[0088] Among them, the enterprise financial and tax relationship network includes graph nodes and edge structure: the graph nodes are used to carry the core subjects and key information in the enterprise financial and tax correlation, the edge structure is used to represent the financial and tax correlation type between the nodes, the dynamic nature is reflected in that the edge weight will be adjusted according to the actual business data update, and the graph database is responsible for storing the network structure and dynamic change data, so as to ensure that the network can reflect the state of enterprise financial and tax correlation in real time.

[0089] In some implementation manners, the extraction of node elements needs to be combined with the field characteristics of standardized financial and tax data to refine the operation: When extracting the enterprise entity as a node, the attributes such as the unified social credit code, enterprise name, industry, registered address and the like of the enterprise need to be obtained synchronously, and the unified social credit code is used as the unique identifier of the node; when extracting the legal representative and shareholder as a node, the personnel information needs to be associated with the enterprise node, and the association basis is the “personnel-enterprise position relationship” field in the standardized data; when extracting the transaction counterpart and fund flow node, the internal associated party and external cooperation party need to be distinguished, the internal associated party is labeled with the “association attribute” label, and the external cooperation party records the cooperation duration, historical transaction total amount and the like auxiliary information.

[0090] The construction of the edge structure needs to follow specific rules defined by the association type: The construction of the equity relationship edge is based on the "shareholding ratio" field in the standardized data, and the edge is only constructed when the shareholding ratio is ≥5%. The construction of the transaction relationship edge takes "monthly cumulative transaction amount ≥100,000 yuan" as the threshold to avoid edge structure redundancy caused by small and sporadic transactions. The construction of the fund flow relationship edge needs to distinguish between "accounts receivable" and "accounts payable" directions, which are marked by the "fund flow type" field. The construction of the personnel association relationship edge is based on the field information of "the same person serving in multiple enterprises".

[0091] The calculation of edge weight uses a multi-dimensional fusion method: The weight of the transaction relationship edge is calculated by combining the monthly transaction amount and the transaction frequency. First, the transaction amount is normalized to the 0-1 interval, and then the transaction frequency is converted to a frequency coefficient in the 0-1 interval. Finally, the edge weight = weight amount × 0.6 + frequency coefficient × 0.4. The weight of the equity relationship edge directly uses the shareholding ratio. The weight of the fund flow relationship edge refers to the fund occupation duration. When the occupation duration is ≥3 months, the weight increases by 0.2, and when it is <1 month, the weight decreases by 0.1.

[0092] It should be noted that the extraction rules of nodes and edge structures are not fixed and can be adjusted according to the business expansion needs of enterprises: when an enterprise adds a "supply chain cooperation partner" node, the node attributes are supplemented with "supply chain role" and "delivery period" fields, and a "supply chain cooperation relationship" edge is added. The construction is based on the "supply chain cooperation agreement number" field in the standardized data. When the tax policy adjusts the association relationship recognition standard, the construction threshold of the edge structure can be directly modified without the need to reconfigure the entire network. Only the edge structure of the historical data that meets the new threshold needs to be supplemented.

[0093] Exemplary, extract nodes from standardized tax data in March 2024: Enterprise A (Unified Social Credit Code 91110000XXXXXX, name "Jia Technology Company", industry "software and information technology services"), Enterprise B (Unified Social Credit Code 91310000XXXXXX, name "B Trade Company"), legal representative C (name "Zhang San", associated enterprise A), shareholder D (name "Li Si", holding enterprise A proportion 15%). Build edge structure: build "equity relationship edge" between enterprise A and shareholder D, weight 0.15; build "transaction relationship edge" between enterprise A and enterprise B, transaction amount 500,000 yuan in March (minimum monthly transaction amount 10,000 yuan, maximum 100,000 yuan, weight amount = (50-10) / (100-10)≈0.44), transaction frequency 8 times (frequency coefficient = 0.8), edge weight = 0.44x0.6 + 0.8x0.4≈0.58; build "personnel association relationship edge" between legal representative C and enterprise A, weight fixed at 1.0 (personnel employment relationship is full-time, weight is set to the highest). Store the above nodes, edge structure and weight to the Neo4j graph database, and set daily fixed point update of edge weight, complete the preliminary construction of dynamic enterprise tax relationship network.

[0094] S302, map risk identification data to network nodes, and simulate risk propagation based on the improved SIR model.

[0095] Among them, the mapping of risk identification data is to associate the risk identification data generated by S2 with the unique identifier of the graph node, ensuring that each risk node can correspond to a specific subject in the network; node risk value initialization is to determine the initial risk level according to the comprehensive risk score in the risk identification data; the improved SIR model adjusts the risk propagation probability by introducing edge weight, and the core is to calculate the risk value change of each node at different times, reflecting the diffusion trend of risk in the network.

[0096] In some implementations, the mapping of risk identification data needs to establish matching rules: taking "node type-unique identifier" as the matching key, the matching key of the enterprise node is "enterprise entity-unified social credit code", and the matching key of the personnel node is "legal representative / shareholder-ID card number", when the "risk subject identifier" field in the standardized data is consistent with the matching key, the mapping of risk identification data and nodes is completed. Node risk value initialization needs to adjust the coefficient combined with node type: The initial risk value of the enterprise node = the comprehensive risk score x 1.0, that is, the enterprise is the core subject of risk propagation, and the coefficient is set to 1.0; the initial risk value of the personnel node such as the legal representative and the shareholder = the comprehensive risk score x 0.8, that is, the risk propagation influence of the personnel node is weaker than that of the enterprise, and the coefficient is lowered; the initial risk value of the external node such as the counterparty = the comprehensive risk score x 0.6, that is, the correlation between the external node and the enterprise is low, and the coefficient is further lowered. The determination basis of the adjustment coefficient is the statistical influence of different node types in the risk propagation cases in the past three years - the risk propagation case proportion of the enterprise node is 65%, the personnel node is 25%, and the external node is 10%, according to which the coefficient is set.

[0097] In some implementations, the improved infectious disease SIR model is used to simulate the risk propagation of the enterprise finance and tax relationship network, and the risk value of node i at time t is calculated as: ; wherein β represents the propagation coefficient (the default value is 0.65), γ represents the recovery coefficient (the default value is 0.2), w ij is the edge weight, and N(i) represents the neighbor set of node i, and i and j represent node indexes.

[0098] Wherein, the propagation coefficient β represents the probability of risk propagation from one node to another node, and the calculation method is "the number of adjacent nodes infected in the risk propagation cases in the past three years / the total number of adjacent nodes", and if the statistical result is 65%, the default value of β is set to 0.65; the recovery coefficient γ represents the probability of the decrease of the node risk value, and the calculation method is "the number of times that the node risk value is reduced below the safety threshold in the risk disposal cases in the past three years / the total number of risk nodes", and if the statistical result is 20%, the default value of γ is set to 0.2; the time step of the risk propagation simulation is set to 1 day, and the basis is the average period of historical risk diffusion - the average time of risk propagation from the initial node to the adjacent node in the past three years is 1.2 days, so the time step is set to 1 day to ensure that the key nodes of risk diffusion can be captured. And the risk value needs to be controlled in the range of 0-1, and if it exceeds 1, it is taken as 1, and if it is lower than 0, it is taken as 0.

[0099] It should be noted that in the risk value calculation formula, the iteration logic of the current risk value + risk transmission increment - risk attenuation amount is used to simulate the dynamic propagation of risk, wherein the risk transmission increment is simulated by the propagation coefficient β, the edge weight w ji , and other parameters, the risk attenuation amount γ x R i (t) simulates the risk reduction brought by the risk disposal of the enterprise, and is used to dynamically track the conduction path and diffusion speed of the risk in the correlation network and mine the key hub of risk propagation.

[0100] For example, in the risk identification data generated by S2, the comprehensive risk score of enterprise B is 0.8, which belongs to the risk node. It is matched with the node of "enterprise entity-91310000XXXXXX" (enterprise B) in the graph network, and the risk value of enterprise B is initialized as 0.8x1.0=0.8; the adjacent nodes of enterprise B are enterprise A (edge weight 0.58) and legal representative E (edge weight 1.0, personnel node), the initial risk value of enterprise A is 0.05, and the initial risk value of legal representative E is 0.05. Set β=0.65, γ=0.2, time step 1 day, calculate the node risk value at t=1: the risk increment of enterprise A is 0.65x0.58x0.8x(1-0.05)≈0.65x0.58x0.8x0.95≈0.28, the risk attenuation amount is 0.2x0.05=0.01, and the risk value of enterprise A is 0.05+0.28-0.01≈0.32; the risk increment of legal representative E is 0.65x1.0x0.8x(1-0.05)≈0.65x1.0x0.8x0.95≈0.494, the risk attenuation amount is 0.2x0.05=0.01, and the risk value of legal representative E is 0.05+0.494-0.01≈0.534; the risk increment of enterprise B is 0, that is, there is no adjacent risk node transmission increment, the risk attenuation amount is 0.2x0.8=0.16, and the risk value of enterprise B is 0.8+0-0.16=0.64. The risk propagation simulation of the first day is completed.

[0101] S303, calculating the risk propagation path by using the path search algorithm, and outputting the financial risk analysis result.

[0102] The path search algorithm is used to filter the key path of risk diffusion from the risk propagation simulation result, and the core is to find the shortest path or high-risk path of risk transmission from the initial node to other nodes; the financial risk analysis result output needs to include path information and node risk change trend, which provides basis for subsequent resource allocation and risk response.

[0103] In some implementations, the path screening needs to set double conditions: one is that the path length (the number of nodes contained in the path) does not exceed 5; the second is that the risk transmission cumulative value of the path (the sum of the edge weight in the path and the corresponding node risk value) ≥0.5, the risk influence degree of the path with risk transmission cumulative value lower than 0.5 is weak, which can be excluded. The path priority sorting adopts the rule of "risk transmission cumulative value descending + path length ascending": first compare the risk transmission cumulative values of different paths, the path with high cumulative value has high priority; if the cumulative values are the same, the path with short length has high priority, the risk diffusion speed of short path is faster, so it needs to be paid attention to first.

[0104] The output of the financial risk analysis results includes three parts: First, a list of risk propagation paths, which marks the names, types, and risk values ​​of nodes in the path; second, a path risk assessment, which explains the scope of risk impact for each path (e.g., "affecting 3 companies and 2 key personnel") and the speed of risk spread (e.g., "expected to spread to all adjacent nodes in 3 days"); and third, a labeling of key risk nodes, which marks nodes in the path with edge weights ≥ 0.6.

[0105] For example, based on the risk propagation simulation results of S302, Dijkstra's algorithm is used to calculate the risk propagation path. Two paths that meet the criteria are selected: Path 1 is "Company B → Company A → Company D", path length 3, cumulative risk propagation value = 0.58 × 0.8 + 0.4 × 0.32 ≈ 0.46 + 0.13 ≈ 0.59 ≥ 0.5; Path 2 is "Company B → Legal Representative E → Company F (another company where Legal Representative E works, edge weight 0.9)", path length 3, cumulative risk propagation value = 1.0 × 0.8 + 0.9 × 0.534 ≈ 0.8 + 0.48 ≈ 1.28 ≥ 0.5. Sorted by priority: the cumulative value of Path 2 is higher than that of Path 1, therefore Path 2 has a higher priority. The risk analysis results are labeled as follows: In path 2, company B (high risk, 0.45), legal representative E (medium risk, 0.534), and company F (medium risk, 0.48). The scope of risk impact is "3 entities, involving the software and trade industries". The risk spread speed is "expected to spread to the adjacent nodes of company F in 2 days". The key risk node is legal representative E (edge ​​weight 0.9, closely related).

[0106] Based on the above technical solutions, S3 constructs a corporate financial and tax relationship network through a dynamic graph structure, ensuring that the network can be adjusted synchronously with business data updates. The improved SIR model introduces edge weights and node type coefficients, which are more in line with the propagation law of corporate financial and tax risks compared with the traditional SIR model, and can improve the accuracy of risk simulation. The path search algorithm combines risk accumulation value and path length to screen key paths, which can accurately locate the core hub of risk propagation, provide enterprises with comprehensive analysis results, and provide direct basis for subsequent allocation of computing resources and formulation of risk response strategies.

[0107] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, the above S4 specifically includes the following S401 to S403: S401. Screen high-risk task queues and critical node queues, and predict future business load based on time series analysis models.

[0108] Among them, the high-risk task queue is the processing task corresponding to the high-risk node extracted from the risk analysis result, and the key node queue is the task corresponding to the core associated node in the risk propagation path, which together constitute the priority basis for resource allocation; Business load prediction is to estimate future demand through historical load data to support resource total planning.

[0109] In some implementations, the screening of the high-risk task queue needs to determine the node high-risk threshold: by counting the risk value distribution of "risk nodes that need emergency intervention" in enterprise financial and tax risk disposal cases in the past three years, taking the 80% quantile of the risk values of all emergency intervention nodes as the high-risk threshold - for example, the risk values of the emergency intervention nodes in the past three years are 0.65, 0.72, 0.68, 0.75, and 0.71, and the 80% quantile is 0.72, then the tasks with node risk values exceeding 0.72 are classified into the high-risk task queue.

[0110] In some implementations, the construction process of the time series analysis model can include: Collect the enterprise's business load historical data for many years, and the data dimensions need to cover "task type, task initiation time, task processing duration, and task occupied computing resource amount", and at the same time, collect the associated business data that affect the load, such as monthly financial and tax reporting period, quarterly closing date, and major transaction time period, to ensure that the model can adapt to the periodic characteristics of financial and tax business.

[0111] Then, the collected historical data is preprocessed to obtain a standardized time series data set.

[0112] In the training phase, the training set in the time series data set is input into the time series model, such as the long short-term memory network LSTM, the optimizer, loss function, training rounds, batch size, learning rate, and other hyperparameters of the model are determined, and then the model capable of predicting the business load in a future period of time is obtained through iterative training. In the training process, after each round of parameter update, the model performance is evaluated using the validation set data, and when the evaluation value meets the verification index requirement, the training is stopped.

[0113] Finally, the model that passes the verification is deployed to the enterprise financial and tax management system, and can also be updated regularly according to the newly added business load data and associated business information.

[0114] S402, divide the elastic resource pool, allocate computing resources according to task priority and trigger elastic scheduling.

[0115] Among them, the elastic resource pool is the standby resource divided from the total load, which is used to cope with the backlog of high-priority tasks; the priority hierarchical allocation is to determine the resource proportion according to the task queue to ensure that high-risk tasks obtain resources first, and the elastic scheduling is to use standby resources when the task backlog is large to ensure processing efficiency.

[0116] In some implementations, the division ratio of the elastic resource pool needs to be determined in combination with the historical load fluctuation: the difference between the peak value of the business load in the past three months and the monthly average load is calculated, and the elastic resource pool ratio is 30% of the difference value. For example, the peak load in the past three months is 1500 tasks, the monthly average load is 1000 tasks, the difference is 500, and the elastic resource pool ratio is 500 x 30% = 15%, that is, 15% of the total load is divided as the elastic resource.

[0117] In some implementations, the computing resources are hierarchically divided and allocated according to the task priority and total load of the node, which can specifically include: The total amount of remaining resources after deducting the elastic resource pool is taken as the basic computing resource; Tasks that belong to both the high-risk and critical node queues are of the first priority, and 40% of the basic computing resources are allocated to them. The average single resource consumption of the tasks in this queue is 1.5 times that of other queues, and they need to be prioritized; Tasks that only belong to the high-risk queue are of the second priority, and 30% are allocated to them; Tasks that only belong to the critical node queue are of the third priority, and 20% are allocated to them; The remaining tasks are of the fourth priority, and 10% are allocated to them.

[0118] When the backlog of nodes of the first or second priority exceeds the preset number, resources are automatically allocated from the elastic resource pool. The preset number of high-priority task backlog is set to 120% of the average daily processing capacity of the priority task: for example, the average daily processing capacity of the first priority task is 100, and the preset number = 100 x 120% = 120. When the real-time backlog exceeds 120, 20% of the elastic resources are automatically allocated to the queue from the elastic resource pool. If the backlog continues to exceed 150, 30% of the elastic resources are allocated, and the backlog is reduced to below the preset number.

[0119] It should be noted that the resource allocation ratio is not fixed: if a priority task has a resource shortage for three consecutive days, resulting in processing delay, the resource ratio of the queue can be temporarily increased by 5%, and then returned to the original ratio when the processing returns to normal. In addition, the maximum calling ratio of the elastic resource pool does not exceed 80%, and 20% of the elastic resources are reserved to deal with the sudden backlog of multiple queues, so as to avoid the depletion of standby resources by a single queue.

[0120] S403, integrate the resource allocation result and the risk analysis information, generate and output a risk assessment report.

[0121] The risk assessment report is a summary of the entire financial and tax risk analysis and resource allocation process, and needs to include key risk information and resource usage to provide decision-making basis for financial and tax management personnel. The output form needs to consider readability and operability.

[0122] In some implementations, the core content of the risk assessment report needs to cover four parts: first, resource allocation details, including resource proportion of each priority task, number and amount of elastic resource calls, current resource remaining amount, and "resource shortage period" and corresponding measures; second, risk propagation analysis, including high-risk propagation path list and risk trend graph of key risk nodes; third, risk level classification, dividing node risk values into "high (> 0.72), medium (0.4-0.72), and low (< 0.4)" three levels, and counting the number and proportion of nodes at each level; fourth, risk response suggestions, such as "complete transaction certificate verification within 3 days" and "suspend large transactions with associated parties" for high-risk nodes, and "pre-process non-urgent tasks 1 day in advance" for resource shortage period. The high-risk propagation path list includes path nodes, node risk values, and edge weights

[0123] For example, the generated risk assessment report shows that the first priority task resource proportion is 40%, the elastic resource is called once with 30 task amount, the resource is tight from the 6th to the 7th day, and it is suggested to pre-process 20% non-urgent tasks in advance; there is one high-risk propagation path (Enterprise B→Legal Representative E→Enterprise F), the risk values of Enterprise B, E, and F are 0.64, 0.534, and 0.48 respectively, and the edge weights are 1.0 and 0.9 respectively; there are 8 high-risk nodes (accounting for 1.5%), 32 medium-risk nodes (6.1%), and 500 low-risk nodes (92.4%); it is suggested to verify 12 large transaction certificates of Enterprise B within 3 days, and suspend the new procurement business of Enterprise F; the report is automatically pushed to the financial and tax management system and the management layer mailbox at 9:00 on the same day.

[0124] Based on the above technical solution, S4 determines the resource allocation priority based on the risk analysis result, avoiding the response lag of high-risk tasks caused by equal allocation; on the other hand, through the elastic resource pool and dynamic scheduling mechanism, it can adapt to business load fluctuations and improve system resource utilization efficiency. Finally, the output risk assessment report integrates risk and resource information, providing support for efficient disposal and management decision-making of financial and tax risks for enterprises.

[0125] The above describes the scheme of the embodiments of the present application mainly from the perspective of device implementation. It can be understood that, in order to implement the above functions, each device, for example, the enterprise financial tax management apparatus, comprises at least one of a hardware structure and a software module for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application of the technical scheme and the design constraints. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0126] The embodiments of the present application can divide the functional units of the enterprise financial tax management apparatus according to the above method examples. For example, each functional unit can be divided according to each function, or two or more functions can be integrated in one processing unit. The integrated unit can be implemented in the form of hardware or software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, and is only a logical functional division. In actual implementation, there can be another division method.

[0127] In the case of using an integrated unit, Figure 5 A possible structure schematic diagram of the enterprise financial tax management apparatus (denoted as enterprise financial tax management apparatus 50) involved in the above embodiments is shown, which comprises a processing unit 501 and a communication unit 502, and can further comprise a storage unit 503. Figure 5 The shown structure schematic diagram can be used to illustrate the structure of the enterprise financial tax management apparatus involved in the above embodiments.

[0128] When Figure 5 When the shown structure schematic diagram is used to illustrate the structure of the enterprise financial tax management apparatus involved in the above embodiments, the processing unit 501 is used to control and manage the actions of the enterprise financial tax management apparatus, the communication unit 502 is used for communication between the enterprise financial tax management apparatus and other devices, and the storage unit 503 is used to store the program code and data of the enterprise financial tax management apparatus.

[0129] For example, the communication unit 502 is used to receive the standardized financial tax data of the enterprise, the historical risk case data and the associated business data of the external system. The processing unit 501 is used to call the preset risk mode hash bucket in the storage unit 503, perform the first risk identification on the received financial tax data, and generate risk identification data.

[0130] In a possible implementation, the processing unit 501 is further configured to construct an enterprise financial tax relationship network of a dynamic graph structure based on the risk identification data, perform risk analysis by using a risk propagation model, and output a risk analysis result.

[0131] In a possible implementation, the communication unit 502 is further configured to acquire enterprise real-time business load data, and the processing unit 501 is further configured to combine the risk analysis result and the business load data, predict future load demand by using a time series model, complete calculation resource allocation, generate a risk assessment report, and send the report to an enterprise financial tax management terminal by using the communication unit 502.

[0132] The processing unit 501 can be a processor or a controller, and the communication unit 502 can be a communication interface, a transceiver, a transceiver, a transceiver circuit, a transceiver device, or the like. The communication interface is a general term, and can include one or more interfaces. The storage unit 503 can be a memory. When the enterprise financial tax management apparatus 50 is a chip, the processing unit 501 can be a processor or a controller, and the communication unit 502 can be an input interface and / or an output interface, a pin, or a circuit, or the like. The storage unit 503 can be a storage unit (for example, a register, a cache, or the like) in the chip, or can be a storage unit (for example, a read-only memory (ROM), a random access memory (RAM), or the like) located outside the chip.

[0133] The communication unit can also be referred to as a transceiving unit. The antenna and the control circuit with a transceiving function in the enterprise financial tax management apparatus 50 can be regarded as the communication unit 502 of the enterprise financial tax management apparatus 50, and the processor with a processing function can be regarded as the processing unit 501 of the enterprise financial tax management apparatus 50. Optionally, the device for implementing the receiving function in the communication unit 502 can be regarded as a communication unit, and the communication unit is configured to perform the receiving steps in the embodiments of the present application, and the communication unit can be a receiver, a receiver, a receiving circuit, or the like. The device for implementing the sending function in the communication unit 502 can be regarded as a sending unit, and the sending unit is configured to perform the sending steps in the embodiments of the present application, and the sending unit can be a transmitter, a sender, a sending circuit, or the like.

[0134] Figure 5If the integrated units in the process are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. Storage media for storing computer software products include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0135] Figure 5 The units in the process can also be called modules; for example, a processing unit can be called a processing module.

[0136] This application embodiment also provides a hardware structure diagram of an enterprise financial and tax management device (referred to as enterprise financial and tax management device 60), see [link to diagram]. Figure 6 The enterprise financial and tax management device 60 includes a processor 601, and optionally, a memory 602 connected to the processor 601.

[0137] In the first possible implementation, see Figure 6 The enterprise financial and tax management device 60 also includes a transceiver 603. The processor 601, memory 602, and transceiver 603 are connected via a bus. The transceiver 603 is used to communicate with other devices or communication networks. Optionally, the transceiver 603 may include a transmitter and a receiver. The device in the transceiver 603 that implements the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of this application. The device in the transceiver 603 that implements the transmitting function can be considered as a transmitter, which is used to perform the transmitting steps in the embodiments of this application.

[0138] Based on the first possible implementation method Figure 6 The structural diagram shown can be used to illustrate the structure of the enterprise financial and tax management device involved in the above embodiments.

[0139] in, Figure 6 This can also be illustrated by the system chip in the enterprise's financial and tax management device. In this case, the actions performed by the aforementioned enterprise financial and tax management device can be implemented by this system chip. For details of the actions performed, please refer to the above text, which will not be repeated here.

[0140] In the implementation process, each step in the method provided by the embodiment can be completed by the integrated logic circuit of hardware in the processor or the instruction in the form of software. The steps of the method disclosed by the embodiment of the present application can be directly embodied as hardware processor execution completion, or execution completion by hardware and software module combination in the processor.

[0141] The processor in the present application can include but is not limited to at least one of the following: central processing unit (CPU), microprocessor, digital signal processor (DSP), microcontroller unit (MCU), or various types of computing devices running software such as artificial intelligence processors, each of which can include one or more cores for executing software instructions to perform operations or processing. The processor can be a separate semiconductor chip, or can be integrated with other circuits as a semiconductor chip, for example, it can form a SoC (system on chip) with other circuits (such as coding and decoding circuits, hardware acceleration circuits, or various bus and interface circuits), or it can be integrated as a built-in processor in the ASIC. The ASIC integrated with the processor can be packaged separately or packaged together with other circuits. In addition to including cores for executing software instructions to perform operations or processing, the processor can further include necessary hardware accelerators, such as field programmable gate arrays (FPGA), PLDs (programmable logic devices), or logic circuits that implement special logic operations.

[0142] The memory in the embodiment of the present application can include at least one of the following types: read-only memory (ROM) or other types of static storage devices that can store static information and instructions, random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, and electrically erasable programmable read-only memory (EEPROM). In some scenarios, the memory can also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited to this.

[0143] The embodiment of the present application further provides a computer readable storage medium, comprising instructions which, when executed on a computer, cause the computer to perform any of the above methods.

[0144] The embodiment of the present application further provides a computer program product comprising instructions which, when executed on a computer, cause the computer to perform any of the above methods.

[0145] The embodiment of the present application further provides a chip, comprising a processor and an interface circuit, wherein the interface circuit is coupled with the processor, the processor is configured to execute a computer program or instructions to implement the above method, and the interface circuit is configured to communicate with other modules outside the chip.

[0146] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or include one or more data storage devices such as servers, data centers, etc. that can be integrated with the medium. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (solid state disk, SSD)) and the like.

[0147] Although the present application is described herein in conjunction with various embodiments, other variations of the disclosed embodiments can be understood and implemented by those skilled in the art through viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. Some measures described in mutually different dependent claims can be combined and produce a good result.

[0148] While the application has been described in connection with specific features thereof, it will be evident that many modifications and variations of the application are possible, and will be evident to those of ordinary skill in the art. Accordingly, it is intended that all such modifications and variations be considered as within the spirit and scope of the application. Other combinations and sub-combinations of features, functions, acts and / or functionalities described herein are also intended to fall within the scope of the application. It will be apparent to one of ordinary skill in the art that features, acts, and / or functions from different aspects of the application can be interchanged lead to still further embodiments.

Claims

1. A big data-based enterprise finance and tax management method, characterized in that, The application relates to a financial risk assessment method and system for enterprises. The method comprises the following steps: Receiving multi-source enterprise financial and tax original data, and performing financial and tax compliance verification and business logic calculation on the pretreated original data to generate standardized financial and tax data; Using a multi-mode matching algorithm to identify risks in the standardized financial and tax data and generating risk identification data; Using a dynamic graph structure to construct an enterprise financial and tax relationship network and performing financial risk analysis on the enterprise financial and tax relationship network based on a risk propagation model; The input data of the risk propagation model is the risk identification data; 2.The enterprise finance and tax management method based on big data according to claim 1, characterized in that, According to the business load data and the risk analysis result, the calculation resource is distributed, and a risk assessment report is output. The generation process of the standardized financial and tax data comprises the following steps: Receiving multi-source original data from enterprise resource planning systems, tax declaration platforms, bank reconciliation systems and electronic invoice systems; Performing data cleaning on missing values, abnormal values and repeated data in the multi-source original data; Performing financial compliance verification on the cleaned data through a rule engine; the financial compliance verification comprises invoice authenticity, tax rate compliance and accounting standards compliance; 3.The enterprise finance and tax management method based on big data according to claim 1, characterized in that, Converting the verified data into a unified format, adding metadata and time stamps, and obtaining standardized financial and tax data; the metadata represents complete traceability information of the data, including data source system, data collection time, data processing process and correlation with other data. The risk identification of the standardized financial and tax data by using the multi-mode matching algorithm comprises the following steps: Performing feature extraction on the standardized financial and tax data to obtain a key feature vector; the feature extraction adopts a principal component analysis-based method; Using a local sensitive hashing algorithm, the key feature vector is mapped into a hash bucket and compared with a pre-set risk pattern hash bucket, and a first risk set is output; Inputting the first risk set into an abnormal rule engine based on a Drools framework to execute a pre-defined IF-THEN rule chain and output a second risk set; According to an established risk confidence fusion calculation mechanism, the number of triggered abnormal rules in the second risk set is converted into a rule confidence score; Performing abnormal detection on the standardized financial and tax data to output an abnormal score for each data; Using a weighted summation method to calculate the rule confidence score and the abnormal score to obtain a comprehensive risk score; 4. The big data-based enterprise finance and tax management method according to claim 3, characterized in that, The standardized financial and tax data with a comprehensive risk score greater than a pre-set risk threshold is marked as risk identification data.

5. The big data-based enterprise finance and tax management method according to claim 3, characterized in that, The abnormal rules comprise invoice consecutive number detection rules, tax number blacklist matching rules, three-flow integration verification rules, tax rate abnormal detection rules and associated transaction cycle detection rules; the three-flow integration verification rules represent rules for verifying the consistency of invoice flow, fund flow and cargo flow; the associated transaction cycle detection rules represent rules for identifying associated party circular transaction behaviors hidden by complex transaction structures. The risk confidence fusion calculation mechanism comprises the following steps: Generating a rule confidence score based on the statistical results of rule triggering events; the generation process of the rule confidence score adopts a hierarchical accumulation strategy, the basic confidence score value is given, the number of triggered abnormal rules is increased, an upper threshold is set to control the score range, and the rule confidence score is obtained.

6. The big data-based enterprise finance and tax management method according to claim 5, characterized in that, The mathematical expression of the risk confidence fusion computer mechanism is: S rule =min(1.0,s0+0.1×(N rules -1)); wherein, S rule represents a rule confidence score, s0 represents a basic confidence score, and N rules represents the number of triggered abnormal rules.

7. The big data-based enterprise finance and tax management method according to claim 3, characterized in that, The abnormality detection on the standardized financial and tax data comprises: Randomly selecting ψ sample subsets from the standardized financial and tax data as training data, and constructing t isolated trees to form an isolated forest, wherein each isolated tree is divided into subspaces by recursively randomly selecting the key feature vectors and split values until the subspace contains only one sample or reaches a tree height limit; the split value is a random division preset in the current data range of the selected feature vector, and the tree height limit is a preset parameter representing the maximum growth depth allowed for a single isolated tree, used to control the model complexity; For the standardized financial and tax data point x to be detected, the path length h(x) of x from the root node to the leaf node is calculated for each isolated tree; According to the path length, an average path length E(h(x)) and an anomaly score of the data point x are calculated; the calculation formula of the anomaly score s(x, n) is: s(x, n)=2 -E(h(x)) / c(n) ; wherein c(n) is a standardized path length under a given standardized financial and tax data sample number n, used to correct the reference value of the path length. 8.The enterprise finance and tax management method based on big data according to claim 1, characterized in that, The enterprise financial and tax relationship network is constructed using a dynamic graph structure, comprising: Extracting enterprise entities, legal representatives, shareholder structures, transaction counterparties, and capital flow elements from the standardized financial and tax data as graph nodes of the enterprise financial and tax relationship network; Based on the equity relationship, transaction relationship, capital flow relationship, and personnel association relationship, the edge structure of the enterprise financial and tax relationship network is constructed; The dynamic changes of the enterprise financial and tax relationship network are stored using a graph database, and the edge weights are updated in real time based on transaction amounts and transaction frequencies. 9.The enterprise finance and tax management method based on big data according to claim 3, characterized in that, The financial risk analysis of the enterprise financial and tax relationship network based on the risk propagation model comprises: mapping risk identification data to a graph database node, initializing a node risk value R according to the composite risk score i (0); The improved SIR model of infectious disease is used to simulate the risk propagation of the enterprise financial and tax relationship network. The risk value of node i at time t is calculated as follows: ; wherein β represents the propagation coefficient (the default value is 0.65), γ represents the recovery coefficient, w ij is the edge weight, N(i) represents the neighbor set of node i, and i and j represent node indexes. Based on the risk value, the shortest path of risk propagation is calculated using a path search algorithm, and multiple risk propagation paths are output.

10. A big data-based enterprise finance and tax management system, characterized in that, Comprise: A data processing module, a risk identification module, and a resource allocation module; wherein, The data processing module is used to receive multi-source enterprise financial and tax original data, and after preprocessing the original data, it carries out financial and tax compliance verification and business logic calculation to generate standardized financial and tax data; The risk identification module is used to identify risks in the standardized financial and tax data using a multi-mode matching algorithm to generate risk identification data; The resource allocation module is used to construct an enterprise financial and tax relationship network using a dynamic graph structure, analyze the financial risk of the enterprise financial and tax relationship network based on a risk propagation model with the risk identification data as input, allocate computing resources according to business load data and risk analysis results, and output a risk assessment report.