Accounting intelligent bookkeeping and anomaly detection method based on deep reinforcement learning
By building an accounting data entropy embedding state and anomaly scoring model through deep reinforcement learning, the structural perception and strategy optimization problems of the accounting system in high-frequency transactions and complex environments are solved, and the efficient and stable operation of automated accounting and anomaly detection is achieved.
Patent Information
- Application Number
- CN202510837499.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-10
AI Technical Summary
Existing accounting systems find it difficult to effectively handle financial events under high-frequency trading and complex policy environments. They lack structural perception capabilities and strategy optimization mechanisms, and are unable to achieve an intelligent feedback chain of prediction-execution-evaluation-correction. In addition, their anomaly detection capabilities are insufficient, resulting in a high false alarm rate and unbalanced model performance in different cycles.
A method based on deep reinforcement learning is used to construct the entropy embedding state of accounting data. Combined with adaptive learning rate and periodic discount factor, automatic voucher generation and dynamic update of account balances are achieved. An anomaly scoring model is introduced to form a closed-loop system of accounting execution and risk control identification.
It achieves high-precision, adaptive accounting processing without human intervention, improves the automation level, risk control capabilities and structural change response efficiency of the financial system, and has strong explainability and system stability.
Smart Images

Figure CN120765404A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of reinforcement learning technology, and in particular to an accounting intelligent bookkeeping and anomaly detection method based on deep reinforcement learning. Background Art
[0002] Against the backdrop of the rapid evolution of automated and intelligent financial management in enterprises, accounting information systems are gradually transforming from traditional voucher entry tools to advanced systems capable of intelligent judgment, automated accounting, risk warning, and adaptive policy. Integrated financial platforms, centered around ERP systems, have been widely deployed across various medium- and large-scale enterprises, supplemented by rule engines, process engines, and data modeling modules to semi-automate certain business processes. However, in environments characterized by high-frequency trading, multi-business collaborative accounting, and complex policy environments, traditional rule-based accounting systems struggle to effectively handle financial events characterized by high structural volatility, high operational frequency, and widespread transaction distribution, thus impacting overall financial transparency and risk control responsiveness. Therefore, key technical challenges in intelligent finance primarily focus on the coordinated design of intelligent decision support, structural awareness, risk identification, and system convergence and explainability.
[0003] Current intelligent accounting systems still have significant shortcomings in the following areas: Most existing systems rely on fixed templates or rules for matching, failing to model the entropy-level changes in the proportional structure of the five major categories on the balance sheet. This results in a lack of awareness of changes in structural dispersion and an inability to provide dynamic feedback on structural mutations. They also lack reinforcement learning-based strategy optimization mechanisms. Traditional machine learning methods rely on historical labeled data for fitting, freezing strategies once deployed and lacking the ability to continuously evolve. Complex financial systems exhibit distinct Markov decision-making characteristics, with correlations between current states and future returns. Reinforcement learning is the most suitable modeling paradigm for such problems. However, there is currently a lack of structured modeling and reliable implementation solutions for deep reinforcement learning in the accounting field. Existing automated accounting systems generally fail to reflect model prediction errors, behavioral biases, or systemic structural changes on accounting strategies. They can only trigger alarms after an anomaly occurs, failing to form an intelligent feedback loop of "prediction-execution-evaluation-correction." Most systems identify anomalies based solely on a single dimension (e.g., monetary thresholds, rule conflicts), failing to integrate three-dimensional metrics: bookkeeping structure, tax deviations, and model prediction residuals. This results in high false positive rates and a lack of multi-dimensional risk characterization. Financial systems inherently fluctuate on a monthly, quarterly, and annual basis, but most strategy learning processes fail to account for the dynamics of time discount factors. This leads to unbalanced model performance at the beginning of the year or during year-end closing, slow convergence, or short-term aggressiveness. Summary of the Invention
[0004] To address these technical issues, we propose an intelligent accounting and anomaly detection method based on deep reinforcement learning. By constructing a state representation that integrates ledger structure, accounting flow, tax information, and structural entropy, combined with a policy optimization mechanism that uses adaptive learning rates and periodic discount factors, we achieve automatic voucher generation and dynamic updating of account balances. We also introduce an anomaly scoring model centered on policy residuals, structural deviations, and tax fluctuations to construct a closed-loop system for accounting execution and risk control identification. This method achieves high-precision, adaptive accounting processing without human intervention, significantly improving the automation level, risk control capabilities, and efficiency of responding to structural changes in financial systems.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0006] An accounting intelligent bookkeeping and anomaly detection method based on deep reinforcement learning, the method comprising:
[0007] Step 1: Obtain the ending balances, cumulative debits, and cumulative credits for all enabled accounting accounts in the current accounting period; map all accounting accounts into five categories: assets, liabilities, owner's equity, income, and expenses, according to enterprise accounting standards, and assign fixed numerical weights to each category; calculate the sum of the absolute values of all current accounting account balances as the state normalization benchmark; calculate the proportion of each category balance to the state normalization benchmark, and use this to calculate the Shannon entropy of the proportions of the five category balances; and construct a unique accounting data entropy embedding state for the current accounting period;
[0008] Step 2: Based on the entropy embedding state of accounting data, a deep reinforcement learning algorithm is used to iterate the policy value and calculate the optimal Q-value prediction for the next accounting period's accounting action in real time. In the next accounting period, the accounting action is executed, and an accounting voucher with multiple debits and credits is generated and automatically recorded.
[0009] Step 3: Calculate the time series difference error after the accounting action is executed; based on the current ending balance of each accounting account, increase the corresponding voucher debit and credit amount difference, and add the product of the absolute value of the time series difference error and the standard deviation of the balance of each accounting account for the past 30 consecutive days as the numerator, and the risk buffer amount of one plus the square root of the number of days in the current fiscal year as the denominator to obtain the updated ending balance of each accounting account; use the updated accounting account balance as the beginning balance of the next accounting period and repeat the calculation of the accounting data entropy embedding state in step 1;
[0010] Step 4: Calculate the anomaly score based on the updated ending balance of the accounting account. When the anomaly score is higher than the standard normal critical threshold corresponding to the two-sided 0.5% significance level, an anomaly alarm is automatically triggered, and the corresponding voucher number, the number of the accounting account involved, and the value of the specific anomaly indicator are immediately recorded for audit and risk control process calls.
[0011] Furthermore, the fixed numerical weight corresponding to assets is 1; the fixed numerical weight corresponding to liabilities is 2; the fixed numerical weight corresponding to owner's equity is 3; the fixed numerical weight corresponding to income is 4; and the fixed numerical weight corresponding to expenses is 5.
[0012] Furthermore, step 3 specifically includes: multiplying the number of unrecorded vouchers in the current accounting period by the Shannon entropy and adding one, and then taking the inverse as the adaptive learning rate; at the same time, dividing the remaining days of the current accounting year by the total number of days in the year (365) to obtain a discount factor; using the adaptive learning rate and the discount factor to perform time-series difference iteration on the reward function and Q value, obtain the optimal Q value prediction for the accounting action in the next accounting period, and determine the accounting action for the next accounting period.
[0013] Furthermore, in step 4, based on the updated ending balance of the accounting account, the ratio of the absolute difference of the ending balance of each accounting account from the average balance of each accounting account for the past 30 consecutive days to the standard deviation of the balance of each accounting account for the past 30 consecutive days is calculated as the first channel indicator; at the same time, the ratio of the absolute difference of the value-added tax amount of the current voucher from the average value-added tax amount of the voucher for the past 90 consecutive days to the standard deviation of the value-added tax amount of the voucher for the past 90 consecutive days is calculated as the second channel indicator; the absolute value of the time series difference error is used as the third channel indicator; the above three channel indicators are summed and divided by the sum of the total number of accounting accounts, the number of vouchers and 1 to obtain the anomaly score.
[0014] Furthermore, the accounting data entropy is embedded in the state S t for:
[0015]
[0016] Among them, S t It represents the entropy embedding state of accounting data at the current accounting period t; It represents the sum of the absolute values of the ending balances of all enabled accounting accounts as of the current accounting period t; N 会计科目 Indicates the total number of accounting subjects enabled in the current account set; K i represents the fixed numerical weight of the category to which the i-th accounting item belongs; It represents the ending balance of the i-th accounting account at the previous accounting period t. A positive value represents a debit balance, and a negative value represents a credit balance. Indicates the cumulative debit amount of the i-th accounting account at the current accounting period t; Indicates the cumulative credit amount of the i-th accounting account in the current accounting period t; M t Indicates the total number of accounting documents that have been posted as of the current accounting period t; VAT j Indicates the value-added tax amount of the j-th voucher; represents the amount of exchange gain or loss generated by the jth voucher; Shannon entropy representing the proportion of balances in the five categories.
[0017] Furthermore, the Shannon entropy of the five categories of balance ratios for: Where g∈{assets, liabilities, owner's equity, income, expenses}; p g,t Indicates the proportion of accounting category g to the total balance in the current accounting period t. The calculation formula is:
[0018] Furthermore, in step 2, based on the entropy embedding state of accounting data, the formula for strategy value iteration using deep reinforcement learning algorithm is:
[0019]
[0020] Among them, Q t+1 (S t ,a t ) is the Q value of the accounting action in the next accounting period; a t Q is the accounting action for the current accounting period; t (S t ,a t ) is the Q value of the accounting action in the current accounting period; Ω t The number of accounting vouchers expected by the current accountant; R t+1 For the next accounting period's reward; a represents an accounting action.
[0021] Furthermore, the updated ending balance for:
[0022]
[0023] in, is the ending balance of the i-th accounting item at the previous accounting period t; is the timing difference error; T is the standard deviation of the balance of each accounting item in the past 30 consecutive days; t The number of days that have passed in the current fiscal year.
[0024] Furthermore, the anomaly score is:
[0025]
[0026] Among them, A t+1 The abnormality score for the next accounting period; The average balance of each accounting item for the past 30 consecutive days; The average value of the VAT amount of the vouchers for 90 consecutive days; sd(VAT90 ) is the standard deviation of the VAT amount of the vouchers for the past 90 consecutive days; M t+1 Indicates the total number of accounting documents that have been posted as of t+1 of the next accounting period.
[0027] Compared with the existing technology, the beneficial effect of the present invention is that it can automatically complete the generation of accounting vouchers, update accounting strategies, adjust account balances, and identify abnormal behaviors without human intervention, significantly improving the automation level and risk control capabilities of the financial system. By introducing a structural entropy modeling mechanism, the system can perceive the proportional structural changes between the five major accounting elements of assets, liabilities, equity, income, and expenses in real time, thereby identifying the degree of discreteness of the account book structure and realizing continuous monitoring of the dynamic stability of the financial system. During the strategy execution process, the system dynamically adjusts the learning rate based on the business processing intensity and the complexity of the financial structure, and implements time discount adjustment of the strategy in combination with the remaining cycles of the fiscal year, thereby improving the stability and response efficiency of the model in different cycles. At the same time, by introducing prediction errors into the balance update process and combining the historical fluctuations of each accounting account to form a risk buffer mechanism, the system has the ability to adaptively correct when facing strategy deviations and structural mutations. In terms of anomaly detection, the present invention designs an anomaly scoring model that integrates three channels of structure, tax, and strategy residuals, which can provide accurate judgment and graded response to systemic errors, business anomalies, and model failures. Overall, the present invention is superior to existing technologies in terms of accounting intelligence, strategy evolution capability, structural perception capability and abnormal response speed. It has strong explainability, system stability and feasibility, and is suitable for various high-frequency, complex and multi-dimensional automatic accounting and risk control scenarios in corporate financial systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a flow chart of the method for intelligent accounting and anomaly detection based on deep reinforcement learning proposed in the present invention. DETAILED DESCRIPTION
[0029] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0030] Reference Figure 1 As shown, the accounting intelligent bookkeeping and anomaly detection method based on deep reinforcement learning includes:
[0031] Step 1: Obtain the ending balances, cumulative debits, and cumulative credits for all enabled accounting accounts in the current accounting period; map all accounting accounts into five categories: assets, liabilities, owner's equity, income, and expenses, according to enterprise accounting standards, and assign fixed numerical weights to each category; calculate the sum of the absolute values of all current accounting account balances as the state normalization benchmark; calculate the proportion of each category balance to the state normalization benchmark, and use this to calculate the Shannon entropy of the proportions of the five category balances; and construct a unique accounting data entropy embedding state for the current accounting period;
[0032] First, during the current accounting period, the system calls the general ledger data interface within the enterprise's financial system to extract the ending balances, cumulative debits, and cumulative credits for all enabled accounts. This data, originally recorded in a multidimensional structure within the financial subsystem, requires unified integration through a data standardization process and conversion to easily calculable book values. The system then categorizes all accounts into five basic categories: assets, liabilities, owner's equity, revenue, and expenses, based on the classification standards predefined in the Accounting Standards for Enterprises. A fixed weight is assigned to each category within the data processing module, enabling subsequent structural embedding modeling of different account types. This process is automated, not relying on manual adjustments but rather through a mapping between standardized account codes and predefined categories. After the category mapping is complete, the system further accumulates the absolute values of all account balances to obtain the total balance size for the current set of accounts, which serves as a normalization factor for subsequent state normalization.
[0033] To construct state characteristics that reflect differences in the internal structure of the balance sheet, the system calculates the proportion of the absolute value of the end-of-period balance of each account category to the total balance, based on the five account categories. Based on this, the system introduces an information entropy model to process the distribution of this proportion and calculates a dispersion index for the current balance sheet structure. This index reflects the degree of balance among the five financial structures in the current set of accounts, namely, structural complexity: if the balance of a certain category of account, such as assets or expenses, is significantly higher than that of other categories, the entropy value is low, indicating a concentrated structure; if the distribution of balances across categories is relatively even, the entropy value is high, indicating a dispersed structure. By incorporating structural entropy into the state representation, the system ensures that the reinforcement learning strategy is not only sensitive to the absolute value of the amount but also has the ability to identify structural evolution trends, thereby avoiding potential risks caused by structural mutations or excessive deviations during the strategy optimization process.
[0034] After the above processing is completed, the system integrates the balance information, period flow information and structural complexity information of each subject to integrate a single accounting state value for representing the financial environment state of the current accounting period under the entire account set dimension. The state value is a key input for the subsequent reinforcement learning model for policy evaluation and updating, has determinism, time sequence and traceability, can dynamically evolve over time, and accurately reflects the changing trajectory of the financial structure of the enterprise in different accounting periods. In this way, the originally scattered accounting data distributed in multiple subsystems is abstracted into a structured, numerically compressed and information-rich state input, so that the deep reinforcement learning model can effectively complete policy learning and decision optimization under the conditions of limited samples, complex structure and nonlinear evolution, and provide a highly reliable basic environment expression for subsequent automatic voucher generation, anomaly identification and risk warning.
[0035] Step 2: Based on the accounting data entropy embedded state, the policy value iteration is performed by the deep reinforcement learning algorithm to calculate the optimal Q value prediction of the next accounting period accounting action in real time; in the next accounting period, the accounting action is executed to generate and automatically account for a multi-borrowing and multi-lending accounting voucher;
[0036] The accounting data entropy embedded state constructed in the first step is taken as the current environment input, and the state-action mapping relationship of reinforcement learning is modeled through a deep neural network, i.e., the value of a plurality of accounting actions executable under a given state is estimated to identify the optimal accounting strategy. The system adopts a policy learning algorithm based on value function approximation, specifically, a multilayer perceptron is used as a function approximator to encode and process the state input, extract deep features such as structure, flow and entropy value, and map them to the evaluation value of each candidate action in the action space. Each action corresponds to a specific accounting behavior, which is manifested as a composite accounting voucher with multiple borrowing and lending entries. The voucher needs to realize the optimal adjustment of the existing accounting state under the premise of structural legality, amount balance and tax compliance. In order to realize the quantitative evaluation of the action effect, the system sets a composite reward function in the reinforcement learning training process, which covers multiple dimensions such as borrowing and lending balance, subject legality, tax stability and automatic voucher generation efficiency, and introduces the historical difference of the dispersion degree of the accounting structure and the volatility of the account as a negative punishment item to measure the influence of a certain accounting action on the stability and risk level of the account book structure. The system prioritizes exploration in the early training stage, guides diversified action attempts through the randomness component or approximate greedy strategy in the policy, and uses the reward value generated by the actual accounting behavior feedback to optimize the estimation function in the reverse direction, constantly adjusting the neural network parameters to better approximate the true mapping relationship between the state and the action value.
[0037] After policy convergence—that is, when the network model's prediction error stabilizes and the action selection demonstrates a clear optimal trend—the system enters the accounting execution phase. At the beginning of each new accounting period, the system first regenerates an accounting state representation based on the current ledger data and feeds this into the trained neural network model. The model then outputs a value assessment of all candidate accounting actions. The system selects the action with the highest assessment value, representing the optimal accounting method determined by the current policy, and generates an accounting voucher with multiple debits and credits based on this value. During voucher generation, the system leverages an existing business rule library and an account validity rule engine to ensure that the generated journal entries are within the permitted range of accounting standards, avoiding illegal account combinations or inconsistent debit and credit relationships across different fields. For each debit and credit entry in the voucher, the system automatically calculates the amount and corresponding tax amount, and assigns a reasonable accounting account based on the current ledger balance and historical accounting trends. This ensures that the generated voucher logically aligns with the evolving trends of actual business behavior while also meeting the accounting system's requirements for continuity and rationality. After the voucher is generated, the system automatically enters it into the financial system and logs metadata such as the execution time of the accounting action, the subjects involved, the tax impact, and whether the debits and credits are fully balanced. This information serves as the basic data source for error assessment and anomaly detection in subsequent steps. At the same time, the system collects real-time feedback from the current execution action in its current state, including indicators such as balance changes after the voucher is generated, structural entropy changes, VAT deviation, and overall voucher production efficiency. This is used as the basis for the reward function in the next iteration to further fine-tune the strategy value function and achieve online continuous optimization of reinforcement learning. In this way, the entire accounting decision-making process no longer relies on traditional fixed rules or static templates. Instead, the system dynamically judges and automatically generates policy models based on historical data training, making accounting behavior more flexible, intelligent, and context-aware. This significantly improves the financial system's adaptability and risk prevention capabilities in scenarios such as processing multi-source heterogeneous data, complex business portfolios, and risk-implicit transactions.
[0038] Step 3: Calculate the time series difference error after the accounting action is executed; based on the current ending balance of each accounting account, increase the corresponding voucher debit and credit amount difference, and add the product of the absolute value of the time series difference error and the standard deviation of the balance of each accounting account for the past 30 consecutive days as the numerator, and the risk buffer amount of one plus the square root of the number of days in the current fiscal year as the denominator to obtain the updated ending balance of each accounting account; use the updated accounting account balance as the beginning balance of the next accounting period and repeat the calculation of the accounting data entropy embedding state in step 1;
[0039] This step begins with the accounting vouchers generated during the previous accounting cycle. Based on the corresponding accounting account for each entry in the voucher, the system extracts the debit and credit amounts and uses these amounts to update the corresponding account's ending balance. The basic update approach is to add the net debit and credit changes caused by the voucher to the original ending balance, thereby reflecting the direct impact of the accounting action on the book value. This process is automatically performed by the system by traversing the voucher entry list and querying the account balance table, ensuring traceability and system consistency for all changes. However, this method goes beyond simply adding the voucher impact results to the ledger. Instead, it further considers the potential volatility risk of financial data over time. Therefore, a standard deviation model based on historical statistics and a dynamic adjustment mechanism based on strategy forecast errors are introduced to provide risk-sensitive buffering for each account update. Specifically, during the update process for each accounting account, the system extracts the end-of-day balance sequence for the account over the past 30 calendar days and calculates its standard deviation, which reflects the recent volatility of the account. At the same time, the system also calculates the temporal difference error (TDE) generated by the current accounting action in real time. This error is derived from the difference between the reinforcement learning model's predicted value of the current period's reward and the actual feedback, and is a quantitative reflection of the effectiveness of the strategy optimization. If the system detects a large TDE, meaning that the strategy fails to effectively predict actual returns under the current state, or the behavior generated by the current strategy deviates significantly from the existing financial structure, it combines this error with the standard deviation to generate a risk buffer factor.
[0040] This factor is constructed based on the principle of dynamic adjustment for time-series risk: the greater the volatility and error, the more conservative the account update should be, thereby reducing the risk of bookkeeping jumps caused by erroneous strategies or data anomalies. To prevent the buffer factor from being overly amplified over time and causing delayed accounting responses, the system also incorporates a fiscal year time parameter, using the number of days elapsed in the current year as a time-decay variable in the buffer value scaling. This provides greater protection at the beginning of the year, when system data is scarce and volatility assessments are unstable. However, as annual operations stabilize and historical data becomes more abundant, the system's confidence in strategy updates increases, and the buffer automatically decreases. By jointly modeling error dynamics, historical volatility, and the time-decay mechanism, the system not only reflects the direct numerical impact of accounting actions when updating account balances, but also dynamically adjusts the correction of bookkeeping values based on strategy maturity and financial structure trends, thereby achieving a risk-controlled and steadily evolving bookkeeping update process. After completing the balance update for all involved accounts, the system writes the generated new balance status to the general ledger module and simultaneously triggers the next round of accounting status construction operations. This means that the updated end-of-period balance in this step serves as the input source for embedding the entropy of the accounting data in the next period, entering the execution process of the first step. This design forms a fully closed-loop accounting system based on reinforcement learning feedback. In this system, accounting behavior is no longer a one-time operation, but is deeply coupled with the entire process of state estimation, strategy feedback, risk adjustment, and structural update. The time-series difference error mechanism and historical volatility correction model introduced in the third step ensure that the system can achieve continuous and stable evolution of the ledger through automatic buffering and flexible adjustment mechanisms in the face of complex scenarios such as structural mutations, abnormal business operations, or policy fluctuations, significantly enhancing the robustness and intelligent adaptability of the financial system.
[0041] Step 4: Calculate the anomaly score based on the updated ending balance of the accounting account. When the anomaly score is higher than the standard normal critical threshold corresponding to the two-sided 0.5% significance level, an anomaly alarm is automatically triggered, and the corresponding voucher number, the number of the accounting account involved, and the value of the specific anomaly indicator are immediately recorded for audit and risk control process calls.
[0042] The implementation of this step begins with the latest ending balance data generated at the end of step 3. The system first extracts the ending balances of all accounting accounts for the current accounting period and compares these values with the mean and standard deviation of the corresponding account's historical ending balances over the past 30 consecutive days. For each account, the system calculates the deviation between the current balance and the historical mean and uses the ratio of this deviation to the historical standard deviation as a normalized risk indicator. This normalization effectively eliminates the amplification effect of absolute value errors caused by differences in the amount of the account itself, ensuring consistent sensitivity in detecting anomalies for both large and small accounts. After calculating the deviation scores for all accounts, the system normalizes and accumulates these scores to form an overall structural volatility indicator at the account level. Simultaneously, the system also performs a tax stability assessment at the document level. This assessment primarily compares the VAT amount in the currently generated document with the average VAT amount for all documents over the past 90 consecutive days and calculates the normalized deviation. Similar to the processing logic at the account level, this deviation is normalized by dividing it by the historical standard deviation to form a tax volatility score. This indicator reflects the stability changes of the current accounting actions in the tax dimension and is used to identify possible tax-related errors, abnormal tax rate application or sharp fluctuations in tax amounts driven by abnormal business.
[0043] In addition to the structural and tax layers, this step also incorporates a feedback mechanism at the strategy level. The difference between the reinforcement learning model's predicted strategy returns in the second step and the actual returns is used as a strategy performance residual, which is directly incorporated into the calculation of the anomaly score. This strategy residual reflects the model's prediction accuracy in the current state. If this error persists, it may indicate generalization failure or insufficient state recognition, necessitating retraining or model reconstruction through subsequent learning mechanisms. Therefore, the system directly incorporates this error into the anomaly detection scoring system, enabling the detection mechanism to not only identify structural anomalies at the data level but also reflect potential issues at the strategy level. After completing the risk scoring for the three dimensions mentioned above, the system weights and sums the account structure score, tax burden score, and strategy residual score according to a unified standardized scale to obtain the total anomaly score for the current accounting period. The system then invokes built-in anomaly detection rules, which establish critical judgment thresholds based on confidence levels. Specifically, the system presets a statistical critical value corresponding to a two-sided 0.5% significance level for a normal distribution as the alarm threshold. When the anomaly score exceeds this threshold, an anomaly signal is identified for the current period. At this time, the system will automatically generate an exception alarm record, which includes the voucher number that triggered the exception, the accounting subject number involved, the specific value of each risk scoring dimension and the comprehensive score, and send the record to the audit interface module for internal auditors to conduct further analysis.
[0044] Furthermore, the fixed numerical weight corresponding to assets is 1; the fixed numerical weight corresponding to liabilities is 2; the fixed numerical weight corresponding to owner's equity is 3; the fixed numerical weight corresponding to income is 4; and the fixed numerical weight corresponding to expenses is 5.
[0045] Assets, as resource-based items, have a direct impact on the scale of a company's resources, so their weight is set to a minimum of 1. Liabilities, representing obligations to external entities, are less volatile than assets and are weighted to 2. Owners' equity, reflecting a company's control over its net assets and its residual value, lies between assets and earnings, and is weighted to 3. Revenue, reflecting the inflow of funds from principal operations or other economic activities, is highly frequent and volatile, and is weighted to 4. Expenses, the primary component of operating expenditures, fluctuate frequently and are highly disruptive, and are weighted to a maximum of 5. This fixed-weight mechanism allows the system to incorporate differential weights of account categories' contributions to the overall structure into state representation and strategy assessment, thereby improving the resolution and risk identification capabilities of structural modeling and enhancing the structural sensitivity of state characteristics over time.
[0046] Furthermore, step 3 specifically includes: multiplying the number of unrecorded vouchers in the current accounting period by the Shannon entropy and adding one, and then taking the inverse as the adaptive learning rate; at the same time, dividing the remaining days of the current accounting year by the total number of days in the year (365) to obtain a discount factor; using the adaptive learning rate and the discount factor to perform time-series difference iteration on the reward function and Q value, obtain the optimal Q value prediction for the accounting action in the next accounting period, and determine the accounting action for the next accounting period.
[0047] Step 3 not only updates accounting account balances and applies a risk buffer mechanism, but also introduces an adaptive learning rate mechanism tied to the current ledger structure complexity and business processing load, as well as a discount factor design dynamically linked to the accounting period. This allows the policy update process in reinforcement learning to respond in real time to the structural characteristics and temporal context of the enterprise's financial environment, thereby achieving a more stable and robust policy evolution process. Specifically, this step begins by extracting information about the document queue length, which represents the current period's business processing intensity, from the number of unrecorded vouchers in the current accounting period. This information is combined with the Shannon entropy of the ledger structure calculated in Step 1 to measure the complexity of the current balance sheet structure. The system multiplies this number of vouchers by the Shannon entropy and adds one to form the denominator of the learning rate. The reciprocal is then taken to obtain the adaptive learning rate for the current period. This learning rate automatically decreases with increasing voucher backlogs or increasing ledger structure dispersion. This reduces the policy update step size under conditions of high business pressure or structural complexity, preventing excessive policy fluctuations and ensuring the stability of policy evolution.
[0048] At the same time, the system also calculates the number of days remaining in the current fiscal year based on the current calendar date and the fiscal year range, and divides this number by the total number of days in the year, 365, to determine the current period's discount factor. This discount factor reflects the current period's position within the entire fiscal year and is used to dynamically adjust the strategy's emphasis on future returns. At the beginning of the year, when there are more days remaining and the discount factor is close to 1, the system places a higher weight on future returns, facilitating long-term strategy planning. Towards the end of the year, when there are fewer days remaining and the discount factor is correspondingly lower, the system's strategy will prioritize immediate feedback from current accounting activities to prevent short-term fluctuations from negatively impacting the year's close.
[0049] After obtaining the aforementioned adaptive learning rate and dynamic discount factor, the system enters the reinforcement learning strategy update phase. Using the reward function value of the current cycle and the maximum Q-value in the next state output by the neural network model, the system calculates the temporal difference error and incrementally updates the Q-value corresponding to the current state and the current action. This update process combines the value estimate of the previous strategy in the current state, the feedback reward obtained from the actual accounting behavior, the optimal return expectation in the next state, and the environmental complexity and time urgency of the current cycle, thereby forming a dynamically adjusted, structurally sensitive, and time-consistent strategy evolution mechanism. After the strategy is updated, the system immediately outputs the latest Q-value predictions for all optional accounting actions in the current state based on the new strategy, and selects the action with the highest Q-value as the optimal accounting decision for the next accounting period, driving subsequent automatic voucher generation and accounting operations.
[0050] By embedding the number of unrecorded vouchers and structural entropy into the learning rate control mechanism and incorporating the time variable into the discount factor design, this step enables reinforcement learning to achieve greater adaptability and policy stability in the face of cyclical fluctuations and structural non-stationarity in accounting business systems. This provides a strategic foundation for highly dynamic response capabilities throughout the accounting optimization and risk control process. This mechanism not only optimizes model training results but also effectively enhances the intelligent accounting system's ability to identify abnormal conditions and the efficiency of intervention by dynamically adjusting the weight distribution. It is one of the key technical links in achieving the deep coupling of reinforcement learning and financial scenarios.
[0051] Furthermore, in step 4, based on the updated ending balance of the accounting account, the ratio of the absolute difference of the ending balance of each accounting account from the average balance of each accounting account for the past 30 consecutive days to the standard deviation of the balance of each accounting account for the past 30 consecutive days is calculated as the first channel indicator; at the same time, the ratio of the absolute difference of the value-added tax amount of the current voucher from the average value-added tax amount of the voucher for the past 90 consecutive days to the standard deviation of the value-added tax amount of the voucher for the past 90 consecutive days is calculated as the second channel indicator; the absolute value of the time series difference error is used as the third channel indicator; the above three channel indicators are summed and divided by the sum of the total number of accounting accounts, the number of vouchers and 1 to obtain the anomaly score.
[0052] During implementation, the system first extracts the current balance for each active account based on the updated ending balances. This balance is then compared to the average ending balance for the account over the past 30 calendar days, calculating the absolute deviation between the two. The system also extracts the standard deviation of the account balance within the same 30-day window and uses the ratio of the current deviation to this standard deviation as a standardized indicator to measure whether the current account balance exhibits statistically significant abnormal fluctuations. This ratio reflects the current balance's relative position within the historical fluctuation range; a larger value indicates that the balance may be outside the normal range, indicating a higher risk level. The standardized deviation ratios for all active accounts are aggregated to form the first channel indicator, which describes structural anomalies. After processing the first channel, the system proceeds to calculate the second channel indicator. This channel focuses on fluctuations in the VAT amount in automatically generated vouchers. The system extracts the VAT amount of the voucher generated by the current accounting action and compares it with the average VAT amount of all vouchers entered over the past 90 calendar days, calculating the absolute deviation between the current value and the historical average. The system also calculates the standard deviation of the 90-day VAT sample and uses the ratio of the deviation to the standard deviation as a secondary indicator to measure the degree of deviation in tax treatment of the current voucher. This indicator is used to identify tax risk behaviors caused by misapplication of tax rates, incorrect tax calculations, or abnormal transaction structures.
[0053] After calculating the structural and tax indicators, the system introduces a third channel indicator: the absolute value of the temporal difference error (TDE) between the execution of the reinforcement learning strategy in the second and third steps, used as a strategy-level anomaly metric. This error reflects the model's ability to predict the subsequent impact of accounting behavior. If this value deviates significantly from the normal range, it indicates that the model is failing to accurately predict the risk-return relationship of the current behavior, possibly due to input distortion, strategy failure, or sudden environmental changes, posing a potential risk of decision-making anomaly. After calculating the three channel indicators, the system aggregates each indicator into a single value and normalizes it. The sum of the first channel indicator, the single value of the second channel indicator, and the absolute value of the third channel indicator are then added together and divided by the total number of currently enabled accounting accounts, the number of vouchers in the current period, and 1 to form the final anomaly score. In this scoring structure, the denominator is designed to be "number of accounts + number of vouchers + 1" to balance the information size of the three channels and avoid score bias caused by excessive data volume in any one channel. The anomaly score reflects the comprehensive degree of deviation from the current accounting cycle in terms of structural rationality, tax compliance, and strategic stability. It quantifies the consistency and risk level of automated accounting behavior across multiple dimensions. After calculating the anomaly score, the system compares it with a preset statistical threshold, typically determined at a 0.5% two-sided significance level under the assumption of a normal distribution. If the anomaly score exceeds this threshold, the system determines that an anomaly exists in the current accounting cycle, automatically triggering the anomaly alarm module. The system then writes detailed data, including the current voucher number, the account number involved, and the three-channel indicators, to the audit log and risk control interface for subsequent internal control processes.
[0054] Furthermore, the accounting data entropy is embedded in the state S t for:
[0055]
[0056] Among them, S t It represents the entropy embedding state of accounting data at the current accounting period t; It represents the sum of the absolute values of the ending balances of all enabled accounting accounts as of the current accounting period t; N 会计科目 Indicates the total number of accounting subjects enabled in the current account set; K i represents the fixed numerical weight of the category to which the i-th accounting item belongs; It represents the ending balance of the i-th accounting account at the previous accounting period t. A positive value represents a debit balance, and a negative value represents a credit balance. Indicates the cumulative debit amount of the i-th accounting account at the current accounting period t; Indicates the cumulative credit amount of the i-th accounting account in the current accounting period t; M tVAT j represents the VAT amount of the jth voucher; represents the exchange gain / loss amount generated by the jth voucher; represents the Shannon entropy of the five-category balance proportion.
[0057] The construction idea of the formula is first reflected in the pursuit of scale independence: by defining Sum the absolute values of the balances of all enabled accounts, and take as the overall scaling factor to ensure that any size of the enterprise falls into the same order of magnitude range at the numerical input level, thereby avoiding numerical instability caused by too large a difference in the amount of magnitude during gradient propagation. Secondly, the formula uses K i The weights applied to the five categories of assets, liabilities, owner's equity, income, and expenses form a monotonically increasing category hierarchy in structure, so that changes on the asset side contribute the least to the state and changes on the expense side contribute the most, indirectly embedding the sensitivity of financial activities to enterprise risk exposure into the state representation. The core of the subject-related part is It combines static balance with this period's cumulative flow into net effect, on the one hand ensuring simultaneous perception of stock and increment, and on the other hand also making the state have an immediate indication function on the direction of cash flow, so that the reinforcement learning model can adjust the strategy in time when cash flow is reduced or expenses surge to maintain the balance of borrowing and lending and the lower limit of risk. Subsequently added injects the profit / loss caused by tax pressure and exchange rate changes into the state at once, reflecting the direct impact of the external economic environment on the cash chain, so that the model not only pays attention to the internal structure of the account when evaluating the value of action, but also can predict the impact of tax burden and foreign exchange fluctuations on future rewards. The above two formulas together constitute two axes of the cuboid: the internal operating aspect and the external environment aspect.
[0058] The second term in the denominator of the formula embodies the idea of information theory, where
[0059] Shannon entropy is used to measure the dispersion of the balance distributions of five categories: assets, liabilities, equity, income, and expenses. A larger entropy value indicates a more uniform balance distribution and a more complex structure. Including it in the denominator along with a constant automatically compresses the state amplitude and reduces gradient jitter when the structure is highly discrete, while maintaining sensitivity to abnormal account expansion when the structure is extremely concentrated, thereby dynamically adjusting the exploration intensity during the policy learning phase. It is important to emphasize that the entropy term does not introduce simple scaling, but rather a form of adaptive regularization: when account balances are overly concentrated in a few categories, it amplifies the gradient signal, urging the model to take balancing actions; when balances tend to be uniform, it reduces the gradient strength, preventing the model from over-adjusting due to minor fluctuations. This reflects the trade-off between structural robustness and operational sensitivity in deep reinforcement learning.
[0060] From the perspective of reinforcement learning, S t The design of is both identifiable and predictable. Identifiable is reflected in the orthogonality of the components: K i Introducing category stretching to ensure that the contribution differences between categories are clear, VAT j and Decouple tax and exchange rate fluctuations from the main accounting events, As a structural regularization term, S is multiplicatively independent of the amount and tax indicators. Predictability comes from the monotonic relationship between the state and subsequent rewards: when a certain type of expense item suddenly increases or the VAT amount deviates from the historical mean, S t The value of is bound to shift significantly, enabling the value network to capture signals of potential high costs or rising tax risks and adjust actions through policy gradients or temporal differences to control risk. Because the state is a single scalar, the network only needs to track the mapping between the scalar and the action value during proximal policy optimization or deep Q-network updates. This significantly reduces the curse of dimensionality common in high-dimensional state spaces and ensures real-time responsiveness even when deployed in resource-constrained financial systems.
[0061] The formula also implies a natural connection between time continuity and historical rolling windows. At the end of each accounting period, a new S t+1 Afterward, the old state is not discarded. Instead, it enters the next cycle with updated balances and incremental flows. Shannon entropy is also continuously calculated along with the category balance distribution, forming a smooth time series. Deep reinforcement learning models can train time-dependent convolutional or recurrent structures based on this sequence to discover inter-period patterns, such as quarterly revenue peaks or cyclical expense declines. More importantly, with the introduction of VAT and exchange gains and losses, the state can simultaneously reflect macroeconomic fluctuations, such as the impact of a stronger exchange rate or policy tax rate adjustments on corporate cash flow and profits. This makes the strategy not only self-consistent at the accounting level but also macroeconomically adaptable.
[0062] Furthermore, the Shannon entropy of the five categories of balance ratios for: Where g∈{assets, liabilities, owner's equity, income, expenses}; p g,t Indicates the proportion of accounting category g to the total balance in the current accounting period t. The calculation formula is:
[0063] The essence of the Shannon entropy index is an information-theoretic measure of the structure of the five categories of accounting balances. Its value range is between zero and logarithmic five, depending on the uniformity of the category distribution. If all balances are concentrated in a certain category, such as expenses or assets, there is only one p g,t ≈1, the rest are close to zero, and the entropy value approaches zero, indicating a highly concentrated structure; if the five types of balances are evenly distributed, that is, p g,t ≈0.2, the entropy value reaches its maximum, which is -5×0.2ln0.2≈ln5, indicating the most dispersed structure. Introducing this indicator into multiple calculation links of the intelligent accounting system can effectively identify the structural tilt caused by the surge or weakening of a certain type of business, and further reflect the volatility and complexity of the organization's financial operations. From a system perspective, the most direct effect of this entropy value is reflected in the state S t In the denominator of , it serves as a structural complexity regularization term. Its inclusion automatically shrinks the system input state when the financial structure becomes more discrete, effectively suppressing drastic adjustments made by aggressive strategies in highly dispersed states, thereby improving the numerical stability and behavioral conservatism of the policy learning process. In learning rate calculation, Shannon entropy is combined with the number of pending documents as an observable indicator of business load, influencing the step size of each policy update in reinforcement learning: the busier the business and the more discrete the structure, the lower the learning rate, preventing the strategy from losing control under highly volatile conditions. Furthermore, this entropy value is used as an indirect reference indicator in the anomaly detection process. In certain scenarios of financial fraud or management manipulation, companies may deliberately adjust their account structures to make their financial statements appear reasonable. However, a short-term, drastic change in structural entropy indicates a sudden abrupt change in the proportional relationship between assets, liabilities, equity, or expense categories. Such a change can conceal significant risks even when the documents themselves are balanced and the accounts are compliant. Therefore, by continuously observing Shannon entropy and analyzing its changing trends, the system can detect early signals of subtle but systemic anomalies, enhancing automated risk control capabilities.
[0064] Furthermore, in step 2, based on the entropy embedding state of accounting data, the formula for strategy value iteration using deep reinforcement learning algorithm is:
[0065]
[0066] Among them, Q t+1 (S t ,a t) is the Q value of the accounting action in the next accounting period; a t Q is the accounting action for the current accounting period; t (S t ,a t ) is the Q value of the accounting action in the current accounting period; Ω t The number of accounting vouchers expected by the current accountant; R t+1 For the next accounting period's reward; a represents an accounting action.
[0067] Q on the left side of the formula t+1 (S t ,a t ) represents the latest value estimate for the same state-action pair after the end of accounting period t, which is determined by the previous period estimate Q t (S t ,a t ) plus a temporal difference term scaled by an adaptive learning rate. Ω in t The number of vouchers currently expected to be posted by the accountant, used to measure the system's immediate concurrent load; Yes The structural Shannon entropy after boundary truncation is used to measure the dispersion of the account book among the five accounting elements. The multiplication of the two, plus one and then taking the inverse, is equivalent to constructing a differentiable learning rate that decreases with the growth of business pressure and structural complexity: when the backlog of vouchers is large or the structural dispersion is high, As the value increases, the learning rate As the learning rate decreases, the update step size decreases accordingly, preventing the model from experiencing value fluctuations in high-noise and highly heterogeneous environments. When business volume decreases or the structure becomes relatively concentrated, the learning rate automatically increases, ensuring accelerated model convergence in low-risk phases. This design replaces traditional fixed or exponentially decreasing learning rates, enabling strategy evolution to be aware of the business's operational rhythm in real time.
[0068] The first term R in the timing difference brackets t+1 The actual income obtained from the reward function in the next accounting period includes multi-dimensional indicators such as credit balance, compliance with regulations, tax stability and closing efficiency, ensuring that the value update is firmly anchored in real business results; the second item Injects discounted expectations of optimal future returns, where Align the discount factor with the financial reporting period: Early in the year, when there are more days remaining, Close to one, the model focuses more on long-term value; at the end of the year, As the value decreases, the strategy focuses more on short-term returns, meeting corporate management requirements for rapid year-end closing and report locking. The entire update formula seamlessly integrates real-time business load, adaptive learning rate, annual time value, and multi-dimensional reward signals into a single-step increment. This enables the deep network to maintain a stable gradient when processing high-frequency discrete actions in accounting scenarios, achieving an optimal compromise between performance goals and internal control requirements.
[0069] It is worth noting that here The solution must be found in a large action space containing multiple borrowing and lending combinations, which is constrained by the legality of accounts and tax rules. Therefore, the system uses the experience replay pool to retain only candidate actions that meet compliance rules and have a balanced debit and credit balance to ensure the feasibility and security of the max operation. As the accounting period iterates, Ω t and The real-time change of allows the learning rate to smoothly transition between different business loads and structural risk levels, avoiding the over-estimation problem exposed by traditional RL in high-noise scenarios; The linear decay of the system naturally suppresses value explosion at the end of the year, effectively preventing the model from ignoring immediate loan balance and tax risks in pursuit of long-term returns. Through this value iteration mechanism driven by quantitative indicators in the accounting environment, the strategy can achieve a dynamic balance in complex financial scenarios: during peak periods of voucher accumulation, the convergence rate of action is automatically slowed to reduce system jitter. When tax policy changes cause a sudden increase in structural entropy, the learning rate is simultaneously reduced to avoid overfitting, and the learning process is quickly resumed after the structure stabilizes. At the same time, the introduction of annual time discounts naturally aligns the strategy with the rhythm of financial settlement, ensuring that the intelligent accounting module can output risk-controlled and cost-optimized voucher production solutions at all stages throughout the year.
[0070] The reward function R t+1 As the core feedback quantity for strategy updating in reinforcement learning, it comprehensively reflects the structural impact, compliance performance, efficiency level, and robustness fluctuations of the automatic accounting behavior on the financial system at the current time t+1. Its structure is as follows:
[0071]
[0072] The function consists of four positive reward terms and two penalty terms, which are defined as follows:
[0073] r 平衡 Indicates the balance degree of the voucher generated by the current action in the debit and credit directions, which is defined as:
[0074]
[0075] in, The debit amount of the first entry in the voucher; The credit amount of the first entry in the voucher. If the debits and credits are completely balanced, this item is 0; if not, it is a negative number, used to penalize accounting errors or logical incompleteness.
[0076] r 准则 It is a binary indicator used to determine whether the automatically generated voucher complies with the accounting system constraints:
[0077]
[0078] r 税负 Indicates the degree of deviation of the VAT amount generated by the current voucher from the historical mean, defined as:
[0079]
[0080] Among them, VAT t+1 : VAT amount of the current document (from the tax account or automatically calculated by the tax rate); The average VAT amount for the voucher over the past 30 days. This is used to encourage the system to generate transaction records with stable tax burdens and no abnormal fluctuations.
[0081] r 关账速 Indicates the time it takes for the system to automatically generate and post a voucher, defined as:
[0082] r 关账速 =-T 生成 ;
[0083] Among them, T 生成 : The time taken to complete the current action (in milliseconds or seconds), recorded by the system log. The shorter the time, the closer the value is to 0, indicating higher accounting efficiency.
[0084] Structural entropy change penalty:
[0085]
[0086] Used to quantify whether the current accounting operation causes a drastic change in the dispersion of the balance sheet structure. Shannon entropy is the ratio of the balances of assets / liabilities / equity / income / expenses at time t. Increasing entropy indicates a more uniform structure, but drastic changes may lead to control risks, so this item is a negative factor.
[0087] Balance fluctuation daily increment penalty:
[0088]
[0089] Indicates the daily increase in the degree of fluctuation of the balance of each accounting account. The standard deviation of the balance of the i-th account in the past 30 days as of the current point in time; Δσ i,t: The growth rate of account balance fluctuations. The greater the fluctuation, the lower the financial stability. The system will deduct points for potential anomalies.
[0090] Furthermore, the updated ending balance for:
[0091]
[0092] in, is the ending balance of the i-th accounting item at the previous accounting period t; is the timing difference error; T is the standard deviation of the balance of each accounting item in the past 30 consecutive days; t The number of days that have passed in the current fiscal year.
[0093] The update formula couples the direct numerical impact of the accounting action and the risk adjustment factor based on the time series difference error into the same expression, thus forming an adaptive and robust update mechanism at the balance level. It is the ending balance of the i-th account after the end of the previous accounting period, representing the stock status of the account before the strategy iteration; This is the net change in the debit and credit direction of the account in the currently generated voucher, reflecting the immediate impact of accounting actions on the flow of funds in this account. If the entry is a pure debit for the account, the net value is positive and the balance increases; conversely, if the net value is negative, the balance decreases. This part follows traditional accounting logic. The key innovation lies in the third term in the formula, which combines the core feedback quantity of deep reinforcement learning—the temporal difference error δ t ── and historical fluctuation characteristics and time decay factor Combined with the above, a subject-level risk buffer is constructed. Measures the strategy's prediction bias for current period rewards: if the model overestimates future returns or underestimates risks, then δ t A larger negative or positive value; |δ t The absolute value of | directly reflects the severity of the strategy mismatch. Using it as a magnification factor ensures that the balance adjustment is proportional to the strategy error. This ensures that when the strategy exhibits significant forecasting bias, the system automatically makes a more conservative adjustment to the balance, reducing the possibility that subsequent decisions will further amplify the error based on the mismatch.
[0094] Is the standard deviation of the account's end-of-day balance series over the past thirty days, representing the intensity of recent fluctuations. The introduction of this term means that the more unstable the account itself is, the more the system tends to leave a larger buffer to resist short-term sharp fluctuations. When the account balance has been stable historically, The lower the time decay factor, the smaller the correction range will be, thus avoiding excessive conservatism and information lag. Introduce the number of days that have passed in the current year into the denominator and achieve sublinear growth through square root form: At the beginning of the year, T t Small, the denominator is close to one, and the buffer item has the greatest impact; as the year progresses, As the denominator increases, the buffer gradually decreases. This design aligns with the financial management principle of "caution at the beginning of the year, lock-in at the end": at the beginning of the year, when data accumulation is insufficient and business plans are still being adjusted, the system tends to be more conservative in updating; at the end of the year, when data is sufficient and quick closing is required, the system reduces the buffer to ensure report accuracy and timeliness.
[0095] In summary, the balance update formula provides a deep reinforcement learning framework with an environmental evolution rule that is deeply integrated with the accounting risk control context. By embedding the strategy prediction residuals, historical fluctuations, and accounting period factors into the change logic of the accounting account balance, the system can instantly evaluate the strategy reliability after each automatic entry and reflect it numerically at the balance level, forming a strategy-account book-risk trinity linkage control mechanism. When the strategy continuously shows high errors in actual business, the balance correction terms will accumulate on multiple accounts, resulting in a systematic offset in the embedded state of accounting data entropy. This offset will affect the strategy evaluation and decision-making through the next round of state input, thereby forcing the model to lower the value estimate of high-risk actions, realizing an endogenous negative feedback correction; on the contrary, when the model prediction is accurate and the account fluctuations are stable, |δ t | and At the same time, it is at a low level, and the balance is almost not affected by the buffer items. The direct effect of the accounting action can be fully reflected, ensuring that the system maintains high-efficiency capital flow and real-time data during normal operation.
[0096] This formula also plays an important supporting role in the anomaly detection module. Because the size of the buffer item is directly related to the time series difference error and the historical standard deviation, if a certain item has a series of abnormally large additional corrections, but the policy error has not decreased accordingly, it may indicate that external anomalies such as forged vouchers or internal policy adjustments have not been identified by the model. The anomaly detection module monitors the buffer ratio in the balance update sequence and its relationship with δ tThe collaborative trend of these factors can capture ambiguous risk signals early, further improving the system's early warning sensitivity. In short, the update formula is not limited to numerical corrections. It assumes the multiple functions of coupling forecast error, volatility significance, and time weight within the deep reinforcement learning architecture, making balance evolution a critical physical layer for the intelligent accounting system to resist strategic deviations, absorb business noise, and self-adjust to changes in the financial cycle. Through this dynamic update rule, the system can manage risk exposure in a real-time, granular, and data-driven manner while maintaining basic accounting principles, providing enterprises with intelligent financial recordkeeping and anomaly prevention capabilities that are both accurate and adaptable.
[0097] Furthermore, the anomaly score is:
[0098]
[0099] Among them, A t+1 The abnormality score for the next accounting period; The average balance of each accounting item for the past 30 consecutive days; The average value of the VAT amount of the vouchers for 90 consecutive days; sd(VAT 90 ) is the standard deviation of the VAT amount of the vouchers in the past 90 consecutive days; M t+1 Indicates the total number of accounting documents that have been posted as of t+1 of the next accounting period.
[0100] Abnormal score A t+1 The design aims to provide a single evaluation indicator for the intelligent accounting framework driven by deep reinforcement learning, which can not only quantitatively measure the health of the account book but also directly trigger risk response. At the methodological level, it undertakes the core functions of transaction posterior monitoring, model residual feedback and risk control early warning linkage. The formula adopts a three-channel fusion approach, uniformly projects the deviations of the structural layer, tax layer and strategy layer to a standard score scale, and then obtains a comprehensive description of the overall abnormality of the current accounting period through normalized averaging. The structural channel uses the ending balances of all enabled accounting accounts as the observation quantity, first calculating the balance of the i-th account and its average value for the past thirty consecutive days. The absolute shift between Dimensionless Z-scores are obtained. This standardization eliminates heteroskedasticity caused by differences in account size, ensuring that the risk contribution of both multi-billion-yuan asset accounts and micro-expense accounts is measured on the same statistical scale. Risk sensitivity is automatically corrected using historical volatility ranges, ensuring that the indicator is equally sensitive to unusual deviations from low-volatility accounts. Summing the absolute Z-scores across all accounts yields the overall structural deviation intensity. This sum will increase significantly if a specific category of accounts exhibits a concentrated pattern of unusual performance, or if multiple accounts experience small fluctuations simultaneously, reflecting the overall imbalance risk of the bookkeeping structure.
[0101] The tax channel builds a risk perspective on the VAT amount at the voucher level by comparing the VAT amount of each voucher in the current accounting period. j Average tax amount over the past 90 days and divided by the corresponding historical standard deviation sd(VAT 90 ), converting possible tax rate misapplication, input / output mismatches, or abnormal discounting into the same Z-score, then summing all vouchers to reflect abnormal tax pressure. Using a 90-day window smooths monthly fluctuations while covering the entire quarterly reporting cycle, capturing both tax burden trends and short-term tax shocks. The third channel introduces strategic layer signals |δ t |, which comes directly from the time series difference error of the reinforcement learning model, reveals the model's prediction accuracy for the current business environment. When the accounting action is misjudged by the system or the environment suddenly changes, causing the model to become inaccurate, the residual is rapidly amplified, adding immediate feedback on the quality of the strategy to the anomaly score. Denominator N 科目 +M t+1 +1 is designed to be the sum of the sample sizes of three channels, corresponding to the number of accounts, the number of vouchers, and a single residual channel. By taking an average rather than a sum, we avoid bias caused by differences in sample bases. This ensures that each structural unit has equal voting power in the scoring, ensuring that the indicator can be compared across accounts and periods. This mathematical construction embodies the principle of fairness in risk measurement and expands the dimension of anomaly detection to include model behavior itself, forming a unified monitoring system for business data and algorithmic behavior.
[0102] Specifically, in the operation process, the system calculates A immediately after completing the accounting period closing. t+1 If the value falls into the tail interval of more than 0.5% on both sides of the normal distribution, it will be immediately regarded as a high-confidence anomaly and trigger an early warning. The audit interface will record the triggering voucher, subject and corresponding three-channel score, which will be used as the data basis for subsequent manual review or strategy retraining. If the anomaly score is in the warning interval but does not reach the alarm threshold, the system will only record A t+1 and δ t The experience replay pool is written together to provide a soft constraint signal for strategy update. It is worth emphasizing that this scoring formula complements the buffer mechanism in the dynamic update of account balance in the previous article: when |δ t Even if large and short-term structural fluctuations have been partially absorbed by the buffer items, the abnormal score will still amplify the residual signal to remind the model of potential systematic biases, preventing the buffer mechanism from covering up deep-seated decision-making errors; and when the abnormal tax amount of a single voucher causes a surge in the second channel score, the system can quickly locate the abnormal voucher and its transaction background through traceability queries, thereby improving risk management efficiency.
[0103] The following example shows anomaly detection for accounting records in accounting period t for a manufacturing company. All amounts are in thousands of yuan, times are in seconds, and statistics such as standard deviation are taken from actual account set snapshots.
[0104] The total absolute value of the ending balance is The proportions of the five categories of balances are assets 0.4138, liabilities 0.2414, owner's equity 0.1724, income 0.1034, and expenses 0.0690, from which we get the Shannon entropy The increment during the period is -18750 after the category weights are aggregated, and the tax impact is 315, so the accounting data entropy is embedded in the state Number of vouchers to be recorded Ω t =20, learning rate The remaining natural days in the year are 165, and the discount factor
[0105] The system's current strategy is in state S t Select action a t : Generate a composite voucher containing Inventory Debit 100, Sales Tax Debit 110, and Bank Deposit Credit 210. Immediately after the action is completed, the reward function is evaluated for six items: First, the balance difference between debit and credit ∑ l (D l -C l )=0, so r 平衡 =-|0|=0. Second, the subject combination has passed the rule engine verification and the legal mark r 准则 = 1. Third, the VAT amount on the voucher is 110, the average value over the past 30 days. r 税负 =-|110-100|=-10. Fourth, the system takes time T 生成 =0.8s, r 关账速 =-0.8. Fifth, after the voucher is recorded, the proportion of the five categories of balances changes to assets 0.4172, liabilities 0.2397, owner's equity 0.1716, income 0.1014, and expenses 0.0701. This difference is directly penalized as a negative penalty of -0.002. Sixth, the rolling standard deviation of the account changes: the fluctuation of cash and inventory increased by 3 and 4 respectively in the past 30 days, and there was no increase in the other 8 accounts. The total ∑ i |Δσ i,t |=3+4=7; penalty -7. Therefore, the final compound reward is R t+1 =0+1-10-0.8-0.002-7≈-16.802.
[0106] The value network gives Q t (S t ,a t )=50, the optimal Q estimate for the next state Timing differential error Q value update Q t+1(S t ,a t )=50+0.0338×(-20.14)≈49.32; Balance update example: Select inventory account: old balance 2700, voucher net increase 100, historical 30-day standard deviation Number of days in the year T t =200.
[0107]
[0108] The remaining accounts are updated in the same way, resulting in a new ledger and entering cycle t+1. The anomaly scoring phase first calculates the Z scores for the ten accounts in the structural channel; in this example, only the significant items are listed: Inventory | (2959.6 - 2600) / 120 | ≈ 3.0, Cash | (3100 - 3000) / 90 | ≈ 1.11, Accounts Receivable | (1450 - 1500) / 70 | ≈ 0.71, with a cumulative sum of approximately 8.00. The VAT amounts for the four vouchers in the tax channel for the current period are 110, 100, 105, and 95, with a historical 90-day average of 100 and a standard deviation of 8, resulting in a total Z score.
[0109] Strategy channel takes |δ t |=20.14. The number of accounts activated in the current period is 10, and the number of vouchers is M t+1 =4, abnormal score The normal two-sided 0.5% critical value is 2.576, which is higher than 2.04. Therefore, the system records "yellow alert" and writes the complete indicators into the audit link. At the same time, δ t = -20.14 replay to force the strategy to adjust its expected value for similar certificates downward in the next round.
[0110] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An accounting intelligent bookkeeping and anomaly detection method based on deep reinforcement learning, characterized by: The method comprises: Step 1: Obtain the ending balances, cumulative debits, and cumulative credits for all enabled accounting accounts in the current accounting period; map all accounting accounts into five categories: assets, liabilities, owner's equity, income, and expenses, according to enterprise accounting standards, and assign fixed numerical weights to each category; calculate the sum of the absolute values of all current accounting account balances as the state normalization benchmark; calculate the proportion of each category balance to the state normalization benchmark, and use this to calculate the Shannon entropy of the proportions of the five category balances; and construct a unique accounting data entropy embedding state for the current accounting period; Step 2: Based on the entropy embedding state of accounting data, a deep reinforcement learning algorithm is used to iterate the policy value and calculate the optimal Q-value prediction for the next accounting period's accounting action in real time. In the next accounting period, the accounting action is executed, and an accounting voucher with multiple debits and credits is generated and automatically recorded. Step 3: Calculate the time series difference error after the accounting action is executed; based on the current ending balance of each accounting account, increase the corresponding voucher debit and credit amount difference, and add the product of the absolute value of the time series difference error and the standard deviation of the balance of each accounting account for the past 30 consecutive days as the numerator, and the risk buffer amount of one plus the square root of the number of days in the current fiscal year as the denominator to obtain the updated ending balance of each accounting account; use the updated accounting account balance as the beginning balance of the next accounting period and repeat the calculation of the accounting data entropy embedding state in step 1; Step 4: Calculate the anomaly score based on the updated ending balance of the accounting account. When the anomaly score is higher than the standard normal critical threshold corresponding to the two-sided 0.5% significance level, an anomaly alarm is automatically triggered, and the corresponding voucher number, the number of the accounting account involved, and the value of the specific anomaly indicator are immediately recorded for audit and risk control process calls.
2. The accounting intelligent bookkeeping and anomaly detection method based on deep reinforcement learning according to claim 1 is characterized in that: The fixed numerical weight corresponding to assets is 1; the fixed numerical weight corresponding to liabilities is 2; the fixed numerical weight corresponding to owner's equity is 3; the fixed numerical weight corresponding to income is 4; and the fixed numerical weight corresponding to expenses is 5.
3. The accounting intelligent bookkeeping and anomaly detection method based on deep reinforcement learning according to claim 2 is characterized in that: Step 3 specifically includes: multiplying the number of unrecorded vouchers in the current accounting period by the Shannon entropy, adding one, and then taking the inverse as the adaptive learning rate; at the same time, dividing the remaining days of the current accounting year by the total number of days in the year (365) to obtain the discount factor; using the adaptive learning rate and discount factor to perform time-series difference iteration on the reward function and Q value to obtain the optimal Q value prediction for the accounting action in the next accounting period, and determine the accounting action for the next accounting period.
4. The method for intelligent accounting and anomaly detection based on deep reinforcement learning according to claim 3, characterized in that: In step 4, based on the updated ending balance of the accounting account, the ratio of the absolute difference between the ending balance of each accounting account and the average balance of each accounting account for the past 30 consecutive days and the standard deviation of the balance of each accounting account for the past 30 consecutive days is calculated as the first channel indicator. At the same time, the ratio of the absolute difference between the value-added tax amount of the current voucher and the average value-added tax amount of the voucher for the past 90 consecutive days and the standard deviation of the value-added tax amount of the voucher for the past 90 consecutive days is calculated as the second channel indicator. The absolute value of the time series difference error is used as the third channel indicator. The sum of the above three channel indicators is divided by the sum of the total number of accounting accounts, the number of vouchers, and 1 to obtain the anomaly score.
5. The method for intelligent accounting and anomaly detection based on deep reinforcement learning according to claim 4, characterized in that: Accounting data entropy embedding state S t for: Among them, S t It represents the entropy embedding state of accounting data at the current accounting period t; It represents the sum of the absolute values of the ending balances of all enabled accounting accounts as of the current accounting period t; N 会计科目 Indicates the total number of accounting subjects enabled in the current account set; k i represents the fixed numerical weight of the category to which the i-th accounting item belongs; It represents the ending balance of the i-th accounting account at the previous accounting period t. A positive value represents a debit balance, and a negative value represents a credit balance. Indicates the cumulative debit amount of the i-th accounting account at the current accounting period t; Indicates the cumulative credit amount of the i-th accounting account in the current accounting period t; M t Indicates the total number of accounting documents that have been posted as of the current accounting period t; VAT j Indicates the value-added tax amount of the j-th voucher; represents the amount of exchange gain or loss generated by the jth voucher; Shannon entropy representing the proportion of balances in the five categories.
6. The accounting intelligent bookkeeping and anomaly detection method based on deep reinforcement learning according to claim 5 is characterized in that: Shannon entropy of the balance ratio of the five categories for: Where g∈{assets, liabilities, owner's equity, income, expenses}; p g,t Indicates the proportion of accounting category g to the total balance in the current accounting period t. The calculation formula is:
7. The method for intelligent accounting and anomaly detection based on deep reinforcement learning according to claim 6, characterized in that: In step 2, the formula for strategy value iteration using deep reinforcement learning algorithm based on the accounting data entropy embedding state is: Among them, Q t+1 (S t ,a t ) is the Q value of the accounting action in the next accounting period; a t Q is the accounting action for the current accounting period; t (S t ,a t ) is the Q value of the accounting action in the current accounting period; Ω t The number of accounting vouchers expected by the current accountant; R t+1 For the next accounting period's reward; a represents an accounting action.
8. The method for intelligent accounting and anomaly detection based on deep reinforcement learning according to claim 7, characterized in that: Updated ending balance for: in, is the ending balance of the i-th accounting item at the previous accounting period t; is the timing difference error; T is the standard deviation of the balance of each accounting item in the past 30 consecutive days; t The number of days that have passed in the current fiscal year.
9. The method for intelligent accounting and anomaly detection based on deep reinforcement learning according to claim 8, characterized in that: The anomaly scores are: Among them, A t+1 The abnormality score for the next accounting period; The average balance of each accounting item for the past 30 consecutive days; The average value of the VAT amount of the vouchers for 90 consecutive days; sd(VAT 90 ) is the standard deviation of the VAT amount of the vouchers in the past 90 consecutive days; M t+1 Indicates the total number of accounting documents that have been posted as of t+1 of the next accounting period.