Food enterprise financial data anomaly detection method based on multi-dimensional feature fusion
Patent Information
- Application Number
- CN202611002817.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]针对现有通用财务异常检测方案应用于食品行业时适配性差、业务正常波动误报率高、异常根因追溯难度大、模型迭代改造成本高的技术缺陷,本发明提供一种基于多维度特征融合的食品企业财务数据异常检测方法,通过行业定制化数据预处理、四维特征体系、双层异构模型架构、贝叶斯风险联动定级、数据分布自适应更新机制,降低误报率、提升异常检出率、简化人工核查流程、降低模型运维迭代成本
本申请全程嵌入食品行业专属业务指标,能够过滤原料季节性损耗、临期商品折价等常规经营波动带来的误判,整体异常识别精准度显著优于市面上通用财务风控工具。依托统一批次编码实现异常业务源头一键追溯,财务人员人工核查工作量可减少六成以上。模型支持增量微调,常规数据变化场景下更新耗时相比全量训练降低七成,行业政策、产线调整时仅需针对性重训,无需整体重构整套识别逻辑。同时同步覆盖成本、税务合规、现金流、库存运营四大食品企业核心财务风险,输出异常明细、业务根因、整改方案与中长期风险预判,形成完整闭环管控,配套加密、分级权限机制保障企业敏感业财数据存储、传输安全。
Smart Images

Figure CN122820355A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent financial risk control technology for enterprises, specifically to a method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion. Background Technology
[0002] The food manufacturing industry's business chain has distinct characteristics compared to general manufacturing. Upstream raw materials rely heavily on agricultural product procurement channels, with corresponding special procurement invoices. Production involves multiple SKUs and batches processed simultaneously, leading to significant fluctuations in raw material losses due to natural depletion and storage conditions, influenced by seasonality and warehousing conditions. Finished products have fixed shelf lives, and discounts nearing expiration or disposal after expiration directly impact accounting data. Sales channels are fragmented, with frequent returns, exchanges, and holiday promotions. Taxation is compounded by value-added tax, consumption tax, and special policies for input tax deduction on agricultural products, resulting in frequent updates to tax and financial regulations.
[0003] Currently, most small and medium-sized food enterprises rely on manual verification of documents one by one for financial risk control. This leads to delays in anomaly detection, frequent omissions, and a heavy workload for finance personnel. Some enterprises have introduced general-purpose financial risk control software, but these tools have not been designed with the specific business characteristics of the food industry in mind, resulting in significant shortcomings after implementation.
[0004] The general system relies solely on standardized financial thresholds to identify anomalies, failing to incorporate industry-specific indicators such as raw material losses, near-expiry inventory, and agricultural product acquisition. Normal fluctuations in production and warehousing are easily misinterpreted as financial anomalies, resulting in a false alarm rate exceeding 25% in industry pilot tests, thus limiting the system's reference value. Existing risk control tools only access financial system data and cannot link to front-end production and inventory business documents. After anomalies occur, finance personnel still need to retrieve data from multiple systems for verification, limiting work efficiency. The detection coverage is narrow, targeting only two general risks: invoices and cash flow. It lacks the ability to identify high-frequency risks in the food industry, such as distorted cost allocation, inventory impairment, and fraudulent agricultural product acquisition. The general model is highly rigid, requiring secondary development and reconstruction once industry tax policies or enterprise production lines change, resulting in long modification cycles and high human resource costs.
[0005] In light of the aforementioned industry pain points, existing general-purpose financial anomaly detection tools cannot meet the specific management and control needs of food companies, and there is an urgent need for a customized risk control and detection solution that fits the entire food business chain. Summary of the Invention
[0006] To address the technical shortcomings of existing general financial anomaly detection solutions when applied to the food industry, such as poor adaptability, high false alarm rate due to normal business fluctuations, difficulty in tracing the root causes of anomalies, and high costs of model iteration and modification, this invention provides a method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion. Through industry-customized data preprocessing, a four-dimensional feature system, a two-layer heterogeneous model architecture, Bayesian risk linkage classification, and an adaptive data distribution update mechanism, this method reduces the false alarm rate, improves the anomaly detection rate, simplifies the manual verification process, and reduces the model operation and maintenance iteration costs.
[0007] The basic technical solution of this invention is as follows: A method for detecting anomalies in the financial data of food enterprises based on multi-dimensional feature fusion includes the following steps: S1. Collect financial data and food industry-specific business data from the entire business process of food enterprises, and obtain a standardized dataset through cleaning, deduplication, and format standardization preprocessing; the business data includes at least one of production batches, raw material losses, inventory shelf life, and agricultural product procurement information; the preprocessing includes weighted KNN missing value imputation based on batch similarity, and establishing a batch coding mapping for the entire chain of procurement, production, inventory, and sales. S2. Based on the standardized dataset, construct a multi-dimensional feature set containing cost features, compliance features, and funding features. Use mutual information method combined with variance analysis to filter out core features, and perform cycle-aware attention-weighted fusion on the core features. S3. Build a two-layer architecture multi-dimensional anomaly detection model, the two-layer architecture including a single-dimensional special detection layer and a cross-dimensional risk linkage and classification layer; The single-dimensional special detection layer sets up independent detection sub-models for each of the four types of features, and outputs the single-item anomaly detection results for each dimension respectively; the four types of sub-models are, in order, the random forest cost anomaly sub-model, the improved FocalLoss logistic regression compliance risk sub-model, the LSTM time series funding anomaly sub-model, and the random forest operation anomaly model. The cross-dimensional risk linkage and rating layer is based on a Bayesian network to construct two directed acyclic graphs of risk transmission: cost → operation → capital and compliance → capital. Based on the risk transmission relationship, it jointly calculates the single abnormal results and outputs the comprehensive risk level. S4. Based on the comprehensive risk level, output the anomaly detection list, root cause analysis, rectification suggestions, and risk trend prediction for the next 1-3 months; S5. Establish a dynamic model update mechanism: periodically detect the degree of distribution deviation of newly added data, and select incremental training or full retraining to update model parameters according to the deviation threshold; trigger emergency full retraining when there are major changes in industry tax policies or enterprise business processes.
[0008] The method also includes data security measures: performing TLS encryption for transmission and AES encryption for storage on all data throughout the process; setting tiered access permissions; and automatically backing up data periodically.
[0009] Step S1: Weighted KNN missing value imputation. Extract raw material category, production season, ambient temperature and humidity, production line number, and raw material shelf life to construct batch feature vectors; calculate the cosine similarity between the batch to be imputed and historical complete batches, selecting the k most similar historical batches as reference samples; calculate the imputed value using cosine similarity as the weight. The imputed value calculation formula is: ; in, Here are the fill values for the fields to be filled, and k is the number of reference samples. For the i-th reference sample, the corresponding field value is... Let be the cosine similarity between the batch to be filled and the i-th reference sample.
[0010] To address the industry characteristics of food companies, such as high batch data missing rates and strong correlation between business data within the same batch, this method replaces traditional mean and median interpolation methods, significantly improving the accuracy of missing data restoration. In this example, k is set to 7, which is an optimal parameter and can be adjusted as needed by those skilled in the art.
[0011] Step S2 involves attention-weighted fusion of periodic perception, dividing the sliding time window by natural month to generate periodic training subsets; and assigning periodic weights based on the detection accuracy of each dimension's features on the corresponding periodic validation set. The weight calculation formula is as follows: ; in, Let be the weight of the d-th type of feature in the t-th period. Let be the anomaly detection accuracy of the d-th feature on the validation set in the t-th period.
[0012] The food industry is characterized by distinct peak and off-peak seasons. This mechanism dynamically adjusts the weights of different cyclical features to eliminate the loss of detection accuracy caused by fixed weights.
[0013] The four sub-models and the improved FocalLoss, the random forest cost anomaly detection sub-model outputs a single batch unit cost benchmark value, and if the deviation between the actual cost and the benchmark value exceeds a preset threshold, the cost is judged to be abnormal. The LSTM time-series funding anomaly detection sub-model predicts funding flows based on time-series data and identifies cash flow gap anomalies. The random forest operation anomaly detection sub-model is used to identify operational and financial anomalies such as near-expiry inventory and gross profit margin fluctuations. The compliance risk detection sub-model is trained using an improved FocalLoss loss function. The loss function is: ; In the formula, N is the total number of samples. For the true labels of the samples, Balance coefficients are used to predict the probability of anomalies in the model. The positive and negative sample balance coefficient has a value of 0.83, and the focusing parameter is... To focus the parameters, a value of 2 is set. This loss function mitigates the missed detection problem caused by imbalanced compliance samples.
[0014] Bayesian network risk linkage classification, Bayesian network calculation of joint probability formula for comprehensive risk: ; In the formula, For the joint probability of comprehensive risks, For the i-th risk node, Let be the set of parent nodes of the i-th risk node. This represents the probability of the current node's risk occurring under the condition that the parent node is abnormal.
[0015] The system establishes a directed acyclic graph of risk transmission between cost, operation, and funding / compliance, quantifying the probability of risk transmission and replacing simple grading rules. The grading results are more in line with actual business practices.
[0016] The KS test adaptive model update mechanism has a regular update cycle of 1-3 months. Each month, the KS test is used to calculate the data distribution shift between the newly added data and the training set. If the shift is below a preset threshold, incremental training is performed; if the shift is above the preset threshold, full retraining is performed. The optimal distribution shift threshold is 0.15, but it can be adjusted based on the company's data situation. In the event of significant changes in financial and tax policies, production lines, or main business operations, emergency full retraining is initiated immediately.
[0017] Batch code mapping and root cause tracing: In the preprocessing stage, a full-link batch code mapping is established, binding the purchase batch, production batch, inventory batch, and sales batch one by one; the root cause location output in step S4 relies on this mapping to directly lock the purchase, production, inventory, or sales business links corresponding to the anomaly, without the need for manual verification of documents across systems.
[0018] Compared with existing general solutions, the technical effects of the present invention are as follows: This application incorporates food industry-specific business indicators throughout the entire process, effectively filtering out misjudgments caused by seasonal raw material losses and price discounts on near-expiry products. Its overall anomaly identification accuracy is significantly superior to general financial risk control tools on the market. Leveraging unified batch coding, it enables one-click tracing of the source of abnormal business activities, reducing the workload of manual verification by financial personnel by more than 60%. The model supports incremental fine-tuning; update time in scenarios of regular data changes is reduced by 70% compared to full training. When industry policies or production lines are adjusted, only targeted retraining is required, without the need to completely reconstruct the entire identification logic. Simultaneously, it covers four core financial risks for food companies: cost, tax compliance, cash flow, and inventory operations. It outputs detailed anomaly information, root causes, rectification plans, and medium- to long-term risk predictions, forming a complete closed-loop management system. An encrypted and tiered access control mechanism ensures the security of sensitive financial data storage and transmission. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the overall process of the present invention: a method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion. Figure 2 This is a hierarchical architecture diagram of the two-layer architecture multi-dimensional anomaly detection model of the present invention. Detailed Implementation
[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0021] Those skilled in the art will understand that the specific parameter values in the following embodiments are preferred options and can be adapted to the actual business scale and data characteristics of the enterprise, all of which fall within the protection scope of the present invention.
[0022] like Figure 1 As shown, the anomaly detection method for financial data of food enterprises based on multi-dimensional feature fusion described in this invention has an overall process divided into five core stages: data collection and preprocessing, feature construction and fusion, model detection and classification, result output, and dynamic updating. Specific implementation details are as follows: Step S1: Full-link business and financial data collection and standardized preprocessing: Establish data collection interfaces covering six major business processes: procurement, production, inventory, sales, taxation, and operations. The interfaces connect to the enterprise's existing ERP system, MES production management system, financial accounting system, and VAT invoice management platform. The data synchronization mode supports two optional modes: daily scheduled batch synchronization and real-time streaming synchronization.
[0023] The detailed financial data and food industry-specific business data collected at each stage are as follows: The procurement process includes accounts payable ledgers, VAT purchase invoices, and bank payment records on the financial side; and supplier qualification files, various raw material categories, purchase prices, complete sets of business documents including agricultural product purchase invoices, and raw material shelf-life information on the business side.
[0024] The financial side of the production process includes manufacturing expense details, labor cost, and monthly production cost transfer vouchers; the business side includes unique production batch numbers, material input for each batch, finished product output, loss classification ledger (distinguishing between natural loss, process loss, and management loss), working hours for each production line, and records of seasonal temporary workers.
[0025] The inventory process includes, on the financial side, the original book value of inventory and records of inventory write-downs; on the operational side, it includes batch inventory entry ledgers, near-expiry warning thresholds, near-expiry item discount disposal documents, and monthly inventory count profit and loss data.
[0026] The sales process includes, on the financial side, main business revenue flow, accounts receivable details, and VAT sales invoices; on the business side, sales orders, channel classification files, product return and exchange records, and promotional documents for near-expiry products.
[0027] The tax and operations process includes monthly / annual tax returns for all tax types, structured information on all invoices, the business scope of the enterprise's food production license, and applicable industry tax and financial policies for the current period.
[0028] The basic standardized preprocessing operations include outlier removal, full data deduplication, and unified field data formatting, which convert unstructured documents such as paper invoices and quality inspection reports after OCR recognition into structured data tables.
[0029] In response to the characteristics of the food industry, such as large amounts of missing batch data and strong correlation between data from the same season and production line, a weighted KNN missing value imputation operation based on batch similarity is performed, thus fully executing all the steps described in claim 2: A five-dimensional batch feature vector is constructed by extracting raw material category, production season, workshop temperature and humidity, production line number, and raw material shelf life. Cosine similarity is used to match complete historical batches. In this embodiment, the highest similarity k=7 groups of historical batches are selected as reference samples. The missing field imputation value is calculated using the following weighted mean formula: ; in, To fill in the values for the field to be filled; k is the number of reference samples. This represents the field value corresponding to the i-th reference sample; Let k be the cosine similarity between the batch to be filled and the i-th reference sample. k=7 is the optimal parameter, which can be adjusted by the enterprise according to its total amount of data.
[0030] Preprocessing synchronously completes the full-link batch coding mapping, establishing a unique binding relationship between procurement batches, production batches, inventory batches, and sales batches, providing a data association basis for subsequent anomaly root cause localization.
[0031] Step S2: Multi-dimensional feature construction and periodic-aware attention-weighted fusion: Based on the preprocessed standardized dataset, four complete feature sets are constructed, with the specific metrics for each feature set as follows: Cost characteristics include raw material procurement cost deviation rate, overall production loss rate, manufacturing overhead allocation deviation rate, seasonal labor cost ratio, and unit cost fluctuation coefficient per batch. Compliance characteristics include invoice compliance rate, matching degree of purchase and sales categories, matching degree of authenticity of agricultural product purchase business, deviation rate between accounting data and tax declaration data, and suitability of business scope and invoice categories; Capital characteristics include capital turnover deviation, accounts receivable collection rate fluctuation coefficient, daily cash flow fluctuation range, accounts payable overdue rate, and operating cash flow deviation. Operational characteristics include the ratio of inventory turnover days to shelf life, the proportion of near-expiry inventory, the order fulfillment anomaly rate, the monthly fluctuation range of gross profit margin, and the profit and loss deviation index per SKU.
[0032] Feature selection employs a combination of mutual information and analysis of variance to eliminate redundant features with a correlation threshold below 0.1, retaining only effective core features. Subsequently, multiple periodic training subsets are generated by dividing the time window into natural months, and the periodic weights are dynamically allocated to the feature weights of each dimension using a periodic weight calculation formula. ; in, The weights of the d-th type of features in the t-th period; Let be the anomaly detection accuracy corresponding to the d-th type of feature in the validation set during the t-th period. The weighted and fused feature set is then input into the lower-level detection model.
[0033] Step S3: Multi-dimensional anomaly detection using a two-layer architecture: This layer consists of two execution logic layers: a single-dimensional special detection layer and a cross-dimensional risk linkage and grading layer.
[0034] Single-dimensional specialized detection layer, four types of sub-model operating rules: Random Forest Cost Anomaly Sub-model: Based on historical complete batch training, the unit cost benchmark ranges for each production line, different seasons, and various raw materials are trained. In this embodiment, a 5% benchmark deviation threshold is set. If the actual cost of a single batch exceeds the threshold, the cost is judged to be abnormal. At the same time, three types of abnormal causes are distinguished: raw material price increase, excessive production loss, and incorrect allocation of manufacturing costs. Improved FocalLoss Logistic Regression Compliance Risk Sub-model: Strictly adopt the loss function described in claim 4 to complete model training. ; Where N is the total number of samples; For the true labels of the samples, Balance coefficients are used to predict the probability of anomalies in the model. The positive and negative sample balance coefficient is 0.83, and the focusing parameter is... To focus on the parameter, the default value is 2. The model outputs a compliance risk score in the range of 0-100, defining the classification standards: 60 points and below is low risk, 60-80 points is medium risk, and above 80 points is high risk, and the corresponding compliance risk points are output simultaneously. LSTM Time Series Cash Flow Anomaly Sub-model: Using historical continuous cash flow as time series input, it predicts the cash inflow and outflow trends for the next 1 to 3 months. If the cash balance at the end of the forecast period is lower than the enterprise's preset safe cash flow threshold, it automatically determines the cash flow gap anomaly and marks the estimated gap period and gap amount. Random Forest Operational Anomaly Sub-model: Identifies financial anomalies derived from inventory and sales operations, such as large amounts of near-expiry inventory and significant fluctuations in monthly gross profit margin.
[0035] Cross-dimensional risk linkage classification layer: Two directed acyclic graphs (DAGs) for risk transmission are pre-built, with transmission paths of cost → operation → funding and compliance → funding. Conditional probability tables for each node of the Bayesian network are trained using historical risk cases of the enterprise. When a single anomaly is detected simultaneously across multiple dimensions, the overall risk occurrence probability is calculated according to the joint probability formula in claim 5. ; in, For the combined probability of comprehensive risks; For the i-th dimension risk node; Let be the set of parent nodes of the i-th risk node; This represents the conditional probability of the current node being at risk when the parent node is abnormal.
[0036] Based on the calculated joint probability interval, three levels of comprehensive risk are mapped: low, medium, and high. High-level risks are simultaneously pushed to the company's financial officer and corresponding business officer for collaborative handling.
[0037] Step S4, output the detection result: The system outputs four categories of standardized content based on the overall risk level, namely: First, an anomaly detection checklist, which specifies the characteristic dimension to which the anomaly belongs, the comprehensive risk level, the unique associated business batch, and the business process in which the anomaly occurred; Second, the root cause localization explanation is that, based on the full-link batch coding mapping established by S1, the abnormal source business node is traced in reverse, corresponding to the technical feature of claim 9. Third, standardized rectification suggestions for different scenarios, providing feasible control and adjustment solutions for various anomalies in cost, compliance, funding, and operations; Fourth, the forecast of medium- to long-term risk trends over the next 1-3 months helps companies adjust their operational and financial management strategies in advance.
[0038] Step S5, Model Dynamic Update Mechanism: The regular update cycle is set to 1-3 months, and the preferred cycle in this embodiment is 2 months. Each month, new business and financial samples are extracted, and the KS test is used to calculate the degree of data distribution deviation between the new samples and the original training set. In this embodiment, the deviation judgment threshold is preferably 0.15. If the calculated deviation is lower than the preset threshold, only incremental training is performed to fine-tune the model parameters. If the deviation is higher than the threshold, the full retraining process is started.
[0039] When the state introduces major new fiscal and tax policies for the food industry, enterprises carry out large-scale transformations of their production lines, or their main business categories change, there is no need to wait for the regular update cycle. Emergency full retraining is triggered immediately, and feature weights and risk assessment rules are adjusted simultaneously.
[0040] The data security measures accompanying this invention are as follows: end-to-end data transmission uses the TLS encryption protocol, and database storage uses AES encryption; hierarchical data access permissions are set, with business personnel only able to view their own business data, finance personnel able to view complete business and financial data and all anomaly detection results, and system administrators having permissions for model configuration and parameter adjustment; the system backend automatically performs full data backups weekly to avoid the risk of data loss or tampering.
[0041] The following examples, using specific enterprise implementations, further illustrate the effectiveness of this solution: Example: A medium-sized domestic condiment and food company was selected as the implementation target. This company mainly produces soy sauce and vinegar, has three standardized food production lines, sells 12 SKUs, and has annual revenue of approximately 120 million yuan. Before implementing this solution, the company relied solely on general financial software and manual verification by financial personnel, resulting in long-standing management pain points such as large cost accounting errors, time-consuming verification of compliance with agricultural product purchase invoices, and delayed prediction of cash flow gaps.
[0042] The overall implementation process is as follows: System integration phase: Integrate the enterprise's four business systems: ERP, MES, financial accounting, and invoice management, and complete the unified mapping of batch codes across the entire chain of procurement, production, inventory, and sales; To address the issue of high missing rates in the raw material loss field, the weighted KNN algorithm of this invention is used to fill in the missing data, achieving an accuracy rate of 92.3%, which is 14.3 percentage points higher than the traditional mean interpolation method; Feature engineering phase: Initially, 28 multi-dimensional features are constructed. After screening using mutual information and variance analysis, 19 core features are retained. Attention weighting is performed monthly according to the natural month window, and the weights of various features are automatically adjusted during peak and off-peak seasons. Model training phase: Using the company's complete labeled business and financial data from the previous two years, four types of special sub-models and a Bayesian network risk transmission model were trained. After manual review and parameter adjustment, the model was tested for 3 months. During routine operation and maintenance: The model adopts a monthly KS distribution test and incremental training update mode, and performs full retraining only when there are major business adjustments.
[0043] After the system has been running stably for 6 months, the actual control effect is as follows: Cost accounting deviation has decreased from 8%~12% to 2.7%, and the cost anomaly detection rate has reached 94%, which can accurately pinpoint the corresponding production batch and the cause of loss; The workload for manual verification of agricultural product purchase invoices and sales invoices has been reduced from 2 person-weeks per month to 0.5 person-days, with a compliance risk false alarm rate of only 3.8% and no compliance omissions. The lead time for predicting cash flow gaps has been increased from 7 days to 45 days, effectively reducing the cost of temporary borrowing and financing for enterprises. The overall financial document verification time for enterprises was reduced by 65%, the workload of manual verification of anomalies was reduced by 62%, and the overall risk control operating cost decreased by 7.2% throughout the year.
[0044] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention; any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion, characterized in that, Includes the following steps: S1. Collect financial and business data from the entire business process of food enterprises, and obtain a standardized dataset after cleaning, deduplication, and format standardization preprocessing; the business data includes at least one of production batches, raw material losses, inventory shelf life, and agricultural product procurement information; the preprocessing includes weighted KNN missing value imputation based on batch similarity, and establishing a batch coding mapping for the entire chain of procurement, production, inventory, and sales. S2. Based on the standardized dataset, construct a multi-dimensional feature set including cost features, compliance features, funding features, and operational features. Use mutual information method combined with variance analysis to filter out core features, and perform cycle-aware attention-weighted fusion on the core features. S3. Construct a two-layer architecture multi-dimensional anomaly detection model, comprising a single-dimensional specialized detection layer and a cross-dimensional risk linkage and rating layer. The single-dimensional specialized detection layer sets up independent detection sub-models for four types of features, and outputs the single-item anomaly detection results for each dimension. The four sub-models are, in order, a random forest cost anomaly sub-model, an improved FocalLoss logistic regression compliance risk sub-model, an LSTM time-series funding anomaly sub-model, and a random forest operation anomaly model. The cross-dimensional risk linkage and rating layer constructs two directed acyclic graphs of risk transmission: cost → operation → funding and compliance → funding, based on a Bayesian network. It jointly calculates the single-item anomaly results based on the risk transmission relationship and outputs the comprehensive risk level. S4. Based on the comprehensive risk level, output an anomaly detection list, root cause analysis, rectification suggestions, and risk trend forecast for the next 1-3 months; S5. Establish a dynamic model update mechanism, periodically detect the degree of data distribution deviation, and select incremental training or full retraining to update model parameters according to the deviation threshold; trigger emergency full retraining when there are major changes in industry tax policies or enterprise business processes.
2. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that... The weighted KNN missing value imputation step in step S1 includes: Extract raw material category, production season, ambient temperature and humidity, production line number, and raw material shelf life to construct batch feature vectors; calculate the cosine similarity between the batch to be filled and historical complete batches, and select the k historical batches with the highest similarity as reference samples; calculate the missing field filling value using cosine similarity as the weight, and the filling value calculation formula is as follows: ; in, Here are the fill values for the fields to be filled, and k is the number of reference samples. Let be the value of the field corresponding to the i-th reference sample. Let be the cosine similarity between the batch to be filled and the i-th reference sample.
3. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that... : Step S2, the attention-weighted fusion for periodic awareness, includes: dividing the sliding time window by natural month to generate a periodic training subset; and assigning periodic weights based on the detection accuracy of each dimension's features on the corresponding periodic validation set. The weight calculation formula is as follows: ; in, Let be the weight of the d-th feature in the t-th period. Let be the anomaly detection accuracy of the d-th feature on the validation set in the t-th period.
4. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that... : In step S3, the compliance risk detection sub-model is trained using an improved FocalLoss loss function. The loss function formula is as follows: ; Where N is the total number of samples, For the true labels of the samples, The balance coefficient represents the anomaly probability predicted by the model. The positive and negative sample balance coefficient has a value of 0.83, and the focusing parameter is... This is the focus parameter, and its value is 2.
5. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that... : In step S3, the Bayesian network calculates the joint probability formula for the overall risk: ; in, For the joint probability of comprehensive risks, For the i-th dimension risk node, Let be the set of parent nodes of the i-th risk node. This represents the probability of the current node's risk occurring under the condition that the parent node is abnormal.
6. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that... : The specific steps of the S5 model dynamic update mechanism are as follows: the regular update cycle is 1-3 months, and the KS test is used every month to calculate the data distribution deviation between the newly added data and the training set; if the deviation is lower than the preset threshold, incremental training is performed, and if the deviation is higher than the preset threshold, full retraining is performed.
7. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that... : In step S1, the preprocessing stage establishes a full-link batch coding mapping, binding the procurement batch, production batch, inventory batch, and sales batch one by one for the purpose of tracing the root cause of anomalies throughout the entire chain.
8. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that... The functions of the four sub-models are as follows: The random forest cost anomaly detection sub-model outputs a single batch unit cost benchmark value. If the deviation between the actual cost and the benchmark value exceeds a preset threshold, the cost is judged to be abnormal. The LSTM time-series funding anomaly detection sub-model predicts funding flows based on time-series data and identifies cash flow gap anomalies. The Random Forest Operational Anomaly Detection Sub-model is used to identify operational and financial anomalies such as near-expiry inventory and gross profit margin fluctuations.
9. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that: The root cause localization description output in step S4 is based on the end-to-end batch coding mapping relationship to locate the procurement, production, inventory or sales business links corresponding to the anomaly.
10. The method for detecting anomalies in financial data of food enterprises based on multi-dimensional feature fusion according to claim 1, characterized in that: The method also includes data security measures: performing TLS encryption for transmission and AES encryption for storage on all data throughout the process; setting tiered access permissions; and automatically backing up data periodically.