Intelligent prediction and control method and system for tobacco financial key indicators based on large models

By constructing a tobacco causal labeling system and decoupling multi-source heterogeneous data with a multi-branch variational autoencoder, and combining tobacco financial knowledge graph and Transformer large language model, the problems of lack of causal mechanism and compliance constraints in financial forecasting in the tobacco industry are solved, and efficient and interpretable forecasting and control of key financial indicators are achieved.

CN122115134APending Publication Date: 2026-05-29GUANGDONG TOBACCO YANGJIANG CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG TOBACCO YANGJIANG CO LTD
Filing Date
2026-02-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing financial forecasting models lack support from the causal mechanisms of the tobacco industry, making it difficult to distinguish between multiple influencing factors such as policy-driven, operational execution, and market feedback. This results in large forecasting biases, poor interpretability, and delayed response. Furthermore, they lack compliance constraints and cannot achieve multi-dimensional verification and confidence quantification.

Method used

A tobacco causal labeling system is constructed to decouple multi-source heterogeneous data from a multi-branch variational autoencoder, generating a causal feature decoupling model. This model is then used to make predictions by combining a tobacco financial knowledge graph and a Transformer large language model. Furthermore, through multi-dimensional compliance verification and closed-loop feedback optimization, the predicted data is dynamically adjusted to achieve compliance and confidence labeling.

Benefits of technology

It significantly improves the interpretability and business consistency of financial indicator forecasts, realizes closed-loop management from single numerical forecasts to credibility perception, enhances resource allocation efficiency and risk response capabilities, and has strong adaptability and long-term robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115134A_ABST
    Figure CN122115134A_ABST
Patent Text Reader

Abstract

The application discloses a kind of tobacco financial key index intelligent prediction and control method and system based on large model, belong to artificial intelligence and tobacco industry financial management cross technical field.The method collects tobacco industry multi-source heterogeneous data, after standardization, based on tobacco causal label system is automatically annotated, constructs causal characteristic decoupling model to separate policy, operation and market factor;Synchronous construction dynamic tobacco financial knowledge graph and generate graph embedding;Decoupling factor and graph embedding are fused and input into the transformer large model adapted to tobacco field, and the prediction data is output under the embedding financial compliance constraint;Further, through the multi-dimensional verification of budget, tariff and cash flow, the results with confidence markers are generated, and the model and graph are optimized based on historical deviation mode feedback. The method realizes the intelligent prediction and dynamic control of high-precision, interpretable and closed-loop controllable tobacco financial indicators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and tobacco industry financial management, and in particular relates to a method and system for intelligent prediction and control of key tobacco financial indicators based on a large model. Background Technology

[0002] In the context of a highly planned and heavily regulated tobacco industry, accurate forecasting and effective control of key financial indicators (such as per-box manufacturing cost, total tax and profit, and inventory turnover rate) are crucial for ensuring national tax revenue and the stable operation of enterprises. However, existing financial forecasting methods largely rely on traditional statistical models or shallow machine learning, making it difficult to effectively integrate heterogeneous data from multiple sources, such as policy documents, production ledgers, and market sentiment. Furthermore, they lack explicit modeling of specific business rules related to the tobacco monopoly system, tobacco leaf purchase quotas, and consumption tax adjustments. Simultaneously, significant data silos exist between industrial companies, commercial companies, and the monopoly management department, making it impossible for forecasting models to distinguish whether indicator fluctuations stem from uncontrollable policy changes, manageable operational anomalies, or lagging market feedback. This results in large forecast biases, poor interpretability, and delayed responses. While large language models have demonstrated strong generalization capabilities in recent years, their direct application to financial scenarios can easily produce "illusory" outputs and are difficult to embed into industry compliance constraints. Furthermore, existing systems generally lack a closed-loop mechanism from prediction to verification to model optimization, and cannot dynamically adjust feature decoupling logic and knowledge representation structure based on actual deviations. This makes them ill-suited to the frequent policy adjustments and rigid planning characteristics of the tobacco industry. Therefore, there is an urgent need for an intelligent prediction and control technology solution that deeply integrates tobacco field knowledge, supports causal factor separation, possesses compliance constraints, and can achieve continuous evolution. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for intelligent prediction and control of key financial indicators in the tobacco industry based on a large model. This addresses the technical problems in existing technologies, such as the lack of causal mechanism support in financial prediction models for the tobacco industry, difficulty in distinguishing multiple influencing factors such as policy-driven, operational execution and market feedback, and the reliance on static rules or isolated data sources, which prevent multi-dimensional collaborative verification and confidence quantification, resulting in inaccurate predictions and lagging control.

[0004] To achieve the above objectives, a first aspect of the present invention provides a method for intelligent prediction and control of key financial indicators in the tobacco industry based on a large model, comprising the following steps: Collect multi-source heterogeneous data from the tobacco industry, preprocess the multi-source heterogeneous data, and obtain standardized data; Automatic causal semantic annotation is performed on the standardized data based on the pre-constructed tobacco causal labeling system to generate training data with causal labels. A causal feature decoupling model is trained and constructed based on the training data with causal labels. The standardized data is input into the causal feature decoupling model. Feature channel filtering and reorganization of the multi-source heterogeneous data are performed through a mask attention mechanism based on causal labels to output policy driving factors, operational execution factors and market feedback factors. A tobacco financial knowledge graph is dynamically constructed based on historical and real-time training data with causal labels. Subgraphs related to the current prediction task are retrieved from the knowledge graph and then encoded by a graph neural network to output graph embedding data. The decoupled policy-driven factors, operational execution factors, and market feedback factors are fused with the graph-embedded data and fed into a prediction model adapted for the tobacco industry. The prediction model is a large language model based on the Transformer architecture and integrating financial compliance constraint mechanisms, and outputs prediction data of key financial indicators. Based on the predicted data, multi-dimensional compliance verification is performed, including budget matching degree, tax compliance and cash flow coverage analysis, to generate predicted data with confidence level labels; The verification results are compared with historical deviation patterns, and the prediction data with confidence level labels are dynamically adjusted to achieve closed-loop feedback optimization. The dynamic adjustment process involves matching the current verification result with a preset historical deviation pattern library for similarity, calculating the deviation correction coefficient, and performing weighted iterative updates on the prediction data with confidence labels until it converges to the optimal prediction interval.

[0005] Furthermore, it also includes, The multi-source heterogeneous data includes structured production and financial data from industrial companies' ERP systems, policy documents issued by national and provincial tobacco monopoly bureaus, inventory ledgers of commercial companies, industry public opinion texts, and external environment data. The multi-source heterogeneous data is standardized with tobacco terminology and spatiotemporally aligned with fiscal year cycles. Based on a pre-built tobacco causal labeling system, each data record is automatically labeled with causal semantic tags, which include policy causes, operational causes, environmental causes, and corresponding financial effect tags, thereby obtaining the standardized data.

[0006] Furthermore, the construction of the causal feature decoupling model includes the following sub-steps: The multi-source heterogeneous data is automatically labeled based on a preset tobacco causal labeling system to generate training data with causal labels. Based on the training data with causal labels, a variational autoencoder with a multi-branch latent representation structure is constructed, and an attention mask is applied to the data corresponding to different causal labels during the encoding process. The training data is then subjected to feature aggregation within the corresponding latent branches. A regularization term is introduced to minimize the mutual information among policy-driven factors, operational execution factors, and market feedback factors; The trained model is used to process standardized data for the current cycle, outputting decoupled policy-driven factors, operational execution factors, and market feedback factors.

[0007] Furthermore, the decoupled policy-driven factors, operational execution factors, and market feedback factors are fused with the graph-embedded data, including: The policy-driven factors, operational execution factors, and market feedback factors are denoted as vectors. , , The graph embedding data is denoted as a vector. A fused input vector is generated through concatenation operations. ; The prediction model uses the following loss function for training or inference constraints when generating output: , in, To predict the loss term, The predicted values ​​of key financial indicators output by the model. For a predefined set of financial compliance rules for the tobacco industry, each rule Corresponding to a constraint function ,when A value greater than 0 indicates that the predicted value violates the rule. , These are non-negative weighting coefficients.

[0008] Furthermore, the causal feature decoupling model includes an encoder and a decoder. The encoder has three latent variable branches, corresponding to policy-driven, operational execution, and market feedback causal labels, respectively. When encoding training data with causal labels, the training data is input to the corresponding latent variable branch according to the causal label attached to the training data, and the key-value pairs of the other two branches are masked in the attention calculation, so that the training data generates a latent representation only within its own branch.

[0009] Furthermore, the decoupled policy-driven factors, operational execution factors, and market feedback factors are input into the prediction model to obtain predicted values ​​for key indicators. The policy-driven factors, operational execution factors, and market feedback factors are mapped to the graph embedding data as vector representations of a unified dimension, and then concatenated to form a fusion input vector, which is fed into the input embedding layer of the prediction model. The fusion input vector is then subjected to contextual semantic encoding by the multi-layer converter encoder of the prediction model and fed into the decoder layer for autoregressive generation. During the generation process of the decoder layer, the embedded rule constraint module is called to obtain the predefined set of financial compliance rules for the tobacco industry in real time, and the currently generated intermediate output words or candidate final output words are substituted into the constraint functions in the set of financial compliance rules for compliance determination. If the current term is determined to violate any financial compliance rule, a negative bias is applied to the logical probability value of the term to suppress its selection; if it is determined to comply with the rule, the original logical probability value is maintained. Based on the distribution of logical probability values ​​after applying a negative bias, or by selecting the next term, the above encoding, compliance determination, and bias correction processes are iteratively executed until a complete sequence is generated as the prediction data for the aforementioned key financial indicators.

[0010] Furthermore, after the predicted data is output, the following verification operations are performed: the predicted data is compared with the preset budget indicator data, the tobacco industry applicable tax rules data, and the enterprise cash flow constraint data respectively to obtain the budget matching degree verification result, the tax rule adaptability verification result, and the cash flow coverage capability verification result; based on whether the three verification results meet their respective preset threshold conditions, three Boolean verification flags are generated. Based on a pre-set confidence mapping table, the combination of the three Boolean verification flags is mapped to a confidence flag, which is an element in a pre-set discrete level set.

[0011] Furthermore, the method includes matching the verification result with a pre-stored historical prediction deviation pattern dataset to obtain a matching deviation pattern identifier; and querying the corresponding model adjustment instruction and graph update instruction from a preset strategy mapping table based on the deviation pattern identifier. The model adjustment instruction is used to modify the training parameters and loss function weights of the causal feature decoupling model, and the graph update instruction is used to adjust the weight update rate and time decay coefficient of at least one causal edge in the tobacco financial knowledge graph.

[0012] Secondly, this invention also provides an intelligent prediction and control system for key financial indicators of tobacco based on a large model, and applies a method for intelligent prediction and control of key financial indicators of tobacco based on a large model, including: The data acquisition module is used to collect multi-source heterogeneous data from the tobacco industry, preprocess the multi-source heterogeneous data, and obtain standardized data. The feature decoupling module is used to automatically perform causal semantic annotation on the standardized data based on a pre-built tobacco causal labeling system, generate training data with causal labels, train and construct a causal feature decoupling model based on the training data with causal labels, input the standardized data into the causal feature decoupling model, and perform feature channel filtering and reorganization on the multi-source heterogeneous data through a mask attention mechanism based on causal labels, and output policy driving factors, operational execution factors and market feedback factors. The graph embedding module is used to dynamically construct a tobacco financial knowledge graph based on historical and real-time training data with causal labels, and to retrieve subgraphs related to the current prediction task from the knowledge graph, and output graph embedding data after encoding by a graph neural network. The intelligent prediction module is used to fuse the decoupled policy-driven factors, operational execution factors, and market feedback factors with the graph-embedded data, and feed them as input into a prediction model adapted for the tobacco industry. The prediction model is a large language model based on the Transformer architecture and integrating financial compliance constraint mechanisms, and outputs prediction data of key financial indicators. The compliance verification module is used to perform multi-dimensional compliance verification based on the predicted data, including budget matching degree, tax compliance and cash flow coverage analysis, and generate predicted data with confidence level labels. The feedback optimization module is used to compare the verification results with historical deviation patterns and dynamically adjust the prediction data with confidence level labels to achieve closed-loop feedback optimization.

[0013] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in any one of the claims.

[0014] The beneficial technical effects of the present invention are at least as follows: This invention provides a method and system for intelligent prediction and control of key financial indicators in the tobacco industry based on a large model, which has the following outstanding advantages: Firstly, by constructing a tobacco causal labeling system and a multi-branch variational autoencoder, the mixed multi-source heterogeneous data is decoupled into policy-driven factors, operational execution factors, and market feedback factors. Simultaneously, a dynamically updated tobacco financial knowledge graph is constructed, transforming unstructured information such as policy documents, production ledgers, and public opinion texts into causal relationship edges with time-weighted attributes. This technical approach enables financial indicator predictions to not only rely on historical time-series statistical patterns but also embed the unique "policy-operation-market" mechanism of the tobacco industry, significantly improving the interpretability and business consistency of the prediction data. It effectively solves the problems of misjudgment or delayed response caused by the neglect of causal paths in traditional models.

[0015] Secondly, the system integrates decoupling factors and graph embedding during the forecasting process, inputting them into a large Transformer model adapted for the tobacco industry. During the inference phase, a multi-dimensional compliance verification mechanism based on budget matching, tax compliance, and cash flow coverage is embedded to generate forecast data with confidence levels. When the verification results trigger different risk levels, the system automatically adjusts the conservatism of subsequent control strategies—supporting refined production scheduling under high confidence and initiating contingency plan reviews under low confidence. This mechanism achieves a shift from "single numerical forecasting" to "closed-loop control based on credibility perception," improving resource allocation efficiency and risk response capabilities while ensuring financial compliance.

[0016] Third, the entire process forms a complete closed loop of "data collection - causal decoupling - graph reasoning - compliance prediction - verification feedback - model optimization". Key parameters such as the loss weight of the causal feature decoupling model and the edge decay coefficient of the knowledge graph can be automatically adjusted according to the verification results and historical deviation patterns. When encountering major policy adjustments or changes in corporate organization, the system can quickly reconstruct the factor distribution and graph structure based on new data, ensuring that the prediction logic is always aligned with the current tobacco business environment, and has strong adaptability and long-term operational robustness. Attached Figure Description

[0017] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the working steps of a method for intelligent prediction and control of key financial indicators in the tobacco industry based on a large model, as disclosed in one embodiment of the present invention. Figure 2 This is a flowchart illustrating the decoupling process steps of the decoupling factor disclosed in one embodiment of the present invention. Figure 3 This is a schematic diagram of a tobacco financial key indicator intelligent prediction and control system based on a large model, disclosed in one embodiment of the present invention. Detailed Implementation

[0019] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0020] Example 1 refer to Figure 1This invention provides an embodiment of a method for intelligent prediction and control of key financial indicators in the tobacco industry based on a large model, comprising the following steps: S1. Collect multi-source heterogeneous data from the tobacco industry, preprocess the multi-source heterogeneous data, and obtain standardized data; S2. Based on the pre-constructed tobacco causal labeling system, the standardized data is automatically labeled with causal semantics to generate training data with causal labels. A causal feature decoupling model is trained and constructed based on the training data with causal labels. The standardized data is input into the causal feature decoupling model. The feature channels of the multi-source heterogeneous data are filtered and reorganized through a mask attention mechanism based on causal labels, and policy driving factors, operational execution factors and market feedback factors are output. S3. Dynamically construct a tobacco financial knowledge graph based on historical and real-time training data with causal labels, and retrieve subgraphs related to the current prediction task from the knowledge graph, and output graph embedding data after encoding by a graph neural network; S4. The decoupled policy-driven factors, operational execution factors, and market feedback factors are fused with the graph-embedded data and fed into a prediction model adapted for the tobacco industry. The prediction model is a large language model based on the Transformer architecture and integrating financial compliance constraint mechanisms, and outputs prediction data of key financial indicators. S5. Perform multi-dimensional compliance verification based on the predicted data, including budget matching degree, tax compliance and cash flow coverage analysis, and generate predicted data with confidence level labels; S6. Compare the verification results with historical deviation patterns, dynamically adjust the prediction data with confidence level labels, and achieve closed-loop feedback optimization; The dynamic adjustment process involves matching the current verification result with a preset historical deviation pattern library for similarity, calculating the deviation correction coefficient, and performing weighted iterative updates on the prediction data with confidence labels until it converges to the optimal prediction interval.

[0021] Based on step S1, it should be noted that in this embodiment, a multi-source data acquisition terminal for the tobacco industry is first established. This terminal uses a combination of interface calls, direct database connections, web crawlers, and text parsing to collect multi-source heterogeneous data from the provincial tobacco commercial company and its municipal branches. Specifically, this includes: structured financial data from the industrial company's ERP system, such as cigarette production output, tobacco leaf purchase amount, and production costs; policy documents issued by the State Tobacco Monopoly Administration and provincial tobacco monopoly bureaus, such as cigarette price control policies and tobacco leaf purchase subsidy policies; cigarette sales and inventory ledgers and retail terminal order data from various municipal branches; public opinion texts published by tobacco industry media and retailer feedback; and external environmental data (such as regional GDP growth rate, per capita disposable income, and market share of cigarette substitutes).

[0022] The collected multi-source heterogeneous data were preprocessed as follows: 1) Data cleaning: missing values ​​and outliers were removed from the structured data (outliers were identified using the Z-score method; when the data deviated from the mean by more than 3 standard deviations, it was determined to be an outlier, and the median of the same type of data was used to fill it in), and text data were deduplicated and denoised. The min-max normalization method is used to map the structured data to the [0,1] interval, eliminating the influence of dimensions. The formula is as follows: ,in The original data, These are the minimum and maximum values ​​of the indicator, respectively; text data is segmented using the BERT word segmenter and converted into word vectors; 3) Data alignment: data from different sources and in different formats are aligned in time and space according to fiscal year cycle and regional dimension (province, city, district / county) to ensure data consistency and comparability, and finally standardized data is obtained and stored in the tobacco finance-specific database for subsequent steps.

[0023] Furthermore, it also includes, The multi-source heterogeneous data includes structured production and financial data from industrial companies' ERP systems, policy documents issued by national and provincial tobacco monopoly bureaus, inventory ledgers of commercial companies, industry public opinion texts, and external environment data. The multi-source heterogeneous data is standardized with tobacco terminology and spatiotemporally aligned with fiscal year cycles. Based on a pre-built tobacco causal labeling system, each data record is automatically labeled with causal semantic tags, which include policy causes, operational causes, environmental causes, and corresponding financial effect tags, thereby obtaining the standardized data.

[0024] Furthermore, the construction of the causal feature decoupling model includes the following sub-steps: The multi-source heterogeneous data is automatically labeled based on a preset tobacco causal labeling system to generate training data with causal labels. Based on the training data with causal labels, a variational autoencoder with a multi-branch latent representation structure is constructed, and an attention mask is applied to the data corresponding to different causal labels during the encoding process. The training data is then subjected to feature aggregation within the corresponding latent branches. A regularization term is introduced to minimize the mutual information among policy-driven factors, operational execution factors, and market feedback factors; The trained model is used to process standardized data for the current cycle, outputting decoupled policy-driven factors, operational execution factors, and market feedback factors.

[0025] Furthermore, the fusion of the decoupled policy-driven factors, operational execution factors, and market feedback factors with graph-embedded data includes: The policy-driven factors, operational execution factors, and market feedback factors are denoted as vectors. , , The graph embedding data is denoted as a vector. A fused input vector is generated through concatenation operations. ;

[0026] The prediction model uses the following loss function for training or inference constraints when generating output: , in, To predict the loss term, The predicted values ​​of key financial indicators output by the model. For a predefined set of financial compliance rules for the tobacco industry, each rule Corresponding to a constraint function ,when A value greater than 0 indicates that the predicted value violates the rule. , These are non-negative weighting coefficients.

[0027] refer to Figure 2 Based on step S2, it should be noted that the following working steps are also included: S21. Generation of training data with causal labels. Based on the preset tobacco causal labeling system, the standardized data after preprocessing in the embodiment is automatically labeled with causal semantics. The labeling method adopts the rule + deep learning hybrid labeling method in embodiment 1. After the labeling is completed, the data with causal labels is divided into training set, validation set and test set. The training set is used for model training, the validation set is used for model parameter tuning, and the test set is used for model performance testing, generating training data with causal labels (i.e. training set data).

[0028] S22, Construction of multi-branch VAE model Based on the training data with causal labels, a variational autoencoder (VAE) with a multi-branch latent representation structure is constructed as the basic framework of the causal feature decoupling model. The VAE model consists of two parts: an encoder and a decoder. The encoder adopts a multi-branch structure with three latent variable branches, corresponding to policy-driven factors, operational execution factors, and market feedback factors, respectively. Each branch corresponds to a causal label type (policy cause, operational cause, and environmental cause, where the market feedback factor corresponds to the market-related label in the environmental cause).

[0029] The encoder's specific structure is as follows: the input layer receives training data with causal labels (feature vectors of structured data and word vectors of text data, which are uniformly converted into 128-dimensional vectors); the input layer is connected to the embedding layer, which maps the input vectors to a high-dimensional feature space (dimensional 256); the embedding layer is connected to three parallel attention mask branches (policy branch, operation branch, and market branch), each branch corresponding to a causal factor. Attention mechanisms and fully connected layers are set within the branches to extract features of the corresponding causal factors.

[0030] During the encoding process, attention masks are applied to the data corresponding to different causal labels: For the input training data, the branch to which it belongs is determined according to its labeled causal label (e.g., data labeled "policy cause" is classified into the policy branch). In the attention calculation, the key-value pairs of the other two branches are masked through the mask operation, so that the training data is only aggregated within its own branch, avoiding feature confusion of different causal factors and achieving the initial decoupling of causal features.

[0031] The decoder is structured to receive the latent variables (policy-driven factors, operational execution factors, and market feedback factors) from the three branches of the encoder output. It then concatenates these three latent variables through a fully connected layer and maps them to a reconstruction vector with the same dimension as the input data. This reconstruction vector is used to calculate the reconstruction error and evaluate the model's feature extraction and decoupling performance.

[0032] S23, Model Regularization and Training Optimization During model training, a regularization term is introduced to minimize the mutual information between policy-driven factors, operational execution factors, and market feedback factors. Mutual information is used to measure the degree of dependence between two random variables. The smaller the mutual information, the stronger the independence of the two factors and the better the decoupling effect.

[0033] The specific formula for calculating the regularization term is: ,in Mutual information between the i-th causal factor and the j-th causal factor , As a policy-driven factor, As an operational execution factor, The market feedback factor is used. The model's total loss function is a weighted sum of the reconstruction loss and the regularization term: ,in The reconstruction loss is calculated using mean squared error loss. The regularization term weight coefficient (in this embodiment) This is used to balance the effects of reconstruction and decoupling.

[0034] During model training, the Adam optimizer was used, with 500 training iterations, a learning rate of 0.001, and a batch size of 32. Every 10 iterations, the model's reconstruction error and mutual information were calculated using validation set data, and model parameters (such as attention mask weights and fully connected layer parameters) were adjusted. When the reconstruction error of the validation set was less than 0.01 and the average mutual information between the three causal factors was less than 0.1, the model was considered to have converged, training was stopped, and the trained causal feature decoupling model was obtained.

[0035] S24, Causal factor decoupling output The standardized data for the current period (e.g., the third quarter of 2025) obtained in the examples are input into the trained causal feature decoupling model. The model uses the attention mask branch of the encoder to filter and reorganize the feature channels of multi-source heterogeneous data: the policy branch filters out policy-related feature channels (e.g., semantic features of policy documents, financial data features of policy impact), and outputs policy driving factors. (128-dimensional feature vector); the operations branch filters out operation-related feature channels (such as operational features of procurement, sales, and inventory), and outputs the operations execution factor. (128-dimensional feature vector); the market branch filters out market-related feature channels (such as retail demand and public opinion influence features), and outputs the market feedback factor $$Z_3$$ (128-dimensional feature vector).

[0036] The decoupling effect was verified using test set data: the mutual information between the three causal factors was calculated. If the average mutual information was less than 0.1 and the matching degree between the feature corresponding to each factor and the causal label was higher than 90%, the decoupling effect was good, and the three causal factors output could be used for subsequent prediction steps. If the decoupling effect was not up to standard (mutual information higher than 0.1 or matching degree lower than 90%), the model parameters (such as regularization term weights and attention masking strategies) were readjusted, and the model was retrained until the decoupling effect met the standard.

[0037] Furthermore, the causal feature decoupling model includes an encoder and a decoder. The encoder has three latent variable branches, corresponding to policy-driven, operational execution, and market feedback causal labels, respectively. When encoding training data with causal labels, the data is input to the corresponding latent variable branch according to the causal label attached to the data, and the key-value pairs of the other two branches are masked in the attention calculation, so that the data generates a latent representation only within its own branch.

[0038] Furthermore, the decoupled policy-driven factors, operational execution factors, and market feedback factors are input into the prediction model to obtain predicted values ​​for key indicators. The policy-driven factors, operational execution factors, and market feedback factors are mapped to the graph embedding data as vector representations of a unified dimension, and then concatenated to form a fusion input vector, which is fed into the input embedding layer of the prediction model. The fusion input vector is then subjected to contextual semantic encoding by the multi-layer converter encoder of the prediction model and fed into the decoder layer for autoregressive generation. During the generation process of the decoder layer, the embedded rule constraint module is called to obtain the predefined set of financial compliance rules for the tobacco industry in real time, and the currently generated intermediate output words or candidate final output words are substituted into the constraint functions in the set of financial compliance rules for compliance determination. If the current term is determined to violate any financial compliance rule, a negative bias is applied to the logical probability value of the term to suppress its selection; if it is determined to comply with the rule, the original logical probability value is maintained. Based on the distribution of logical probability values ​​after applying a negative bias, or by selecting the next term, the above encoding, compliance determination, and bias correction processes are iteratively executed until a complete sequence is generated as the prediction data for the aforementioned key financial indicators.

[0039] The prediction model includes a multi-layer converter encoder-decoder structure. Its input embedding layer receives the policy-driven factors, operational execution factors, market feedback factors, and graph embedding data extracted from the tobacco industry knowledge graph. In the multi-layer converter encoder-decoder structure, a rule constraint module is embedded in at least one decoder layer. The rule constraint module, based on a predefined set of tobacco industry financial compliance rules, applies a negative bias to the logical value of any intermediate or final output term that violates any rule in the compliance rule set during the generation of prediction values, thereby generating prediction data for the key financial indicators.

[0040] Furthermore, after the predicted data is output, the following verification operation is performed: The predicted data are compared with the preset budget indicator data, the tobacco industry applicable tax rules data, and the enterprise cash flow constraint data to obtain the budget matching degree verification result, tax suitability verification result, and cash flow coverage capability verification result. Based on whether the three verification results meet their respective preset threshold conditions, three Boolean verification flags are generated. Based on a predefined confidence level mapping table, the combination of the three Boolean verification flags is mapped to a confidence level label, which is an element in a preset discrete level set.

[0041] Furthermore, the method includes matching the verification result with a pre-stored historical prediction deviation pattern dataset to obtain a matching deviation pattern identifier; and querying the corresponding model adjustment instruction and graph update instruction from a preset strategy mapping table based on the deviation pattern identifier. The model adjustment instruction is used to modify the training parameters and loss function weights of the causal feature decoupling model, and the graph update instruction is used to adjust the weight update rate and time decay coefficient of at least one causal edge in the tobacco financial knowledge graph.

[0042] refer to Figure 3 Secondly, this invention also provides an intelligent prediction and control system for key tobacco financial indicators based on a large model, and applies a method for intelligent prediction and control of key tobacco financial indicators based on a large model, including: The data acquisition module is used to collect multi-source heterogeneous data from the tobacco industry, preprocess the multi-source heterogeneous data, and obtain standardized data. The feature decoupling module is used to automatically perform causal semantic annotation on the standardized data based on a pre-built tobacco causal labeling system, generate training data with causal labels, train and construct a causal feature decoupling model based on the training data with causal labels, input the standardized data into the causal feature decoupling model, and perform feature channel filtering and reorganization on the multi-source heterogeneous data through a mask attention mechanism based on causal labels, and output policy driving factors, operational execution factors and market feedback factors. The graph embedding module is used to dynamically construct a tobacco financial knowledge graph based on historical and real-time training data with causal labels, and to retrieve subgraphs related to the current prediction task from the knowledge graph, and output graph embedding data after encoding by a graph neural network. The intelligent prediction module is used to fuse the decoupled policy-driven factors, operational execution factors, and market feedback factors with the graph-embedded data, and feed them as input into a prediction model adapted for the tobacco industry. The prediction model is a large language model based on the Transformer architecture and integrating financial compliance constraint mechanisms, and outputs prediction data of key financial indicators. The compliance verification module is used to perform multi-dimensional compliance verification based on the predicted data, including budget matching degree, tax compliance and cash flow coverage analysis, and generate predicted data with confidence level labels. The feedback optimization module is used to compare the verification results with historical deviation patterns and dynamically adjust the prediction data with confidence level labels to achieve closed-loop feedback optimization.

[0043] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in any one of the claims.

[0044] In the description of this specification, the references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0045] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for intelligent prediction and control of key financial indicators in the tobacco industry based on a large model, characterized in that, Includes the following steps: S1. Collect multi-source heterogeneous data from the tobacco industry, preprocess the multi-source heterogeneous data, and obtain standardized data; S2. Based on the pre-constructed tobacco causal labeling system, the standardized data is automatically labeled with causal semantics to generate training data with causal labels. A causal feature decoupling model is trained and constructed based on the training data with causal labels. The standardized data is input into the causal feature decoupling model. The feature channels of the multi-source heterogeneous data are filtered and reorganized through a mask attention mechanism based on causal labels, and policy driving factors, operational execution factors and market feedback factors are output. S3. Dynamically construct a tobacco financial knowledge graph based on historical and real-time training data with causal labels, and retrieve subgraphs related to the current prediction task from the knowledge graph, and output graph embedding data after encoding by a graph neural network; S4. The decoupled policy-driven factors, operational execution factors, and market feedback factors are fused with the graph-embedded data and fed into a prediction model adapted for the tobacco industry. The prediction model is a large language model based on the Transformer architecture and integrating financial compliance constraint mechanisms, and outputs prediction data of key financial indicators. S5. Perform multi-dimensional compliance verification based on the predicted data, including budget matching degree, tax compliance and cash flow coverage analysis, and generate predicted data with confidence level labels; S6. Compare the verification results with historical deviation patterns, dynamically adjust the prediction data with confidence level labels, and achieve closed-loop feedback optimization; The dynamic adjustment process involves matching the current verification result with a preset historical deviation pattern library for similarity, calculating the deviation correction coefficient, and performing weighted iterative updates on the prediction data with confidence labels until it converges to the optimal prediction interval.

2. The intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in claim 1, characterized in that, It also includes, The multi-source heterogeneous data includes structured production and financial data from industrial companies' ERP systems, policy documents issued by national and provincial tobacco monopoly bureaus, inventory ledgers of commercial companies, industry public opinion texts, and external environment data. The multi-source heterogeneous data is standardized with tobacco terminology and spatiotemporally aligned with fiscal year cycles. Based on a pre-built tobacco causal labeling system, each data record is automatically labeled with causal semantic tags, which include policy causes, operational causes, environmental causes, and corresponding financial effect tags, thereby obtaining the standardized data.

3. The intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in claim 1, characterized in that, The construction of the causal feature decoupling model includes the following sub-steps: The multi-source heterogeneous data is automatically labeled based on a preset tobacco causal labeling system to generate training data with causal labels. Based on the training data with causal labels, a variational autoencoder with a multi-branch latent representation structure is constructed, and an attention mask is applied to the data corresponding to different causal labels during the encoding process. The training data is then subjected to feature aggregation within the corresponding latent branches. A regularization term is introduced to minimize the mutual information among policy-driven factors, operational execution factors, and market feedback factors; The trained model is used to process standardized data for the current cycle, outputting decoupled policy-driven factors, operational execution factors, and market feedback factors.

4. The intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in claim 1, characterized in that, The decoupled policy-driven factors, operational execution factors, and market feedback factors are fused with the graph-embedded data, including: The policy-driven factors, operational execution factors, and market feedback factors are denoted as vectors. The graph embedding data is denoted as a vector. A fused input vector is generated through concatenation operations. ; The prediction model uses the following loss function for training or inference constraints when generating output: ; in, To predict the loss term, The predicted values ​​of key financial indicators output by the model. For a predefined set of financial compliance rules for the tobacco industry, each rule Corresponding to a constraint function when This indicates that the predicted value violates the rule. , These are non-negative weighting coefficients.

5. The intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in claim 3, characterized in that, The causal feature decoupling model includes an encoder and a decoder. The encoder has three latent variable branches, which correspond to policy-driven, operational execution, and market feedback causal labels, respectively. When encoding training data with causal labels, the training data is input to the corresponding latent variable branch according to the causal label attached to the training data, and the key-value pairs of the other two branches are masked in the attention calculation, so that the training data generates a latent representation only within its own branch.

6. The intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in claim 1, characterized in that, The decoupled policy-driven factors, operational execution factors, and market feedback factors are input into the prediction model to obtain the predicted values ​​of key indicators. The policy-driven factors, operational execution factors, and market feedback factors are mapped to the graph embedding data into vector representations of a unified dimension, and then concatenated to form a fused input vector, which is then fed into the input embedding layer of the prediction model. The fused input vector is then subjected to contextual semantic encoding by the multi-layer converter encoder of the prediction model and then fed into the decoder layer for autoregressive generation. During the generation process of the decoder layer, the embedded rule constraint module is called to obtain the predefined set of financial compliance rules for the tobacco industry in real time, and the currently generated intermediate output words or candidate final output words are substituted into the constraint functions in the set of financial compliance rules for compliance determination. If the current term is determined to violate any financial compliance rule, a negative bias is applied to the logical probability value of the term to suppress its selection; if it is determined to comply with the rule, the original logical probability value is maintained. Based on the distribution of logical probability values ​​after applying a negative bias, or by selecting the next term, the above encoding, compliance determination, and bias correction processes are iteratively executed until a complete sequence is generated as the prediction data for the aforementioned key financial indicators.

7. The intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in claim 1, characterized in that, After the predicted data is output, the following verification operation is performed: The predicted data are compared with the preset budget indicator data, the tobacco industry applicable tax rules data, and the enterprise cash flow constraint data to obtain the budget matching degree verification result, tax suitability verification result, and cash flow coverage capability verification result. Based on whether the verification results meet their respective preset threshold conditions, three Boolean verification flags are generated. Based on a pre-set confidence mapping table, the combination of the three Boolean verification flags is mapped to a confidence flag, which is an element in a pre-set discrete level set.

8. The intelligent prediction and control method for key financial indicators of tobacco based on a large model as described in claim 1, characterized in that, Also includes: The verification result is matched with the pre-stored historical prediction deviation pattern dataset to obtain the matching deviation pattern identifier. Based on the deviation mode identifier, the corresponding model adjustment instruction and map update instruction are queried from the preset strategy mapping table; The model adjustment instruction is used to modify the training parameters and loss function weights of the causal feature decoupling model, and the graph update instruction is used to adjust the weight update rate and time decay coefficient of at least one causal edge in the tobacco financial knowledge graph.

9. A large-scale model-based intelligent prediction and control system for key tobacco financial indicators, employing the large-scale model-based intelligent prediction and control method for key tobacco financial indicators as described in any one of claims 1-8, characterized in that, include: The data acquisition module is used to collect multi-source heterogeneous data from the tobacco industry, preprocess the multi-source heterogeneous data, and obtain standardized data. The feature decoupling module is used to automatically perform causal semantic annotation on the standardized data based on a pre-built tobacco causal labeling system, generate training data with causal labels, train and construct a causal feature decoupling model based on the training data with causal labels, input the standardized data into the causal feature decoupling model, and perform feature channel filtering and reorganization on the multi-source heterogeneous data through a mask attention mechanism based on causal labels, and output policy driving factors, operational execution factors and market feedback factors. The graph embedding module is used to dynamically construct a tobacco financial knowledge graph based on historical and real-time training data with causal labels, and to retrieve subgraphs related to the current prediction task from the knowledge graph, and output graph embedding data after encoding by a graph neural network. The intelligent prediction module is used to fuse the decoupled policy-driven factors, operational execution factors, and market feedback factors with the graph-embedded data, and feed them as input into a prediction model adapted for the tobacco industry. The prediction model is a large language model based on the Transformer architecture and integrating financial compliance constraint mechanisms, and outputs prediction data of key financial indicators. The compliance verification module is used to perform multi-dimensional compliance verification based on the predicted data, including budget matching degree, tax compliance and cash flow coverage analysis, and generate predicted data with confidence level labels. The feedback optimization module is used to compare the verification results with historical deviation patterns and dynamically adjust the prediction data with confidence level labels to achieve closed-loop feedback optimization.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the intelligent prediction and control method for key tobacco financial indicators based on a large model as described in any one of claims 1-8.