Addition and deduction intelligent management method based on machine learning algorithm
By building a multi-level technical architecture, the automated collection of R&D expense data and feature dimensionality optimization are realized, combined with the random forest classification model and gradient descent optimization algorithm, the shortcomings in R&D expense data fusion, feature selection and model timeliness in the existing technology are solved, and the ability to efficiently process data and quickly respond to policy changes is achieved.
Patent Information
- Application Number
- CN202510386411.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The prior art has significant shortcomings in the fusion of multi-source heterogeneous data, feature selection and model timeliness of R&D cost data, and it is difficult to effectively utilize time series characteristics and quickly respond to policy changes.
By building a multi-level technical architecture, we can realize automatic collection of R&D expense data, feature dimensionality reduction optimization and dynamic parameter adjustment of classification models, combine information entropy and statistical learning methods to select features, and use random forest classification model and gradient descent optimization algorithm to achieve efficient data processing and rapid iteration of the model.
It improves the efficiency of the integration of R&D expense data and the feature selection effect, enhances the timeliness and policy adaptability of the model, and significantly improves the efficiency and accuracy of R&D expense review.
Smart Images

Figure CN120235718A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of tax informatization, and particularly relates to an intelligent management system integrating multi-modal data processing, adaptive feature engineering, and dynamic model optimization. Background Art
[0002] There are three major technical pain points in the implementation process of the current additional deduction policy:
[0003] Data heterogeneity: R & D expense data is scattered in multiple data sources such as ERP systems (such as SAP, Oracle), project management platforms (such as Jira), and financial systems (such as Kingdee). Format differences make it difficult to fuse data.
[0004] Feature complexity: A typical R & D expense data set contains more than 80 feature dimensions. Traditional feature selection methods (such as chi-square test) cannot effectively extract policy-sensitive features.
[0005] Model timeliness: Policy adjustments (such as the addition of the clause "additional deduction for R & D expenses in artificial intelligence" in 2024) require the model to have the ability to quickly iterate. Traditional batch training methods are difficult to meet the requirements.
[0006] The existing technologies have significant deficiencies: The hybrid model proposed in Rule-based Tax Compliance with Machine Learning (2022) still relies on manual rule annotation. The deep model in Deep Learning for Tax Fraud Detection (2023) is prone to overfitting in small sample scenarios. Existing systems lack effective utilization of the time series characteristics of R & D expense data. Summary of the Invention
[0007] Based on the above deficiencies of the existing technologies, this method realizes the automatic collection of R & D expense data (including structured database interfaces and unstructured document parsing), the dimensionality reduction and optimization of high-dimensional features (combining information theory and statistical learning methods), and the dynamic tuning of parameters of the classification model (parameter update mechanism based on the gradient descent algorithm) by constructing a multi-level technical architecture, and finally outputs the compliance determination result of R & D expenses that conforms to the latest tax policy.
[0008] The technical solution of the present invention is: A method for intelligent management of additional deductions based on machine learning algorithms, including the following steps:
[0009] Step 1: Establish an enterprise R & D activity data collection module. Obtain multi-source heterogeneous data through a structured database interface and a distributed file system, and perform data cleaning and standard preprocessing. This module establishes a connection with a relational database through the JDBC / ODBC protocol, uses the Flume data flow framework to collect unstructured data such as R & D project documents, financial vouchers, and patent application forms in real time, uses regular expression matching technology to clean missing values and outliers, and converts numerical data into a standard normal distribution with a mean of 0 and a standard deviation of 1 through the Z-score standardization method.
[0010] Step 2: Build a feature engineering module. Use a feature selection algorithm based on information entropy to screen key R & D expense indicators, and perform feature space dimensionality reduction through linear discriminant analysis. This algorithm evaluates the information content of each feature by calculating the Shannon entropy formula, selects the top k features with an information gain greater than the threshold, then constructs between-class scatter matrices and within-class scatter matrices, obtains the optimal projection direction by solving the generalized eigenvalue problem, and maps the original feature space to a d-dimensional subspace.
[0011] Step 3: Deploy a random forest classification model. Set the number of decision trees to N, use the Gini coefficient as the node splitting criterion, and establish a rule for judging the compliance of R & D expenses. This model generates M training subsets through the bootstrap sampling method. Each decision tree randomly selects mtry features for node splitting during training, evaluates the splitting effect based on the Gini impurity formula, and finally integrates the prediction results of all trees through the majority voting mechanism.
[0012] Step 4: Dynamically adjust the parameters of the classification model based on the gradient descent optimization algorithm, and output a list of R & D expenses that meet the additional deduction policy and risk warning results. This algorithm defines a cross-entropy loss function, calculates the gradient through backpropagation, uses the stochastic gradient descent update rule with momentum to dynamically adjust the learning rate and momentum parameters, and finally outputs a structured list containing expense item codes and policy matching score ratings, as well as red / yellow warning indicators based on the confidence threshold.
[0013] Furthermore, the data preprocessing in Step 1 includes:
[0014] Step a: Filling missing values
[0015] When preprocessing enterprise R & D activity data, for the missing values in the data, use the method of filling with the moving window mean of the historical data of the same enterprise. Specifically, with the time point where the current missing value is located as the center, select a certain number of data points before and after to form a moving window, calculate the mean of the historical data of the same enterprise within this window, and use this mean as the filling value for the missing value.
[0016] Step b: Standardization of numerical fields
[0017] For numerical fields, Z-score standardization is implemented to convert the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. The calculation formula is as follows:
[0018]
[0019] Where:
[0020] x′ is the data value after standardization;
[0021] x is the original data value;
[0022] μ is the mean of all data in this numerical field;
[0023] σ is the standard deviation of all data in this numerical field;
[0024] Step c: Categorical variable encoding
[0025] For categorical variables, One-Hot encoding is used to generate a sparse matrix; One-Hot encoding is a method of converting categorical variables into numerical variables, converting the values of each categorical variable into a binary vector, the length of the vector is equal to the number of values of the categorical variable, and only the position corresponding to the value is 1, and the rest are 0;
[0026] Specifically, for each categorical variable, determine all its possible values, and then create a new binary feature for each value; in the original dataset, for each sample, if the value of its categorical variable is a certain specific value, the corresponding new feature is 1, and the rest of the new features are 0.
[0027] Furthermore, the feature selection algorithm based on information entropy in step 2 is specifically as follows:
[0028] In the additional deduction intelligent management system based on machine learning algorithms, for a given feature set
[0029] F = {f1, f2,..., f n}
[0030] It is necessary to calculate the information gain value of each feature to evaluate the importance of this feature for classification; the calculation formula of the information gain value is:
[0031]
[0032] Where:
[0033] IG(F|C): represents the information gain of feature F relative to class C;
[0034] H(C): is the entropy of class C, also known as class entropy, and its calculation formula is;
[0035]
[0036] Among them, C is the set of categories, that is, the set composed of all possible categories;
[0037] c is a specific category in the set C;
[0038] p(c) is the probability that category c appears in the sample set S, and the calculation method is where |s c | is the number of samples with category c, and |S| is the total number of samples in the sample set S; Entropy(C) is the category entropy, which measures the degree of uncertainty of category C without any feature information;
[0039] V: is the value set of feature F, that is, the set composed of all possible values of feature F;
[0040] S: represents the sample set, which is the set of all samples used for analyzing and training the model;
[0041] S v : is the sample subset in the sample set S where the value of feature F is v,
[0042] |S v | represents the number of samples in this subset;
[0043] H(C|S v ) is the conditional entropy of category C in the sample subset where the value of feature F is v
[0044] S v The calculation formula is
[0045]
[0046] where
[0047] p(c|S v ) is the probability that category c appears in the sample subset S v The calculation method is
[0048] |S cv | is the number of samples with category c in the sample subset S v ;
[0049] The conditional entropy H(C|S v ) represents the degree of uncertainty of category C when the value of feature F is known to be v;
[0050] is the weighted sum of the conditional entropies H(C|S v ) corresponding to all feature values v, and the weight is That is, the sample subset S vThe proportion in the sample set S;
[0051] After calculating the information gain value of each feature, select the features that satisfy IG(F|C)≥θ, and these features constitute the optimized feature subset; where θ is a preset threshold, and the setting of this threshold needs to be determined according to specific business requirements and data characteristics.
[0052] Further, the method for generating decision trees in the random forest classification model in step 3 is as follows:
[0053] In step 3 of the additional deduction intelligent management system based on the machine learning algorithm, the random forest classification model generates decision trees through a specific method; perform Bootstrap sampling on the original training set, and this sampling method is random sampling with replacement to generate N sub-training sets; then, for each sub-training set, recursively perform the following operations until the node purity reaches the preset threshold or the sample quantity is insufficient;
[0054] Step a: Randomly select candidate features
[0055] Randomly select candidate features from d features; in actual operation, select k features from the entire feature set through a random number generator as the candidate features when splitting the current node;
[0056] Step b: Calculate the Gini coefficient of the candidate features
[0057] For each candidate feature, calculate its Gini coefficient; the calculation formula of the Gini coefficient is:
[0058]
[0059] Where:
[0060] Gini(D): Represents the Gini coefficient of the data set D, which measures the impurity of the data set;
[0061] c: Is the number of categories, that is, the total number of all possible categories in the data set;
[0062] p i : Is the sample proportion of category i, and the calculation method is Where |D i | is the number of samples in the data set with category i, and |D| is the total number of samples in the data set D;
[0063] Step c: Select features and splitting points for node splitting
[0064] After calculating the Gini coefficient of each candidate feature, select the feature and splitting point that minimize Gini(D A ) for node splitting; here DA It is a subset obtained by dividing according to a certain feature and splitting point; for continuous features, all possible splitting points will be traversed, and the Gini coefficient of the subset after division under each splitting point will be calculated;
[0065] For discrete features, all possible value combinations will be considered for division; by selecting the feature and splitting point that minimize Gini(D A ), the purity of the subset after division can be made as high as possible, thus constructing a more effective decision tree;
[0066] Through the above method, a decision tree is generated for the random forest classification model, and finally the classification prediction result is obtained by synthesizing the results of multiple decision trees, which is used for the determination of the compliance of R & D expenses in the intelligent management system for additional deductions.
[0067] Furthermore, the parameter update formula of the gradient descent optimization algorithm in step 4 is:
[0068] In step 4 of the intelligent management system for additional deductions based on machine learning algorithms, the gradient descent optimization algorithm is used to update the model parameters; this algorithm continuously iterates to find the model parameter values that minimize the objective function;
[0069] The parameter update formula is:
[0070]
[0071] Where:
[0072] W t+1 : represents the model parameters at the (t + 1)-th iteration; the model parameters w at the current iteration number t are updated in combination with the gradient of the objective function, gradually approaching the parameter values that minimize the objective function; t t
[0073] w t : are the model parameters at the t-th iteration; in each iteration process, they are adjusted according to the gradient information of the objective function to achieve the purpose of optimizing the model;
[0074] η: is the learning rate, which controls the step size of parameter update in each iteration;
[0075] is the gradient of the objective function J(w) at w = w t ; the gradient is a vector, and its direction represents the direction in which the objective function value increases fastest, and the opposite direction is the direction in which the objective function value decreases fastest;
[0076] The calculation formula of the objective function J(w) is:
[0077]
[0078] Wherein:
[0079] m: is the number of samples, i.e., the total number of samples used to train the model;
[0080] h w (x): is the model prediction value, which is the result obtained by predicting the input x based on the current model parameters w; for different machine learning models, the specific form of h w (x) will be different;
[0081] x (i) : represents the input feature vector of the i-th sample;
[0082] y (i) : is the true label value of the i-th sample;
[0083] (h w (x (i) ) - y (i) ) 2 measures the squared error between the predicted value and the true value of the model for the i-th sample;
[0084] This part is the mean squared error (MSE) loss function, which calculates the average of the squared prediction errors of all samples and is used to measure the fitting degree of the model on the training data;
[0085] λ: is the L2 regularization coefficient, which controls the weight of the regularization term;
[0086] ||w||2: is the L2 norm of the model parameter w, and the calculation formula is Where is the j-th component of the parameter vector w;
[0087] By continuously iteratively updating the model parameter w, the objective function J(w) is gradually reduced, and finally a set of optimal model parameters is obtained.
[0088] Furthermore, the projection matrix of the linear discriminant analysis is calculated as:
[0089] In the additional deduction intelligent management system based on the machine learning algorithm, when step 2 uses linear discriminant analysis for feature space dimensionality reduction, it is necessary to calculate the projection matrix W; the goal of linear discriminant analysis is to find a projection matrix W such that the data of different classes are separated as much as possible after projection, and the data of the same class are aggregated as much as possible;
[0090] The calculation of the projection matrix W is achieved by maximizing the following objective function:
[0091]
[0092] Wherein:
[0093] W is a projection matrix that projects the original high-dimensional feature space into a low-dimensional space;
[0094] S b is the between-class scatter matrix, which is used to measure the dispersion between different classes; its calculation formula is:
[0095]
[0096] c is the number of classes, that is, the total number of all possible classes in the dataset;
[0097] n i is the number of samples in the i-th class, which reflects the scale of the samples of this class in the dataset;
[0098] \(\mu_i\) is the mean vector of the samples in the i-th class,
[0099] \(\mu\) is the global mean vector,
[0100] The between-class scatter matrix S b is derived based on the measurement of the mean differences between different classes. By calculating the deviation of the mean of each class from the global mean and considering the weights of the sample numbers, a quantitative representation of the between-class dispersion is obtained;
[0101] s w is the within-class scatter matrix, which is used to measure the dispersion of data within the same class; its calculation formula is For each class of samples, calculate the deviation of each sample from the mean of this class, and then sum over all classes to obtain the within-class scatter matrix; the within-class scatter matrix reflects the tightness of the samples within the same class;
[0102] The objective function The numerator |W T S b W| represents the dispersion between different classes after projection, and the denominator |W T S w W| represents the dispersion within the same class after projection; by maximizing this objective function, a projection matrix W can be found such that the separation between different classes after projection is maximized and the aggregation within the same class is maximized, thus achieving effective dimensionality reduction of the feature space;
[0103] In actual calculations, usually by solving the generalized eigenvalue problem S b w = λS w w to obtain the column vectors of the projection matrix W, and these column vectors are the eigenvectors corresponding to the largest eigenvalues.
[0104] The technical innovation of the present invention is embodied in the following key modules:
[0105] Data Acquisition and Preprocessing Subsystem
[0106] Multi-source Data Integration:
[0107] Structured Data: Connect to the relational database through the JDBC 4.2 protocol and implement incremental extraction using Sqoop 1.4.7
[0108] Unstructured Data: Use Apache Tika 2.9.0 to parse PDF contracts and extract key information such as "R & D Personnel List" and "Equipment Lease Term" in combination with regular expressions
[0109] Real-time Data Stream: Build a message queue based on Kafka 2.8.1 to achieve a processing throughput of 5000 data items per minute
[0110] Data Cleaning Technology:
[0111] Sliding Window Filling Algorithm
[0112] The window size is automatically adjusted according to the data frequency (w = 7 for daily data, w = 3 for monthly data)
[0113] Outlier Detection: Based on the IQR method, the calculation formula is:
[0114] Q1 = 25 th percentile, Q3 = 75 th
[0115] percentile IQR = Q3 - Q1, Lower bound = Q1 - 1.5 × IQR
[0116] Upper bound = Q3 + 1.5 × IQR
[0117] Feature Engineering Optimization Subsystem
[0118] Information Entropy Feature Selection:
[0119] Implementation Steps:
[0120] Calculate the class entropy: H(C) = -∑ c∈C p(c)log2p(c)
[0121] Calculate the conditional entropy: H(C|S v ) = -∑ c∈C P(c|S , )log2p(c|S v )
[0122] Information Gain Calculation:
[0123] Optimize the threshold θ using the genetic algorithm, with the value range [0.1, 0.3]
[0124] Implementation of LDA dimensionality reduction:
[0125] Between-class scatter matrix:
[0126] Within-class scatter matrix:
[0127] Generalized eigenvalue solution: S b w = γS w Take the eigenvectors corresponding to the first d largest eigenvalues of w to form the projection matrix
[0128] Classification model construction subsystem
[0129] Random forest parameter configuration:
[0130] Boot strap sampling: Each sub-training set contains 60% of the original samples
[0131] Feature subset size: (d is the number of features after dimensionality reduction)
[0132] Decision tree depth: Adopt a dynamic adjustment strategy, with an initial depth of 5, increasing by 1 layer each time until the Gini gain < 0.03
[0133] Node splitting criterion:
[0134] Gini coefficient calculation:
[0135]
[0136] Splitting condition:
[0137]
[0138] Dynamic optimization subsystem
[0139] Objective function design:
[0140] The regularization term coefficient λ is determined by Bayesian optimization, with the search space [0.001, 0.1]
[0141] Loss function:
[0142]
[0143] Parameter update algorithm:
[0144] The momentum term coefficient γ = 0.9, and the learning rate adopts exponential decay: η t = η0 × 0.95 t / 100
[0145] Gradient calculation:
[0146]
[0147] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0148] I. Improving the R & D cost audit efficiency through the multi-source heterogeneous data fusion ability
[0149] By constructing a data acquisition system with a hybrid architecture, the present invention realizes the real-time integration of structured financial data and unstructured R & D documents. The specific advantages include:
[0150] Cross-system data connectivity: By using the JDBC 4.2 protocol to connect to mainstream ERP systems (such as SAP S / 4HANA), combined with the distributed log collection framework of Apache Flume, it can simultaneously process relational databases such as MySQL and Oracle, as well as unstructured data such as R & D contract PDFs and patent applications in HDFS, solving the limitation that traditional systems only support a single data source.
[0151] Real-time data stream processing: Based on the Kafka message queue to build a data pipeline, achieving a throughput of 5000 data items per minute, with an efficiency improvement of 60% compared to traditional ETL tools (such as Informatica), ensuring the timeliness of R & D cost data. For example, an electronics enterprise shortened the cycle of cost data from generation to analysis from 72 hours to 2 hours through this system.
[0152] Multi-modal data parsing: Using Apache Tika 2.9.0 to parse PDF contracts, combined with regular expressions to extract key information such as "social insurance payment certificates for R & D personnel" and "equipment lease term", breaking through the bottleneck of existing systems relying on manual entry of unstructured data. In the application of a biopharmaceutical enterprise, the utilization rate of unstructured data increased from 35% to 89%.
[0153] II. Enhancing policy adaptability through the dynamic model optimization mechanism
[0154] The parameter update strategy innovatively designed in the present invention enables the model to quickly respond to policy changes. The specific advantages are as follows:
[0155] Adaptive learning rate scheduling: By using an exponentially decaying learning rate (η t = η0 × 0.95 ^ (t / 100)) combined with a momentum term (γ = 0.9), when the policy is adjusted (such as a change in the additional deduction ratio), the model can complete re-training within 48 hours, with a response speed improvement of 85% compared to the traditional batch training method. After the policy update in 2024, the model accuracy of a manufacturing enterprise only decreased by 0.8%, while the average decrease of similar systems was 5.2%.
[0156] Combination of Regularization and Gradient Optimization: Through the combination of L2 regularization (λ = 0.01) and the AdamW optimizer, the model complexity is effectively controlled. In the tests of a semiconductor company, the overfitting risk of the model in the small-sample scenario (500 pieces of data) is reduced by 63%, and the false alarm rate is reduced from the industry average of 8.6% to 3.7%.
[0157] Continuous Feedback Optimization: Establish a model performance monitoring dashboard to track indicators such as the change rate of the Gini coefficient and the loss function value in real time. When the classification accuracy of policy-sensitive features (such as "artificial intelligence R & D expenses") is detected to be lower than the threshold for three consecutive days, the system automatically triggers incremental learning to ensure that the model always maintains the best state. Description of the Drawings
[0158] Figure 1 is the overall system architecture diagram
[0159] Figure 2 is the data preprocessing flow chart
[0160] Figure 3 is the random forest model construction diagram Detailed Implementation Manner
[0161] As Figures 1-3 shown, the method of the present invention includes the following steps:
[0162] Step 1, establish an enterprise R & D activity data collection module, obtain multi-source heterogeneous data through a structured database interface and a distributed file system, and perform data cleaning and standardization preprocessing; this module establishes a connection with a relational database through the JDBC / ODBC protocol, uses the Flume data flow framework to collect unstructured data such as R & D project documents, financial vouchers, and patent applications in real time, uses regular expression matching technology to clean missing values and outliers, and converts numerical data into a standard normal distribution with a mean of 0 and a standard deviation of 1 through the Z-score standardization method;
[0163] Step 2, construct a feature engineering module, use a feature selection algorithm based on information entropy to screen key R & D expense indicators, and perform feature space dimensionality reduction through linear discriminant analysis; this algorithm evaluates the information content of each feature by calculating the Shannon entropy formula, selects the top k features with an information gain greater than the threshold, then constructs between-class scatter matrices and within-class scatter matrices, obtains the optimal projection direction by solving the generalized eigenvalue problem, and maps the original feature space to a d-dimensional subspace;
[0164] Step 3: Deploy a random forest classification model, set the number of decision trees to N, use the Gini coefficient as the node splitting criterion, and establish a compliance determination rule for R & D expenses; this model generates M training subsets through the bootstrap sampling method. Each decision tree randomly selects mtry features for node splitting during training, evaluates the splitting effect based on the Gini impurity formula, and finally integrates the prediction results of all trees through a majority voting mechanism;
[0165] Step 4: Dynamically adjust the classification model parameters based on the gradient descent optimization algorithm, and output a list of R & D expenses that meet the additional deduction policy and risk warning results; this algorithm defines a cross-entropy loss function, calculates the gradient through backpropagation, and uses the update rule of stochastic gradient descent with momentum to dynamically adjust the learning rate and momentum parameters. Finally, it outputs a structured list containing expense item codes and policy matching score, as well as red / yellow warning indicators based on the confidence threshold.
[0166] This invention takes the R & D expense audit of a semiconductor enterprise in 2024 as an example:
[0167] Data collection configuration:
[0168] Structured data source: MySQL 8.0.28, including a R & D project table (1200 entries) and an expense details table (35000 entries)
[0169] Unstructured data source: 500 PDF contracts, and the field of "social insurance payment certificates for R & D personnel" is extracted through Tika parsing
[0170] Data pipeline: The Kafka cluster is configured with 3 partitions and the number of ZooKeeper nodes is 3
[0171] Preprocessing parameters:
[0172] Sliding window filling: Use a window with w = 7 for the field of "usage duration of R & D equipment"
[0173] Standardization: For the field of "unit price of R & D material procurement", the mean μ = 850 and the standard deviation σ = 200
[0174] Encoding: Convert "patent type" (invention patent / utility model / external design) into a 3-dimensional sparse vector
[0175] Feature engineering process:
[0176] Initial features: 42 expense-related indicators
[0177] Information gain screening: Retain 18 features with IG ≥ 0.25 (such as average salary of R & D personnel, equipment depreciation rate)
[0178] LDA dimensionality reduction: Project the 18-dimensional features into an 8-dimensional space, and the cumulative contribution rate reaches 92.3%
[0179] Model training parameters:
[0180] Number of decision trees: 150, maximum depth of each tree: 12
[0181] Subsampling ratio: 70%, size of feature subset k = 4
[0182] Node splitting threshold: Gini gain ≥ 0.04, minimum number of samples in leaf nodes: 3
[0183] Dynamic optimization results:
[0184] Optimization algorithm: AdamW (weight decay coefficient 0.01)
[0185] Learning rate scheduling: Cosine annealing strategy, initial n = 0.001, T _ max = 500
[0186] Convergence curve of loss function: Reached the optimal value of 0.028 at the 320th iteration
[0187] System performance metrics:
[0188] Accuracy: 95.2% (2000 data in the test set)
[0189] Response time: Processing of a single piece of data < 50ms (based on NVIDIA A100 GPU)
[0190] Model update period: Retraining is completed within 48 hours after policy adjustment
[0191] This embodiment verifies the feasibility of the system in a real industrial scenario. Feature optimization improves the model training speed by 40%, and the dynamic parameter adjustment mechanism ensures that the model performance drops by no more than 1.5% after policy update.
[0192] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent management method for additional deductions based on machine learning algorithms, characterized in that: The following steps are involved: Step 1: Establish an enterprise R&D activity data collection module, obtain multi-source heterogeneous data through structured database interfaces and distributed file systems, and perform data cleaning and standardization preprocessing; this module establishes a connection with a relational database through the JDBC / ODBC protocol, uses the Flume data flow framework to collect R&D project documents, financial vouchers, and patent application documents in real time, uses regular expression matching technology to clean missing values and outliers, and converts numerical data into a standard normal distribution with a mean of 0 and a standard deviation of 1 through the Z-score standardization method; Step 2: Build a feature engineering module, use a feature selection algorithm based on information entropy to screen key R&D expense indicators, and perform feature space dimensionality reduction through linear discriminant analysis; The algorithm evaluates the information content of each feature by calculating the Shannon entropy formula, selects the first k features whose information gain is greater than the threshold, and then constructs the inter-class scatter matrix and the intra-class scatter matrix. The optimal projection direction is obtained by solving the generalized eigenvalue problem, and the original feature space is mapped to the d-dimensional subspace. Step 3: Deploy the random forest classification model, set the number of decision trees to N, use the Gini coefficient as the node splitting criterion, and establish the R&D expense compliance judgment rule; the model generates M training subsets through the self-service sampling method, and each decision tree randomly selects mtry features for node splitting during training. The splitting effect is evaluated based on the Gini impurity formula, and finally the prediction results of all trees are integrated through the majority voting mechanism; Step 4: Dynamically adjust the classification model parameters based on the gradient descent optimization algorithm to output the R&D expense list and risk warning results that meet the additional deduction policy; The algorithm defines the cross entropy loss function, calculates the gradient through back propagation, and uses the stochastic gradient descent update rule with momentum to dynamically adjust the learning rate and momentum parameters. The final output is a structured list containing cost item codes, policy matching scores, and red / yellow warning signs based on confidence thresholds.
2. The intelligent management method for additional deductions based on machine learning algorithm according to claim 1 is characterized by: The data preprocessing in step 1 includes: Step a: Missing value filling When preprocessing the enterprise R&D activity data, for the missing values in the data, the method of filling the missing values with the mean of the sliding window of the historical data of the same enterprise is adopted; specifically, taking the time point where the current missing value is located as the center, a certain number of data points before and after are selected to form a sliding window, and the mean of the historical data of the same enterprise in the window is calculated, and this mean is used as the filling value of the missing value; Step b: Standardize numeric fields For numeric fields, Z-score standardization is implemented to convert the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. The calculation formula is: in: x′ is the standardized data value; x is the original data value; Hong is the mean of all data in the numeric field; σ is the standard deviation of all data in this numeric field; Step c: Encoding categorical variables For categorical variables, One-Hot encoding is used to generate a sparse matrix. One-Hot encoding is a method for converting categorical variables into numerical variables. Each categorical variable value is converted into a binary vector. The length of the vector is equal to the number of categorical variable values, in which only the corresponding value position is 1 and the rest of the positions are 0. In the specific operation, for each categorical variable, determine all its possible values, and then create a new binary feature for each value; in the original data set, for each sample, if its categorical variable takes a certain value, the corresponding new feature is 1, and the remaining new features are 0.
3. The intelligent management method for additional deductions based on machine learning algorithm according to claim 1 is characterized by: The feature selection algorithm based on information entropy in step 2 is specifically: In the intelligent management system for additional deductions based on machine learning algorithms, for a given feature set F={f1,f2,…,f n } It is necessary to calculate the information gain value for each feature to evaluate the importance of the feature for classification; the calculation formula for the information gain value is: in: IG(F|C): represents the information gain of feature F relative to category C; Piece (C): is the entropy of category C, also known as category entropy, and its calculation formula is; Among them, C is the category set, that is, the set of all possible categories; c is a specific category in set C; p(c) is the probability that category c appears in sample set S, which is calculated as Where |S c | is the number of samples of category c, |S| is the total number of samples in sample set S; H(C) is the category entropy, which measures the degree of uncertainty of category C in the absence of any feature information; V: is the value set of feature F, that is, the set of all possible values of feature F; S: stands for sample set, which is the collection of all samples used to analyze and train the model; S v : is the sample subset in the sample set S whose feature F value is v, |S v | represents the number of samples in the subset; H(C|S v ): is a subset of samples where the feature F takes the value v S v The conditional entropy of category C in is calculated as in p(c|S v ) is in the sample subset S v The probability of category c appearing in is calculated as |S cv | is the sample subset S v The number of samples of category c in ; Conditional entropy H(C|S v ) represents the degree of uncertainty of category C when the value of feature F is known to be v; is the conditional entropy H(C|S v ) is weighted summed, and the weight is That is, the sample subset S v The proportion of the sample set S; After calculating the information gain value of each feature, the features that satisfy IG(F|C)≥θ are screened out. These features constitute the optimized feature subset; where θ is the preset threshold, and the setting of this threshold needs to be determined according to specific business needs and data characteristics.
4. The intelligent management method for additional deductions based on machine learning algorithm according to claim 1 is characterized by: The decision tree generation method of the random forest classification model in step 3 is: In step 3 of the intelligent management system for additional deductions based on machine learning algorithms, the random forest classification model generates a decision tree through a specific method; Bootstrap sampling is performed on the original training set, which is a random sampling method with replacement, to generate N sub-training sets; then, for each sub-training set, the following operations are recursively performed until the node purity reaches the preset threshold or the number of samples is insufficient; Step a: Randomly select candidate features Randomly select from d features candidate features; in actual operation, a random number generator is used to select k features from the entire feature set as candidate features for the current node split; Step b: Calculate the Gini coefficient of the candidate features For each candidate feature, calculate its Gini coefficient; the calculation formula of the Gini coefficient is: in: Gini(D): represents the Gini coefficient of the data set D, which measures the impurity of the data set; c: is the number of categories, that is, the total number of all possible categories in the data set; p i : is the sample proportion of category i, calculated as Where |D i | is the number of samples of category i in the data set, |D| is the total number of samples in data set D; Step c: Select features and split points for node splitting After calculating the Gini coefficient of each candidate feature, select Gini(D A ) The smallest feature and segmentation point are used for node splitting; here D A It is a subset obtained by dividing according to a certain feature and split point. For continuous features, all possible split points will be traversed to calculate the Gini coefficient of the subset divided at each split point. For discrete features, all possible value combinations are considered for partitioning; by selecting Gini (D A ) The smallest features and split points can make the purity of the divided subsets as high as possible, thus building a more effective decision tree; Through the above method, a decision tree is generated for the random forest classification model, and finally the classification prediction result is obtained by combining the results of multiple decision trees, which is used to determine the compliance of R&D expenses in the additional deduction intelligent management system.
5. The intelligent management method for additional deductions based on machine learning algorithm according to claim 1 is characterized by: The parameter update formula of the gradient descent optimization algorithm in step 4 is: In step 4 of the intelligent management system for additional deductions based on machine learning algorithms, the model parameters are updated using a gradient descent optimization algorithm; the algorithm continuously iterates to find the model parameter values that minimize the objective function; The parameter update formula is: in: w t+1 : represents the model parameters at the t+1th iteration; the model parameters w at the current iteration number t t Update the objective function based on its gradient, and gradually approach the parameter value that minimizes the objective function. w t : is the model parameter of the tth iteration; in each iteration, it is adjusted according to the gradient information of the objective function to achieve the purpose of optimizing the model; η: is the learning rate, which controls the step size of parameter update at each iteration; is the objective function J(w) when w=w t The gradient at; the gradient is a vector, its direction indicates the direction in which the objective function value increases fastest, and its opposite direction indicates the direction in which the objective function value decreases fastest; The calculation formula of the objective function J(w) is: in: m: is the number of samples, that is, the total number of samples used to train the model; h w (x): is the model prediction value, which is the result of predicting the input x based on the current model parameters w; for different machine learning models, h w The specific form of (x) will vary; x (i) : represents the input feature vector of the i-th sample; y (i) : is the true label value of the i-th sample; (h w (x (i) )-y (i) ) 2 Measures the square error between the model's predicted value and the true value for the i-th sample; This part is the mean square error (MSE) loss function, which calculates the average of the squared prediction errors of all samples and is used to measure how well the model fits the training data; λ: L2 regularization coefficient, which controls the weight of the regularization term; ||w||2: is the L2 norm of the model parameter w, calculated as in is the jth component of the parameter vector w; By continuously iteratively updating the model parameters w, the objective function J(w) is gradually reduced, and finally a set of optimal model parameters is obtained.
6. The intelligent management method for additional deductions based on machine learning algorithm according to claim 3 is characterized by: The projection matrix of the linear discriminant analysis is calculated as: In the intelligent management system for additional deductions based on machine learning algorithms, when using linear discriminant analysis to reduce the dimension of feature space in step 2, it is necessary to calculate the projection matrix W; The goal of linear discriminant analysis is to find a projection matrix W so that data of different categories are separated as much as possible after projection, and data of the same category are clustered as much as possible; The calculation of the projection matrix W is achieved by maximizing the following objective function: in: W is the projection matrix, which projects the original high-dimensional feature space into a low-dimensional space; s b It is the inter-class scatter matrix, which is used to measure the degree of dispersion between different categories; its calculation formula is: c is the number of categories, that is, the total number of all possible categories in the data set; n i is the number of samples in the i-th category, reflecting the size of samples of this category in the data set; μ i is the mean vector of the i-th class of samples, μ is the global mean vector, The inter-class divergence matrix S b The derivation of is based on the measurement of the difference in the mean values of different categories. By calculating the deviation of the mean of each category from the global mean and considering the weight of the number of samples, a quantitative representation of the degree of dispersion between categories is obtained. S w It is the intra-class scatter matrix, which is used to measure the degree of dispersion of data within the same category; its calculation formula is For each class of samples, the deviation of each sample from the class mean is calculated, and then the sum is taken for all classes to obtain the intra-class scatter matrix; the intra-class scatter matrix reflects the closeness of samples in the same class; Objective Function The molecule |W T S b W| represents the degree of discreteness between different categories after projection, and the denominator |W T S w W| represents the degree of discreteness within the same category after projection; by maximizing this objective function, a projection matrix W can be found so that the separation between different categories after projection is maximized and the aggregation within the same category is maximized, thereby achieving effective feature space dimensionality reduction; In practical calculations, it is usually solved by solving the generalized eigenvalue problem S b w=λS w w is used to obtain the column vectors of the projection matrix W, which are the eigenvectors corresponding to the largest eigenvalues.
Citation Information
Patent Citations
Software defect prediction method based on a hybrid active learning strategy
CN109656808A
Method and system for calculating enterprise income tax
CN111899081A
Task quality inspection method and device and computer storage medium
CN114330536A
Intelligent express cost approval method and device based on machine learning
CN117952547A
Research and development cost addition deduction risk detection method and device, terminal and medium
CN118037031A