Enterprise benefit prediction analysis method, system and equipment based on big data model, and medium
By constructing a multi-dimensional enterprise-benefit label system and a dynamic matching model using big data models, the problems of low efficiency in manual screening and large budget forecasting deviations in traditional enterprise-benefit forecasting have been solved, achieving high-precision enterprise screening and subsidy forecasting, and improving decision-making efficiency.
Patent Information
- Application Number
- CN202510963780.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional methods for formulating and predicting the effectiveness of subsidies for businesses rely on manual experience, resulting in low efficiency in matching businesses, distorted subsidy budget predictions, and difficulty in identifying hidden clauses and integrating multi-dimensional data.
We employ a big data model-based predictive analysis method for enterprise benefits. We use an improved BERT model for semantic parsing to construct a multi-dimensional enterprise benefit label system. We combine the random forest algorithm to calculate the fit score and use the Monte Carlo simulation algorithm to predict the subsidy payment amount, generating a three-dimensional visualization report.
It achieved an increase in the accuracy of enterprise screening to 92.3%, a reduction in the error rate of subsidy payment prediction to 7.5%, and a response time shortened to within 10 minutes, thereby improving decision-making efficiency and budget control accuracy.
Smart Images

Figure CN120875138A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence and public welfare for enterprises, specifically a predictive analysis method, system, device and medium for enterprise welfare based on big data models. Background Technology
[0002] The traditional methods for formulating and predicting the effectiveness of preferential policies for businesses have the following technical shortcomings:
[0003] (1) Analysis of preferential policies for enterprises relies on human experience: Existing technologies extract preferential policies for enterprises through keyword matching, but it is difficult to identify implicit clauses (such as qualification relevance and nested industry restrictions);
[0004] (2) Low efficiency in enterprise matching: Relying on a single dimension of business registration information for screening, without integrating dynamic indicators such as financial and tax data and R&D investment;
[0005] (3) Distortion in subsidy budget forecast: The number of applicant enterprises was estimated using a linear regression model, without considering the nonlinear impact of the adjustment of preferential policies on the willingness to apply. Summary of the Invention
[0006] The technical objective of this invention is to provide a method, system, device, and medium for predictive analysis of enterprise benefits based on big data models, in order to solve the problems of low efficiency and large deviations in traditional enterprise benefit matching due to reliance on manual screening.
[0007] The technical objective of this invention is achieved as follows: a predictive analysis method for benefiting enterprises based on a big data model, the specific method of which is as follows:
[0008] By using natural language processing technology to semantically analyze preferential policies for enterprises, a multi-dimensional preferential policy label system is constructed, including the level of effectiveness, the target beneficiaries, and the application conditions.
[0009] Construct a multi-dimensional feature library for enterprises and use random forest to calculate the fit score between preferential enterprise tags and enterprise features;
[0010] A time series prediction model was trained based on historical declaration data, and the Monte Carlo simulation algorithm was used to simulate the predicted range of subsidy payment amounts in Shanghai.
[0011] Generate a 3D visualization evaluation report that includes a heat map of the industry's driving effect.
[0012] As a preliminary step, semantic parsing uses an improved BERT model to identify the intent of preferential policies for enterprises. Through the Attention mechanism, hierarchical features including the level of effectiveness, the target beneficiaries, and the application conditions are extracted to generate a preferential knowledge graph for enterprises that includes industry codes, qualification requirements, and subsidy types.
[0013] Furthermore, the improved BERT model performs the following specific actions to identify the intent behind preferential policies for businesses:
[0014] Collect genuine documents benefiting businesses;
[0015] The improved BERT model repeatedly learns from real-world enterprise benefit documents and masters the document's specific expression methods;
[0016] The improved BERT model extracts and optimizes key information;
[0017] Generate a business-benefit relationship map using key information.
[0018] As a preferred option, during the fit scoring process, the enterprise feature vector is represented as Xi = [x1, x2, ..., x20]; where x1, x2, ..., x20 represent 20-dimensional features related to the enterprise.
[0019] As a preferred option, during the fit scoring process, the enterprise benefit label vector is represented as: Tj=[t1,t2,...,tm]; where t1,t2,...,tm represent the key information extracted by the improved BERT model.
[0020] A predictive analysis system for business benefit based on a big data model, the system comprising:
[0021] The data acquisition layer is used to connect to government data sharing platform APIs to various levels of government service websites that release information on preferential policies for enterprises.
[0022] The algorithm model layer is used to build a multi-dimensional feature library for enterprises, use random forest to calculate the fit score between preferential enterprise tags and enterprise features, train a time series prediction model based on historical application data, and combine Monte Carlo simulation algorithm to simulate the predicted range of subsidy payment amounts in Shanghai.
[0023] The application service layer is used to generate a 3D visualization evaluation report, including a heat map of the industry's driving effect.
[0024] As a preferred approach, the data acquisition layer collects information from the preferential policies webpage in real time by establishing dynamic field mapping rules;
[0025] The application service layer includes:
[0026] The front end is used to generate 3D heatmaps using ECharts.
[0027] The backend is used to provide real-time prediction services via the Flask API.
[0028] More preferably, the algorithm model layer includes:
[0029] The Enterprise Benefit Semantic Parsing Engine is used to identify the intent of enterprise benefit documents using an improved BERT model. It extracts hierarchical features (effectiveness level, support targets, application conditions) through the Attention mechanism and generates an enterprise benefit knowledge graph containing industry codes, qualification requirements, and subsidy types.
[0030] The enterprise profile matching engine integrates over 20 dimensions of enterprise-related features, uses the random forest algorithm to calculate the fit score between enterprise features and preferential enterprise tags, supports dynamic weight adjustment, and leverages Spark MLlib for distributed training of the random forest.
[0031] The Monte Carlo prediction module is designed to accelerate tens of thousands of iterations using the MPI parallel computing framework.
[0032] An electronic device includes: a memory and at least one processor;
[0033] The memory contains computer programs;
[0034] The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the enterprise-benefiting predictive analysis method based on the big data model described above.
[0035] A computer-readable storage medium storing a computer program that can be executed by a processor to implement the above-described method for predictive analysis of business benefits based on a big data model.
[0036] The enterprise benefit prediction and analysis method, system, equipment, and medium based on big data models of the present invention have the following advantages:
[0037] (I) This invention uses natural language processing technology to semantically analyze preferential enterprise texts and construct a multi-dimensional preferential enterprise label system; it combines an enterprise feature database to construct a dynamic matching model to achieve intelligent adaptation between preferential enterprise conditions and enterprise profiles; it uses a Monte Carlo simulation algorithm to predict the number of eligible enterprises and the range of subsidy payment amounts, and generates a three-dimensional visualization effect evaluation report. This solves the problems of low efficiency and large deviation in subsidy budget prediction caused by traditional preferential enterprise matching relying on manual screening. It achieves an enterprise screening accuracy rate of 92.3% and a subsidy payment amount prediction error rate of ≤7.5%. It is applicable to preferential enterprise formulation departments, industrial park management agencies, and enterprise preferential enterprise application service platforms.
[0038] (ii) This invention improves the accuracy of enterprise screening from 68% in traditional methods to 92.3% through multi-source data fusion and dynamic weight adjustment, thereby enhancing matching precision;
[0039] (iii) This invention uses Monte Carlo simulation to make the prediction error rate of subsidy payment amount ≤7.5%, which is better than the 15%-20% error of traditional linear models, thus optimizing budget control;
[0040] (iv) This invention supports real-time simulation and deduction of preferential policies for enterprises through a visual dashboard, shortening the response time from 3-5 days for manual analysis to within 10 minutes, thereby enhancing decision-making efficiency. Attached Figure Description
[0041] The invention will be further described below with reference to the accompanying drawings.
[0042] Appendix Figure 1 This is a flowchart of a predictive analysis method for benefiting enterprises based on big data models. Detailed Implementation
[0043] The following detailed description of the enterprise-benefiting predictive analysis method, system, equipment, and medium based on big data models of the present invention is provided with reference to the accompanying drawings and specific embodiments.
[0044] Example 1:
[0045] As attached Figure 1 As shown in the figure, this embodiment provides a predictive analysis method for benefiting enterprises based on a big data model. The method is as follows:
[0046] S1. Use natural language processing technology to semantically analyze the preferential policies for enterprises and construct a multi-dimensional preferential policy label system that includes the level of effectiveness, the target beneficiaries, and the application conditions.
[0047] S2. Construct a multi-dimensional feature library for enterprises and use random forest to calculate the fit score between the preferential enterprise tags and enterprise features;
[0048] S3. Train a time series prediction model based on historical declaration data, and combine the Monte Carlo simulation algorithm to simulate the predicted range of subsidy payment amounts in Shanghai.
[0049] S4. Generate a 3D visualization effect evaluation report including a heat map of the industry's driving effect.
[0050] In step S1 of this embodiment, the semantic parsing uses an improved BERT model to identify the intent of preferential policies for enterprises. The Attention mechanism is used to extract hierarchical features including the level of effectiveness, the target beneficiaries, and the application conditions, and to generate a preferential knowledge graph for enterprises that includes industry codes, qualification requirements, and subsidy types.
[0051] In this embodiment, the improved BERT model performs intent recognition for preferential policies for businesses as follows:
[0052] ① Collect genuine documents benefiting businesses;
[0053] ② The improved BERT model repeatedly learns from real enterprise benefit documents and masters the document-specific expression methods;
[0054] ③ The improved BERT model extracts and optimizes key information;
[0055] ④ Generate a business-benefit relationship map using key information.
[0056] In the fit scoring process of step S2 in this embodiment, the enterprise feature vector is represented as Xi = [x1, x2, ..., x20]; where x1, x2, ..., x20 represent 20-dimensional features related to the enterprise.
[0057] In the fit scoring process of step S2 in this embodiment, the enterprise benefit label vector is represented as: Tj=[t1,t2,...,tm]; where t1,t2,...,tm represent the key information extracted by the improved BERT model.
[0058] Example 2:
[0059] This embodiment provides a business benefit prediction and analysis system based on a big data model. The system includes:
[0060] The data acquisition layer is used to connect to government data sharing platform APIs to various levels of government service websites that release information on preferential policies for enterprises.
[0061] The algorithm model layer is used to build a multi-dimensional feature library for enterprises, use random forest to calculate the fit score between preferential enterprise tags and enterprise features, train a time series prediction model based on historical application data, and combine Monte Carlo simulation algorithm to simulate the predicted range of subsidy payment amounts in Shanghai.
[0062] The application service layer is used to generate a 3D visualization evaluation report, including a heat map of the industry's driving effect.
[0063] In this embodiment, the data acquisition layer collects information from the enterprise benefit webpage in real time by establishing dynamic field mapping rules.
[0064] The application service layer in this embodiment includes:
[0065] The front end is used to generate 3D heatmaps using ECharts.
[0066] The backend is used to provide real-time prediction services via the Flask API.
[0067] The algorithm model layer in this embodiment includes:
[0068] The Enterprise Benefit Semantic Parsing Engine is used to identify the intent of enterprise benefit documents using an improved BERT model. It extracts hierarchical features (effectiveness level, support targets, application conditions) through the Attention mechanism and generates an enterprise benefit knowledge graph containing industry codes, qualification requirements, and subsidy types.
[0069] The enterprise profile matching engine integrates over 20 dimensions of enterprise-related features, uses the random forest algorithm to calculate the fit score between enterprise features and preferential enterprise tags, supports dynamic weight adjustment, and leverages Spark MLlib for distributed training of the random forest.
[0070] The Monte Carlo prediction module is designed to accelerate tens of thousands of iterations using the MPI parallel computing framework.
[0071] Example 3:
[0072] Calculation example:
[0073] Company A characteristics: [Number of patents = 8, R&D expenditure = 6%, Industry = Integrated Circuits]; Preferential treatment label for enterprises: [Requires ≥ 5 patents, R&D expenditure > 5%];
[0074] Out of 1000 trees, 920 were deemed suitable → Fit score = 0.92;
[0075] Dynamic weight adjustment mechanism
[0076] Step 1: Initial weight allocation, as shown in the table below:
[0077] Feature type Initial weights Allocation basis Patent holdings 0.25 Frequent Requirements Characteristics of Documents Promoting Business Development R&D expenditure ratio 0.20 Affecting the amount of subsidies Industry matching 0.18 Relevance of support direction ... ... Ranking of Gini coefficient importance
[0078] Step 2: Dynamically adjust the triggering conditions, as follows:
[0079] When the enterprise benefit document library adds ≥3 new similar enterprise benefit documents;
[0080] Or the historical matching error rate is >10%;
[0081] Step 3: Weight update algorithm, as follows:
[0082] # Sliding window update formula (α = forgetting factor)
[0083] new_weight = 0.7 * old_weight + 0.3 * real_time_weight # Real-time weight calculation logic if the new policy document for supporting enterprises emphasizes "green manufacturing":
[0084] Environmental certification weighting += 0.15 #Increase focus on emerging policies
[0085] Patent weight -= 0.05# Reduce the weight of non-core features;
[0086] The technological advantages are compared in the table below.
[0087]
[0088]
[0089] Application prediction model:
[0090] An LSTM time series model is trained based on historical declaration data to predict the base number of eligible enterprises.
[0091] LSTM time series prediction workflow;
[0092] The data preparation phase includes the following steps:
[0093] Data source: Anonymized historical data (2019-2023) obtained from the enterprise benefit application platform.
[0094] Date Enterprise Benefit Document ID Number of companies applying Pass rate 202301 GOV086 1,240 82.3% 202302 GOV092 986 79.1%
[0095] Feature engineering, as detailed below:
[0096] Sliding window generation: The sequence is cut into 12-month cycles;
[0097] Normalization: The number of firms is scaled to [0,1] using the Min-Max method;
[0098] The key parameters for model construction and training are shown in the table below:
[0099]
[0100] Enterprise Base Forecast: Input the declaration data for the most recent 12 months → Output the predicted number of enterprises for future months. Example:
[0101] Input full-year 2023 data → Predicted base number for Q1 2024: 3,850 companies;
[0102] (Measured error rate ≤ 4.2%, 11.8% lower than the traditional ARIMA model);
[0103] Example 4:
[0104] This embodiment also provides an electronic device, including: a memory and a processor;
[0105] The memory stores the instructions executed by the computer.
[0106] The processor executes computer execution instructions stored in the memory, causing the processor to execute the enterprise-benefiting predictive analysis method based on a big data model in any embodiment of the present invention.
[0107] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.
[0108] Memory is used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory devices, or other volatile solid-state storage devices.
[0109] Example 5:
[0110] This embodiment also provides a computer-readable storage medium storing multiple instructions, which are loaded by a processor to cause the processor to execute the enterprise-benefiting predictive analysis method based on a big data model according to any embodiment of the present invention. Specifically, a system or apparatus equipped with a storage medium may be provided, on which software program code implementing the functions of any of the above embodiments is stored, and the computer (or CPU or MPU) of the system or apparatus may read and execute the program code stored in the storage medium.
[0111] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.
[0112] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.
[0113] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0114] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A predictive analysis method for enterprise benefits based on a big data model, characterized in that, The method is as follows: By using natural language processing technology to semantically analyze preferential policies for enterprises, a multi-dimensional preferential policy label system is constructed, including the level of effectiveness, the target beneficiaries, and the application conditions. Construct a multi-dimensional feature library for enterprises and use random forest to calculate the fit score between preferential enterprise tags and enterprise features; A time series prediction model was trained based on historical declaration data, and the Monte Carlo simulation algorithm was used to simulate the predicted range of subsidy payment amounts in Shanghai. Generate a 3D visualization evaluation report that includes a heat map of the industry's driving effect.
2. The enterprise benefit prediction and analysis method based on big data model according to claim 1, characterized in that, Semantic parsing employs an improved BERT model to identify the intent of preferential policies for enterprises. Through the Attention mechanism, hierarchical features including the level of effectiveness, the target beneficiaries, and the application conditions are extracted to generate a preferential knowledge graph for enterprises that includes industry codes, qualification requirements, and subsidy types.
3. The enterprise benefit prediction and analysis method based on big data model according to claim 2, characterized in that, The improved BERT model for identifying the intent behind preferential policies for businesses is as follows: Collect genuine documents benefiting businesses; The improved BERT model repeatedly learns from real-world enterprise benefit documents and masters the document's specific expression methods; The improved BERT model extracts and optimizes key information; Generate a business-benefit relationship map using key information.
4. The enterprise benefit prediction and analysis method based on big data model according to claim 1, characterized in that, During the fit scoring process, the enterprise feature vector is represented as Xi = [x1, x2, ..., x20]; where x1, x2, ..., x20 represent 20-dimensional features related to the enterprise.
5. The enterprise benefit prediction and analysis method based on big data model according to claim 1, characterized in that, During the fit scoring process, the enterprise benefit label vector is represented as: Tj=[t1,t2,...,tm]; where t1,t2,...,tm represent the key information extracted by the improved BERT model.
6. A predictive analysis system for enterprise benefits based on a big data model, characterized in that, The system includes: The data acquisition layer is used to connect to government data sharing platform APIs to various levels of government service websites that release information on preferential policies for enterprises. The algorithm model layer is used to build a multi-dimensional feature library for enterprises, use random forest to calculate the fit score between preferential enterprise tags and enterprise features, train a time series prediction model based on historical application data, and combine Monte Carlo simulation algorithm to simulate the predicted range of subsidy payment amounts in Shanghai. The application service layer is used to generate a 3D visualization evaluation report, including a heat map of the industry's driving effect.
7. The enterprise-benefit prediction and analysis system based on a big data model according to claim 6, characterized in that, The data acquisition layer collects information from the preferential policies webpage in real time by establishing dynamic field mapping rules; The application service layer includes: The front end is used to generate 3D heatmaps using ECharts. The backend is used to provide real-time prediction services via the Flask API.
8. The enterprise-benefit prediction and analysis system based on a big data model according to claim 6 or 7, characterized in that, The algorithm model layer includes: The Enterprise Benefit Semantic Parsing Engine is used to identify the intent of enterprise benefit documents using an improved BERT model. It extracts hierarchical features through the Attention mechanism and generates an enterprise benefit knowledge graph that includes industry codes, qualification requirements, and subsidy types. The enterprise profile matching engine integrates over 20 dimensions of enterprise-related features, uses the random forest algorithm to calculate the fit score between enterprise features and preferential enterprise tags, supports dynamic weight adjustment, and leverages Spark MLlib for distributed training of the random forest. The Monte Carlo prediction module is designed to accelerate tens of thousands of iterations using the MPI parallel computing framework.
9. An electronic device, characterized in that, include: Memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the enterprise-benefiting predictive analysis method based on a big data model as described in any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the enterprise-benefiting predictive analysis method based on a big data model as described in any one of claims 1 to 5.