Target fund total expenditure prediction method and system under public health emergencies
By collecting multi-source heterogeneous data and utilizing a bidirectional cross-attention mechanism and meta-learning to optimize the temporal large language model, the accuracy and timeliness issues of fund expenditure forecasting under public health emergencies were solved, achieving efficient and accurate forecasting and emergency response.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CAPINFO CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-24
Smart Images

Figure CN122453528A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical insurance fund forecasting technology, and in particular to a method and system for forecasting total expenditure of target funds under public health emergencies. Background Technology
[0002] In the emergency response to public health emergencies (such as major epidemics and natural disasters), the target fund, as a core payment tool of the medical security system, plays a crucial role in ensuring the rational allocation of medical resources and preventing the risk of fund depletion through accurate prediction of its expenditure scale. Existing technologies mainly rely on traditional time series models or deep learning models for fund expenditure prediction, but these methods have significant limitations, as they struggle to effectively integrate dynamic correlation information from multi-source heterogeneous data. Specifically, target fund expenditures under public health emergencies are influenced by multiple coupled factors, including the intensity of public opinion dissemination, the extent of policy adjustments, and historical expenditure patterns. Existing models typically only convert textual data into structured indicators through simple feature engineering and then concatenate them with the fund's time series data, failing to establish a two-way dynamic correlation mechanism between the textual modality and the structured time series modality. This shallow integration method prevents the model from quantifying the implicit transmission path of public opinion events, policy responses, and fund expenditures. Especially in data-scarce emergency scenarios, the prediction error increases significantly, making it difficult to meet the 24-hour emergency decision-making response timeliness and prediction accuracy requirements stipulated in the "National Emergency Response Plan for Public Emergencies." Summary of the Invention
[0003] In view of this, the present invention proposes a method to improve the accuracy and response speed of total expenditure forecasting of target funds during public health emergencies. The present invention provides the following technical solution: A method for forecasting total expenditure of a target fund under a public health emergency includes: Collect multi-source heterogeneous data, including at least time-series data of the target fund, textual data of public health events, and textual data of the target policy, and construct a multimodal feature matrix; The multimodal feature matrix is fused using a bidirectional cross-attention mechanism to obtain a fused standard feature set; The standard feature set is input into a pre-trained temporal large language model through a domain knowledge-guided adaptation method; The temporal large language model is adaptively optimized based on the corresponding historical scene data using a meta-learning mechanism. Based on the optimized temporal large language model and the standard feature set, the predicted results of the target total fund expenditure under the public health emergency scenario are output.
[0004] Optionally, the data collection includes at least multi-source heterogeneous data such as time-series data of the target fund, textual data of public health events, and textual data of the target policy, and the construction of a multimodal feature matrix includes: Credibility assessment and duplication detection are performed on textual data of public health events and target policies, and textual data is screened based on credibility weights and duplication thresholds; A pre-trained language model is used to extract features from the filtered text data, and text feature vectors are generated by combining term frequency-inverse document frequency (TF-IDF) weighting. Missing values are filled and outliers are corrected in the time series data of the target fund, and the time granularity is unified to daily to generate a structured data vector. The multimodal feature matrix is constructed using daily timestamps as the horizontal dimension and text feature vectors and structured data vectors as the vertical dimensions.
[0005] Optionally, the step of fusing the multimodal feature matrix through a bidirectional cross-attention mechanism to obtain a fused standard feature set includes: Based on the text feature vector and structured data vector in the multimodal feature matrix, a bidirectional cross-attention association calculation is performed to calculate the positive influence weight of the text modality on the structured data modality and the reverse association strength of the structured data modality on the text modality. The positive impact weight and negative correlation strength are dynamically adjusted based on the product of the public health event level and the effectiveness of the target policy. The event level is determined according to the classification standard of the emergency response plan for public health emergencies, and the policy effectiveness is determined according to the level of the policy issuing body. Multi-head attention fusion is performed based on the adjusted weights and correlation strength to output the standard feature set.
[0006] Optionally, the step of inputting the standard feature set into the pre-trained temporal large language model through a domain knowledge-guided adaptation method includes: The standard feature set is dynamically divided into time patches according to the target business cycle; The time patch is mapped to the word embedding space of a pre-trained temporal large language model by linear projection, generating a sequence of temporal feature vectors. The prompt text containing target domain knowledge is converted into a prompt vector sequence, and then concatenated with the temporal feature vector sequence in the vector space to form a complete input sequence, which is then input into the temporal big language model.
[0007] Optionally, the adaptive optimization of the temporal large language model using corresponding historical scene data based on the meta-learning mechanism includes: Based on the event type, scope of impact, and policy response characteristics of public health emergencies, historical scenario data corresponding to the current scenario are retrieved from historical data to construct a meta-training set; Meta-training is performed on the temporal large language model based on the meta-training set to obtain initial parameters with cross-scene adaptability. Using the limited samples of the public health emergency scenario, the temporal large language model is fine-tuned based on the initial parameters to achieve adaptive optimization for the current scenario.
[0008] Optionally, the prediction results of the target total fund expenditure under the public health emergency scenario, based on the optimized temporal large language model and the standard feature set, include: Based on the standard feature set, the time-series big language model with adaptive optimization outputs the predicted total expenditure of the target fund at multiple preset time scales in the future. Generate and output interpretable analysis data associated with the prediction results, the interpretable analysis data including at least a fund expenditure trend chart and a feature importance analysis chart reflecting the degree of influence of different data modalities on the prediction results.
[0009] This invention further discloses a target fund total expenditure prediction system under a public health emergency, comprising: The data acquisition and feature construction module is used to collect multi-source heterogeneous data, including at least time-series data of the target fund, textual data of public health events, and textual data of the target policy, and to construct a multimodal feature matrix. The feature fusion module is used to fuse the multimodal feature matrix through a bidirectional cross-attention mechanism to obtain a fused standard feature set; The domain adaptation module is used to input the standard feature set into the pre-trained temporal big language model through an adaptation method guided by domain knowledge. The meta-learning optimization module is used to adaptively optimize the temporal large language model based on the meta-learning mechanism and corresponding historical scene data. The prediction output module is used to output the prediction results of the target total expenditure of the fund under the public health emergency scenario based on the optimized temporal large language model and the standard feature set.
[0010] The present invention further discloses a computer-readable storage medium, characterized in that the storage medium stores a computer program, which, when executed by a processor, implements the above-described method.
[0011] The present invention further discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the above-described method when executing the program.
[0012] The present invention further discloses a computer program product, including a computer program, characterized in that the computer program implements a method when executed by a processor.
[0013] According to the technical solution of this invention, by collecting multi-source heterogeneous data such as target funds, public health events, and policy texts, and using a bidirectional cross-attention mechanism for deep fusion, a dynamic correlation between "public opinion-policy-fund expenditure" is constructed, thereby significantly improving prediction accuracy. Furthermore, through a domain knowledge-guided adaptation method, the fused features are input into a time-series large language model, and combined with a meta-learning mechanism, the model is rapidly and adaptively optimized in low-sample emergency scenarios. This enables the model to not only have a deep understanding of the patterns in areas such as target payment cycles and policy impacts, but also to achieve rapid and accurate prediction responses to rare emergencies. Finally, multi-period predicted values and interpretable analysis charts are output, providing efficient and transparent intelligent support for emergency decision-making by target funds under major health events while ensuring prediction accuracy. Attached Figure Description
[0014] For illustrative and not limiting purposes, the present invention will now be described in conjunction with embodiments and accompanying drawings, wherein: Figure 1 This is a schematic diagram of the target fund total expenditure prediction method under a public health emergency according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the constituent modules of the target fund total expenditure prediction system under a public health emergency in an embodiment of the present invention; Figure 3 This is a schematic diagram of the composition structure of the electronic device in an embodiment of the present invention. Detailed Implementation
[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application.
[0016] It should be noted that, where there is no conflict, the embodiments and features of the embodiments in this application can be combined with each other. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0017] refer to Figure 1 This embodiment discloses a method for predicting total expenditure of a target fund under a public health emergency, including the following steps: S100: Collect multi-source heterogeneous data, including at least time-series data of the target fund, textual data of public health events, and textual data of the target policy, and construct a multimodal feature matrix.
[0018] First, for the collection of target fund time-series data, a structured data table containing fields such as daily fund expenditure total and settlement number is collected within a preset time period through the target business data interface. For the collection of public health event text data, this implementation method uses the Scrapy framework to write a targeted web crawler to automatically crawl daily news reports on emergencies published by authoritative media such as Xinhua News Agency and Health World, as well as the official websites of local health commissions. The crawled content must at least include the event title, body text, publication time, event type, scope of impact, and source media. For the collection of target policy text data, this implementation method uses the PyPDF2 text parsing tool to batch parse policy documents published by the target management agency or relevant government platform official websites, extracting structured information such as policy title, issuing unit, publication date, effective date, and core adjustment content (such as payment method reform and reimbursement ratio adjustment).
[0019] Subsequently, text data credibility assessment and duplication detection were performed, and the captured public health event text data and target policy text data were cleaned. Specifically, credibility weights were assigned based on the authority of the source (for example, official media had a weight of 1.0, and non-authoritative sources had a weight of 0.3). Semantic similarity between texts was calculated, and when the similarity was ≥0.85, it was determined to be duplicate content, and only the one with the highest weight or the earliest publication time was retained, thus achieving text data filtering based on credibility weights and duplication thresholds. Further, a BERT pre-trained language model was used to segment and encode the filtered text data to obtain deep semantic features. Simultaneously, the statistical features of keywords related to target fund expenditures (such as "confirmed cases," "isolation," and "reimbursement adjustment") were extracted using the TF-IDF weighted method. The BERT semantic vectors and TF-IDF weighted features were fused to generate the final text feature vector. At the same time, the time-series data of the target fund was preprocessed. For missing values, linear interpolation (continuous time series) or K-nearest neighbor imputation (discrete classification) was used to fill in the missing values. Outliers are identified using the 3σ criterion and confirmed in conjunction with business system logs. If the outlier is due to data collection anomalies, the average of the next three days is used for correction. Finally, all data is standardized to a daily time granularity, and the monetary unit is standardized to "ten thousand yuan," generating a well-structured data vector.
[0020] Finally, using the "daily" timestamp as a unified horizontal dimension, the generated text feature vectors and structured data vectors are precisely aligned according to their timestamps. For a given daily timestamp, if there is no corresponding text data, its text feature vector is set to zero. Ultimately, the feature vectors from all time steps are stacked vertically to construct the multimodal feature matrix. The horizontal dimension of this matrix represents the time series, while the vertical dimension contains the fused features of the text and structured data modalities, providing standardized input for subsequent feature fusion steps.
[0021] S200: The multimodal feature matrix is fused through a bidirectional cross-attention mechanism to obtain a fused standard feature set.
[0022] Specifically, the multimodal feature matrix constructed in step S100 is split into text feature sequences (containing feature vectors of public health events and target policies) and structured data feature sequences (containing feature vectors of target fund time-series data) according to their feature sources. The two sequences are then processed through an independent linear transformation layer, mapping them to the same feature dimension to prepare for subsequent attention calculations.
[0023] Subsequently, a bidirectional cross-attention correlation calculation is performed. The forward cross-attention calculation uses text feature sequences as the query source and structured data feature sequences as the key and value source. This is used to quantify the potential positive impact weight of textual information (such as pandemic news and policy provisions) on the target fund data. An attention weight matrix is further calculated, where the magnitude of each element represents the degree of influence of textual information at one time on fund expenditure data at another time. The reverse cross-attention calculation uses structured data feature sequences as the query source and text feature sequences as the key and value source. This is used to explore the strength of inverse correlations that may be associated with fluctuations in the target fund data (such as abnormal expenditure peaks) but are not explicitly described in the text. Another attention weight matrix is further calculated to uncover potential event or policy clues behind changes in fund data.
[0024] Furthermore, to enhance key information and suppress redundancy, a dynamic adjustment mechanism is introduced after obtaining the two attention weights mentioned above. First, based on predefined rules, each piece of text data is assigned two attribute values: the event level of the public health event (divided into levels 1-4, referring to national emergency response standards) and the policy effectiveness of the target policy (divided into levels 1-3, determined according to the issuing agency level). Then, the contribution of the data at each time step to the attention calculation is dynamically adjusted based on the product of the event level and policy effectiveness. Specifically, this is achieved through a predefined weight mapping rule: the larger the product, the more significant the event or the more important the policy, and the corresponding positive impact weight and negative correlation strength will be amplified (e.g., increased to 1.8 to 2.5 times the normal level); for data with small products (such as duplicate news or local temporary notices), their weight is correspondingly reduced (e.g., reduced to 0.1 to 0.3 times). This process dynamically adjusts the positive impact weight and negative correlation strength based on the product of the public health event level and the target policy effectiveness.
[0025] The dynamically adjusted forward and backward attention calculation results are integrated and input into a multi-head attention fusion layer (e.g., set to 8 heads). This layer can capture complex intermodal dependencies in parallel from multiple different representation subspaces, achieving a deeper level of fusion.
[0026] Finally, the fused features undergo layer normalization to stabilize the learning process, and then are transformed nonlinearly through a feedforward neural network. A linear projection layer adjusts the feature dimensions to a preset uniform size, and the entire output sequence is standardized to obtain the final fused standard feature set. This feature set is a high-dimensional temporal feature representation that integrates dynamic associations of multi-source information, providing high-quality input for subsequent temporal large-scale language models.
[0027] S300: Input the standard feature set into the pre-trained temporal large language model through a domain knowledge-guided adaptation method.
[0028] Specifically, firstly, dynamic time patching is implemented. Based on the inherent cyclical characteristics of the target business and the current forecast scenario, the standard feature set fused in step S200 is intelligently segmented on the time axis. In this embodiment, the specific segmentation rules are as follows: under normal business scenarios, a 7-day time patch is used to align with the inherent cycle of weekly settlement of target expenses and medical treatment; during the emergency response phase of a public health emergency, the patch length is shortened to 1 day to accurately capture short-term drastic fluctuations in fund expenditures; when conducting medium- to long-term trend forecasting or policy impact assessment, a 30-day time patch is used to cover the complete monthly management cycle or a typical policy impact cycle. This dynamic segmentation mechanism ensures that the organization of time-series data matches the actual rhythm of the target business.
[0029] Secondly, input reprogramming is performed. The feature data within each time patch is mapped through a trainable fully connected linear layer to the word embedding space dimension (e.g., 4096 dimensions) pre-set by the pre-trained TIME-LLM model. This transforms the numerical temporal feature vectors into a sequence of temporal feature vectors that the model can process, thereby enabling it to understand temporal information without modifying the core architecture of the model.
[0030] Finally, the embedding of domain knowledge guidance is completed. A structured prompt text containing target domain knowledge is constructed and converted into a corresponding prompt vector sequence using TIME-LLM's own word segmenter. In this embodiment, the prompt text adopts the following exemplary layered design: The first layer is the task instruction layer, which clearly defines the prediction target, such as "Please predict the total expenditure of the target fund in the next 14 days based on the following time series characteristics."; the second layer is the domain pattern layer, which injects prior knowledge specific to the target business, such as "Note: Target expenditure usually shows a seasonal peak in winter and at the end of the year and the beginning of the new year, and the full impact of the DRG / DIP payment method reform has a lag of about 6 months."; the third layer is the scenario identification layer, which provides the context of the current prediction, such as "Current prediction background: This city is in the peak season for influenza, and a new round of target protection policies has recently been implemented." In the model input stage, this "prompt vector sequence" is used as a prefix and directly concatenated with the above time series feature vector sequence in the sequence dimension to form a complete input sequence, which is then fed into the TIME-LLM model for forward computation. This enables the model to efficiently activate its ability to recognize and apply complex laws in the target domain through external guidance information.
[0031] S400: Based on a meta-learning mechanism, the temporal large language model is adaptively optimized using relevant historical scene data. This addresses the challenge of data scarcity in the early stages of public health emergencies. The optimization process is divided into an offline meta-training phase and an online rapid fine-tuning phase.
[0032] Specifically, the first step involves retrieving historical similar scenarios and constructing a meta-training set. When a new public health emergency (e.g., a new localized outbreak) is identified as corresponding to the current prediction task, all historical scenario data matching the key features are automatically retrieved from the historical database based on the event type, scope of impact, and initial policy response characteristics of that scenario. For example, for a new event of "a localized outbreak of respiratory infectious disease in City A," the database will retrieve all historical event cycles that occurred in any region, were caused by respiratory infectious diseases, and initially involved similar control measures. The complete data for each historical scenario (from its occurrence to its end) is treated as an independent learning task, and all these tasks together constitute the meta-training set used for meta-learning.
[0033] Secondly, Model-Independent Meta-Learning (MAML) training is performed. In this offline phase, the temporal large language model is meta-trained using the MAML algorithm framework. This optimizes the model's initial parameters by repeatedly simulating a small number of samples learning across multiple tasks (i.e., multiple historical scenarios). Specifically, in each training iteration, the algorithm samples a support set (simulating "small samples") and a query set from a historical scenario task. The model performs one or more gradient updates based on the support set, then calculates the loss on the query set, and uses this loss for backpropagation to adjust the model's initial parameters. After multiple iterations, the model is trained with a set of highly generalizable initial parameters, enabling it to achieve good performance with minimal gradient update steps when facing a completely new task with very little data. This allows the model to quickly adapt to new tasks.
[0034] Finally, a rapid fine-tuning process is triggered for the current emergency scenario. Once a new public health emergency actually occurs and a small amount of sample data containing actual target fund expenditure labels has accumulated over the first few days (e.g., 5-10 days), the online optimization process is initiated. At this point, the model, trained with initial parameters obtained through meta-training, is loaded, and gradient descent fine-tuning is performed iteratively (usually less than 10 iterations) using only this small amount of sample data from the current scenario. This process is computationally inefficient and time-saving (usually completed within minutes), enabling the model parameters to be quickly and accurately adapted to the specific patterns of the new event, thus completing the final adaptive optimization. Subsequently, the optimized model can be used to accurately predict the subsequent development of the event. This mechanism ensures that the model maintains high predictive accuracy even in the early stages of an emergency response when data is extremely limited.
[0035] To address the severe imbalance between the number of regular scenario samples and emergency scenario samples in the training data for the training process of the temporal large language model, this implementation method designs a dedicated loss function to optimize model performance. A weighted loss function is used in both the main training phase and the rapid fine-tuning phase for the current scenario. This function, based on the standard mean squared error (MSE), assigns a higher penalty weight to the prediction errors of samples belonging to the "emergency public health scenario" in the training batch. Specifically, the loss function takes the following form: in This is a sample collection for public health emergencies. This is a collection of samples from typical scenarios. This is the weighted mean square error loss value. This represents the total number of samples in the current training batch. For the first The true label or observation value of a sample For the model to the first The predicted value for each sample, For set or The penalty weight coefficient for the i-th sample.
[0036] This mechanism forces the model to focus more on learning rare but critical emergency scenario patterns during training, thereby significantly improving the model's generalization ability and robustness in emergency prediction tasks.
[0037] S500: Based on the optimized temporal large language model and the standard feature set, output the prediction results of the target total fund expenditure under the public health emergency scenario.
[0038] During the forecasting calculation, the standard feature set constructed under the current public health emergency scenario is input into the aforementioned adaptively optimized Time-Large Language Model (TIME-LLM). Guided by prior domain knowledge prompts, the model performs deep reasoning on the input features and outputs multi-timescale forecasts covering different decision-making perspectives in one go. Specifically, the model generates short-term rolling forecasts for the next 7 days, medium-term trend forecasts for the next 14 days, and long-term trend forecasts for the next 30 days in parallel. The forecast results directly correspond to the daily estimated value of the target fund's total expenditure, providing core quantitative basis for the target management department's decisions on fund allocation and emergency plan activation.
[0039] Furthermore, during the interpretability analysis output, supplementary analysis content closely related to the aforementioned predicted values is generated and output to enhance decision-makers' understanding of the model's predictive logic and key driving factors. The generated interpretability analysis data should include at least the following two types of visualizations: Fund Expenditure Trend Chart: This chart, presented as a time-series line graph, visually compares and displays historical actual expenditures, current data, and model-predicted future multi-scale expenditure trends. The chart clearly marks the prediction intervals and confidence ranges, making the direction and magnitude of expenditure fluctuations readily apparent.
[0040] Feature Importance Analysis Chart: This chart, in the form of a heatmap or bar chart, quantifies the relative contribution or influence weight of different input features (e.g., the intensity of COVID-19 news on a specific date, the effectiveness of a policy, historical fund expenditures, etc.) to the final prediction result during the model's prediction process. This allows policymakers to trace the main driving sources of the prediction results; for example, identifying whether recent news of an escalating epidemic in a high-risk area or a new reimbursement policy became the most influential factor in this prediction.
[0041] Ultimately, all predicted values and analytical charts are rendered and published in real time through an integrated web visualization platform. From the moment a user submits a prediction request to the presentation of the complete results page, the end-to-end response time of the entire system is kept within 5 minutes, thus meeting the stringent timeliness requirements for emergency decision-making during major public health events.
[0042] To verify the comprehensive performance of the target fund prediction method constructed in this invention based on a multi-source fusion framework guided by cross-attention and dynamic weight allocation, and a temporal large language model, a model evaluation stage was established. This stage was executed on an independent test set, and a multi-dimensional evaluation index was used to comprehensively and quantitatively analyze the accuracy, error scale, bias direction, and fitting stability of the prediction results. Specifically: We used complete multi-source datasets, divided chronologically into training, validation, and test sets in a 7:2:1 ratio. The test set was completely independent and did not participate in any model training or hyperparameter tuning process. On the test set, we calculated and analyzed the following key metrics to evaluate model performance: Regarding the overall accuracy metrics, the overall accuracy rate of the model predictions reached 88.55%, approaching the industry standard of ">90% for excellent," indicating a strong ability of the model to grasp the overall trend of target fund expenditures. Specifically, the Mean Absolute Percentage Error (MAPE) and Symmetrical MAPE (sMAPE) were as follows: MAPE was 11.45%, close to the threshold of "<10% for excellent"; sMAPE was 12.55%, effectively avoiding the problem of percentage error distortion caused by the actual value being close to zero. These two metrics together confirm that the percentage deviation between the model's predicted values and the actual values is controlled at a low level.
[0043] Regarding the error magnitude indicators, the root mean square error (RMSE) and mean absolute error (MAE) are as follows: RMSE is 63,870,542.9, and MAE is 44,705,064.78. These absolute error values are consistent with the target fund's expenditures, which are in the hundreds of millions, and are within an acceptable business range. MAE is not sensitive to extreme values, and its relatively low value indicates that the overall deviation of the model's predictions is controllable. The mean square error (MSE) is 40,794,462,504,357,433.5. Its relatively large value is mainly due to the squared penalty applied to individual large error points. This reflects the model's efforts to avoid extreme prediction errors, meeting the risk management requirements of strictly avoiding significant deviations in the target fund's predictions.
[0044] Regarding the direction of bias and systematic analysis, the mean bias rate and mean percentage error (MPE) are as follows: the mean bias rate is -11.13%, and the MPE is 9.6%. The values are close and their signs are opposite, indicating a slight, non-systematic tendency to underestimate the model's predictions, but overall there is no serious systematic bias. The mean residual is 41,453,305.2717, which is close to zero under the target prediction, further statistically demonstrating that the model does not have significant systematic prediction bias.
[0045] Regarding the stability and explanatory power indices, the coefficient of determination (R²) and adjusted R² are shown: R² is 0.7639, and adjusted R² is 0.7606. This indicates that the model can explain approximately 76.39% of the variation in target fund expenditure data, demonstrating strong explanatory power. The small difference between adjusted R² and R² indicates appropriate model complexity, sufficient sample size, and robust evaluation results. The residual standard deviation is 48,590,840, a relatively low value, reflecting the model's stable predictive performance at different time points with low volatility. Therefore, this prediction method demonstrates good overall accuracy, controllable error scale, weak non-systematic bias, and excellent stability and explanatory power on the test set. All indicators meet or approach the preset excellent standards, validating the effectiveness and advancement of the complete technical solution of "cross-attention-guided multi-source fusion," "domain knowledge adaptation," and "meta-learning few-shot optimization," providing reliable technical support for emergency decision-making regarding target funds during major public health events. The evaluation results support the practical application of this model.
[0046] In summary, the complete technical solution constructed in this specific implementation method achieves dynamic correlation between events, policies, and expenditures by collecting and deeply fusing multi-source data, including target funds, event texts, and policy texts, based on bidirectional cross-attention, thereby improving the quality of feature representation and prediction accuracy. Furthermore, through dynamic time patches and target-specific hint prefixes, domain knowledge is seamlessly embedded into the input space of the time-series large language model, achieving deep adaptation of the model to the periodicity of targets and the patterns of policy impact. Addressing the challenge of scarce data in emergency scenarios, a meta-learning mechanism is employed, utilizing historically similar scenarios to train the model's rapid adaptability, and fine-tuning it online with only a very small number of samples from the current scenario. This ensures prediction accuracy while enabling rapid emergency response to major health events. Finally, the system outputs multi-timescale predicted values and interpretable analysis charts, providing decision-makers with intelligent support that combines accurate quantitative results with transparent decision-making basis, comprehensively achieving high precision, strong adaptability, and rapid response technical effects.
[0047] refer to Figure 2 This embodiment further discloses a target fund total expenditure prediction system under a public health emergency, including: The data acquisition and feature construction module 21 is used to collect multi-source heterogeneous data, including at least target fund time-series data, public health event text data, and target policy text data, and to construct a multimodal feature matrix. This module is responsible for the automated acquisition, cleaning, and feature engineering of multi-source heterogeneous data. Specifically, it periodically calls the target data platform API through the built-in data interface unit to obtain structured target fund time-series data; it automatically crawls public health event news reports from designated authoritative media and government websites through a targeted crawling unit based on the Scrapy framework; and it parses target policy PDF files through a document parsing unit (integrating tools such as PyPDF2). After acquiring the raw data, its built-in preprocessing engine performs credibility assessment, duplication detection, missing value imputation, outlier correction, and normalization operations, and calls pre-trained language models (such as BERT) and the TF-IDF algorithm to generate text feature vectors. Finally, this module aligns all features according to a unified timestamp and outputs the constructed multimodal feature matrix to the downstream feature fusion module 22.
[0048] The feature fusion module 22, connected to the data acquisition and feature construction module 21, is used to fuse the multimodal feature matrix through a bidirectional cross-attention mechanism to obtain a fused standard feature set. Specifically, it receives the multimodal feature matrix output by the data acquisition and feature construction module 21 and performs deep cross-modal fusion. Its core is a bidirectional cross-attention network. This network first splits the input matrix into text and structured data streams, calculating positive attention from text to structure (quantifying event impact) and negative attention from structure to text (mining implicit associations) respectively. The module's embedded dynamic weight adjuster amplifies the attention weight of key information in real time based on the event level and policy effectiveness labeled in the input data. Specifically: Dynamic weighting mechanism: A scenario-driven and cross-attention-guided dynamic weighting mechanism is designed to achieve adaptive weighted penalty for errors in key samples, thereby improving the model's predictive robustness under sudden events.
[0049] For each sample in the training batch Define binary scenario indicator variables Used to distinguish sample types: ; Set base weights for the two types of samples, ensuring that the weight for critical scenarios is higher than that for ordinary scenarios: Basic weighting for public health emergency scenarios: (Values: 1.5, 2, 2.5, 5) Basic weights for typical scenarios: =1 Dynamic weighted coefficient calculation: The weighted coefficient integrates three types of information: scene label, cross-attention importance, and business rule strength. The calculation formula is as follows: ,in For the first Dynamic weighting coefficients for each sample; , , Assign coefficients to the weights, satisfying This can be determined through experimental optimization; This is the label for the sample scene, and its value can be 0 or 1. The importance of the sample output by the cross-attention module represents the strength of the association between the sample and external public opinion and policy information; The business rule strength coefficient is quantified by the level of the emergency, the policy level, and the degree of risk, and its value ranges from [0,1].
[0050] Weight calculation: The final weights of the samples are obtained by adaptively adjusting the basic weights based on the weighting coefficients. .
[0051] When the sample belongs to a public health emergency scenario =1, Too big Approaching This significantly increases the penalty for errors; When the sample is a typical scenario =0, Smaller than average Approaching =1, maintain the normal level of punishment.
[0052] Weight normalization processing To ensure the stability of loss values across different batches, the weights of all samples within a batch are normalized: ,in, This represents the total number of samples in the current training batch. This represents the final weight after normalization.
[0053] Finally, a multi-head attention fusion layer is used to aggregate bidirectional information, and after layer normalization and linear projection, a unified, high-dimensional standard feature set is output to the domain adaptation module 23.
[0054] Domain adaptation module 23 is used to input the standard feature set into a pre-trained temporal large language model through a domain knowledge-guided adaptation method. It is used to adapt temporal features to the input space of a temporal large language model (such as TIME-LLM). Specifically, it includes: The dynamic patch partitioning unit is used to divide the standard feature set into time patches of different lengths based on the scenario (regular / sudden / monthly). Input reprogramming units are used to map each patch to the word embedding space of the model through trainable linear layers; The prompt construction and concatenation unit is used to generate structured target domain prompt text (including task instructions, domain rules and scene identifiers) according to the task. After converting it into a vector, it is used as a prefix and concatenated with the reprogrammed temporal feature vector to form a complete model input sequence, which is then fed into the pre-trained temporal large language model.
[0055] The meta-learning optimization module 24 is used to adaptively optimize the temporal large language model based on relevant historical scene data using a meta-learning mechanism. A model-independent meta-learning framework is employed. It includes a scene retrieval unit, used to quickly retrieve similar scene data from the historical database based on the characteristics of new events to construct a meta-training task set. A meta-training engine trains the temporal large language model offline using these historical task sets, optimizing its initial parameters to achieve rapid adaptation. When a new sudden event occurs and a small amount of labeled data accumulates, an online fine-tuner is triggered, using this small amount of data to fine-tune the model, completing the adaptive optimization for the current scene, and synchronizing the optimized model parameters to the prediction process.
[0056] The prediction output module 25 is used to output the predicted total expenditure of the target fund under the public health emergency scenario based on the optimized time-series large language model and the standard feature set. The prediction output module 25 is used to perform the final prediction and generate a visualization report. Specifically, it includes: The prediction calculation unit is used to call the temporal large language model configured by the domain adaptation module and optimized by the meta-learning optimization module, process the input standard feature set, and directly calculate the predicted fund expenditure values for the next 7 days, 14 days and 30 days. The interpretability analysis unit is used to calculate feature importance using algorithms such as integrated gradients, and drives the visualization rendering unit to automatically generate analysis reports that include trend line charts and feature importance heatmaps.
[0057] All results are output in real time via API interface or web visualization interface to ensure that the end-to-end response time meets the emergency requirement of 5 minutes.
[0058] Figure 3 A schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as... Figure 3 As shown, the electronic device 50 includes: a processor 501, a memory 502, and a bus 503; The processor 501 and the memory 502 communicate with each other via the bus 503; the processor 501 is used to call the program instructions in the memory 502 to execute the methods provided in the above-described embodiments.
[0059] This embodiment provides a non-transitory computer-readable storage medium that stores computer instructions that cause a computer to execute the methods provided in the above-described embodiments.
[0060] Those skilled in the art will understand that all or part of the steps of the above-described method implementation can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above-described method implementation. The aforementioned storage medium includes various storage media capable of storing program code, such as ROM, RAM, magnetic disk, or optical disk.
[0061] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0062] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0063] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for predicting total expenditure of a target fund under a public health emergency, characterized in that, include: Collect multi-source heterogeneous data, including at least time-series data of the target fund, textual data of public health events, and textual data of the target policy, and construct a multimodal feature matrix; The multimodal feature matrix is fused using a bidirectional cross-attention mechanism to obtain a fused standard feature set; The standard feature set is input into a pre-trained temporal large language model through a domain knowledge-guided adaptation method; The temporal large language model is adaptively optimized based on the corresponding historical scene data using a meta-learning mechanism. Based on the optimized temporal large language model and the standard feature set, the predicted results of the target total fund expenditure under the public health emergency scenario are output.
2. The method for predicting total expenditure of the target fund according to claim 1, characterized in that, The data collection includes at least multi-source heterogeneous data such as time-series data of the target fund, textual data of public health events, and textual data of the target policy, and the construction of a multimodal feature matrix includes: Credibility assessment and duplication detection are performed on textual data of public health events and target policies, and textual data is screened based on credibility weights and duplication thresholds; A pre-trained language model is used to extract features from the filtered text data, and text feature vectors are generated by combining term frequency-inverse document frequency (TF-IDF) weighting. Missing values are filled and outliers are corrected in the time series data of the target fund, and the time granularity is unified to daily to generate a structured data vector. The multimodal feature matrix is constructed using daily timestamps as the horizontal dimension and text feature vectors and structured data vectors as the vertical dimensions.
3. The method for predicting total expenditure of the target fund according to claim 1, characterized in that, The multimodal feature matrix is fused using a bidirectional cross-attention mechanism to obtain a fused standard feature set, including: Based on the text feature vector and structured data vector in the multimodal feature matrix, a bidirectional cross-attention association calculation is performed to calculate the positive influence weight of the text modality on the structured data modality and the reverse association strength of the structured data modality on the text modality. The positive impact weight and negative correlation strength are dynamically adjusted based on the product of the public health event level and the effectiveness of the target policy. The event level is determined according to the classification standard of the emergency response plan for public health emergencies, and the policy effectiveness is determined according to the level of the policy issuing body. Multi-head attention fusion is performed based on the adjusted weights and correlation strength to output the standard feature set.
4. The method for predicting total expenditure of the target fund according to claim 1, characterized in that, The step of inputting the standard feature set into a pre-trained temporal large language model through a domain knowledge-guided adaptation method includes: The standard feature set is dynamically divided into time patches according to the target business cycle; The time patch is mapped to the word embedding space of a pre-trained temporal large language model by linear projection, generating a sequence of temporal feature vectors. The prompt text containing target domain knowledge is converted into a prompt vector sequence, and then concatenated with the temporal feature vector sequence in the vector space to form a complete input sequence, which is then input into the temporal big language model.
5. The method for predicting total expenditure of the target fund according to claim 1, characterized in that, The adaptive optimization of the temporal large language model based on the meta-learning mechanism using relevant historical scene data includes: Based on the event type, scope of impact, and policy response characteristics of public health emergencies, historical scenario data corresponding to the current scenario are retrieved from historical data to construct a meta-training set; Based on the model-independent meta-learning framework, the temporal large language model is meta-trained using the meta-training set to obtain initial parameters with cross-scene adaptability. Using the limited samples of the public health emergency scenario, the temporal large language model is fine-tuned based on the initial parameters to achieve adaptive optimization for the current scenario.
6. The method for predicting total expenditure of the target fund according to claim 1, characterized in that, The prediction results of the target total fund expenditure under the public health emergency scenario, based on the optimized temporal large language model and the standard feature set, include: Based on the standard feature set, the time-series big language model with adaptive optimization outputs the predicted total expenditure of the target fund at multiple preset time scales in the future. Generate and output interpretable analysis data associated with the prediction results, wherein the interpretable analysis data includes at least a fund expenditure trend chart and a feature importance analysis chart reflecting the degree of influence of different data modalities on the prediction results.
7. A target fund total expenditure forecasting system under a public health emergency, characterized in that, include: The data acquisition and feature construction module is used to collect multi-source heterogeneous data, including at least time-series data of the target fund, textual data of public health events, and textual data of the target policy, and to construct a multimodal feature matrix; The feature fusion module is used to fuse the multimodal feature matrix through a bidirectional cross-attention mechanism to obtain a fused standard feature set; The domain adaptation module is used to input the standard feature set into the pre-trained temporal big language model through an adaptation method guided by domain knowledge. The meta-learning optimization module is used to adaptively optimize the temporal large language model based on the meta-learning mechanism and corresponding historical scene data. The prediction output module is used to output the prediction results of the target total expenditure of the fund under the public health emergency scenario based on the optimized temporal large language model and the standard feature set.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method of any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method of any one of claims 1-6.