Multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion
Patent Information
- Application Number
- CN202610965203.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-29
AI Technical Summary
[0010]本发明所要解决的技术问题在于针对上述现有技术中的不足,提供一种基于闭环知识过滤与跨模态融合的多变量时序预测方法,用于解决现有多变量时序预测方法缺少局部时序片段与文本语义之间的细粒度对应关系,且文本知识筛选主要依赖时间关联、文本相关性或预设过滤规则,无法根据各文本知识对预测误差的实际影响迭代更新过滤逻辑,导致冗余或负向文本知识参与预测的技术问题
一种基于闭环知识过滤与跨模态融合的多变量时序预测方法,通过将原型级跨模态语义融合与闭环文本知识过滤结合,形成从多变量时序数据输入、跨模态时序表征生成、文本知识筛选、过滤逻辑更新到预测输出的完整技术链。先利用时序片段原型与文本嵌入原型之间的一一对应关系,为输入时序片段建立可供预测大语言模型利用的文本语义映射,再通过多头交叉注意力将时序片段嵌入与文本语义映射进行融合,使局部时序变化能够获得相应的文本语义表征;同时,基于候选文本知识消融前后的预测误差差值确定知识边际贡献,并据此迭代更新过滤逻辑,使文本知识筛选不再仅依赖时间关联或文本相关性。由此,可减少冗余或产生负向影响的文本知识进入预测过程,并使筛选后的文本知识与跨模态时序表征共同参与预测,形成与现有静态知识注入方式不同的闭环预测机制。
Smart Images

Figure CN122838985A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multivariate time series prediction technology, specifically involving a multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion. Background Technology
[0002] Multivariate time series forecasting, based on historical observation sequences of multiple related variables, predicts the changes of one or more variables in subsequent periods and has been applied in scenarios such as meteorology, finance, power, and wind power. Existing methods typically use numerical time series data as input and build predictive models by utilizing the temporal correlation and inter-variable correlation of variables.
[0003] In real-world operating environments, multivariate time series data may be affected by external disturbances and sudden events, leading to abnormal fluctuations or distribution shifts. Relevant external information often exists in textual form, and relying solely on numerical time series data is insufficient to reflect the impact of these external factors on time series changes.
[0004] To utilize textual information, existing methods typically map temporal representations to similar representation spaces using linear mapping, contrastive learning, or knowledge distillation. These methods primarily achieve alignment at the embedding or distribution level, but have not yet established a correspondence between local temporal segments and specific textual semantics.
[0005] Furthermore, time-series data often lacks text annotations that correspond one-to-one with local time-series segments. Manual annotation would require configuring text descriptions for a large number of time-series segments, resulting in high implementation costs. Therefore, existing methods struggle to obtain the textual semantics corresponding to local time-series changes in the absence of text annotations.
[0006] Regarding the utilization of textual knowledge, existing methods mostly employ static prompts or timestamp-based retrieval, inputting textual knowledge related to the time period to be predicted into the prediction model; some methods also filter based on preset filtering rules or textual relevance. These filtering methods primarily examine temporal or semantic relevance, without evaluating the actual impact of individual textual knowledge on prediction error.
[0007] Therefore, the screening results may simultaneously include textual knowledge that has a positive effect on prediction, no significant effect, or a negative impact. Existing methods typically do not compare the prediction results before and after the introduction of knowledge items, nor do they quantify the prediction contribution of individual knowledge items, making it difficult to identify and exclude invalid or negative knowledge.
[0008] Meanwhile, existing filtering rules are usually pre-set before prediction and remain unchanged in subsequent processes. Even if the filtered textual knowledge leads to an increase in prediction error, there is a lack of a processing mechanism to correct the filtering rules based on the prediction results.
[0009] Therefore, existing technologies still need to address the following issues: how to establish the correspondence between local temporal segments and text semantics in the absence of temporal text annotations; how to evaluate knowledge items based on the actual impact of textual knowledge on prediction errors, and how to feed the evaluation results back to the filtering logic to reduce the interference of invalid or negative knowledge on the prediction process. Summary of the Invention
[0010] The technical problem to be solved by this invention is to address the shortcomings of the prior art by providing a multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion. This method addresses the lack of fine-grained correspondence between local time series segments and text semantics in existing multivariate time series prediction methods. Furthermore, the text knowledge filtering mainly relies on time correlation, text relevance, or preset filtering rules, and cannot iteratively update the filtering logic according to the actual impact of each text knowledge on the prediction error, resulting in redundant or negative text knowledge participating in the prediction.
[0011] The present invention adopts the following technical solution: A multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion includes the following steps: S1. Obtain multivariate time-series data and construct a multi-source text knowledge base, wherein the knowledge base includes time-series variable definition knowledge; S2. Divide and cluster the multivariate time series data into time series segment prototypes, generate text embedding prototypes for each time series segment prototype, match the input time series segment with the time series segment prototype, generate text semantic mapping based on the matching results, and perform multi-head cross-attention fusion of the input time series segment embedding and the text semantic mapping to obtain cross-modal time series representation. S3. Using the initial filtering logic as the current filtering logic, the knowledge filtering agent filters the corresponding text knowledge in the verification set. Based on the difference in prediction error before and after the text knowledge item ablation, the marginal contribution of knowledge is determined and a marginal contribution label is generated. The knowledge evaluation agent updates the current filtering logic according to the text knowledge item, the marginal contribution label and the current filtering logic. After a preset iteration, the updated filtering logic is obtained. S4. Based on the updated filtering logic, filter the text knowledge corresponding to the input time series samples, form a prefix hint with the time series variable definition knowledge, and embed it into the cross-modal time series representation before inputting it into the prediction large language model to obtain the multivariate time series prediction results.
[0012] Furthermore, in S1: The C-dimensional multivariate time series data is split into C univariate time series along the variable dimension, and the input-output sample pairs for the prediction task are constructed based on each of the univariate time series. The multi-source text knowledge base also includes covariate knowledge with timestamps and news event knowledge with timestamps. The text knowledge entries in the multi-source text knowledge base are stored in JSON format and labeled with their respective fields, geographical locations, knowledge types, and timestamps.
[0013] Furthermore, in S2, the multivariate time series data is divided and clustered into time series segment prototypes, including: The multivariate time-series data is segmented into time-series segments of a preset length; k-means clustering is performed on the time-series segments to obtain M time-series segment prototypes; and text embedding prototypes are generated for each time-series segment prototype, including: A large language model is used to generate text descriptions for each of the time-series segment prototypes; the word segmenter of the predictive large language model is used to encode the text descriptions to obtain text embedding prototypes that correspond one-to-one with each of the time-series segment prototypes.
[0014] Furthermore, in S2, the input timing segment is matched with the timing segment prototype, including: The input time series samples are segmented into multiple non-overlapping input time series segments; Calculate the cosine similarity between each input time segment and its prototype; The prototype of the time segment with the highest cosine similarity is determined as the prototype of the time segment that matches the corresponding input time segment; Based on the time segment prototype matched by each input time segment, the text embedding prototype corresponding to the time segment prototype is called to obtain the text semantic mapping.
[0015] Furthermore, in S2, the input temporal segment embeddings and text semantic mappings are fused using multi-head cross-attention, including: The input temporal segment is projected onto the hidden dimension of the prediction large language model through a linear layer to obtain the input temporal segment embedding; the text semantic map is flattened; and multi-head cross-attention processing is performed using the input temporal segment embedding as the query and the flattened text semantic map as the key and value to obtain the cross-modal temporal representation.
[0016] Furthermore, in S3, the initial filtering logic is constructed based on expert knowledge; the knowledge filtering agent retrieves text knowledge from the multi-source text knowledge base within the last S time steps of the input time series samples located in the validation set, and filters the retrieved text knowledge according to the current filtering logic. The knowledge filtering agent outputs the filtered text knowledge entries, the knowledge types corresponding to the text knowledge entries, and the reasons for the filtering.
[0017] Furthermore, in S3, the knowledge marginal contribution is determined based on the difference in prediction error before and after the text knowledge item ablation, including: Based on the complete candidate text knowledge set, the input time series samples in the validation set are predicted to obtain the mean squared error of the complete knowledge prediction. Set the text embedding position corresponding to each text knowledge entry to zero to obtain the ablation text knowledge set corresponding to each text knowledge entry; Predictions were made based on each ablation text knowledge set, and the corresponding mean squared error of ablation knowledge prediction was obtained. The knowledge margin contribution of the corresponding text knowledge item is determined based on the difference between the mean squared error of the complete knowledge prediction and the mean squared error of the ablation knowledge prediction.
[0018] Furthermore, in S3, each knowledge marginal contribution is discretized into a corresponding marginal contribution label. The marginal contribution label includes a strong negative label, a weak negative label, a neutral label, and a positive label. The knowledge evaluation agent performs multi-step reasoning based on the text knowledge item, the marginal contribution label corresponding to the text knowledge item, the current filtering logic, and the filtering reason to determine the filtering constraints that need to be tightened or added, and updates the current filtering logic according to the filtering constraints.
[0019] Furthermore, in S4, the knowledge filtering agent filters text knowledge corresponding to the input time series sample from the multi-source text knowledge base according to the updated filtering logic to obtain the filtered text knowledge set; the filtered text knowledge set and the time series variable definition knowledge corresponding to the input time series sample are organized into a prefix hint. The prefix hint embedding and the cross-modal temporal representation are concatenated in the embedding layer of the large language prediction model to obtain the predicted input; The predicted input is fed into the prediction large language model, and the output of the prediction large language model is mapped through the output linear layer to obtain the multivariate time series prediction result. The backbone parameters of the prediction large language model are efficiently fine-tuned using a low-rank adaptation method, and the model parameters of the prediction large language model are optimized using L2 loss.
[0020] Furthermore, the multivariate time-series data is wind turbine operation monitoring data, which includes at least one operating parameter among active power output and rotor speed. The multi-source text knowledge base includes wind power time-series variable definition knowledge, wind turbine external covariate data, and wind power news event knowledge. The wind turbine external covariate data includes at least one of wind speed, wind direction, ambient temperature, and tower base humidity.
[0021] A second aspect of the present invention is a multivariate time series prediction system based on closed-loop knowledge filtering and cross-modal fusion, used to perform the method, comprising: The data and knowledge construction module is used to acquire multivariate time-series data and construct a multi-source text knowledge base, which includes time-series variable definition knowledge. The text semantic fusion module is used to divide and cluster multivariate time series data into time series segment prototypes, generate text embedding prototypes for each time series segment prototype, match the input time series segment with the time series segment prototype, generate text semantic mapping based on the matching results, and perform multi-head cross-attention fusion of the input time series segment embedding and the text semantic mapping to obtain cross-modal time series representation. The closed-loop text knowledge filtering module is used to use the initial filtering logic as the current filtering logic, filter the corresponding text knowledge in the verification set through the knowledge filtering agent, determine the marginal contribution of knowledge based on the difference in prediction error before and after the text knowledge item ablation and generate a marginal contribution label, and update the current filtering logic through the knowledge evaluation agent according to the text knowledge item, the marginal contribution label and the current filtering logic. After a preset iteration, the updated filtering logic is obtained. The text knowledge enhancement prediction module is used to filter the text knowledge corresponding to the input time series samples based on the updated filtering logic. The filtered text knowledge and the time series variable definition knowledge are combined to form a prefix hint, which is then embedded and concatenated with the cross-modal time series representation before being input into the prediction large language model to obtain multivariate time series prediction results.
[0022] Furthermore, the data and knowledge construction module is used to split the C-dimensional multivariate time series data into C univariate time series along the variable dimension, and construct input-output sample pairs for the prediction task based on each of the univariate time series; The data and knowledge construction module is also used to construct a multi-source text knowledge base that includes knowledge of time-series variable definitions, knowledge of covariates with timestamps, and knowledge of news events with timestamps.
[0023] Furthermore, the text semantic fusion module is used to perform k-means clustering on the time-series segments to obtain M time-series segment prototypes, generate text descriptions for each time-series segment prototype using a large language model, and encode the text descriptions to obtain text embedding prototypes.
[0024] Furthermore, the closed-loop text knowledge filtering module is used to set the text embedding position corresponding to the text knowledge item to zero, forming an ablation text knowledge set, and to determine the knowledge marginal contribution based on the difference between the mean square error of the complete knowledge prediction and the mean square error of the ablation knowledge prediction.
[0025] Furthermore, the text knowledge enhancement prediction module is used to organize the filtered text knowledge set and the temporal variable definition knowledge corresponding to the input temporal sample into prefix hints, and to concatenate the embedding of the prefix hints with the cross-modal temporal representation in the embedding layer of the prediction large language model.
[0026] Thirdly, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the aforementioned multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion.
[0027] Fourthly, embodiments of the present invention provide an electronic device, including a computer program, which, when executed by the electronic device, implements the steps of the above-described multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion.
[0028] Compared with the prior art, the present invention has at least the following beneficial effects: A multivariate temporal prediction method based on closed-loop knowledge filtering and cross-modal fusion is proposed. This method combines prototype-level cross-modal semantic fusion with closed-loop textual knowledge filtering to form a complete technology chain from multivariate temporal data input, cross-modal temporal representation generation, textual knowledge filtering, filtering logic update to prediction output. First, a one-to-one correspondence between temporal fragment prototypes and text embedding prototypes is used to establish a textual semantic mapping for the input temporal fragments that can be used by a large-scale predictive language model. Then, multi-head cross-attention is used to fuse the temporal fragment embeddings and textual semantic mapping, enabling local temporal changes to obtain corresponding textual semantic representations. Simultaneously, the marginal contribution of knowledge is determined based on the difference in prediction error before and after candidate textual knowledge ablation, and the filtering logic is iteratively updated accordingly, so that textual knowledge filtering no longer relies solely on temporal correlation or textual relevance. This reduces redundant or negatively impactful textual knowledge from entering the prediction process and allows the filtered textual knowledge and cross-modal temporal representations to participate in the prediction together, forming a closed-loop prediction mechanism different from existing static knowledge injection methods.
[0029] Furthermore, the C-dimensional multivariate time series data is split into C univariate time series along the variable dimension, which facilitates the construction of input-output sample pairs in a unified manner, and performs time series segmentation and prototype matching for local changes of different variables. The multi-source text knowledge base further incorporates time-stamped covariate knowledge and news event knowledge, enabling the prediction model to combine external information related to the period to be predicted.
[0030] Furthermore, textual knowledge entries are stored in JSON format and labeled with their domain, geographical location, knowledge type, and timestamp, enabling structured management of knowledge from different sources and of different types. The knowledge filtering agent can perform retrieval and filtering based on these attributes, and the knowledge evaluation agent can also form or adjust filtering constraints accordingly, providing a clear data basis for the knowledge filtering process.
[0031] Furthermore, k-means clustering is performed on time series segments of a preset length to extract M representative time series segment prototypes from a large number of local time series segments. After generating text descriptions for each prototype and encoding them as text embedding prototypes, it is not necessary to annotate each input time series segment individually; the corresponding text semantics can be obtained through prototype matching, reducing the need for segment-by-segment text annotation.
[0032] Furthermore, the input temporal samples are divided into non-overlapping temporal segments, and cosine similarity is used to determine the closest temporal segment prototype. A correspondence between the input segments and the prototypes can be established according to a unified similarity metric. With the help of the one-to-one correspondence between the temporal segment prototypes and the text embedding prototypes, the text semantic mapping of the input segments can be further obtained.
[0033] Furthermore, the linear layer projects the input temporal segments onto the hidden dimensions of the large language prediction model, enabling the temporal segment embeddings to participate in subsequent attention calculations. By performing multi-head cross-attention with the temporal segment embeddings as queries and the flattened text semantic mappings as keys and values, numerical temporal information and text semantic information can be fused at the segment level to obtain cross-modal temporal representations.
[0034] Furthermore, the initial filtering logic, constructed based on expert knowledge, provides explicit rules for the initial knowledge screening. The knowledge filtering agent retrieves textual knowledge within the last S time steps of the input time-series sample and outputs knowledge entries, knowledge types, and screening reasons, preserving necessary contextual information for subsequent evaluation of knowledge marginal contributions and updates to the filtering logic.
[0035] Furthermore, based on the complete candidate text knowledge set, the text embedding position of each individual knowledge entry is set to zero, which can form a corresponding ablation text knowledge set while keeping other knowledge entries unchanged. Comparing the mean squared error of the complete knowledge prediction with the mean squared error of the ablation knowledge prediction can quantify the impact of a single knowledge entry on the prediction error, providing a numerical basis for updating the filtering logic.
[0036] Furthermore, the filtered text knowledge set and the corresponding temporal variables of the input temporal samples are defined and organized as prefix hints. The prefix hints are then embedded and concatenated with the cross-modal temporal representations at the embedding layer. This allows the large-scale predictive language model to simultaneously receive external text knowledge, variable semantics, and numerical temporal representations. After mapping through the output linear layer, the model output can be converted into multivariate temporal prediction results.
[0037] Furthermore, the filtered text knowledge set and the corresponding temporal variables of the input temporal samples are defined and organized as prefix hints. The prefix hints are then embedded and concatenated with the cross-modal temporal representations at the embedding layer. This allows the large-scale predictive language model to simultaneously receive external text knowledge, variable semantics, and numerical temporal representations. After mapping through the output linear layer, the model output can be converted into multivariate temporal prediction results.
[0038] Furthermore, LoRA is employed to efficiently fine-tune the parameters of the backbone of the large-scale language prediction model, enabling adaptation to multivariate time-series prediction tasks without updating all model parameters. L2 loss uses the error between the predicted and actual results as the optimization basis to update model parameters, allowing the large-scale language prediction model to learn the mapping relationship between the predicted input and the multivariate time-series results.
[0039] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0040] In summary, this invention establishes a correspondence between temporal segments and textual semantics through prototype matching, and uses the marginal contribution of knowledge as the feedback basis for filtering logic. The filtered textual knowledge and cross-modal temporal representations are used together for prediction, thereby reducing the interference of invalid or negative knowledge on multivariate temporal prediction.
[0041] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0042] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a text knowledge illustration used for time series sample prediction of the machine-side current (Ic) of the wind power converter 1 in this invention; Figure 3 The time-series sample prediction curve of the current (Ic) on the machine side of wind power converter 1 is shown. Figure 4 A schematic diagram of a computer device provided in an embodiment of the present invention; Figure 5 This is a block diagram of a chip provided according to an embodiment of the present invention.
[0043] Among them, 60. Computer equipment; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access memory unit; 6202. Cache memory unit; 6203. Read-only memory unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0046] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0047] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.
[0048] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0049] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0050] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.
[0051] This invention provides a multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion. First, a prototype-level cross-modal semantic association is established through a text semantic fusion module, generating a cross-modal time series representation that can be understood by a large language model. Second, a multi-source text knowledge base is constructed, relying on an agent-driven closed-loop text knowledge filtering framework. The filtering logic is iteratively optimized using knowledge margin contribution as feedback, dynamically selecting non-redundant text knowledge with practical value for prediction. Finally, the filtered text knowledge is concatenated with the time series representation to form a text knowledge-enhanced prediction input. The large language model, after efficient parameter fine-tuning, completes the multivariate time series prediction output. This invention effectively bridges the gap between time series and text modalities, improves the efficiency of text knowledge utilization, enables the large language model to establish a correspondence between time series segments and text semantics in complex time series fluctuation scenarios, and updates the filtering logic based on knowledge margin contribution feedback, allowing the filtered text knowledge and cross-modal time series representation to be used together for multivariate time series prediction.
[0052] Please see Figure 1 This invention discloses a multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion, comprising the following steps: S1. Preprocess the multivariate time series data, converting the C-dimensional multivariate time series data into C-dimensional multivariate time series data. The system is split into C univariate sequences along the variable dimension to construct input-output sample pairs for the prediction task. At the same time, a multi-source textual knowledge base (MS-TKB) is constructed. The MS-TKB contains multivariate time series variable definition knowledge, time-stamped covariate knowledge, and time-stamped news event knowledge compiled by experts. All knowledge entries are stored in JSON format and labeled with their domain, geographical location, knowledge type, and timestamp attribute. S2. Generate cross-modal temporal representations that can be understood by large language models through the Prototype-based Textual Semantic Fusion Module (ProtoTSFM); S201. Divide the multivariate time series data into segments of length [length missing]. The patch is obtained by k-means clustering, resulting in M discriminative temporal patch prototypes. The large language model is used to generate text descriptions for each temporal patch prototype, which are then encoded by the word segmenter of the large language model used for prediction to obtain a set of text embedding prototypes. ,in It is the length of the text embedding sequence after uniform padding. It predicts the hidden dimensions of large language models, thereby establishing a one-to-one semantic correspondence between temporal sequences and texts at the prototype level. ; S202. Segment the input time series samples to be predicted into non-overlapping patches. Number of patches For each patch, the closest temporal patch prototype is matched using cosine similarity. ,in This represents the cosine similarity. Because this patch-level prototype matching method accurately approximates the input temporal sequence at a local scale, it matches the prototype... Related text embedding prototype Can be used as A reliable textual semantic representation is obtained. Corresponding text semantic mapping ; S203. Project the input temporal patches through a linear layer to the hidden dimension of the large language prediction model to obtain the patch embedding. ,by For querying and flattening Using keys and values, multi-head cross-attention (MHCA) is used to achieve cross-modal semantic fusion, generating temporal representations that can be understood by large language models. ,Right now ,in It is Flatten to match attention input.
[0053] S3. Useful text knowledge is filtered and predicted using an agent-driven closed-loop textual knowledge filtering framework (CKFF). S301. Based on expert knowledge, an initial filtering logic is constructed. The knowledge filtering agent retrieves text knowledge from the MS-TKB within the last S time steps of the input time series sample to be predicted, and obtains the text knowledge set according to the initial filtering logic. ,in, Indicates the first The filtered text knowledge, It is a type of knowledge. These are the corresponding selection reasons; S302, Text knowledge set Used for prediction, to obtain prediction results; only for the time series samples of the validation set, for each time series sample in the validation set... Calculate each knowledge entry The marginal contribution to prediction, to quantify its actual effect on prediction; within a complete text knowledge set. Lieutenant General The corresponding text embedding positions are set to zero, forming the ablation knowledge set. Will use and The difference in the mean square error of the predictions is defined as knowledge. Marginal contribution: ; To reduce the interference of minute numerical fluctuations and facilitate stable reasoning by the knowledge evaluation agent, continuous marginal contributions are discretized into four types of text labels: strongly negative, weakly negative, neutral, and positive.
[0054] in, For knowledge entries Marginal contribution label A strong negative label. For knowledge entries The corresponding marginal contribution of knowledge, It is a weak negative label. This is a neutral label. This is a positive label.
[0055] S303: The knowledge evaluation agent combines knowledge items, marginal contribution labels, and the current filtering logic to obtain, through multi-step reasoning, the filtering constraints that need to be tightened or added to exclude invalid knowledge. Then, it updates the filtering logic and returns to S301, until the preset number of iterations is completed. After completing the preset number of iterations, the filtering logic obtained from the last update is used as the updated filtering logic, and the knowledge filtering agent filters the text knowledge corresponding to the input time series sample to be predicted according to the updated filtering logic to obtain the filtered text knowledge set.
[0056] S4. Text knowledge-enhanced prediction based on large language models; S401, Filter the obtained text knowledge set The input temporal variables are defined and organized as prefix prompts, and the embeddings of these prefix prompts are integrated with the temporal representation that the large language model can understand. By concatenating the data at the embedding layer, a text-based knowledge-enhanced input is formed; S402. LoRA is used to efficiently fine-tune the parameters of the large language prediction model backbone. The text knowledge-enhanced input is fed into the large language prediction model backbone, and the final prediction result is obtained through the output linear layer mapping. The model parameters are optimized with L2 loss as the objective.
[0057] In another embodiment of the present invention, a multivariate time series prediction system based on closed-loop knowledge filtering and cross-modal fusion is provided. This system can be used to implement the above-mentioned multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion. Specifically, the multivariate time series prediction system based on closed-loop knowledge filtering and cross-modal fusion includes a data and knowledge construction module, a text semantic fusion module, a closed-loop text knowledge filtering module, and a text knowledge enhancement prediction module.
[0058] The data and knowledge construction module is used to acquire multivariate time-series data and construct a multi-source text knowledge base, which includes time-series variable definition knowledge. The text semantic fusion module is used to divide and cluster multivariate time series data into time series segment prototypes, generate text embedding prototypes for each time series segment prototype, match the input time series segment with the time series segment prototype, generate text semantic mapping based on the matching results, and perform multi-head cross-attention fusion of the input time series segment embedding and the text semantic mapping to obtain cross-modal time series representation. The closed-loop text knowledge filtering module is used to use the initial filtering logic as the current filtering logic, filter the corresponding text knowledge in the verification set through the knowledge filtering agent, determine the marginal contribution of knowledge based on the difference in prediction error before and after the text knowledge item ablation and generate a marginal contribution label, and update the current filtering logic through the knowledge evaluation agent according to the text knowledge item, the marginal contribution label and the current filtering logic. After a preset iteration, the updated filtering logic is obtained. The text knowledge enhancement prediction module is used to filter the text knowledge corresponding to the input time series samples based on the updated filtering logic. The filtered text knowledge and the time series variable definition knowledge are combined to form a prefix hint, which is then embedded and concatenated with the cross-modal time series representation before being input into the prediction large language model to obtain multivariate time series prediction results.
[0059] The data and knowledge construction module is used to split C-dimensional multivariate time series data into C univariate time series along the variable dimension, and construct input-output sample pairs for the prediction task based on each univariate time series; The data and knowledge construction module is also used to construct a multi-source text knowledge base that includes knowledge of time-series variable definitions, knowledge of covariates with timestamps, and knowledge of news events with timestamps.
[0060] The text semantic fusion module is used to perform k-means clustering on time-series segments to obtain M time-series segment prototypes. A large language model is used to generate text descriptions for each time-series segment prototype, and the text descriptions are encoded to obtain text embedding prototypes.
[0061] The closed-loop text knowledge filtering module is used to set the text embedding position corresponding to the text knowledge entry to zero, forming an ablation text knowledge set, and to determine the knowledge marginal contribution based on the difference between the mean square error of the complete knowledge prediction and the mean square error of the ablation knowledge prediction.
[0062] The text knowledge enhancement prediction module is used to organize the filtered text knowledge set and the time-series variable definition knowledge corresponding to the input time-series samples into prefix hints, and to concatenate the embedding of the prefix hints with the cross-modal time-series representation in the embedding layer of the prediction large language model.
[0063] This invention provides a terminal device comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve corresponding method flows or corresponding functions. The processor described in this embodiment can be used for the operation of a multivariate time-series prediction method based on closed-loop knowledge filtering and cross-modal fusion, including: A multivariate time-series data set is acquired and a multi-source text knowledge base is constructed, including time-series variable definition knowledge. The multivariate time-series data is divided and clustered into time-series segment prototypes, and text embedding prototypes are generated for each time-series segment prototype. Input time-series segments are matched with time-series segment prototypes, and text semantic mappings are generated based on the matching results. The input time-series segment embeddings and text semantic mappings are fused using multi-head cross-attention to obtain cross-modal time-series representations. Using the initial filtering logic as the current filtering logic, a knowledge filtering agent filters the corresponding text knowledge in the validation set. The marginal contribution of knowledge is determined based on the difference in prediction error before and after the text knowledge entries are ablated, and a marginal contribution label is generated. The knowledge evaluation agent updates the current filtering logic based on the text knowledge entries, marginal contribution labels, and the current filtering logic. After a preset iteration, the updated filtering logic is obtained. Based on the updated filtering logic, the text knowledge corresponding to the input time-series samples is filtered. The filtered text knowledge and the time-series variable definition knowledge are combined to form a prefix hint, which is then embedded and concatenated with the cross-modal time-series representation and input into the prediction large language model to obtain multivariate time-series prediction results.
[0064] Please see Figure 4 The terminal device is a computer device. In this embodiment, the computer device 60 includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When executed by the processor 61, the computer program 63 implements the multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion in this embodiment. To avoid repetition, these details are not elaborated here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the multivariate time series prediction system based on closed-loop knowledge filtering and cross-modal fusion in this embodiment. To avoid repetition, these details are not elaborated here.
[0065] Computer device 60 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will understand that... Figure 4 This is merely an example of computer device 60 and does not constitute a limitation on computer device 60. It may include more or fewer components than shown, or combine certain components, or different components. For example, computer device may also include input / output devices, network access devices, buses, etc.
[0066] The processor 61 may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0067] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computer device 60.
[0068] Furthermore, the memory 62 may include both internal storage units of the computer device 60 and external storage devices. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.
[0069] Please see Figure 5 The terminal device is an electronic device 600, which is manifested in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.
[0070] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the method section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.
[0071] Storage unit 620 may include readable media in the form of volatile storage units, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include read-only memory (ROM) 6203.
[0072] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0073] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.
[0074] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem). This communication can be performed via input / output interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network, wide area network, and / or public network, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.
[0075] This invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the terminal device and extended storage media supported by the terminal device; it can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). More specific examples of the computer-readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical fiber, portable compact disk read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.
[0076] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency, etc., or any suitable combination thereof.
[0077] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0078] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the multivariate time-series prediction method based on closed-loop knowledge filtering and cross-modal fusion in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps: A multivariate time-series data set is acquired and a multi-source text knowledge base is constructed, including time-series variable definition knowledge. The multivariate time-series data is divided and clustered into time-series segment prototypes, and text embedding prototypes are generated for each time-series segment prototype. Input time-series segments are matched with time-series segment prototypes, and text semantic mappings are generated based on the matching results. The input time-series segment embeddings and text semantic mappings are fused using multi-head cross-attention to obtain cross-modal time-series representations. Using the initial filtering logic as the current filtering logic, a knowledge filtering agent filters the corresponding text knowledge in the validation set. The marginal contribution of knowledge is determined based on the difference in prediction error before and after the text knowledge entries are ablated, and a marginal contribution label is generated. The knowledge evaluation agent updates the current filtering logic based on the text knowledge entries, marginal contribution labels, and the current filtering logic. After a preset iteration, the updated filtering logic is obtained. Based on the updated filtering logic, the text knowledge corresponding to the input time-series samples is filtered. The filtered text knowledge and the time-series variable definition knowledge are combined to form a prefix hint, which is then embedded and concatenated with the cross-modal time-series representation and input into the prediction large language model to obtain multivariate time-series prediction results.
[0079] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0080] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0081] To verify the effectiveness of this invention, actual operation monitoring data of a 3.5MW wind turbine generator set from a domestic wind farm were selected for model verification. This data was automatically collected by the wind turbine condition monitoring system at 10-minute intervals, covering a total of 42 key operating parameters, including active power output and rotor speed, which can truly reflect the operating characteristics of the wind turbine under all operating conditions.
[0082] Step 1: Wind power data preprocessing and construction of a dedicated multi-source text knowledge base Normalization preprocessing is performed on the multivariate time series data of wind power, and the 42-dimensional time series data is split into 42 single-variable time series along the variable dimension. A parameter sharing mechanism between variables is adopted to complete the unified modeling. A multi-source text knowledge base specifically for wind power scenarios is constructed, comprising three categories of wind power-related knowledge: first, expert-compiled definitions of wind power time-series variables, including variable definitions and their relationships with covariates; second, timestamped external covariate data for wind turbines, including wind speed, wind direction, ambient temperature, and tower base humidity; and third, timestamped news events in the wind power field. Each knowledge entry is stored in JSON format, with timestamped knowledge entries labeled with a "time" field for time-aligned retrieval with the time-series data.
[0083] Step 2: The text semantic fusion module generates a wind power time series representation. Wind power time-series data is processed using a text semantic fusion module, and the patch length is set. The number of prototypes is M=100; 100 discriminative wind power time-series patch prototypes are obtained through k-means clustering, and DeepSeek is used for further analysis. The V3.2 large language model generates text descriptions for each temporal patch prototype and encodes them into text embedding prototypes, establishing a one-to-one correspondence between temporal patch prototypes and text embedding prototypes.
[0084] Input wind power time sequence Perform non-overlapping patch segmentation to obtain The patch is matched with the optimal prototype using cosine similarity to obtain the corresponding text embedding prototype sequence. ;Will The patch embeddings are obtained by projecting them onto the hidden dimension of the large language model through a linear layer. ,by For querying and flattening Using a key-value pair, a multi-head cross-attention mechanism is employed to achieve cross-modal semantic fusion, generating a wind power time-series representation adapted to the understanding of large language models. .
[0085] Step 3: Use a closed-loop text knowledge filtering framework to screen wind power forecasting knowledge. A closed-loop text knowledge filtering framework driven by intelligent agents is used to process the wind power-specific knowledge base, where both the knowledge filtering agent and the knowledge evaluation agent adopt DeepSeek. V3.2 large language model implementation, setting knowledge retrieval time step S=3 and closed-loop iteration count. The initial filtering logic is constructed based on the knowledge of wind turbine time series variable definitions compiled by experts. The knowledge filtering agent performs preliminary screening of wind power text knowledge in the input time series for the next three time steps.
[0086] The marginal contribution of knowledge is calculated based on time-series samples of the validation set. The marginal contribution is discretized into four types of labels. The knowledge evaluation agent combines the contribution labels to iteratively optimize the filtering logic. The updated filtering logic is used by the knowledge filtering agent to filter out a text knowledge set that is more suitable for wind power forecasting.
[0087] Step 4: Text-based knowledge-enhanced wind power time-series forecasting The filtered wind power text knowledge and the input wind power time series variables are organized into prefix hints. The prefix hints are embedded and concatenated with the wind power time series representation in the embedding layer to form the prediction input. LoRA is used to efficiently fine-tune the parameters of the LLaMA2-7B large language model, and the wind power multivariate time series prediction results are obtained by mapping through the output linear layer.
[0088] To illustrate the application of this invention in wind power scenarios, Figure 2 The textual knowledge used in the time-series sample prediction of the machine-side current (Ic) of wind power converter 1 is shown. Figure 3 The corresponding predicted curves and measured curves are shown. Figure 2 The textual knowledge is obtained by filtering through the updated filtering logic. Figure 3 This reflects the relationship between the predicted results and the measured results obtained based on the textual knowledge and cross-modal temporal representation.
[0089] This embodiment establishes a correspondence between local temporal changes and text semantics by using temporal segment prototypes and text embedding prototypes, and generates cross-modal temporal representations through prototype matching and multi-head cross-attention. At the same time, the filtering logic is iteratively updated based on the marginal contribution of knowledge items, so that the selected text knowledge and cross-modal temporal representations participate in prediction together.
[0090] In summary, this invention presents a multivariate temporal prediction method based on closed-loop knowledge filtering and cross-modal fusion. It fragments multivariate temporal data, obtains temporal fragment prototypes based on clustering, and generates corresponding text embedding prototypes for each temporal fragment prototype, establishing a one-to-one correspondence between local temporal changes and text semantics. Furthermore, through prototype matching and multi-head cross-attention processing, it forms a cross-modal temporal representation that combines numerical temporal features with text semantic information. Regarding text knowledge utilization, this invention uses a knowledge filtering agent to screen candidate text knowledge and determines the marginal contribution of knowledge by using the difference in prediction error before and after knowledge item ablation. Then, a knowledge evaluation agent iteratively updates the filtering logic based on the marginal contribution label, allowing text knowledge filtering to be adjusted according to its actual impact on the prediction results. Thus, this invention reduces the interference of redundant and negative knowledge on the prediction process and inputs the filtered text knowledge and cross-modal temporal representations into a large-scale predictive language model, achieving joint modeling and prediction of multivariate temporal data.
[0091] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.
Claims
1. A multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion, characterized in that, Includes the following steps: S1. Obtain multivariate time-series data and construct a multi-source text knowledge base, wherein the knowledge base includes time-series variable definition knowledge; S2. Divide and cluster the multivariate time series data into time series segment prototypes, generate text embedding prototypes for each time series segment prototype, match the input time series segment with the time series segment prototype, generate text semantic mapping based on the matching results, and perform multi-head cross-attention fusion of the input time series segment embedding and the text semantic mapping to obtain cross-modal time series representation. S3. Using the initial filtering logic as the current filtering logic, the knowledge filtering agent filters the corresponding text knowledge in the verification set. Based on the difference in prediction error before and after the text knowledge item ablation, the marginal contribution of knowledge is determined and a marginal contribution label is generated. The knowledge evaluation agent updates the current filtering logic according to the text knowledge item, the marginal contribution label and the current filtering logic. After a preset iteration, the updated filtering logic is obtained. S4. Based on the updated filtering logic, filter the text knowledge corresponding to the input time series samples, form a prefix hint with the time series variable definition knowledge, and embed it into the cross-modal time series representation before inputting it into the prediction large language model to obtain the multivariate time series prediction results.
2. The multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion according to claim 1, characterized in that, In S1: The C-dimensional multivariate time series data is split into C univariate time series along the variable dimension, and the input-output sample pairs for the prediction task are constructed based on each of the univariate time series. The multi-source text knowledge base also includes covariate knowledge with timestamps and news event knowledge with timestamps. The text knowledge entries in the multi-source text knowledge base are stored in JSON format and labeled with their respective fields, geographical locations, knowledge types, and timestamps.
3. The multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion according to claim 1, characterized in that, In S2, multivariate time series data are divided and clustered into time series segment prototypes, including: The multivariate time-series data is segmented into time-series segments of a preset length; k-means clustering is performed on the time-series segments to obtain M time-series segment prototypes; and text embedding prototypes are generated for each time-series segment prototype, including: A large language model is used to generate text descriptions for each of the time-series segment prototypes; the word segmenter of the predictive large language model is used to encode the text descriptions to obtain text embedding prototypes that correspond one-to-one with each of the time-series segment prototypes.
4. The multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion according to claim 1, characterized in that, In S2, the input timing segment is matched with the timing segment prototype, including: The input time series samples are segmented into multiple non-overlapping input time series segments; Calculate the cosine similarity between each input time segment and its prototype; The prototype of the time segment with the highest cosine similarity is determined as the prototype of the time segment that matches the corresponding input time segment; Based on the time segment prototype matched by each input time segment, the text embedding prototype corresponding to the time segment prototype is called to obtain the text semantic mapping.
5. The multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion according to claim 1, characterized in that, In S2, the input temporal segment embeddings and text semantic mappings are fused using multi-head cross-attention, including: The input temporal segment is projected onto the hidden dimension of the prediction large language model through a linear layer to obtain the input temporal segment embedding; the text semantic map is flattened; and multi-head cross-attention processing is performed using the input temporal segment embedding as the query and the flattened text semantic map as the key and value to obtain the cross-modal temporal representation.
6. The multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion according to claim 1, characterized in that, In S3, the initial filtering logic is constructed based on expert knowledge; the knowledge filtering agent retrieves text knowledge within the last S time steps of the input time series samples located in the validation set from the multi-source text knowledge base, and filters the retrieved text knowledge according to the current filtering logic. The knowledge filtering agent outputs the filtered text knowledge items, the knowledge type corresponding to the text knowledge items, and the reasons for the filtering.
7. The multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion according to claim 1, characterized in that, In S3, the knowledge marginal contribution is determined based on the difference in prediction error before and after the text knowledge item ablation, including: Based on the complete candidate text knowledge set, the input time series samples in the validation set are predicted to obtain the mean squared error of the complete knowledge prediction. Set the text embedding position corresponding to each text knowledge entry to zero to obtain the ablation text knowledge set corresponding to each text knowledge entry; Predictions were made based on each ablation text knowledge set, and the corresponding mean squared error of ablation knowledge prediction was obtained. The knowledge margin contribution of the corresponding text knowledge item is determined based on the difference between the mean squared error of the complete knowledge prediction and the mean squared error of the ablation knowledge prediction.
8. The multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion according to claim 1, characterized in that, In S3, each knowledge marginal contribution is discretized into a corresponding marginal contribution label. The marginal contribution label includes a strong negative label, a weak negative label, a neutral label, and a positive label. The knowledge evaluation agent performs multi-step reasoning based on the text knowledge item, the marginal contribution label corresponding to the text knowledge item, the current filtering logic, and the filtering reason to determine the filtering constraints that need to be tightened or added, and updates the current filtering logic according to the filtering constraints.
9. The multivariate time series prediction method based on closed-loop knowledge filtering and cross-modal fusion according to claim 1, characterized in that, In S4, the knowledge filtering agent filters text knowledge corresponding to the input time series sample from the multi-source text knowledge base according to the updated filtering logic, and obtains the filtered text knowledge set; the filtered text knowledge set and the time series variable definition corresponding to the input time series sample are organized into a prefix hint. The prefix hint embedding and the cross-modal temporal representation are concatenated in the embedding layer of the large language prediction model to obtain the predicted input; The predicted input is fed into the prediction large language model, and the output of the prediction large language model is mapped through the output linear layer to obtain the multivariate time series prediction result. The backbone parameters of the prediction large language model are efficiently fine-tuned using a low-rank adaptation method, and the model parameters of the prediction large language model are optimized using L2 loss.
10. The method for multivariate time series prediction using a large language model according to claim 1, characterized in that, The multivariate time-series data is wind turbine operation monitoring data, which includes at least one operating parameter among active power output and rotor speed. The multi-source text knowledge base includes wind power time-series variable definition knowledge, wind turbine external covariate data, and wind power news event knowledge. The wind turbine external covariate data includes at least one of wind speed, wind direction, ambient temperature, and tower base humidity.
11. A multivariate time series prediction system based on closed-loop knowledge filtering and cross-modal fusion, characterized in that, For performing the method according to any one of claims 1 to 10, comprising: The data and knowledge construction module is used to acquire multivariate time-series data and construct a multi-source text knowledge base, which includes time-series variable definition knowledge. The text semantic fusion module is used to divide and cluster multivariate time series data into time series segment prototypes, generate text embedding prototypes for each time series segment prototype, match the input time series segment with the time series segment prototype, generate text semantic mapping based on the matching results, and perform multi-head cross-attention fusion of the input time series segment embedding and the text semantic mapping to obtain cross-modal time series representation. The closed-loop text knowledge filtering module is used to use the initial filtering logic as the current filtering logic, filter the corresponding text knowledge in the verification set through the knowledge filtering agent, determine the marginal contribution of knowledge based on the difference in prediction error before and after the text knowledge item ablation and generate a marginal contribution label, and update the current filtering logic through the knowledge evaluation agent according to the text knowledge item, the marginal contribution label and the current filtering logic. After a preset iteration, the updated filtering logic is obtained. The text knowledge enhancement prediction module is used to filter the text knowledge corresponding to the input time series samples based on the updated filtering logic. The filtered text knowledge and the time series variable definition knowledge are combined to form a prefix hint, which is then embedded and concatenated with the cross-modal time series representation before being input into the prediction large language model to obtain multivariate time series prediction results.
12. The large language model multivariate time series prediction system according to claim 11, characterized in that, The data and knowledge construction module is used to split C-dimensional multivariate time series data into C univariate time series along the variable dimension, and construct input-output sample pairs for the prediction task based on each univariate time series; The data and knowledge construction module is also used to construct a multi-source text knowledge base that includes knowledge of time-series variable definitions, knowledge of covariates with timestamps, and knowledge of news events with timestamps.
13. The large language model multivariate time series prediction system according to claim 11, characterized in that, The text semantic fusion module is used to perform k-means clustering on time-series segments to obtain M time-series segment prototypes. A large language model is used to generate text descriptions for each time-series segment prototype, and the text descriptions are encoded to obtain text embedding prototypes.
14. The large language model multivariate time series prediction system according to claim 11, characterized in that, The closed-loop text knowledge filtering module is used to set the text embedding position corresponding to the text knowledge entry to zero, forming an ablation text knowledge set, and to determine the knowledge marginal contribution based on the difference between the mean square error of the complete knowledge prediction and the mean square error of the ablation knowledge prediction.
15. The large language model multivariate time series prediction system according to claim 11, characterized in that, The text knowledge enhancement prediction module is used to organize the filtered text knowledge set and the time-series variable definition knowledge corresponding to the input time-series samples into prefix hints, and to concatenate the embedding of the prefix hints with the cross-modal time-series representation in the embedding layer of the prediction large language model.