Sensor time series data processing method and system based on large language model
Patent Information
- Application Number
- CN202610938368.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-25
AI Technical Summary
现有技术缺乏一种能够统一处理各类异构传感器信号且保持其物理语义完整性的分词框架
[0065]1.高泛化性与通用性:本发明提供了一种专为通用传感器时间序列信号设计的分词策略与系统架构,成功将各种异构的物理量时间序列(包括但不限于压力、振动、温度、流量、加速度等)转化为大语言模型可以理解的嵌入表示,不局限于特定的传感器类型或特定的下游任务。与现有技术主要面向通用时间序列预测任务(如电力负荷、天气温度等)不同,本发明专门针对传感器数据的异构性、物理量纲多样性和多源融合场景进行了系统性设计,具有更强的跨传感器类型泛化能力。
Smart Images

Figure CN122819221A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to the processing and analysis technology of sensor time series data. Specifically, this invention discloses a sensor time series data processing method and system based on a large language model. It is a method and system specifically designed to perform tokenization processing on multidimensional time series data generated by various sensors (including but not limited to physical sensors, chemical sensors, mechanical sensors, industrial sensors, and environmental sensors) to adapt to downstream inference tasks using a large language model (LLM). Background Technology
[0002] Large language models (such as the GPT series, LLaMA series, and DeepSeek series) are pre-trained on massive text corpora based on the Transformer architecture, demonstrating powerful sequence modeling capabilities, contextual understanding capabilities, and few-shot / zero-shot reasoning capabilities. In recent years, the industry has begun to explore extending the application boundaries of LLM from natural language processing to the field of time series analysis, attempting to leverage the general reasoning capabilities of LLM to process time-series data generated by various sensor systems (such as wind turbine vibration sensors, industrial oil pipeline pressure sensors, and environmental monitoring sensors) to achieve downstream tasks such as equipment status monitoring, anomaly detection, and predictive maintenance. However, the essence of LLM is modeling discrete text tokens, while sensor time-series data are continuous numerical signals; a natural "modal gap" exists between the two in terms of data form and semantic representation. Therefore, how to effectively "translate" or "tokenize" sensor time series data into a sequence form that LLM can understand and process has become a key prerequisite for realizing the above-mentioned technical vision.
[0003] In existing technologies, the processing of sensor time series data has mainly developed along two technical routes:
[0004] Route 1: Traditional Machine Learning and Dedicated Deep Learning Models. This method trains specialized models (such as Support Vector Machines, Random Forests, LSTMs, Transformer prediction models, etc.) for specific sensor tasks (e.g., anomaly detection, fault classification). These methods require separate data collection, feature design, and model training for each task. Once trained, these models lack generalization capabilities across tasks and sensor types. When the sensor type or downstream task changes, a complete retraining process is typically required, resulting in high data and computational costs.
[0005] Route Two: Temporal Data Adaptation Methods Based on Large Language Models. Recently, the academic community has proposed several technical solutions for applying LLM to general time series prediction. For example, Time-LLM (arXiv:2310.01728) proposes a "reprogramming" framework that uses a cross-attention mechanism to map time series data to a text prototype space, then inputs it into a frozen LLM for processing, and uses a Prompt-as-Prefix (PaP) strategy to achieve few-shot prediction. GPT4TS (One Fits All) adopts a similar approach, mapping time series data to the embedding dimension of the LLM through data patching and linear projection layers, freezing most of the LLM parameters (only fine-tuning the positional encoding and layer normalization layers). Subsequent works such as PatchInstruct further explore time series data segmentation methods based on structured instructions.
[0006] Although the aforementioned prior art (especially Time-LLM and GPT4TS) has provided useful explorations for solving the modal alignment problem between time-series data and LLM, the inventors have determined that the prior art still has the following technical defects or unresolved technical problems:
[0007] First, existing solutions are not specifically designed for the particular data format of "sensor time series," lacking a systematic consideration of the heterogeneity of sensor data, the diversity of physical dimensions, and multi-source sensor fusion scenarios. Solutions like Time-LLM primarily target general time series forecasting tasks (such as power load, weather temperature, traffic flow, etc.), where these data often have relatively regular sampling frequencies and uniform physical meanings. However, in real-world sensor applications, different types of sensors (such as pressure sensors, vibration sensors, and temperature sensors) have different sampling rates, dimensions (Pascals, degrees Celsius, millimeters / second, etc.), noise characteristics, and physical semantics. Existing technologies lack a unified word segmentation framework capable of processing various heterogeneous sensor signals while maintaining the integrity of their physical semantics.
[0008] Second, existing technologies, when performing modal alignment, either require the introduction of additional complex modules (such as the cross-attention reprogramming module of Time-LLM) or fine-tuning of some LLM parameters (such as fine-tuning position encoding and layer normalization in GPT4TS), which increases the complexity and deployment cost of the system to some extent. More importantly, when multiple different downstream tasks need to be processed simultaneously (e.g., simultaneously determining whether a bearing is abnormal, identifying the type of fault, and predicting the remaining life), existing technologies usually need to maintain a complete copy of the model for each task or perform task-specific parameter updates, lacking an efficient multi-task adaptation architecture of "one frozen LLM core + multiple lightweight task heads".
[0009] Third, while existing technologies have initially proposed prompt templates (such as PaP in Time-LLM) that include context, task instructions, and statistical information in the construction of prompt words, their prompt content is often static or semi-static. They fail to fully utilize the real-time statistical characteristics of sensor time series data to dynamically adjust the prompt content and lack structured and tagged encapsulation methods for sensor domain knowledge. This limits LLM's accurate understanding of the physical meaning of sensor signals in low-sample scenarios.
[0010] Fourth, existing technologies lack standardized solutions for handling the boundary conditions at the end of the sequence in the crucial data patching stage. When the length of the sensor time series cannot be divided evenly by the patching step size, the way the tail data is processed directly affects the sequence integrity and inference accuracy of the input LLM, and existing technologies have not given sufficient attention to this issue or provided a systematic solution.
[0011] In summary, existing technologies lack a unified word segmentation method and system architecture specifically designed for various types of sensor time-series data, capable of low-cost modality alignment and adaptation, supporting efficient multi-task expansion, and independent of LLM parameter fine-tuning. This invention aims to solve the aforementioned technical problems. Summary of the Invention
[0012] To address the aforementioned deficiencies in existing technologies, this invention aims to provide a sensor time series data processing method and system based on a large language model. It is a unified word segmentation method and system architecture specifically designed for various types of sensor time series data, capable of low-cost modal alignment and adaptation, supporting efficient multi-task expansion, and independent of large language model parameter fine-tuning.
[0013] Specifically, the technical problems to be solved by this invention include:
[0014] First, existing LLM-based time-series data processing solutions are not specifically designed for the "sensor time series" data format, and lack systematic consideration of the heterogeneity of sensor data, the diversity of physical dimensions, and multi-source sensor fusion scenarios. Different types of sensors (such as pressure sensors, vibration sensors, and temperature sensors) have different sampling rates, dimensions, noise characteristics, and physical semantics. Existing technologies lack a word segmentation framework that can uniformly process various heterogeneous sensor signals while maintaining the integrity of their physical semantics.
[0015] Second, existing technologies, when performing modal alignment, either require the introduction of additional complex modules (such as the cross-attention reprogramming module of Time-LLM) or fine-tuning of some LLM parameters (such as fine-tuning position encoding and layer normalization in GPT4TS), increasing system complexity and deployment costs. When multiple different downstream tasks need to be processed simultaneously, existing technologies lack an efficient multi-task adaptation architecture of "one frozen LLM core + multiple lightweight task heads".
[0016] Third, while existing technologies have initially proposed prompt templates that include context, task instructions, and statistical information (such as Prompt-as-Prefix in Time-LLM) for constructing prompt words, their prompt content is often static or semi-static. They fail to fully utilize the real-time statistical characteristics of sensor time series to dynamically adjust the prompt content, and also lack a structured and labeled encapsulation method for sensor domain knowledge.
[0017] Fourth, existing technologies lack standardized solutions for handling the boundaries of the sequence tail in the data segmentation process. When the length of the sensor time series cannot be divided by the segmentation step size, the way the tail data is processed directly affects the sequence integrity and inference accuracy of the input LLM.
[0018] The purpose of this invention is to overcome at least some of the defects of the prior art and provide a unified technical solution that can efficiently and cost-effectively adapt various types of sensor time series data to large language models.
[0019] To achieve the above-mentioned objectives, this invention provides a sensor time series data processing system and method based on a large language model.
[0020] In a first aspect, the present invention provides a sensor time series data processing system based on a large language model, comprising:
[0021] The data acquisition module is used to acquire raw time-series data collected by at least one sensor. The sensors include, but are not limited to, physical sensors (such as pressure sensors, vibration sensors, and temperature sensors), chemical sensors, mechanical sensors, industrial sensors, or environmental sensors.
[0022] A data blocking module, configured to preprocess the original time series data and then divide the preprocessed data into a plurality of consecutive data blocks according to a predetermined block length and a sliding step size.
[0023] A feature projection module, configured to map each of the data blocks to a latent space of a large language model through a trainable feature projection layer to obtain data block embedding features.
[0024] A prompt construction module, configured to construct a structured text prompt containing domain knowledge and / or task information according to the type of the sensor and / or the current downstream task to be executed.
[0025] An embedding fusion module, configured to convert the structured text prompt through a tokenizer and an embedding layer of the large language model to obtain text embedding features, and fuse the text embedding features with the data block embedding features to generate a unified multi-modal embedding sequence.
[0026] A large language model inference module, which adopts a pre-trained large language model with parameters kept frozen, and is configured to perform sequence modeling and inference on the unified multi-modal embedding sequence, and output a general feature representation.
[0027] A downstream task adaptation module, which includes at least one independently trainable task-specific projection head, and is configured to map the general feature representation into an output result corresponding to a specific downstream task.
[0028] Further, the data blocking module first normalizes the original time series data to make it have zero mean and unit standard deviation. Preferably, the normalization adopts a Z-score normalization method: calculating the mean μ and standard deviation σ of the segment of sensor time series data, and processing each data point x in the original data according to the formula x norm =(x-μ) / σ for conversion, so that the entire data sequence conforms to the distribution of zero mean and unit standard deviation.
[0029] Subsequently, the data blocking module divides the normalized time series into N data blocks according to a predetermined block length P and a sliding step size S. Wherein, when S=P, it is non-overlapping division; when S<P, it is overlapping division, and the overlapping length of adjacent data blocks is P-S. For a time series with a length of L, the extraction interval of the i-th data block is [i×S, i×S+P], where i∈[0, N-1].
[0030] Specifically, for the portion of the sequence that is less than a complete data block at the end, the data segmentation module employs padding to ensure complete division. When (LP) is not divisible by S, data of length Pad = S - ((LP) mod S) is padded at the end of the sequence. Padding methods include, but are not limited to, zero padding, edge duplication padding, or mean padding. The padded sequence is divided into N = ⌈(LP) / S⌉ + 1 complete data blocks. Each data block is vectorized and combined into a two-dimensional matrix of dimension N×P, which serves as the input to the feature projection module.
[0031] The above block processing has dual technical effects: (1) by aggregating local information into each data block, it better preserves the local semantic and trend information of the time series; (2) as a coarse-grained word segmentation method, it forms a compact input token sequence, which significantly reduces the computational burden of LLM.
[0032] Furthermore, the feature projection module maps each data block to the latent space of the large language model in the following way: Let the dimension of the input embedding layer of the large language model be D. For N data blocks of length P, a trainable linear mapping matrix W is used... proj Using a ∈R^(P×D) or a small multilayer perceptron (MLP), each data block is projected from the temporal signal space to the latent space of the large language model to obtain the data block embedding features E of dimension N×D. patch =Patch×W proj .
[0033] After the projection is completed, the temporal features are mathematically perfectly aligned with the text embedding vectors of the large language model, achieving cross-modal alignment between the two different modalities of "time series signals" and "natural language text." This lays the foundation for subsequently activating the time series understanding and reasoning capabilities of the large language model. Unlike existing alignment methods that require additional cross-attention modules (such as Time-LLM) or fine-tuning of some LLM parameters (such as GPT4TS), this invention achieves modal alignment through a lightweight, trainable projection layer, significantly reducing system complexity and training costs.
[0034] Furthermore, the prompt construction module constructs the structured text prompt through a structured template splicing mechanism. Specifically:
[0035] A predefined structured template containing multiple marker bits is defined, each marker bit corresponding to a sensor description component, a task instruction component, and a data statistics component. Upon receiving sensor time-series data to be processed, the prompt construction module performs the following operations:
[0036] (1) Static text loading: Read static description information that matches the current sensor type from the system configuration library and fill it into the corresponding flag bit. The static description information includes the sensor type, deployment environment, sampling frequency and / or domain prior knowledge.
[0037] (2) Task instruction generation: Generate natural language instructions based on the current downstream task and fill them into the corresponding marker bits. The natural language instructions include inference objectives and / or output constraints.
[0038] (3) Dynamic statistical calculation and textification: Statistical features are calculated in real time for the currently input time series data and formatted into text and then filled into the corresponding marker positions. The statistical features include at least one of mean, standard deviation, maximum value, minimum value, kurtosis and skewness.
[0039] The completed structured template is output as the structured text prompt.
[0040] The structured prompt construction method described above differs from existing technologies (such as Time-LLM's Prompt-as-Prefix) in that: (1) it adopts a structured template with clear marker bits, which realizes the standardized encapsulation and dynamic splicing of prompt components; (2) the statistical features are calculated in real time based on the current input window and automatically formatted, rather than using static or semi-static statistical information; (3) the domain knowledge is dynamically read from the system configuration library in a tagged manner, which supports flexible adaptation to different sensor types.
[0041] Furthermore, the embedding fusion module generates the unified multimodal embedding sequence in the following manner:
[0042] First, the structured text prompts are converted into a Token ID sequence using the built-in word segmenter of the large language model. Then, the Token ID sequence is converted into text embedding features E_text∈R^(M×D) through the text embedding layer of the large language model, where M is the length of the text embedding sequence.
[0043] Secondly, the text embedding feature E_text and the data block embedding feature E_patch are concatenated and concatenated along the sequence length dimension to obtain the joint multimodal feature matrix E_concat∈R^((M+N)×D).
[0044] Then, a position encoding matrix P∈R^((M+N)×D) is generated based on the concatenated total index position. The position encoding matrix is then added element-wise to the joint multimodal feature matrix so that the large language model can recognize the sequential order in the sequence.
[0045] Furthermore, a learnable modality type encoding is introduced, which superimposes a text type feature T_text on the position corresponding to the text embedding feature and a sensor type feature T_sensor on the position corresponding to the data block embedding feature, thereby explicitly marking the modality boundary in the latent space.
[0046] Finally, the fused multimodal embedding sequence is represented as E_final = E_concat + P + T, where T = [T_text; T_sensor]. This unified multimodal embedding sequence with (M+N)×D dimensions is fed as a whole into a pre-trained large language model that remains frozen.
[0047] The difference between the above fusion method and the existing technology is that the existing technology (such as Time-LLM) usually only splices the prompts before the temporal embedding, while the present invention further superimposes learnable modality type encoding on the basis of splicing, explicitly distinguishing between text modality and sensor modality in the latent space, so that the self-attention mechanism of the large language model can process information of different modalities more specifically.
[0048] Furthermore, the large language model inference module employs a pre-trained large language model with parameters frozen to perform sequence modeling and inference on the unified multimodal embedding sequence. The pre-trained large language model includes, but is not limited to, models based on the Transformer architecture such as the GPT series, LLaMA series, ChatGLM series, DeepSeek series, or Qianwen series.
[0049] During sequence modeling, the multi-head self-attention layer of the large language model automatically establishes cross-modal attention weights between text embeddings and sensor data block embeddings when calculating the correlation between the query matrix and the key matrix. The text semantics at the front end of the sequence (such as task descriptions and domain rules) act as "semantic anchors" of the context in the calculation, guiding and standardizing the parsing logic of the attention heads at the back end of the model for sensor data features (such as trends, periods, and impulses). This activates the sequential understanding, logical deduction, and pattern matching capabilities that the original large language model has solidified during massive text training, enabling temporal data reasoning without comprehensive fine-tuning.
[0050] Throughout the process, all parameters of the large language model remain frozen, requiring no fine-tuning or updates to the model parameters based on sensor data.
[0051] Furthermore, the downstream task adaptation module includes multiple task-specific projection heads configured in parallel, each with a different output dimension, corresponding to different types of downstream tasks. The downstream tasks include at least two of the following: classification tasks, multi-classification tasks, and regression tasks.
[0052] The projection head corresponding to the binary classification task includes a linear mapping layer and a softmax activation function, used to map the general feature representation to a probability distribution of two categories; the projection head corresponding to the multi-class classification task includes a linear mapping layer and a softmax activation function, used to map the general feature representation to a probability distribution of multiple categories; the projection head corresponding to the regression task includes a linear mapping layer, used to map the general feature representation to continuous numerical values.
[0053] When switching to a new downstream task, all parameters of the pre-trained large language model are frozen, and only a new task-specific projection head corresponding to the new downstream task is trained. The number of parameters in the new task-specific projection head is much smaller than the number of parameters in the pre-trained large language model (usually less than one-thousandth of the latter), thereby achieving extremely low-cost business expansion.
[0054] Secondly, the present invention also provides a sensor time series data processing method based on a large language model, comprising the following steps:
[0055] Step 1: Acquire raw time-series data from at least one sensor;
[0056] Step 2: Perform normalization preprocessing on the original time series data;
[0057] Step 3: Divide the normalized time series into multiple consecutive data blocks according to the predetermined block length and sliding step size;
[0058] Step 4: Map each data block to the latent space of the large language model through a trainable feature projection layer to obtain data block embedding features;
[0059] Step 5: Based on the type of the sensor and / or the downstream task to be performed, construct a structured text prompt containing domain knowledge and / or task information;
[0060] Step six: The structured text prompts are converted into text embedding features by the word segmenter and embedding layer of the large language model;
[0061] Step 7: Fuse the text embedding features with the data block embedding features to generate a unified multimodal embedding sequence;
[0062] Step 8: Perform sequence modeling and inference on the pre-trained large language model with the unified multimodal embedding sequence input parameters kept frozen, and output a general feature representation;
[0063] Step nine: Map the general feature representation to the output corresponding to the specific downstream task using an independently trainable task-specific projector.
[0064] Compared with the prior art, the present invention has the following beneficial effects:
[0065] 1. High Generalization and Versatility: This invention provides a word segmentation strategy and system architecture specifically designed for general-purpose sensor time series signals. It successfully transforms various heterogeneous physical quantity time series (including but not limited to pressure, vibration, temperature, flow rate, acceleration, etc.) into embedded representations that can be understood by large language models, without being limited to specific sensor types or specific downstream tasks. Unlike existing technologies that mainly target general time series prediction tasks (such as power load, weather temperature, etc.), this invention is specifically designed for the heterogeneity of sensor data, the diversity of physical dimensions, and multi-source fusion scenarios, exhibiting stronger cross-sensor type generalization capabilities.
[0066] 2. Extremely low model adaptation and deployment costs: This invention achieves extremely low adaptation costs through the following mechanisms:
[0067] (1) Freeze LLM strategy: All parameters of the large language model itself are completely frozen during inference and task adaptation, without the need for expensive, time-consuming and extremely resource-intensive full or partial model fine-tuning for new sensor datasets;
[0068] (2) Lightweight projection layer: Modal alignment is achieved by a single trainable linear projection matrix or a small MLP, with a parameter count much smaller than that of the LLM itself;
[0069] (3) Independently trainable task projection head: When it is necessary to develop new downstream sensor business, the R&D personnel only need to train a linear projection head with a very small number of parameters (usually less than one-thousandth of the number of LLM parameters) to directly connect to the system, which greatly reduces the development and computing costs of the system.
[0070] (4) Structured prompts guide LLM inference by declarative prompts containing domain knowledge, task instructions and statistical features. This eliminates the need to retrain or fine-tune the model for each new task and can efficiently activate the zero-shot or few-shot learning capabilities of pre-trained LLMs.
[0071] 3. Preserving local semantics and reducing computational overhead: This invention uses a data segmentation mechanism to divide long-sequence sensor time data into multiple consecutive data blocks, with each data block serving as a basic token unit input into the LLM. This mechanism has dual technical benefits:
[0072] (1) By aggregating information within blocks, key local features and trend information of the time series are effectively preserved;
[0073] (2) As a coarse-grained word segmentation method, it significantly reduces the computational burden caused by long sequence inputs, enabling LLM to efficiently process long sensor time series windows. In particular, the present invention adopts normalized tail boundary padding (zero padding, edge duplication, or mean padding) during the block segmentation process, ensuring the integrity of the input sequence and the accuracy of inference, while the prior art lacks a systematic solution for this.
[0074] 4. Explicit Cross-Modal Boundary Marking and Controllable Inference: This invention introduces learnable modality type encoding during the embedding and fusion stage, explicitly marking the boundaries between text modalities and sensor modalities in the latent space. Compared to existing technologies that achieve modality fusion solely through sequence concatenation, the modality type encoding of this invention enables the self-attention mechanism of large language models to more accurately identify and process information from different modalities, enhancing the controllability and accuracy of cross-modal inference.
[0075] 5. Highly Efficient Multi-Task Expansion Capability: The downstream task adaptation module of this invention supports multiple parallel-configured task-specific projection heads, each with different output dimensions, corresponding to different types of downstream tasks such as classification, multi-classification, and regression. When a new task needs to be added, only a new lightweight projection head needs to be trained, without any modification to the LLM or existing projection heads. This achieves a highly efficient multi-task architecture of "one frozen LLM core + multiple lightweight task heads," greatly improving the flexibility of sensor data processing and business expansion in fields such as industrial IoT, smart manufacturing, and environmental monitoring.
[0076] 6. Expanding the application boundaries of large language models: This invention can directly utilize advanced large language models to perform a wide range of downstream sensor tasks (including but not limited to sensor event classification and identification, anomaly detection, equipment status monitoring, predictive maintenance, environmental or traffic trend prediction, etc.), introducing the powerful sequence modeling and generalization reasoning capabilities of LLM into the field of sensor data analysis, and greatly expanding the applicability of large language models in industrial applications. Attached Figure Description
[0077] Figure 1 This is a flowchart illustrating the overall architecture of the sensor time series data processing system based on a large language model, as described in this invention.
[0078] Figure 2 This is a schematic diagram of the sensor time series data patching process in this invention.
[0079] Figure 3 This is a schematic diagram of the data block word segmentation and modal alignment process in this invention.
[0080] Figure 4 This is a schematic diagram of the structured domain knowledge prompt construction process in this invention.
[0081] Figure 5 This is a schematic diagram of the multimodal embedding and fusion process in this invention.
[0082] Figure 6 This is a schematic diagram of the downstream task multi-projection head architecture in this invention.
[0083] Figure 7 This is the overall flowchart of the sensor time series data processing method based on a large language model according to the present invention. Detailed Implementation
[0084] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0085] like Figure 1 As shown, the sensor time series data processing system based on a large language model provided in this embodiment of the invention includes two parallel processing paths: the left path is the text feature extraction path, and the right path is the sensor data word segmentation path. The outputs of the two paths converge at the embedding layer, forming a unified multimodal embedding sequence, which is then fed into a pre-trained large language model with its parameters frozen for sequence modeling and inference. Finally, the downstream projection head outputs the specific task results.
[0086] In one specific implementation, the sensors include, but are not limited to, physical sensors (such as pressure sensors, vibration sensors, acceleration sensors, temperature sensors, and flow sensors), chemical sensors, mechanical sensors, industrial process sensors, or environmental monitoring sensors. The sensor can be a single sensor, or a sensor array or sensor network composed of multiple sensors. The raw time series data can be a univariate time series or a multivariate time series. When it is a multivariate time series, the system processes each variable separately in blocks, and then concatenates or processes the data blocks of each variable independently along the feature dimensions.
[0087] The pre-trained large language model can be any commercial or open-source large language model based on the Transformer architecture, including but not limited to the GPT series (such as GPT-2, GPT-3, GPT-4), the LLaMA series (such as LLaMA-2, LLaMA-3), the ChatGLM series, the DeepSeek series, the Qwen series, and the Baichuan series. The choice of large language model can be determined based on the computational resource constraints and performance requirements of the specific application scenario. For example, in resource-constrained edge computing scenarios, a model with a smaller number of parameters (such as GPT-2 or LLaMA-2-7B) can be selected; in high-performance cloud computing scenarios, a model with a larger number of parameters (such as LLaMA-3-70B or GPT-4) can be selected. Regardless of the model selected, its parameters are kept frozen in this invention, that is, no parameter updates or fine-tuning are performed based on sensor data.
[0088] like Figure 2 As shown, the block processing of sensor time series data is a key step in achieving coarse-grained word segmentation in this invention. In one specific embodiment, this process includes sub-steps such as data acquisition, normalization processing, block parameter configuration, slice index calculation, tail boundary processing, and output mapping.
[0089] (a) Data acquisition and normalization processing
[0090] The system receives raw time-series data from the sensors via a data acquisition module. Let the received raw time series be X = {x1, x2, ..., x...}. L}, where L is the sequence length, x t This represents the sensor reading at time step t.
[0091] The system first performs Z-score standardization (mean-variance normalization) on this data segment. Specifically, the system calculates the mean μ and standard deviation σ of this segment of sensor time series data:
[0092] μ=(1 / L)×Σ {t=1} ^{L} x t
[0093] σ = sqrt{(1 / L)×Σ {t=1} ^{L} (x t -μ)²}
[0094] Then, for each data point x in the original data t Convert using the following formula:
[0095] X norm,t =(x t -μ) / σ
[0096] After the above transformation, the entire data sequence X norm = {x norm,1 , x norm,2 , ..., x norm,L} conforms to a distribution with zero mean and unit standard deviation. The purpose of normalization is to eliminate the influence of differences in dimension and magnitude of different sensors on model inference, so that data from different types of sensors (e.g., pressure sensors in pascals, temperature sensors in degrees Celsius) are comparable before being input into the large language model.
[0097] The system divides the normalized time series into N consecutive data blocks according to a predetermined block length P and a sliding step size S. Wherein, P represents the number of time steps included in each data block, and S represents the number of interval steps between the starting positions of adjacent data blocks.
[0098] When S=P, the division is non-overlapping between data blocks, that is, each time step belongs to only one data block. When S<P, the division is overlapping between data blocks, and the overlapping length of adjacent data blocks is P-S. The advantage of overlapping division is that it can enhance the continuity of local context and avoid losing cross-boundary temporal pattern information due to fixed boundary segmentation. In a preferred embodiment, for sensor signals with strong periodicity (such as vibration signals), an overlapping division method is adopted, and the overlapping length is set to P / 2; for sensor signals with relatively gentle changes (such as temperature signals), a non-overlapping division method can be adopted to further reduce the amount of calculation.
[0099] The specific values of P and S can be determined according to the sampling frequency of the sensor signal and the time scale of the temporal pattern to be analyzed. For example, for an industrial oil pipeline pressure sensor with a sampling rate of 500Hz, if it is necessary to analyze the pressure fluctuation pattern on the second-level scale, P can be set to 500 (corresponding to 1 second of data) and S to 250 (50% overlap); for an ambient temperature sensor with a sampling rate of 1Hz, P can be set to 60 (corresponding to 1 minute of data) and S to 60 (non-overlapping division).
[0100] For a normalized time series with a length of L, the extraction interval of the i-th data block is [i×S, i×S+P], where i∈[0, N-1]. The i-th data block can be expressed as:
[0101] Patch i ={x norm,{i×S} , x norm,{i×S+1} , ..., x norm,{i×S+P-1}}
[0102] When (L - P) is not divisible by S, the remaining data at the end of the sequence is insufficient to form a complete data block. In this case, the system pads the end of the sequence with data of length Pad, where:
[0103] Pad = S - ((LP) mod S)
[0104] The total length of the padded sequence is L+Pad, and it can be completely divided into N=⌈(LP) / S⌉+1 complete data blocks.
[0105] The filling method can be any of the following three methods:
[0106] (1) Zero padding: Fill the end of the sequence with the value 0. This method is the simplest to implement and is suitable for normalized signals with a mean close to 0.
[0107] (2) Edge copy padding: The last valid data value of the sequence is copied Pad times and padded to the end of the sequence. This method can maintain the numerical continuity at the end of the sequence and is suitable for scenarios where there is no obvious change in the signal at the end.
[0108] (3) Mean padding: Calculate the mean of the last few time steps of the sequence and use this mean as the padding value. This method can smoothly expand the sequence and is suitable for sensor signals with high noise.
[0109] In a preferred embodiment, the system adaptively selects the filling method based on the characteristics of different sensor signals. For example, for impact signals (such as the pulse signal from a vibration sensor during bearing failure), edge replication filling is preferred to preserve the impact characteristics; for stable signals (such as temperature sensor signals), mean filling is preferred.
[0110] After completing the above processing, the system vectorizes the N data blocks and combines them into a two-dimensional matrix of dimension N×P, which serves as the input to the subsequent data block word segmentation module. Each row of this two-dimensional matrix corresponds to a data block, and each column corresponds to a time step within the data block.
[0111] Through the block-based processing described above, the original long sequence of length L is compressed into N data blocks (N is much smaller than L), with each data block serving as a basic token unit. This preserves the local semantic information of the time series (each data block contains information from P consecutive time steps) while significantly reducing the sequence length input to a large language model, thereby reducing the computational burden.
[0112] like Figure 3 As shown, data block segmentation and modal alignment are key steps in mapping features from the sensor signal space to the text space of a large language model.
[0113] In one specific implementation, the dimension of the input embedding layer of the pre-trained large language model is set to D (e.g., D=4096 for LLaMA-2-7B, D=768 for GPT-2). For N sensor data blocks of length P (i.e., matrices of dimension N×P) obtained after data segmentation, the system projects each data block from the temporal signal space to the latent space of the large language model through a trainable feature projection layer.
[0114] The feature projection layer can be implemented in the following two ways:
[0115] Method 1: Linear Mapping. The system constructs a learnable linear mapping matrix W. proj ∈R^(P×D). For the i-th data block Patch i ∈R^P, its projected embedding vector is:
[0116] E patch,i = Patch i ×W proj
[0117] Where E patch,i ∈R^D. Stack the projection results of all N data blocks to obtain the data block embedding features E. patch ∈R^(N×D).
[0118] Method 2: Small Multilayer Perceptron (MLP) Mapping. The system constructs a small MLP network containing one or more hidden layers, and each data block is mapped to a D-dimensional space through this MLP network. The MLP method can capture more complex nonlinear relationships within the data block, but the number of parameters is slightly larger than that of the linear mapping method. In a preferred embodiment, the MLP mapping method is used for sensor signals with more complex data patterns (such as multi-component vibration signals); the linear mapping method is used for sensor signals with relatively simple data patterns (such as single temperature signals) to reduce training costs.
[0119] After the above projection is completed, the temporal features are mathematically perfectly aligned with the embedding vector of the standard text—both are D-dimensional continuous vectors. However, this alignment is only numerical. To truly achieve semantic alignment between the two different modalities of "time series signals" and "natural language text," it is necessary to use subsequent structured prompts and self-attention mechanisms to complete cross-modal semantic association.
[0120] It is worth noting that, unlike existing modality alignment methods that require the introduction of additional cross-attention modules (such as reprogramming the text prototype in Time-LLM) or fine-tuning of some parameters of the large language model (such as fine-tuning positional encoding and layer normalization in GPT4TS), this invention achieves mathematical modality alignment using only a lightweight trainable projection layer. This design significantly reduces the system's complexity and training cost, while retaining ample flexibility for subsequent semantic alignment—the self-attention mechanism of the large language model will automatically establish cross-modal semantic associations between the text and sensor data during inference.
[0121] like Figure 4 As shown, the construction of structured domain knowledge prompts is an important means for guiding the large language model to understand the physical meaning of sensor data. Unlike directly inputting abstract time-series numerical values into the large language model, this invention injects the prior knowledge of domain experts, the reasoning objectives of specific tasks, and the statistical characteristics of the current data into the model in the form of natural language through carefully designed structured text prompts, providing rich semantic context for the cross-modal reasoning of the large language model.
[0122] In one specific implementation, the system predefines a structured template containing multiple marker bits. This template includes the following three core components:
[0123] (a) Sensor description component (corresponding to the tags <dataset description> and <domain knowledge>)
[0124] This component describes the sensor type, deployment environment, sampling parameters, and relevant domain prior knowledge. This information is static, stored in the system configuration library, and is read and populated based on the type of sensor currently being processed.
[0125] For example, for a high-frequency pressure sensor deployed on an industrial oil pipeline, the component's content could be: "This dataset was acquired by a high-frequency pressure sensor deployed on an industrial oil pipeline, with a sampling rate of 500Hz. Prior knowledge in this field indicates that the pressure signal exhibits stable random fluctuations under normal pipeline conditions; when a valve is closed, it triggers an instantaneous pressure shock wave (water hammer effect), which manifests as a sudden increase in amplitude followed by decaying oscillations; when a pipeline leaks, the pressure signal shows a sustained decrease in mean and an increase in variance."
[0126] For example, for a vibration sensor deployed on the main bearing of a wind turbine, the populated content of this component could be: "This dataset was acquired by an acceleration vibration sensor deployed on the main bearing housing of a wind turbine, with a sampling rate of 2560Hz. Prior knowledge in this field indicates that the energy of the vibration signal during normal bearing operation is mainly distributed in the low-frequency band; inner ring wear generates impact pulses characterized by the bearing passage frequency and its harmonics; outer ring wear generates an impact sequence modulated by the cage rotation frequency; and rolling element spalling generates high-frequency impacts at random intervals."
[0127] (ii) Task instruction component (corresponding to the <task information> flag)
[0128] This component describes the specific goals and output constraints of the downstream task currently to be executed. This information is dynamically generated by the system based on the downstream scenario of the current business call.
[0129] For example, for an anomaly detection task, the task instruction could be: "Based on the characteristics of the time-series data block on the right, analyze whether there are any abnormal operating conditions in the current pipeline. If an anomaly exists, determine whether the anomaly type is 'water hammer effect' or 'pipeline leakage'; if no anomaly exists, output 'normal operation'. The returned result must be given in JSON format, including event category labels and confidence scores."
[0130] For example, in a multi-task joint reasoning scenario (simultaneously performing fault binary classification, fault type multi-classification, and remaining life regression prediction), the task instruction can be: "Based on the bearing vibration signal data block on the right, please complete the following three tasks in sequence: (1) Determine whether the bearing is currently in a 'normal' or 'abnormal' state; (2) If it is abnormal, further determine whether the specific fault type is 'inner ring wear,' 'outer ring wear,' or 'rolling element peeling'; (3) Predict the remaining service life of the bearing (in days)."
[0131] (III) Data statistics component (corresponding marker bit <sensor time series statistics x>)
[0132] This component describes the real-time statistical characteristics of the current input time series window. After receiving the data to be processed, the system calls statistical algorithms to perform real-time calculations on the current time series window, extracts key feature indicators, and automatically formats them into text lines.
[0133] The statistical characteristics include, but are not limited to, the following indicators:
[0134] Mean: Reflects the overall level of the signal;
[0135] Standard deviation: reflects the degree of dispersion of a signal;
[0136] Maximum and minimum values: reflect the amplitude range of the signal;
[0137] Kurtosis: Reflects the sharpness of the signal distribution and is of great value in detecting shocking abnormal events;
[0138] Skewness: Reflects the asymmetry of signal distribution;
[0139] Crest Factor: The ratio of the maximum value to the effective value, often used to detect mechanical faults;
[0140] Impulse Factor: The ratio of the maximum value to the absolute mean.
[0141] For example, for a time window of the pressure sensor mentioned above, the data statistics component could be populated with the following:
[0142] "Window signal mean: 2.34 Pa;"
[0143] Window signal standard deviation: 0.87 Pa;
[0144] Maximum window signal value: 12.45 Pa;
[0145] Minimum window signal value: 1.23 Pa;
[0146] Window signal kurtosis: 4.82;
[0147] Window signal skewness: 1.56.
[0148] The specific values of the above statistics will change dynamically depending on each input window, enabling the large language model to obtain real-time feature information of the current data, thereby making more accurate inferences and judgments.
[0149] After the above three components are filled, the system outputs the complete structured plain text string to the standard large language model tokenizer. The tokenizer then segments it into discrete text token sequences and converts them into continuous vectors that can be fused with the sensor signal matrix through the text embedding layer.
[0150] like Figure 5 As shown, multimodal embedding fusion is a key step in unifying and fusing the outputs of text paths and sensor paths at the embedding layer. This process includes four sub-steps: text feature extraction, sequence concatenation, positional encoding overlay, and modality type encoding overlay.
[0151] (I) Text Feature Extraction
[0152] The system first uses the built-in tokenizer of the pre-trained large language model to segment the structured text prompts. The tokenizer is based on sub-word segmentation algorithms fixed during the pre-training phase of the large language model, including but not limited to BPE (Byte Pair Encoding), SentencePiece, or WordPiece algorithms. The tokenizer compares the text prompts to a pre-defined vocabulary, dividing the plain text into the smallest tokens, and converting each token into a unique integer index (Token ID) in the vocabulary, thus transforming the text string into a one-dimensional integer sequence.
[0153] Let the total length of the resulting token ID sequence after segmentation be M. The system inputs these M integer IDs into the text embedding layer of the large language model. Based on these integer IDs, the text embedding layer retrieves the corresponding D-dimensional dense vector (D is the hidden feature dimension of the large language model) from the pre-trained lookup table matrix, and finally outputs a text embedding feature matrix E of shape M×D. text .
[0154] It should be noted that this invention directly reuses the matching word segmentation component that comes with the selected pre-trained large language model when it is released. There is no need to perform secondary fine-tuning or retraining of the word segmenter's vocabulary and text segmentation weights based on sensor data, thereby ensuring the efficiency and low computational consumption of the text extraction path.
[0155] (ii) Sequence cascading splicing
[0156] The system embeds the text into the matrix E along the sequence length dimension. text ∈R^(M×D) and data block embedding matrix E patch ∈R^(N×D) are concatenated and assembled into a unified joint multimodal feature matrix E. concat ∈R^((M+N)×D):
[0157] E concat =[E text E patch ]
[0158] That is, the embedding vectors of N sensor data blocks are concatenated after the embedding vectors of M text tokens to form a unified sequence of total length (M+N). This concatenation order has specific technical significance: placing the text prompts at the beginning of the sequence allows the large language model to use text semantics as an "anchor" in self-attention calculation, guiding subsequent understanding and reasoning of sensor data features.
[0159] (iii) Timing position coding superposition
[0160] To enable the large language model to recognize the order of tokens in a sequence, the system generates a corresponding two-dimensional positional encoding matrix P∈R^((M+N)×D) based on the concatenated total index position. Positional encoding can employ absolute positional encoding (such as the sine-cosine positional encoding in the original Transformer paper) or relative positional encoding (such as RoPE rotational positional encoding). The specific positional encoding method used depends on the encoding scheme of the pre-trained large language model itself—the system directly reuses the positional encoding mechanism built into the large language model to ensure consistency with the configuration during the pre-training phase.
[0161] The system performs element-wise addition on the location encoding matrix and the aforementioned joint multimodal feature matrix to incorporate absolute and relative temporal location information into the multimodal features.
[0162] (iv) Modal type encoding superposition
[0163] Furthermore, the system introduces a one-dimensional learnable modality type encoding (Token Type Embedding). This encoding appends an identifier vector indicating the source modality to each sequence position.
[0164] Specifically, text type feature T is superimposed on the M position vectors belonging to the text path. text , which is the superposition of N position vectors belonging to the sensor path and the sensor type feature T. sensor :
[0165] T=[T text ;T sensor ]
[0166] Where T text ∈R^D,T sensor ∈R^D, where T is a learnable parameter vector. During model training, T... text and T sensor They are jointly optimized during the training of the feature projection layer.
[0167] The system adds the modality type encoding to the joint multimodal feature matrix and the position encoding matrix element-wise to obtain the final multimodal embedding sequence:
[0168] E final =E concat +P+T
[0169] The E finalThe shape is (M+N)×D, which contains textual semantic information, sensor signal characteristics, temporal location information and modal type information.
[0170] Compared to existing technologies (such as Time-LLM's Prompt-as-Prefix, which achieves modality fusion solely through sequence concatenation), this invention introduces learnable modality type encoding, which explicitly marks the boundaries between text and sensor modalities in the latent space. This enables the multi-head self-attention mechanism of large language models to more accurately identify and process information from different modalities—when calculating attention weights, the model can selectively focus on feature interactions within the same modality and semantic associations across modalities, thereby enhancing the controllability and accuracy of cross-modal reasoning.
[0171] After completing multimodal embedding and fusion, the system will E final The entire sequence is fed into a pre-trained large language model with its parameters frozen for sequence modeling and inference.
[0172] In one specific implementation, the large language model is a model based on a Transformer decoder architecture, which includes multiple Transformer layers stacked sequentially. Each Transformer layer includes a multi-head self-attention sub-layer and a feed-forward network sub-layer, and is equipped with residual connections and layer normalization.
[0173] During sequence modeling, when the multi-head self-attention layer of a large language model calculates the relevance between the query matrix and the key matrix, it automatically establishes the text embedding E. text Embedded with sensor data blocks E patch Cross-modal attention weights between them.
[0174] Specifically, in the self-attention calculation, the text token at the front of the sequence (containing domain knowledge, task instructions, and statistical information) serves as the query and can be used to calculate an attention score with the key of the sensor data block token at the back of the sequence. This calculation process enables the model to "pay attention" to the sensor signal features most relevant to the current text semantics. Simultaneously, the sensor data block token, as the query, can also pay attention to the text token—this bidirectional, cross-modal attention mechanism enables deep semantic interaction between text and sensor signals.
[0175] More importantly, the semantic text at the beginning of the sequence (such as the domain knowledge that "closing the valve will cause an instantaneous pressure shock wave" or the task instruction that "analyze whether there is a leak in the current pipeline") acts as "semantic anchors" in self-attention computation. These semantic anchors guide and regulate the parsing logic of the attention head at the back end of the model for sensor data features—for example, the model tends to look for patterns in pressure signals that match the semantics of "water hammer effect" (sudden increase in amplitude followed by decaying oscillation) or "leakage" (continuous decrease in mean) that match the semantics of "leakage".
[0176] Through the above mechanism, this invention successfully activates the sequential understanding, logical deduction, and pattern matching capabilities that the original large language model has solidified in the training of massive amounts of text, enabling it to identify meaningful physical event patterns from sensor time series without any form of parameter fine-tuning for sensor data.
[0177] Throughout the process, all parameters of the large language model remain frozen—that is, no gradient updates or parameter adjustments are performed. This design brings significant technical advantages: (1) it avoids expensive and time-consuming large-scale model training or fine-tuning for each new sensor dataset; (2) it preserves the general reasoning capabilities of the original large language model; and (3) it enables the system to be deployed in resource-constrained edge computing environments at extremely low computational cost.
[0178] like Figure 6 As shown, the general feature representations (Hidden States) output by the large language model are passed to the downstream task adaptation module. This module includes one or more parallel task-specific projection heads, each used to map the general features to the output space of a specific downstream task.
[0179] In one specific implementation, the large language model outputs a feature vector of shape [1, D] (assuming the hidden layer dimension D=4096). This vector is obtained by appropriately pooling the output of the last Transformer layer (e.g., taking the output of the last token, or performing average pooling on the sequence dimension). This feature vector contains both temporal features such as the frequency, amplitude, and trend of the sensor signal, and integrates domain background knowledge and task context information injected into the prompt words.
[0180] The system can deploy multiple task-specific projection heads with different structures simultaneously, each targeting a specific type of downstream task. The following example, using wind turbine bearing anomaly monitoring, illustrates the implementation methods of three typical projection heads:
[0181] (a) Two-class projection head (Task A: Determine whether the bearing is "normal" or "abnormal")
[0182] The structure of the projection head consists of a lightweight linear mapping layer (matrix size D×2, i.e., 4096×2) followed by a Softmax activation function.
[0183] Its working mechanism is as follows: the 1×D feature vector output by LLM is multiplied by the 4096×2 weight matrix to reduce the dimension to a 1×2 vector (corresponding to the logical values of the "normal" and "abnormal" categories respectively), and then transformed into a probability distribution through the Softmax function. For example, the output may be [normal: 12%, abnormal: 88%], and the system determines that the bearing is in an abnormal state based on this.
[0184] (ii) Multi-class projection head (Task B: Determine the specific fault type as "inner ring wear", "outer ring wear" or "rolling element peeling")
[0185] The structure of the projector consists of an independently trained linear mapping layer (matrix size D×3, i.e., 4096×3) followed by a Softmax activation function.
[0186] Its working mechanism is as follows: the same 1×D feature vector is multiplied by the 4096×3 weight matrix, reducing the dimension to a 1×3 vector, and then transformed into a probability distribution of three fault categories through Softmax. For example, the output may be [inner ring wear: 75%, outer ring wear: 15%, rolling element spalling: 10%], and the system determines that the most likely fault type is inner ring wear.
[0187] (III) Regression Prediction Projection Head (Task C: Predicting Bearing Remaining Service RUL)
[0188] The structure of this projection head is a linear mapping layer (matrix size D×1, i.e., 4096×1), and no classification activation function is added to the back end.
[0189] Its working mechanism is as follows: after the feature vector is multiplied by the 4096×1 weight matrix, a continuous mathematical scalar is directly output. For example, the output is [14.5], which means that the remaining service life of the bearing is predicted to be 14.5 days.
[0190] The three projection heads mentioned above can be deployed in parallel within the same system. When the system receives a segment of bearing vibration sensor data, the large language model outputs a general feature vector, which simultaneously flows into the three projection heads to complete three tasks: "normal / abnormal binary classification," "fault type multi-classification," and "remaining life regression prediction," respectively, thus achieving an efficient inference architecture of "one feature, multiple task outputs."
[0191] More importantly, the large language model itself (the core architecture of the D dimension) remains completely frozen throughout the entire process described above. When a factory needs to develop new sensor diagnostic services—for example, adding a gearbox that needs to monitor its gear wear—R&D personnel do not need to fine-tune or retrain the massive large language model. They only need to: (1) write a structured prompt template for the gearbox vibration signal (containing domain knowledge of the gearbox and new task instructions); (2) train a new linear projection head of several hundred to several thousand bytes (such as a 4096×k matrix, where k is the number of task categories). The computational and time costs of the entire adaptation process are far lower than those of traditional methods, greatly reducing the development and deployment threshold of sensor data analysis services in industrial IoT scenarios.
[0192] like Figure 7 As shown, the sensor time series data processing method based on a large language model provided by the present invention includes the following steps:
[0193] Step S1: Acquire raw time series data collected by at least one sensor.
[0194] Step S2: Perform Z-score normalization preprocessing on the original time series data to give it zero mean and unit standard deviation.
[0195] Step S3: Divide the normalized time series into multiple consecutive data blocks according to the predetermined block length P and sliding step size S. When the end of the sequence is less than a complete data block, use zero padding, edge copy padding, or mean padding to fill the end.
[0196] Step S4: Map each data block to the latent space of the large language model through a trainable feature projection layer (linear mapping matrix or small MLP) to obtain the data block embedding features E. patch .
[0197] Step S5: Based on the type of the sensor and / or the downstream task to be executed, construct a structured text prompt containing a sensor description component, a task instruction component, and a data statistics component through a structured template splicing mechanism.
[0198] Step S6: The structured text prompt is converted into text embedding features E by the word segmenter and embedding layer of the large language model. text .
[0199] Step S7: Embed the text into feature E text With the data block embedding feature E patch The fusion process involves concatenating and splicing sequences along their length to obtain E. concatSuperimpose positional encoding matrices P; superimpose learnable modality type encodings T to generate a unified multimodal embedding sequence E. final =E concat +P+T.
[0200] Step S8: Perform sequence modeling and inference on the pre-trained large language model with the unified multimodal embedding sequence input parameters kept frozen, and output a general feature representation.
[0201] Step S9: Map the general feature representation to the output corresponding to a specific downstream task using one or more independently trainable task-specific projectors. When switching to a new downstream task, keep all parameters of the large language model frozen and train only the new task-specific projector corresponding to the new downstream task.
[0202] The technical effects of the present invention will be verified through several specific embodiments below.
[0203] Example 1: Anomaly Detection of Pressure Sensors in Industrial Oil Pipelines
[0204] This embodiment uses a high-frequency pressure sensor (sampling rate 500Hz) deployed on an industrial oil pipeline as an example to illustrate the specific application of the present invention in pipeline anomaly detection scenarios.
[0205] The system receives time-series data collected by pressure sensors in real time. The length of each processing window is L = 2000 (corresponding to 4 seconds of data). The system first performs Z-score normalization on the data in this window, and then sets the block length P = 500 (corresponding to 1 second) and the sliding step size S = 250 (50% overlap), dividing the 2000 time steps into N = ⌈(2000-500) / 250⌉+1 = 7 data blocks, with each data block containing 500 time steps.
[0206] The system reads the static description information of the pressure sensor from the configuration library: "This dataset was collected by a high-frequency pressure sensor deployed on an industrial oil pipeline, with a sampling rate of 500Hz. Domain prior knowledge: Valve closure causes water hammer effect (pressure shock wave followed by decaying oscillation); pipeline leakage causes a continuous decrease in the average pressure."
[0207] The system generates a task instruction based on the current business scenario: "Based on the characteristics of the time-series data block on the right, analyze whether there are any abnormal operating conditions in the current pipeline. If there is an abnormality, determine it as 'water hammer effect' or 'pipeline leakage'; if normal, output 'normal operation'. Return the result in JSON format."
[0208] The system calculates the statistical characteristics of the current window data in real time: "Mean: 2.34 Pa, Standard deviation: 0.87 Pa, Maximum value: 12.45 Pa, Kurtosis: 4.82".
[0209] The three components mentioned above are concatenated to form a structured text prompt, which is then converted into a text embedding through a word segmenter and embedding layer. Sensor data is converted into data block embeddings through a block segmentation and projection layer. The two are then fused at the embedding layer and input into the frozen LLaMA-2-7B model for inference. The general feature vector output by the model is then processed by a binary classification projector (4096×3) to output the probability distributions of the three categories.
[0210] In the practical verification of this embodiment, data was collected over a period of three months on an industrial oil pipeline testing platform, covering different combinations of operating conditions with varying flow velocities (1 m / s to 5 m / s) and different leakage orifice diameters (3 mm to 10 mm). The test dataset contained 20,000 samples, of which 15% were abnormal. Experimental results show that the overall accuracy of anomaly detection in the system of this invention reaches 98.7%, with a false positive rate (FPR) of only 0.8% and a false negative rate (FNR) of 0.5%. Compared with the traditional baseline method based on support vector machine (SVM), the accuracy is improved by approximately 12%, and it still exhibits extremely high robustness under the strong noise background caused by the start-up and shutdown of water pumps.
[0211] Example 2: Multi-task monitoring of wind turbine bearings
[0212] This embodiment takes an acceleration vibration sensor (sampling rate 2560Hz) deployed on the main bearing of a wind turbine as an example to illustrate the specific application of the present invention in the scenario of "one feature, multiple task outputs".
[0213] The system receives data from vibration sensors, with each processing window having a length L = 5120 (corresponding to 2 seconds). The system sets the block length P = 512 and the sliding step size S = 256, dividing the 5120 time steps into N = ⌈(5120-512) / 256⌉+1 = 19 data blocks.
[0214] The structured prompts of the system construction include domain knowledge of bearings ("Inner ring wear generates impact pulses characterized by the bearing passing frequency; outer ring wear generates impact sequences modulated by the cage rotation frequency; rolling element spalling generates high-frequency impacts at random intervals") and multi-task instructions ("Complete in sequence: (1) Determine normal / abnormal; (2) If abnormal, determine the fault type; (3) Predict the remaining life").
[0215] The system deploys three projection heads simultaneously: a binary classification head (4096×2+Softmax), a multi-classification head (4096×3+Softmax), and a regression head (4096×1). A single common feature vector is fed into all three projection heads simultaneously, and the results of the three tasks are output in parallel.
[0216] Experimental results demonstrate that, when using the publicly available bearing dataset from Case Western Reserve University (CWRU) for fault classification tasks (covering normal operation, inner ring wear, outer ring wear, rolling element failure, and various load conditions), our system achieves a classification accuracy of 99.2% and an F1-score of 0.99 without fine-tuning the large language model. Furthermore, in tests on the IMS bearing degradation dataset for predicting remaining useful life (RUL), our system outputs a root mean square error (RMSE) as low as 0.083 and a mean absolute percentage error (MAPE) of 6.5%, significantly outperforming existing dedicated Transformer time-series prediction models in both prediction accuracy and generalization ability.
[0217] Example 3: Multi-source sensor fusion monitoring
[0218] This embodiment takes multiple heterogeneous sensors (vibration sensor, temperature sensor, and pressure sensor) deployed simultaneously on industrial equipment as an example to illustrate the extended application of the present invention in multi-source sensor data fusion scenarios.
[0219] The system receives data streams from three different sensors, performs normalization, segmentation, and feature projection on each independently, and obtains three sets of data block embedding features. The system then concatenates these three sets of data block embedding features along the sequence length dimension (or first fuses them with text embeddings before concatenation) to form a unified multimodal input sequence.
[0220] The structured prompts incorporate multi-source fusion of domain knowledge (e.g., "the high-frequency component of the vibration signal and the gradual change trend of the temperature signal jointly indicate the bearing lubrication status") to guide the large language model in establishing cross-sensor type association reasoning. Experiments show that multi-source fusion significantly improves the accuracy and robustness of anomaly detection compared to single-sensor fusion.
[0221] Specific performance comparison data shows that on a large industrial gearbox test bench, when using only a single-axis vibration signal as input, the system's fault identification accuracy rate is 89.5%. However, when using a multi-source heterogeneous signal fusion input combining vibration, temperature, and hydraulic pressure, the system accuracy rate significantly jumps to 97.8% thanks to structured prompts guiding semantic association of cross-modal features. Furthermore, in the feature extraction of early, subtle faults, the multi-source fusion scheme advances the system's fault warning time by an average of 45 hours, greatly enhancing the reliability of monitoring.
[0222] In other embodiments of the present invention, in addition to the above fixed-length division method, adaptive-length division or multi-scale division may also be adopted for dividing data blocks. Adaptive-length division means dynamically adjusting the length of data blocks according to the local change rate of the sensor signal—shorter data blocks are used in regions with drastic signal changes to retain details, and longer data blocks are used in regions with gentle signal changes to reduce the amount of calculation. Multi-scale division means dividing the same sequence with multiple different block lengths simultaneously to form a multi-scale collection of data blocks, thereby enabling the large language model to perceive local details and global trends of the time sequence at the same time.
[0223] Overlapping or non-overlapping methods can be adopted between data blocks. The overlapping method (S<P) can enhance the continuity of local context, which is suitable for scenarios where signal patterns may cross block boundaries; the non-overlapping method (S=P) has higher computational efficiency, which is suitable for scenarios where signal patterns are well aligned with block boundaries.
[0224] In addition to the above tokenization method based on continuous data blocks, discrete vector coding (quantizing each data block to the nearest vector in a predefined codebook), clustering coding (clustering data blocks into K clusters and using the cluster index as the Token) or other sequence mapping methods can also be adopted for tokenization of time series data.
[0225] The large language model adopted in the present invention is not limited to the Transformer decoder architecture, and other structures of pre-trained models can also be adopted, including but not limited to Transformer encoder architecture (such as BERT), encoder-decoder architecture (such as T5), state space model (such as Mamba), etc. No matter what architecture is adopted, the core is to keep model parameters frozen, and the adaptation of sensor data is realized only through the projection layer and the prompting mechanism.
[0226] The content of the prompt can be flexibly adjusted according to different application scenarios, including but not limited to adding domain rules (such as "it is abnormal if the pressure exceeds the threshold X and lasts for Y seconds"), expert knowledge (such as "the fault characteristic frequency of this model of bearing is XXXHz") or more detailed task description information.
[0227] In addition to the above fusion method of sequence cascading concatenation plus position encoding plus modality type encoding, the following alternative methods can also be adopted for fusion of multimodal embeddings: (1) weighted sum fusion—performing element-wise weighted summation on text embeddings and data block embeddings according to learnable weights; (2) cross-attention fusion—realizing deep interaction between text and sensor data through a cross-attention layer; (3) gated fusion—adaptively controlling the proportion of text information and sensor information in fusion through a gating mechanism.
[0228] Compared with the prior art, the present invention has the following verifiable technical effects:
[0229] (1) Verification of high generalization capability: Under the same system architecture, the present invention can process time series data from different types of sensors (pressure, vibration, temperature, acceleration, etc.) without replacing or retraining the large language model.
[0230] (2) Verification of low adaptation cost: For adaptation to new sensor tasks, the present invention only requires training a lightweight projection head (with a parameter quantity of D×k, usually at the level of thousands to tens of thousands), without any fine-tuning of the large language model itself (the parameter quantity of a large language model is usually at the level of billions to hundreds of billions). The reduction in adaptation cost can reach several orders of magnitude.
[0231] (3) Verification of inference performance: The present invention compresses long sequences into short sequences through a data chunking mechanism (the sequence length is reduced from L to N, and usually N<<L), which significantly reduces the computational complexity during inference of the large language model (the computational complexity of the self-attention mechanism is reduced from O(L²) to O(N²)).
[0232] (4) Verification of multi-task parallel inference: The present invention realizes an efficient inference architecture of "one set of features, multi-task output" through a plurality of parallel task-specific projection heads.
[0233] To further verify the technical effects of the present invention, the following experimental settings are adopted: The basic pre-trained large language model adopts the LLaMA-2-7B model with frozen parameters; the hardware platform is a single NVIDIA RTX 4090 GPU. LSTM, the dedicated time series pre-training model PatchTST, and the existing time series alignment solution Time-LLM are selected as comparative methods. Comprehensive experimental results show that:
[0234] (1) Verification of high generalization capability and accuracy: In the above public datasets, the average classification accuracy of the present invention reaches 98.5%, which is significantly improved compared with LSTM (85.2%) and PatchTST (94.1%); in the few-shot scenario (only 10% of the training data is used), the accuracy still remains above 92%.
[0235] (2) Verification of low adaptation cost: For adaptation to new sensor tasks, the present invention only needs to train a linear projection head with a parameter scale of thousands, and the single-task adaptation training time on RTX 4090 does not exceed 15 minutes; while the fine-tuning time of the Time-LLM solution for the same task is as long as more than 6 hours.
[0236] (3) Inference performance and multi-task parallel verification: In the three-task parallel inference of Example 2, the single inference latency of this system is controlled within 50 milliseconds by sharing and distributing a single general feature vector; compared with the traditional method of deploying an independent dedicated model for each task, the memory usage is reduced by about 65% and the inference throughput is increased by nearly 3 times.
[0237] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A sensor time series data processing system based on a large language model, characterized in that, Comprising: a data acquisition module, configured to acquire raw time series data collected by at least one sensor; a data chunking module, configured to, after preprocessing the raw time series data, divide the preprocessed data into a plurality of consecutive data chunks according to a predetermined chunk length and a sliding step size; a feature projection module, configured to map each of the data chunks to the latent space of a large language model through a trainable feature projection layer to obtain data chunk embedding features; a prompt construction module, configured to construct a structured text prompt containing domain knowledge and / or task information according to the type of the sensor and / or the current downstream task to be performed; an embedding fusion module, configured to convert the structured text prompt through a tokenizer and an embedding layer of the large language model to obtain text embedding features, and fuse the text embedding features with the data chunk embedding features to generate a unified multi-modal embedding sequence; a large language model inference module, which adopts a pre-trained large language model with parameters kept frozen, configured to perform sequence modeling and inference on the unified multi-modal embedding sequence, and output a general feature representation; and a downstream task adaptation module, which comprises at least one independently trainable task-specific projection head, configured to map the general feature representation into an output result corresponding to a specific downstream task.
2. The system according to claim 1, characterized in that, The preprocessing performed on the raw time series data by the data chunking module comprises: normalizing the raw time series data to have zero mean and unit standard deviation; dividing the normalized time series into N data chunks according to a predetermined chunk length P and a sliding step size S, wherein non-overlapping division is implemented when S=P, overlapping division is implemented when S<P, and the overlapping length of adjacent data chunks is P-S; for the part at the end of the sequence that is less than one complete data chunk, padding processing is adopted to enable complete division, and the padding processing is any one of zero padding, edge replication padding or mean padding.
3. The system according to claim 2, characterized in that, When the data chunking module performs padding processing on the end of the sequence: for a time series with a length of L, when (L-P) is not divisible by S, data with a length of Pad=S-((L-P) mod S) is padded at the end of the sequence; the padded sequence is divided into N=⌈(L-P) / S⌉+1 complete data chunks; each of the data chunks is vectorized separately and then combined into a two-dimensional matrix with a dimension of N×P, which serves as the input of the feature projection module.
4. The system according to claim 1, characterized in that, The feature projection module maps each of the data chunks to the latent space of the large language model in the following manner: Let the dimension of the input embedding layer of the large language model be D. For N data blocks of length P, a trainable linear mapping matrix W is used. proj Using an ∈R^(P×D) or a small multilayer perceptron, each data block is projected from the temporal signal space to the latent space of the large language model to obtain the data block embedding features E of dimension N×D. patch .
5. The system according to claim 1, characterized in that, the structured text prompt constructed by the prompt construction module comprises at least two of the following three components: a sensor description component, configured to describe the type, deployment environment and / or domain prior knowledge of the sensor; a task instruction component, configured to describe the objective and / or output constraints of the current downstream task to be performed; a data statistics component, configured to describe the statistical features of the currently input time series data, wherein the statistical features comprise at least one of mean, standard deviation, maximum value, minimum value, kurtosis and skewness.
6. The system according to claim 5, characterized in that, the prompt construction module constructs the structured text prompt through a structured template splicing mechanism: A predefined structured template containing multiple marker bits, the marker bits corresponding to the sensor description component, the task instruction component, and the data statistics component, respectively; When receiving sensor time series data to be processed, static description information matching the current sensor type is read from the system configuration library and filled into the corresponding tag position. Natural language instructions are generated according to the current downstream task and filled into the corresponding tag position. Statistical features of the currently input time series data are calculated in real time, formatted into text, and then filled into the corresponding tag position. The completed structured template is output as the structured text prompt.
7. The system according to claim 1, characterized in that, The embedding fusion module generates the unified multimodal embedding sequence in the following manner: The text embedding feature E is applied along the sequence length dimension. text ∈R^(M×D) and the data block embedding feature E patch Concatenating and splicing ∈R^(N×D) yields the joint multimodal feature matrix E. concat ∈R^((M+N)×D), where M is the sequence length of the text embedding and N is the sequence length of the data block embedding; A position encoding matrix P∈R^((M+N)×D) is generated based on the concatenated total index position, and the position encoding matrix is added element by element to the joint multimodal feature matrix; A learnable modality type encoding is introduced, and text type feature T is superimposed on the position corresponding to the text embedding feature. text The carrier type feature T is superimposed on the position corresponding to the data block embedding feature. sensor ; Output the fused multimodal embedding sequence E final =E concat +P+T, where T=[T text ; T sensor ].
8. The system according to claim 1, characterized in that, The downstream task adaptation module includes multiple task-specific projection heads set in parallel, each with a different output dimension, corresponding to different types of downstream tasks. The downstream tasks include at least two of the following: classification tasks, multi-classification tasks, and regression tasks. The projection head corresponding to the classification task includes a linear mapping layer and a softmax activation function, which are used to output the classification probability distribution; The projection head corresponding to the regression task includes a linear mapping layer for outputting continuous numerical values.
9. A method for processing sensor time series data based on a large language model, characterized in that, Includes the following steps: Acquire raw time-series data from at least one sensor; The original time series data is preprocessed by normalization; The normalized time series is divided into multiple consecutive data blocks according to the predetermined block length and sliding step size; Each data block is mapped to the latent space of a large language model through a trainable feature projection layer to obtain data block embedding features; Based on the type of the sensor and / or the downstream task to be performed, construct a structured text prompt that includes domain knowledge and / or task information; The structured text prompts are converted into text embedding features by the word segmenter and embedding layer of the large language model; The text embedding features are fused with the data block embedding features to generate a unified multimodal embedding sequence; The pre-trained large language model with its unified multimodal embedding sequence input parameters kept frozen is used for sequence modeling and inference to output a general feature representation; The general feature representation is mapped to the output corresponding to a specific downstream task through an independently trainable task-specific projector.
10. The method according to claim 9, characterized in that, The method further includes: When it is necessary to switch to a new downstream task, keep all parameters of the pre-trained large language model frozen and only train a new task-specific projection head corresponding to the new downstream task. The number of parameters of the new task-specific projection head is less than the number of parameters of the pre-trained large language model.