A fast adaptation method based on edge side timing model

CN122818043APending Publication Date: 2026-09-25UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611010734.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本申请的目的在于提供一种基于边缘侧时序模型快速适配方法,以解决工业设备工况变化导致预训练时序模型预测性能下降,以及传统在线微调方式在边缘节点中计算开销大、适配速度慢、关键结构信息利用不足的问题

Benefits of technology

[0057]其一,本申请通过对多源运行时序数据进行对齐、标记、补齐和分段,得到能够被边缘节点统一处理的当前时序样本窗口,使不同通道、不同采样周期和不同质量状态的数据可以进入同一结构分解处理链。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application relates to the technical field of industrial equipment predictive maintenance, and discloses a rapid adaptation method based on an edge side time sequence model, which comprises the following steps: an edge node acquires multi-source operation time sequence data, equipment working condition information and resource constraint information of an industrial equipment, and obtains a current time sequence sample window through time alignment, quality marking, missing data completion and window segmentation; the window is subjected to statistical, trend, periodic and time sequence dependent structure decomposition to generate a structure prompt vector; the structure contribution degree on a verification window and edge resource constraints are used to dynamically allocate a prompt dimension; in the case that the main parameters of a pre-trained time sequence model are frozen, only the structure prompt parameters are subjected to module perception update; the updated structure prompt vector is fused with a window input representation to output equipment state prediction, abnormal scoring and predictive maintenance prompts; the application can reduce the online adaptation overhead on the edge side, improve the real-time performance of industrial equipment abnormal early warning and the accuracy of predictive maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of predictive maintenance technology for industrial equipment, and more specifically, to a fast adaptation method based on edge-side time series models. Background Technology

[0002] Industrial equipment generates multi-source runtime sequence data during continuous operation, including temperature, current, voltage, pressure, vibration, speed, flow rate, and process stage indicators. By analyzing this multi-source runtime sequence data to predict equipment status, score anomalies, and generate predictive maintenance prompts, potential degradation trends can be identified before significant failures occur in industrial equipment, thereby reducing the risk of unplanned downtime.

[0003] With the application of pre-trained time series models in industrial monitoring, edge nodes can perform real-time inference near industrial equipment, reducing data transmission latency. However, different equipment models, process stages, load levels, and equipment aging conditions can cause changes in the statistical structure, trend structure, periodic structure, and time-series dependency structure of runtime time series data. If pre-trained time series models are used directly, problems such as increased prediction residuals, delayed anomaly scoring, and unstable maintenance prompts can easily occur.

[0004] Existing model adaptation methods typically include full parameter fine-tuning, low-rank parameter updates, or fixed cue vector updates. Full parameter fine-tuning requires significant GPU memory and long training time, making it difficult to apply to the low latency requirements of edge nodes in industrial settings. While fixed cue vector updates can reduce the amount of parameter updates, they usually do not dynamically adjust based on the current operating stage, differences in structural feature contributions, and edge node resource constraints, resulting in limited cue parameters failing to prioritize serving the current prediction and maintenance task.

[0005] Therefore, there is a need for a technical solution that can quickly generate a structure hint vector to be updated based on the current operating stage and runtime sequence data structure characteristics of industrial equipment under the condition of limited edge node resources, and only update trainable hint parameters, thereby improving the adaptation efficiency and prediction stability of the pre-trained time series model in the current industrial equipment predictive maintenance scenario. Summary of the Invention

[0006] The purpose of this application is to provide a fast adaptation method based on edge-side temporal models to solve the problems of declining prediction performance of pre-trained temporal models caused by changes in the operating conditions of industrial equipment, as well as the high computational cost, slow adaptation speed, and insufficient utilization of key structural information in traditional online fine-tuning methods at edge nodes.

[0007] To achieve the above objectives, this application adopts the following technical solution.

[0008] A fast adaptation method based on edge-side temporal models, applied to edge nodes deployed on industrial equipment, includes the following steps:

[0009] Step S101: Obtain multi-source runtime sequence data, equipment operating condition information and resource constraint information of industrial equipment in the current maintenance cycle; align, mark, complete and segment the multi-source runtime sequence data to obtain the current time series sample window.

[0010] Step S102: Perform statistical, trend, periodic, and time-dependent structural decomposition on the current time series sample window to obtain a set of structural features;

[0011] Step S103: Determine the current operating stage and structural prompt template based on the equipment operating condition information, write the structural feature set into the structural prompt template, and generate the initial structural prompt vector;

[0012] Step S104: Filter the verification window that is consistent with the current operating stage according to the equipment operating condition information, and determine the prompt dimension according to the structural contribution and resource constraint information of the initial structural prompt vector on the verification window, and form the structural prompt vector to be updated according to the prompt dimension.

[0013] Step S105: Freeze the main parameters of the pre-trained temporal model, update the trainable cue parameters in the structure cue vector to be updated, and generate the updated structure cue vector.

[0014] Step S106: Convert the current time series sample window into a time series input vector, concatenate the updated structure prompt vector to the front of the time series input vector to form an enhanced model input and input it into the pre-trained time series model to obtain the equipment status prediction, anomaly score and predicted maintenance prompt results.

[0015] Further, in step S101, the multi-source runtime sequence data, equipment condition information, and resource constraint information of the industrial equipment during the current maintenance cycle are obtained, including: reading temperature, current, voltage, pressure, vibration, speed, and flow channel data through the equipment controller, sensor acquisition interface, or industrial communication bus, and recording each channel data as a data item including channel identifier, sampling timestamp, sampling value, and communication status; reading the equipment number, process stage, start / stop status, and load level from the equipment controller to generate equipment condition information; reading available memory, available computing power, maximum adaptation latency, and cache capacity from the edge node runtime monitoring program to generate resource constraint information; and associating the multi-source runtime sequence data, equipment condition information, and resource constraint information to form an edge cache record for use in the window generation of step S101 and for further calls in steps S103 to S105.

[0016] Furthermore, step S101 aligns, marks, pads, and segments the multi-source runtime sequence data, including:

[0017] The edge nodes determine a unified time granularity based on the device control cycle and the effective sampling cycle of each channel in the edge cache records, and generate a unified time axis based on the unified time granularity;

[0018] The data from each channel is mapped to a unified time axis according to the sampling timestamp, and quality markers are generated based on whether the sampled value is empty, whether it exceeds the channel range, whether the communication status is interrupted, and the start / stop status.

[0019] When the quality marker indicates a missing value and the duration of the missing value does not exceed the upper limit of missing value completion, the missing sample value is completed by forward hold or the nearest mean.

[0020] Then, based on the cache capacity and maximum adaptation latency in the resource constraint information, determine the window length and window step size, and truncate continuous windows according to the window length and window step size to generate the current time series sample window.

[0021] Furthermore, methods for obtaining the set of structural features include:

[0022] The edge node reads the sampled values ​​and quality markers corresponding to each sensor channel in the current time-series sample window, first removes the stop marker positions and unusable fill positions, and then calculates the mean, variance, skewness and kurtosis of the remaining sampled values ​​to form statistical structure features;

[0023] Perform moving average, linear regression, or first-order differencing on the retained sample values ​​to form trend structure characteristics;

[0024] After removing the trend component from each channel window sequence, frequency domain transformation or periodic peak search is performed on the residual sequence to form periodic structure features;

[0025] The autocorrelation coefficient is calculated according to the preset lag order and combined with the cross-channel correlation coefficient to form the time-dependent structural characteristics.

[0026] Each structural feature is combined with the structural type, channel identifier, window number, and feature value to form a set of structural features, including statistical structural features, trend structural features, and periodic structural features.

[0027] Further, in step S103, the current operating stage and structural prompt template are determined based on the equipment operating condition information, including:

[0028] Read the process stage, start / stop status and load level from the equipment operating condition information, and determine the current operating condition stage based on the process stage;

[0029] Select the structural prompt template that corresponds to the current working condition stage and structural type from the structural prompt template table;

[0030] Following the order of data items in the structure prompt template, write the structure type, channel identifier, feature name and feature value in the structure feature set into the corresponding positions, and add the equipment number, current operating condition stage, load level and prediction task identifier into the structure prompt template to obtain the structure prompt record;

[0031] Perform a hint encoding mapping on the structure hint record to generate an initial structure hint vector.

[0032] Furthermore, step S104 involves filtering verification windows that match the current operating condition stage based on the equipment operating condition information, including:

[0033] Read the historical sample window within the current maintenance cycle and compare the operating condition stage corresponding to the historical sample window with the current operating condition stage;

[0034] The validation window is obtained by retaining historical sample windows that are consistent with the operating conditions and have a window missing ratio less than the missing ratio threshold, a noise amplitude less than the noise amplitude threshold, and a number of key channels not less than the number of key channels threshold.

[0035] Generate status labels or weak labels for the verification window based on equipment maintenance records, process alarm records, or predicted residual thresholds;

[0036] The verification window and its status label or weak label are used to calculate the structural contribution corresponding to the initial structural cue vector.

[0037] Further, step S104, which involves determining the cue dimension and forming the structure cue vector to be updated according to the cue dimension, includes:

[0038] The upper limit of the total dimension of the prompt is determined based on the input dimension of the pre-trained temporal model, available memory, available computing power, and maximum adaptation latency;

[0039] Assign initial cue dimensions to the initial structure cue vectors corresponding to statistical structure, trend structure, periodic structure, and time-dependent structure;

[0040] Each initial structural cue vector is appended with a unit dimension increment to form a candidate cue dimension scheme;

[0041] Calculate the candidate validation loss corresponding to the candidate cue dimension scheme on the validation window, and take the difference between the baseline validation loss and the candidate validation loss as the corresponding structural contribution.

[0042] The cue dimension of each initial structural cue vector is determined based on the structural contribution, and the initial structural cue vectors are truncated, linearly compressed, or linearly expanded according to the cue dimension to form the structural cue vector to be updated.

[0043] Further, step S105, which involves freezing the main parameters of the pre-trained temporal model and updating the trainable cue parameters in the structure cue vector to be updated, includes:

[0044] Edge nodes set the corresponding parameters of the input embedding layer, temporal coding layer, and prediction output layer of the pre-trained temporal model to an untrainable state;

[0045] The parameters in the structure cue vector to be updated that are limited by the cue dimension are determined as trainable cue parameters; the cue learning rate and structure importance coefficient are read from the parameter configuration table according to the structure type;

[0046] Calculate the adaptation loss based on at least one of the prediction loss, anomaly recognition loss, or cue parameter regularization term on the verification window;

[0047] The trainable cue parameters are updated based on the adaptation loss, cue learning rate, and structural importance coefficient, generating an updated structural cue vector while keeping the main parameters of the pre-trained temporal model unchanged.

[0048] Further, step S106, which involves forming the input to the enhanced model and inputting it into the pre-trained temporal model, includes:

[0049] The sampled values ​​of each channel in the current time series sample window are normalized and vectorized to generate a time series input vector;

[0050] The updated structure hint vectors are arranged in a fixed order of statistical structure, trend structure, periodic structure, and time-dependent structure.

[0051] The arranged structural cue vector is concatenated to the beginning of the temporal input vector to obtain the enhanced model input;

[0052] The augmented model input is used as the input data for the pre-trained time series model. The pre-trained time series model outputs at least one of the following: sensor prediction values, equipment health, or failure probability within the future target time period. Based on the prediction residual between the sensor prediction values ​​and the actual observation values, the equipment health, or the failure probability, the equipment status prediction and anomaly score are generated.

[0053] Furthermore, after receiving the prediction and maintenance prompts, the edge nodes continuously receive new multi-source runtime sequence data and write the new window prediction residuals, the current operating stage, and the verification loss into the adaptation log.

[0054] When the current working condition changes, the prediction residuals of multiple consecutive windows reach the refit threshold, or the validation loss increases continuously, the edge node uses the structural cue vector and cue dimension obtained from the previous fit as the initial result and re-executes steps S102 to S105.

[0055] When the verification loss after re-adaptation is less than the rollback threshold, the re-adapted structure hint vector is retained; when the verification loss after re-adaptation reaches the rollback threshold, the previous version's structure hint vector is restored, and the triggering reason, rollback time, and verification loss before and after the update are recorded in the adaptation log.

[0056] Compared with related technologies, this application has the following advantages:

[0057] Firstly, this application aligns, marks, completes, and segments multi-source runtime sequence data to obtain a current time-series sample window that can be uniformly processed by edge nodes, enabling data from different channels, different sampling periods, and different quality states to enter the same structural decomposition processing chain.

[0058] Secondly, this application performs statistical, trend, periodic and time-dependent structural decomposition on the current time-series sample window, and writes the resulting set of structural features into the structural prompt template determined by the equipment operating condition information, so that the structural prompt vector reflects both the current operating condition stage and the current window structure change.

[0059] Third, this application calculates the structural contribution based on a verification window consistent with the current working stage, and determines the hint dimension in conjunction with resource constraint information, so that limited edge resources are preferentially allocated to structural hint vectors that contribute more to the current predictive maintenance task, thereby reducing the amount of invalid updates.

[0060] Fourth, this application freezes the main parameters of the pre-trained temporal model, only updates the trainable cue parameters in the structure cue vector to be updated, and concatenates the updated structure cue vector to the front of the temporal input vector to form an enhanced model input, which can improve the adaptability of equipment status prediction, anomaly scoring and predictive maintenance cue with a small amount of parameter updates. Attached Figure Description

[0061] Figure 1 This is an overall flowchart of a fast adaptation method based on edge-side temporal models proposed in this application.

[0062] Figure 2 This is a scenario diagram of edge-side predictive maintenance implementation in this application.

[0063] Figure 3 This is a flowchart of the multi-source runtime sequence data alignment, marking, completion, and segmentation process in this application.

[0064] Figure 4 This is a flowchart of the decomposition of the statistical, trend, periodic, and time-series dependency structures in this application.

[0065] Figure 5 This is a flowchart of the process for writing the structure hint template and generating the initial structure hint vector in this application.

[0066] Figure 6This is a flowchart of the verification window screening, structural contribution calculation, and prompt dimension determination process for this application.

[0067] Figure 7 This is a flowchart of the pre-trained temporal model training and structure hint parameter update process in this application.

[0068] Figure 8 This is a flowchart of the process for constructing the input of the enhanced model and generating the prediction and maintenance results in this application.

[0069] Figure 9 This is a flowchart of the re-adaptation condition judgment and structural hint vector rollback process for this application.

[0070] Figure 10 This is a schematic diagram of the edge-side rapid adaptation functional unit structure of this application. Detailed Implementation

[0071] The technical solution of this application will be further described below with reference to the accompanying drawings. The following embodiments are used to illustrate this application and are not intended to limit the scope of protection of this application.

[0072] In this embodiment, the current maintenance cycle refers to the continuous operating time period from the time of the most recent equipment maintenance completion, equipment startup, or process batch start to the current data acquisition time. The start time of the current maintenance cycle is determined by the edge node reading equipment maintenance records, equipment start / stop records, or production batch records, and the current data acquisition time is used as the end time. The edge node writes the start and end times of the current maintenance cycle into the maintenance cycle record table, and determines the range of data to be processed based on the maintenance cycle record table when executing step S101. In an optional implementation, when equipment maintenance records are missing, the edge node uses the most recent startup time as the start time of the current maintenance cycle.

[0073] In this embodiment, the edge node is an industrial computer, embedded computing unit, or controller with edge inference capability deployed on the industrial equipment side. The edge node stores at least a parameter configuration table, a structure prompt template table, a warning level table, a model version table, and an adaptation log table. The parameter configuration table includes at least the parameter name, parameter value, setting basis, applicable equipment type, applicable operating condition stage, update time, and configuration source. All content in this specification involving thresholds, coefficients, rules, templates, and tables is read by the edge node from the aforementioned tables and then processed accordingly.

[0074] See Figure 1 and Figure 2The edge-side timing model fast adaptation method provided in this embodiment is executed by edge nodes deployed on the industrial equipment side; the industrial equipment can be semiconductor manufacturing equipment, CNC machining equipment, pump sets, compressors, fans, motors or production line transmission equipment; the sensor group is used to collect multi-source runtime timing data such as temperature, current, voltage, pressure, vibration, speed, flow rate, etc.; the edge node is deployed with a pre-trained timing model and a structure prompt adaptation program; the maintenance terminal receives the predicted maintenance prompt results output by the edge node.

[0075] See Figure 3 In step S101, the edge node acquires multi-source runtime sequence data through the device controller, sensor acquisition interface, or industrial communication bus. The industrial communication bus can be Ethernet, CAN bus, Modbus bus, OPC-UA interface, or a data interface provided by the device manufacturer. Each acquisition record includes channel identifier, sampling timestamp, sampled value, unit information, and communication status. The edge node writes the acquisition record into the edge buffer queue and uses the edge buffer queue as input for alignment, marking, padding, and segmentation processing. In the basic implementation, the sampling timestamp is reported by the device controller. In the optional implementation, when the device controller does not report the sampling timestamp, the sampling timestamp is determined by the receiving time of the edge node.

[0076] Channel identifiers are used to uniquely identify sensor channels; for example, a temperature channel can be represented as TEMP_01, a vibration channel as VIB_01, and a pressure channel as PRES_01; the sampled value is the measurement value of the corresponding channel at the sampling timestamp; unit information is used to determine the physical dimensions of different channels. Communication status is used to determine whether there is a communication interruption for the corresponding sampling record.

[0077] Edge nodes read equipment status information from the device controller. This information includes the device number, process stage, start / stop status, and load level. The process stage is used to determine the current status stage and filter verification window. The device number and load level are used to write the structure prompt template. The start / stop status is used for quality marking and missing data completion judgment. Edge nodes also read resource constraint information from the runtime monitoring program. This information includes available memory, available computing power, maximum adaptation latency, and cache capacity. Available memory is obtained by subtracting safety reserve memory from the current free memory read by the edge node's operating system. Available computing power is calculated by reading the processor utilization rate, graphics processor utilization rate, or neural network accelerator utilization rate from the edge node. Maximum adaptation latency is determined by maintenance response requirements or process control allowable delay. Cache capacity is determined by the maximum length of the edge cache queue or the remaining storage space. This resource constraint information is used to determine the prompt dimension and the number of fast update rounds.

[0078] Edge nodes generate a unified time axis with a unified time granularity. The basic implementation of the unified time granularity is as follows: when the sampling periods of each sensor channel are consistent, the sampling period is taken as the unified time granularity; when the sampling periods of each sensor channel are inconsistent, the larger value between the device control period and the minimum effective sampling period of each channel is taken as the unified time granularity. The effective sampling period is obtained by the edge node by calculating the difference between the nearest preset number of adjacent effective sampling timestamps and taking the median. In the optional implementation, when the channel sampling period is registered in the sensor parameter table, the edge node prioritizes reading the sampling period in the sensor parameter table to reduce online statistical overhead.

[0079] Edge nodes map data from each channel to a unified time axis based on sampling timestamps. In the basic implementation, for multiple sampled values ​​falling within the same time granularity interval, the edge node selects the sampled value closest to the corresponding time point. In the optional implementation, the edge node selects the mean, median, or last sampled value of multiple sampled values ​​within the corresponding time granularity interval, and records the mapping method used in the parameter configuration table.

[0080] After alignment, the edge nodes generate quality markers for each sampling point. In the basic implementation, when a sampled value exists and is within the corresponding channel range, and the communication status is normal, the quality marker is "normal". When a sampled value is empty or no sampling record is received within the corresponding time granularity, the quality marker is "missing". When a sampled value is less than the minimum channel range or greater than the maximum channel range, the quality marker is "out of bounds". When the communication interface does not return data within the communication waiting time, the quality marker is "communication interrupted". When the start / stop status is "shutdown", the quality marker is "shutdown". The channel range is read from the sensor parameter table, and the communication waiting time is determined by the communication protocol response time limit or a manual configuration file. In the optional implementation, the edge nodes can also mark short-term spike sampling points as suspected abnormal sampling points and reduce the weight of these sampling points during subsequent structural decomposition.

[0081] For channel values ​​missing at the same time point, edge nodes first read the quality marker; if the quality marker indicates a single missing point, forward hold is used for filling; if the duration of consecutive missing values ​​does not exceed the missing value filling limit, the average of the adjacent valid values ​​before and after the missing point is used for filling; if the start / stop status is shutdown, no numerical filling is performed, and a shutdown marker is written. The missing value filling limit is obtained by multiplying the corresponding channel's effective sampling period by the missing value filling factor; the missing value filling factor can be determined based on the historical distribution of missing durations, for example, by taking the 95th percentile of the missing duration during historical normal acquisition, dividing it by the effective sampling period, and rounding it up, or by having equipment maintenance personnel write it into the parameter configuration table based on the channel sampling stability.

[0082] The current time series sample window after alignment, labeling, and padding can be represented as:

[0083]

[0084] in, Indicates the current time The current time series sample window at the end point, Indicates the current data collection time. Indicates the window length. Indicates the number of sensor channels. Indicates the current time 3D channel sampling vector, Indicates by a time point and A real number matrix space consisting of channels.

[0085] Window length The setting method is as follows: when the shortest duration of the device fault precursor is known, the corresponding shortest duration is divided by a uniform time granularity to obtain the window length; when the duration of the fault precursor is unknown, the window length is set according to the response cycle of the state prediction task; the window step size is set according to the maximum adaptation latency and cache capacity; when the maximum adaptation latency is less than the latency threshold, the window step size is set to a value that allows adjacent windows to overlap; when the cache capacity is lower than the cache capacity threshold, the window step size is set to a value that reduces cache usage; the latency threshold and cache capacity threshold are read from the parameter configuration table.

[0086] The percentage of abnormally missing items within a window can be expressed as:

[0087]

[0088] in, Indicates the percentage of abnormal missing items within the window. This indicates the number of sample points in the window that are marked as missing, out of bounds, or communication interrupted. Indicates the window length. This indicates the number of sensor channels. This formula is used to illustrate a quantitative rule for window quality screening.

[0089] when Less than the missing percentage threshold When, edge nodes retain the current time-series sample window; when Not less than When this happens, edge nodes discard the corresponding window or generate a data quality alarm; missing proportion threshold. You can take the 95th percentile of the historical missing window ratio, or the equipment maintenance personnel can configure it according to the stability of the data collection system.

[0090] In some implementations, the following example can be used to illustrate the construction process of the current timing sample window; for example, the industrial equipment is a semiconductor etching equipment. The edge node reads the process stage and start / stop status from the equipment controller, and reads multi-source runtime timing data from pressure sensors, temperature sensors, current sensors, voltage sensors, and vibration sensors; the effective sampling period of the pressure channel PRES_01, temperature channel TEMP_01, current channel CUR_01, and voltage channel VOLT_01 is 1 second, the effective sampling period of the vibration channel VIB_01 is 0.5 seconds, and the equipment control cycle is 1 second; the edge node determines a uniform time granularity of 1 second, and takes the current acquisition time 10:00:00 as the end time of the window, and extracts 128 time points backward to obtain the current timing sample window with a window length of 128 and a number of channels of 5.

[0091] In the current time-series sample window, if the pressure channel PRES_01 does not receive sampled values ​​at 09:59:12 and 09:59:13, and the missing completion factor for the pressure channel is set to 3, then the upper limit for missing completion is 3 seconds. Since the continuous missing duration of the pressure channel is 2 seconds, which does not exceed the upper limit for missing completion, the edge node reads two adjacent valid sampled values ​​at 09:59:11 and 09:59:14, and uses the average of the two values ​​to complete the corresponding positions at 09:59:12 and 09:59:13, while retaining the completion position marker. If the vibration channel VIB_01 shows a communication interruption marker in the same window, and the number of abnormal sampling points is 18, then the abnormal missing ratio in the window is 18 / (128×5)=0.028. When the missing ratio threshold is set to 0.05, 0.028 is less than 0.05, the edge node retains the current time-series sample window, and outputs the current time-series sample window to step S102.

[0092] See Figure 4 In step S102, the edge node performs structural decomposition on the current time series sample window. The edge node takes the current time series sample window output in step S101 as input, extracts statistical structural features, trend structural features, periodic structural features and time-dependent structural features for each sensor channel, and combines each structural feature into a structural feature set.

[0093] For the For each sensor channel, edge nodes first filter out stop marker positions and unusable fill positions based on quality markers, and then calculate statistical structure characteristics for the remaining valid sampled values. Statistical structure characteristics include mean, variance, skewness, and kurtosis; as an example, the statistical structure characteristics are:

[0094]

[0095] in, Indicates the first Statistical structural characteristics of each channel Indicates the first in the window The average of each channel, Indicates the first in the window The variance of each channel, Indicates the first in the window The skewness of each channel, Indicates the first in the window The kurtosis of each channel, This represents the transpose of a vector.

[0096] Optionally, statistical structural features may also include peak values, ranges, or quantiles, but these optional indicators are not essential components of statistical structural features in this embodiment.

[0097] Trend structure features are used to represent the direction and magnitude of the channel sample values ​​changing over time within the current time series sample window; in the basic implementation, edge nodes perform linear regression on the window data to obtain the trend slope; simultaneously, a moving average is performed to obtain a smoothed trend value. The trend slope of each channel can be expressed as:

[0098]

[0099] in, Indicates the first The trend slope of each channel, Indicates the window length. Indicates the time sequence number within the window. This represents the average of the time sequence numbers. Indicates the first in the window The time point, the first The sampled values ​​of each channel, Indicates the first The average value of each channel within the window. This represents the smallest positive number that prevents the denominator from being zero.

[0100] In the optional implementation, the edge nodes use an exponential moving average to extract the trend structure; the smoothing coefficient of the exponential moving average is set according to the sampling period and the allowable response delay; the shorter the sampling period or the greater the noise, the smaller the smoothing coefficient value; the smoothing coefficient value is written into the parameter configuration table and updated according to the operating conditions.

[0101] Periodic structure features are used to represent repetitive changes caused by process cycles, rotating parts, or periodic control actions during equipment operation. In the basic implementation, edge nodes first remove trend components from the original sequence to obtain a residual sequence, then perform a frequency domain transformation on the residual sequence, and select one or more periodic components with the highest amplitude as periodic structure features. Periodic structure features include period amplitude and period length, and can be expressed as:

[0102]

[0103] in, Indicates the first The periodic structural characteristics of each channel Indicates the first The period amplitude corresponding to each period component Indicates the first The period length corresponding to each periodic component. This indicates the number of periodic components selected.

[0104] If the first is obtained through frequency domain transformation One main frequency Then the period length can be expressed as:

[0105]

[0106] in, Indicates the first The first channel The period length corresponding to each periodic component. Indicates the first The first channel The main frequency corresponding to each period component.

[0107] Temporal dependency structure features are used to represent the dependency relationships between different time points within the same channel, as well as the linkage relationships between different channels; in the basic implementation, edge nodes calculate the autocorrelation coefficient under a preset lag order and calculate the cross-channel correlation coefficient. Each channel has a lag order. The autocorrelation coefficient under these conditions can be expressed as:

[0108]

[0109] in, Indicates the first Each channel has a lag order. The autocorrelation coefficient under the following conditions Indicates the lag order. Indicates the window length. Indicates the first in the window The time point, the first The sampled values ​​of each channel, Indicates the first The average of each channel, This represents the smallest positive number that prevents the denominator from being zero.

[0110] The preset hysteresis order is determined based on the device control cycle, sampling cycle, and fault precursor response time; for example, when the sampling cycle is 1 second and it is necessary to identify dependency changes within 2 seconds, 4 seconds, and 8 seconds, the hysteresis order is set to 2, 4, and 8, respectively. In an optional implementation, edge nodes can also calculate cross-channel Pearson correlation coefficients and record channel pairs whose absolute correlation coefficient values ​​reach a correlation threshold as cross-channel dependency structures. The correlation threshold can be taken as the 90th percentile of the absolute value of the cross-channel correlation coefficients in a historical stable window, or it can be manually configured.

[0111] Furthermore, to illustrate the generation process of the structural feature set, the following example can be used. For instance, for the pressure channel PRES_01, the edge node calculates a mean of 0.42, a variance of 0.018, a skewness of 0.36, and a kurtosis of 2.91, thus forming a statistical structural feature. The edge node further performs linear regression using time series 1 to 128 as the horizontal sequence and the normalized pressure sample values ​​as the vertical sequence, obtaining a trend slope of 0.0048. When the trend slope threshold is 0.0030, the edge node records the trend structural feature corresponding to the pressure channel PRES_01 as a slow increase in pressure.

[0112] For the periodic structure feature, the edge node removes the moving average trend component from the original sequence of the pressure channel PRES_01 to obtain the residual sequence, and performs a frequency domain transformation on the residual sequence. If the two main frequencies with the first amplitude are 0.125Hz and 0.250Hz, and the corresponding amplitudes are 0.23 and 0.08, the corresponding period lengths are 8 seconds and 4 seconds, respectively. The edge node combines the period amplitude of 0.23 and the period length of 8 seconds, and the period amplitude of 0.08 and the period length of 4 seconds as the periodic structure feature of the pressure channel PRES_01.

[0113] For the time-dependent structural features, the edge nodes calculate the autocorrelation coefficients of the pressure channel PRES_01 at lag orders of 2 and 8, respectively. If the autocorrelation coefficient for lag order 2 is 0.74 and the autocorrelation coefficient for lag order 8 is 0.39, the edge nodes consider both as the time-dependent structural features of the pressure channel PRES_01. If the cross-channel correlation coefficient between the pressure channel PRES_01 and the temperature channel TEMP_01 is 0.68 and the correlation threshold is 0.60, the edge nodes further record that there is a cross-channel dependency structure between PRES_01 and TEMP_01.

[0114] Step S102 outputs a set of structural features. The set of structural features includes statistical structural features, trend structural features, periodic structural features, and time-dependent structural features. Each structural feature contains a structural type, channel identifier, window number, feature name, and feature value. This set of structural features serves as the input for step S103.

[0115] See Figure 5 In step S103, the edge node receives the set of structural features output in step S102 and reads the process stage, start / stop status, and load level from the equipment operating condition information obtained in step S101. The edge node determines the current operating condition stage based on the process stage and selects the structural prompt template corresponding to the current operating condition stage and structural type from the structural prompt template table.

[0116] The structure prompt template table includes template number, structure type, operating condition stage, data item order, data item type, default value, and applicable equipment type. In the basic implementation, the structure prompt template includes structure name, equipment number, current operating condition stage, channel identifier, feature name, feature value, and prediction task identifier. In the optional implementation, the structure prompt template can also include load level and start / stop status to further describe the current operating conditions of the industrial equipment.

[0117] The edge nodes, following the data item order in the structure prompt template, write the structure type, channel identifier, feature name, and feature value from the structure feature set into the corresponding positions, and add the equipment number, current operating condition stage, load level, and prediction task identifier into the structure prompt template to obtain the structure prompt record. Since the equipment operating condition information is involved in determining the current operating condition stage and selecting the structure prompt template, the equipment operating condition information is actually invoked in step S103.

[0118] Numerical structural features can be normalized before being written into the template; in one optional implementation, edge nodes read the mean and standard deviation of the corresponding channel and structural feature in the historical stable window, and then perform standardization; the standardization process can be expressed as:

[0119]

[0120] in, Represents the standardized first Type of structure, first The structural characteristics of each channel Indicates the structural features before standardization. This represents the mean of the corresponding structural feature within the historical stable window. This represents the standard deviation of the corresponding structural feature within the historical stable window. This indicates the prevention of extremely small positive numbers with a denominator of zero. This normalization process is used to reduce the impact of dimensional differences between different sensor channels on cue encoding.

[0121] In the basic implementation, the edge nodes call the local cue encoder to encode the structural cue records into fixed-length vectors. The local cue encoder can be a text encoding layer accompanying a pre-trained temporal model, or a lightweight encoding network deployed in the edge nodes; the initial structural cue vector can be represented as:

[0122]

[0123] in, Indicates the first The initial structure hint vector corresponding to each structure type. This indicates a local encoder prompt. Indicates the first This is a record of structural prompts.

[0124] In an optional implementation, the edge node does not call the text encoder, but instead uses structure type embedding, channel embedding, working condition embedding and numerical projection network to generate the initial structure cue vector. Specifically, the edge node looks up the structure type, channel identifier and current working condition stage to obtain discrete embeddings. The normalized numerical structure features are input into a lightweight multilayer perceptron to obtain numerical embeddings. The discrete embeddings and numerical embeddings are then concatenated and mapped to the initial structure cue vector through a linear projection layer.

[0125] In some implementations, the following example can be used to illustrate the generation process of the structure hint record and the initial structure hint vector; for example, for the trend structure feature of pressure channel PRES_01, the edge node reads the process stage as steady-state etching from the equipment operating condition information, and selects the trend structure hint template corresponding to the steady-state etching stage accordingly; the edge node writes the structure name as trend structure, the equipment number as EQP_01, the current operating condition stage as steady-state etching, the channel identifier as PRES_01, the feature name as trend slope, the feature value as 0.0048, the load level as medium load, and the prediction task identifier as pressure prediction for the next 30 minutes, thus obtaining the trend structure hint record.

[0126] The edge node calls the local cue encoder to encode the trend structure cue records and outputs a 128-dimensional initial trend structure cue vector. In an optional implementation, the edge node reads the structure type embedding, channel embedding, and working condition embedding respectively, and inputs the standardized trend slope into the numerical projection network. Then, the above embedding results are concatenated and mapped to a 128-dimensional initial trend structure cue vector through a linear projection layer.

[0127] Step S103 outputs the initial structural cue vector. The initial structural cue vector serves as the input for step S104 to calculate the structural contribution and determine the cue dimension.

[0128] See Figure 6 In step S104, the edge node first filters verification windows that match the current operating condition stage based on the equipment operating condition information. In the basic implementation, the edge node reads historical sample windows from the current maintenance cycle and compares the operating condition stage corresponding to the historical sample window with the current operating condition stage. If the operating condition stage corresponding to the historical sample window matches the current operating condition stage, and the missing proportion of the historical sample window is less than the missing proportion threshold, the noise amplitude is less than the noise amplitude threshold, and the number of critical channels is not less than the number of critical channels threshold, then the edge node retains the historical sample window as a verification window.

[0129] The missing proportion threshold can be taken as the 95th percentile of the missing proportion of historical normal operating windows. The noise amplitude can be represented by the mean of the absolute values ​​of the differences between adjacent samples, and the noise amplitude threshold can be taken as the 95th percentile of the noise amplitude of historical stable windows. The critical channel number threshold is determined by the number of channels required for the prediction task. For example, a pressure prediction task requires at least three critical channels: pressure, current, and temperature, so the critical channel number threshold is set to 3. In an optional implementation, the edge node can further filter the verification window based on the load level, specifically retaining only historical sample windows that are consistent with the current load level or located within the same load level range.

[0130] Edge nodes determine the cue dimension based on the structural contribution of the initial cue vector to the validation window and resource constraints. In the basic implementation, edge nodes determine the upper limit of the total cue dimension based on the input dimension of the pre-trained temporal model, available memory, available computing power, and maximum adaptation latency. The upper limit of the total cue dimension does not exceed the input dimension of the pre-trained temporal model, and also does not exceed the size of the cue parameters that the edge node is allowed to update within the maximum adaptation latency.

[0131] In one alternative implementation, the upper limit of the total dimensions can be expressed as:

[0132]

[0133] in, This indicates the maximum number of dimensions. This represents the input dimension of the pre-trained temporal model. Indicates the available memory for edge nodes. This represents the base memory used for model inference and caching. This indicates the number of bytes occupied by a single prompt parameter. This represents the memory amplification factor for the optimizer state and gradient state. After obtaining candidate values ​​according to this formula, the edge nodes round down the candidate values ​​and use the rounded result as the upper limit of the total available cue dimension.

[0134] Let the number of structure types be In this embodiment These correspond to statistical structure, trend structure, periodic structure, and time-dependent structure, respectively; edge nodes assign initial cue dimensions to the cue vector for each initial structure and set unit dimension increments. The dimension allocation vector before round allocation can be represented as:

[0135]

[0136] in, Indicates the first Dimensional allocation vector before round allocation. This indicates the dimensions of the statistical structure. This indicates the dimension of the trend structure. The periodic structure indicates the dimension. This indicates the dimension of the temporal dependency structure.

[0137] In each round of allocation, the edge nodes append a unit dimension increment to an initial structure cue vector, forming candidate cue dimension schemes. For the The candidate suggestion dimension scheme can be represented as follows:

[0138]

[0139] in, Indicates the first Wheel to the First Candidate suggestion dimension schemes after adding a unit dimension increment to the structure. Indicates the increment per unit dimension. Indicates the first A unit vector with 1 at one position and 0 at the other positions.

[0140] Edge nodes calculate the candidate validation loss corresponding to the candidate cue dimension scheme on the validation window, and determine the difference between the baseline validation loss and the candidate validation loss as the structural contribution; the structural contribution can be expressed as:

[0141]

[0142] in, Indicates the first The structure in the first Structural contribution in the wheel Indicates the first The benchmark validation loss when no unit dimension increment is added in the round. Indicates the first Wheel direction The candidate validation loss after adding a unit dimension increment to the structure. If A positive value indicates that the validation loss decreases after adding dimensions; the larger the value, the higher the validation loss. The greater the contribution of a structure after it gains additional hint dimensions.

[0143] Edge nodes determine the cue dimension of each initial structural cue vector based on structural contribution and resource constraint information, and then truncate, expand, or remap the initial structural cue vectors according to the cue dimension to form the structural cue vector to be updated. If the cue dimension of a structure is greater than its initial vector dimension, the edge node can expand the vector dimension through linear projection; if the cue dimension of a structure is less than its initial vector dimension, the edge node can truncate the beginning of the vector according to the cue dimension or obtain the corresponding dimension through linear compression. The structural cue vector to be updated serves as the input for step S105.

[0144] In some implementations, the following example may be used to illustrate the process of determining the prompt dimension.

[0145] For example, the edge node reads the pre-trained temporal model input dimension as 1024 and the number of structure types as 4, namely statistical structure, trend structure, periodic structure, and temporal dependency structure. The edge node assigns an initial cue dimension of 128 to each of the four initial structure cue vectors and sets the unit dimension increment to 32.

[0146] In the first round of dimension allocation, if the baseline validation loss without adding a unit dimension increment is 0.126, the validation loss for statistical structures is 0.122, for trend structures it is 0.112, for periodic structures it is 0.118, and for time-dependent structures it is 0.120, then the contribution of statistical structures is 0.004, the contribution of trend structures is 0.014, the contribution of periodic structures is 0.008, and the contribution of time-dependent structures is 0.006. Since the trend structure has the largest contribution, the edge nodes allocate the current unit dimension increment of 32 dimensions to the trend structure hint vector, updating the dimension allocation result from [128, 128, 128, 128] to [128, 160, 128, 128]. Subsequently, the edge nodes form the structure hint vector to be updated according to this hint dimension.

[0147] See Figure 7In this embodiment, the pre-trained time series model refers to a time series model that has been pre-trained on historical operating data of industrial equipment and deployed to the edge nodes on the industrial equipment side to receive inputs from the augmentation model and output equipment state prediction results. In the basic implementation, the pre-trained time series model includes an input embedding layer, a time series coding layer, a structure cue access layer, and a prediction output layer. The input embedding layer is used to convert the current time series sample window into a time series input vector; the time series coding layer is used to extract the time dependencies in the multi-channel time series; the structure cue access layer is used to receive the updated structure cue vector and form the augmentation model input with the time series input vector; the prediction output layer is used to output the sensor prediction values ​​within the future target time period; in an optional implementation, the pre-trained time series model also includes a state output layer, which is used to output the equipment health status or failure probability.

[0148] In the basic implementation, the temporal coding layer can be formed by concatenating a temporal convolutional coding layer and a multi-head self-attention coding layer. The temporal convolutional coding layer is used to extract local change features between adjacent sampling points, and the multi-head self-attention coding layer is used to extract temporal dependencies between different time points and different channels. In the optional implementation, the temporal coding layer can also be composed of a gated recurrent unit coding layer, a long short-term memory coding layer, or only a multi-head self-attention coding layer, as long as it can receive multi-channel temporal input and output the corresponding time series representation.

[0149] The training data for the pre-trained time series model is formed from historical industrial equipment operation data. The training data acquisition method is consistent with the acquisition method in step S101, that is, temperature, current, voltage, pressure, vibration, speed and flow channel data are read through the equipment controller, sensor acquisition interface or industrial communication bus, and the equipment number, process stage, start and stop status and load level are read simultaneously. The edge node or offline training server performs time alignment, quality marking, missing data filling and sliding window segmentation on the above historical operation data according to a unified time granularity to obtain historical time series sample windows. Each historical time series sample window includes multi-channel sampled values ​​at multiple time points and carries the corresponding equipment operating condition information.

[0150] Training labels can be determined based on the prediction task. For sensor prediction tasks, training labels are the actual sensor sampling values ​​within the target time period after the historical time series sample window. For anomaly identification tasks, training labels can come from equipment maintenance records, process alarm records, downtime records, or manually labeled records. If there are no manually labeled anomalies, edge nodes or offline training servers can generate weak labels based on the distribution of prediction residuals of historical normal windows. For example, windows with prediction residuals reaching the residual threshold can be used as weak anomaly samples, and windows with prediction residuals below the residual threshold can be used as weak normal samples.

[0151] In the basic implementation, the training input of the pre-trained time series model is the time series input vector corresponding to the historical time series sample window, and the training output is the sensor prediction value within the future target time period. If the pre-trained time series model includes a state output layer, the training output also includes the device health or failure probability. The device health can be generated based on maintenance records, failure occurrence time, and the interval between the current window and the failure occurrence time. The failure probability can be trained based on anomaly labels or weak labels. After pre-training is completed, the main parameters of the pre-trained time series model are solidified and deployed to the edge nodes.

[0152] As an example, a historical time series sample window can be represented as:

[0153]

[0154] in, Indicates the first A historical time series sample window, Indicates the historical time series sample window number. Indicates the window length. Indicates the number of sensor channels. Indicates the first The first historical time series sample window Multi-channel sampling vectors at various time points Indicates the time point number within the window. Indicates by a time point and A real number matrix space composed of sensor channels.

[0155] The prediction output of the pre-trained time series model for a window of historical time series samples can be expressed as:

[0156]

[0157] in, This indicates that the pre-trained time series model is for the first... The sensor prediction values ​​for the future target duration are output from a historical time-series sample window. This represents a pre-trained temporal model. Represents the main parameters of the pre-trained time series model. Indicates the first A historical time series sample window.

[0158] When the training task only includes sensor predictions, the pre-training loss function can be the prediction mean square error, expressed as:

[0159]

[0160] in, Indicates pre-training loss, Indicates the number of training samples. Indicates the training sample number. This represents the predicted value output by the pre-trained time series model. Indicates the first The true future sampled value corresponding to each historical time series sample window This represents the square of the L2 norm.

[0161] When the training task includes both sensor prediction and anomaly recognition, the pre-training loss function can be expressed as:

[0162]

[0163] in, Indicates the weight of the predicted loss. Indicates the anomaly identification loss weight. Indicates the number of training samples. Indicates the training sample number. Indicates the predicted value. Represents the true future sampled value. This represents the anomaly category prediction or fault probability output by the pre-trained time series model. This indicates an abnormal or weak tag. This represents the cross-entropy loss function.

[0164] The convergence conditions for a pre-trained temporal model include at least one of the following: the number of training epochs reaches the upper limit; the decrease in validation set loss over multiple consecutive training epochs is less than the pre-training convergence threshold; the validation set loss reaches the target loss threshold; and the validation set anomaly detection accuracy, recall, or prediction error reaches the task metrics. In the basic implementation, if the decrease in validation set loss over five consecutive training epochs is less than the pre-training convergence threshold, training is stopped and the main parameters of the current pre-trained temporal model are saved. In an optional implementation, the main parameters of the pre-trained temporal model can be saved in the training epoch with the lowest validation set loss, and these main parameters can be deployed to edge nodes.

[0165] After pre-training, the edge nodes freeze the main parameters of the pre-trained temporal model during step S105, updating only the trainable cue parameters in the structure cue vector to be updated. The training data for the edge-side fast adaptation phase is the validation window selected in step S104. The training input is the augmented model input formed by the unupdated structure cue vector and the validation window. The training output is the actual sensor sampling value, state label, or weak label corresponding to the validation window. The edge-side fast adaptation phase does not change the main parameters of the pre-trained temporal model; it only changes the trainable cue parameters in the structure cue vector to be updated.

[0166] The loss function for the edge-side fast adaptation phase can be expressed as:

[0167]

[0168] in, This indicates the loss from fast adaptation at the edge. Indicates the weight of the predicted loss. Indicates the anomaly identification loss weight. This indicates the weight of the regularization term as a prompt parameter. This represents the mean square error. This represents the sensor predictions of the pre-trained temporal model during the validation window. This represents the true future sampled value corresponding to the verification window. Represents the cross-entropy loss function. This indicates the anomaly category prediction result or failure probability on the verification window. This indicates the status label or weak label corresponding to the verification window. Indicates the first The trainable cue parameters corresponding to this structure Indicates the structure type number. Indicates the number of structure types. This represents the squared L2 norm of the trainable cue parameters; when performing only sensor prediction tasks. It can be set to 0.

[0169] The convergence conditions for the edge-side fast adaptation phase include at least one of the following: the number of fast update rounds reaches the upper limit determined by the maximum adaptation delay; the decrease in verification window loss is less than the fast adaptation convergence threshold; the verification window loss is lower than the adaptation target loss threshold; and the update amount of structural cue parameters is lower than the cue parameter update threshold. In the basic implementation, the edge node determines the upper limit of the number of fast update rounds based on the maximum adaptation delay and the time consumed in a single update round, and stops updating when the upper limit of the number of fast update rounds is reached. In the optional implementation, if the decrease in verification window loss in two consecutive fast update rounds is less than the fast adaptation convergence threshold, the update is stopped early.

[0170] In some implementations, the following example may be used to illustrate the training process of a pre-trained temporal model.

[0171] Edge nodes or offline training servers collect 30 consecutive days of temperature, current, voltage, cavity pressure, and vibration channel data from the same type of semiconductor etching equipment, and simultaneously read the process stage, start / stop status, and load level. The historical operating data is aligned to a 1-second time granularity, and a historical time-series sample window is formed with 128 sampling points as the window length and 16 sampling points as the window step size. The actual sampled values ​​of cavity pressure, temperature, and vibration within 30 minutes after each historical time-series sample window are used as sensor prediction labels. Anomaly labels are generated from process alarm records and maintenance records. The historical time-series sample window is input into the pre-trained time-series model, and the pre-training loss is composed of the sensor prediction mean square error and anomaly identification cross-entropy. When the decrease in the validation set loss is less than the pre-training convergence threshold for five consecutive training rounds, training is stopped, and the main parameters of the pre-trained time-series model are saved.

[0172] In other implementations, the following example can be used to illustrate the fast adaptation training process on the edge side. After deploying the main parameters of the pre-trained temporal model, the edge node selects 32 verification windows under the same operating conditions during the current steady-state etch phase. The edge node determines the cue dimension based on the structural contribution of the initial structural cue vector on the verification windows and forms a structural cue vector to be updated. Subsequently, the edge node freezes the main parameters of the pre-trained temporal model and only updates the trainable cue parameters in the structural cue vector to be updated. If the maximum adaptation latency is 10 seconds and the single-round update time is 2 seconds, the maximum number of fast update rounds is 5 rounds. If the decrease in verification window loss in the 3rd and 4th rounds is less than the fast adaptation convergence threshold, the edge node stops updating early and outputs the updated structural cue vector.

[0173] See Figure 7 In step S105, the edge node receives the structure cue vector to be updated output in step S104 and freezes the main parameters of the pre-trained temporal model. The main parameters of the pre-trained temporal model include the input embedding layer parameters, the temporal coding layer parameters, the parameters in the structure cue access layer excluding the structure cue parameters, and the prediction output layer parameters. Freezing the main parameters of the pre-trained temporal model means setting the above parameters to an untrainable state so that they are not updated during backpropagation.

[0174] Edge nodes set the parameters corresponding to the cue dimension in the cue vector of the structure to be updated as trainable cue parameters. In the basic implementation, all parameters in the cue vector of the structure to be updated are used as trainable cue parameters. In an optional implementation, the edge nodes set only the cue parameters corresponding to the structures with the highest structural contribution as trainable cue parameters, while the cue parameters corresponding to the remaining structures remain unchanged. Thus, the cue dimension is further used in step S105 to limit the range of trainable cue parameters.

[0175] Edge nodes read the learning rate and structural importance coefficient based on the structure type. Statistical structure hints describe overall distribution changes and can have corresponding learning rates set; trend and periodic structure hints describe slow changes and rhythmic patterns and can also have corresponding learning rates set; temporal dependency structure hints describe dependencies between time points and channels and can be combined with the structural importance coefficient to suppress over-updates. The learning rates and structural importance coefficients are stored in a parameter configuration table.

[0176] No. The gradient ratio suggested by this structure can be expressed as:

[0177]

[0178] in, Indicates the first The gradient ratio suggested by the structure, Indicates the current structure type number. Indicates the structure type number participating in the summation. The loss function represents the loss function on the th The structure suggests the gradient of the parameters. The loss function represents the loss function on the th The structure suggests the gradient of the parameters. Indicates the number of structure types. express The 2-norm, express The 2-norm, This represents the smallest positive number that prevents the denominator from being zero.

[0179] No. The structural hint vector in the first The updated parameters can be represented as follows:

[0180]

[0181] in, Indicates the first The first update before Such structural prompt parameters, Indicates the first The 1st update Such structural prompt parameters, This indicates an update to the round number. Indicates the first The corresponding cue learning rate for each structure Indicates the first The structural importance coefficient corresponding to each structure Indicates the first The gradient ratio suggested by the structure, Indicates the first The gradient of the structure suggests the parameters.

[0182] In some implementations, the following example can be used to illustrate the process of updating trainable cue parameters. For instance, the edge node sets the learning rate for statistical structure cue to 1×10⁻⁴, the learning rate for trend structure cue to 5×10⁻⁵, the learning rate for periodic structure cue to 5×10⁻⁵, and the learning rate for time-dependent structure cue to 3×10⁻⁵. Simultaneously, the edge node sets the structure importance coefficient for statistical structure cue to 1.0, the structure importance coefficients for trend and periodic structure cue to 0.8, and the structure importance coefficient for time-dependent structure cue to 0.6.

[0183] If the gradient ratio of the trend structure hint in a single update is 0.35, the learning rate of the trend structure hint is 5 × 10^-5, and the structure importance coefficient is 0.8, then the edge nodes update the trend structure hint parameters according to the corresponding update formula. After the update, the edge nodes compare the parameter hash values ​​of the pre-trained temporal model's main parameters before and after the update. If the main parameter hash values ​​have not changed, the updated structure hint parameters are retained; if the main parameter hash values ​​have changed, the parameters before the update are restored, and an update anomaly log is recorded.

[0184] Step S105 outputs the updated structure cue vector. The updated structure cue vector serves as the input for step S106, which constructs the augmented model.

[0185] See Figure 8 In step S106, the edge node receives the current temporal sample window output in step S101 and the updated structure cue vector output in step S105. The edge node first converts the current temporal sample window into a temporal input vector, and then concatenates the updated structure cue vector to the front of the temporal input vector to form the enhanced model input.

[0186] In the basic implementation, edge nodes normalize and vectorize the sampled values ​​of each channel in the current time-series sample window to generate a time-series input vector. Normalization can use the mean and standard deviation from a historical stable window, or the minimum and maximum values ​​from a historical normal window. Vectorization can be performed by the input embedding layer of the pre-trained time-series model.

[0187] The input to the augmented model can be represented as:

[0188]

[0189] in, Indicates the input to the augmentation model. This represents a statistical structure hint vector. This indicates a trend structure cue vector. Represents a periodic structure hint vector. Represents the temporal dependency structure hint vector. Indicates the current time series sample window The corresponding time-series input vector, This indicates a sequential splicing operation.

[0190] In optional implementations, edge nodes project each structural cue vector to the same dimension as the temporal input vector and then add it element-wise to the temporal input vector; alternatively, the structural cue vector is used as an attention key prefix, enabling the pre-trained temporal model to call upon prior structural information when encoding the current temporal sample window. If the element-wise addition method is used, the edge nodes first perform dimensionality correction on the structural cue vectors using a linear projection matrix to ensure that the vector dimensions are consistent before addition.

[0191] After receiving input from the augmented model, the pre-trained time series model outputs sensor predictions for a future target duration. The target duration is determined based on the equipment maintenance response time and the rate of fault evolution. For example, for equipment requiring a 30-minute advance warning, the target duration can be set to 30 minutes; for equipment with a fast production cycle, the target duration can be set to several production cycles. Sensor predictions serve as one form of equipment status prediction.

[0192] When the pre-trained temporal model has a state output header, the device health and failure probability are directly output from the state output header. When the pre-trained temporal model only has sensor prediction outputs, the edge nodes generate the device health and failure probability based on the prediction residuals. The prediction residuals can be represented as the absolute difference or mean square error between the actual observations and the predicted values. The health status can be calculated based on the normalized prediction residuals, and the failure probability can be obtained from the anomaly scoring mapping function.

[0193] In one alternative implementation, the anomaly score can be represented as:

[0194]

[0195] in, This indicates the abnormal score at the current moment. This represents the score weight corresponding to the predicted residual. This indicates the scoring weight corresponding to the amount of decline in equipment health. This represents the scoring weight corresponding to the probability of failure. Represents the normalization function. Represents the actual observed value. Indicates the predicted value. Indicates the health status of the device. This represents the probability of failure. This formula is used to illustrate one quantitative implementation of anomaly scoring, but it does not require anomaly scoring to adopt this weighted form.

[0196] normalization function The scoring weights can be determined based on the minimum and maximum residuals of historical normal windows, or based on the mean and standard deviation of the residuals of historical normal windows. The scoring weights can be determined based on the impact of historical fault samples on the early warning accuracy, or configured by operations and maintenance personnel. To facilitate edge computing, the three scoring weights can be normalized. .

[0197] In some implementations, the following example may be used to illustrate the process of generating anomaly scores and predictive maintenance prompts.

[0198] For example, the edge node concatenates the updated structure cue vector with the temporal input vector corresponding to the current temporal sample window and inputs it into the pre-trained temporal model. The pre-trained temporal model outputs the predicted value of pressure channel PRES_01 for the next 30 minutes. If the actual observed value of pressure channel PRES_01 at a certain prediction time is 73.2 kPa and the predicted value is 69.6 kPa, then the prediction residual is 3.6 kPa. The edge node reads the minimum and maximum values ​​of the pressure prediction residual in the historical normal window and normalizes 3.6 kPa to 0.72.

[0199] If the pre-trained time series model outputs a device health score of 0.76 and a failure probability of 0.55, and the scoring weights are set as follows: , , The anomaly score at the current moment is If the warning threshold in the warning level table is 0.35 for Level 1, 0.50 for Level 2, and 0.70 for Level 3, then 0.573 reaches the Level 2 warning threshold but does not reach the Level 3 warning threshold, and the edge node generates a Level 2 warning result.

[0200] The edge node further calculates the predicted residual contribution for each channel. If the residual contribution of pressure channel PRES_01 is ranked first, and the trend structure contribution is ranked first, then the edge node identifies PRES_01 as the main abnormal channel, determines the source of the abnormality as the increase in prediction deviation caused by the rising pressure trend, and outputs predicted maintenance prompts based on the maintenance suggestion mapping table to prompt inspection of the chamber pressure control loop, pressure valve opening feedback, and related actuators. The predicted maintenance prompts are sent to the maintenance terminal and simultaneously written to the adaptation log table.

[0201] See Figure 9After the prediction maintenance prompt results are generated, the edge node continuously receives new multi-source runtime sequence data and writes the new window prediction residual, the current operating condition stage, and the verification loss into the adaptation log. When the current operating condition stage changes, the prediction residual of multiple consecutive windows reaches the re-adaptation threshold, or the verification loss increases continuously, the edge node triggers the re-adaptation process.

[0202] The refit threshold can be determined based on the distribution of predicted residuals from historical normal windows. For example, the refit threshold can be taken as the 97th percentile of the predicted residuals from historical normal windows. The number of consecutive windows is set according to the equipment control cycle and false alarm tolerance, for example, 3 to 10 windows. Changes in the current operating condition stage can be determined by the process stage field, load level field, or production batch field reported by the equipment controller. A continuous increase in verification loss can be represented by the verification loss of the most recent verification windows being greater than the verification loss of the previous verification window.

[0203] In the refitting process, the edge nodes use the structural cue vector and cue dimension obtained from the previous fit as the initial result and re-execute steps S102 to S105. Specifically, the edge nodes re-perform structural decomposition on the sample window of the most recent time period to obtain a new set of structural features; they re-select the structural cue template according to the current working condition stage to generate a new initial structural cue vector; they redetermine the cue dimension according to the structural contribution on the same working condition verification window to form a new structural cue vector to be updated; and then they only update the trainable cue parameters in the new structural cue vector to be updated to obtain the refitted structural cue vector.

[0204] If the verification loss after re-adaptation is less than the rollback threshold, the edge node retains the re-adapted structure hint vector; if the verification loss after re-adaptation reaches the rollback threshold, the edge node restores the structure hint vector of the previous version. The rollback threshold can be set to 1.05 to 1.20 times the verification loss of the previous version. The edge node also records the working stage, triggering reason, verification loss before and after the update, rollback time, and hint parameter version, forming a re-adaptation log.

[0205] In some implementations, the following example can be used to illustrate the refitting process. For instance, if an edge node detects that the prediction residuals have all reached the refitting threshold within the last 5 sample windows, or if the device controller switches from the steady-state etching stage to the cleaning stage in its reporting phase, the edge node triggers the refitting process. The edge node uses the previous version's structure cue vector and cue dimension as initial results and selects the last 48 sample windows backward from the current time to perform refitting.

[0206] If the verification loss of the previous version was 0.118, and the rollback threshold is set to 1.10 times the verification loss of the previous version, then the rollback threshold is 0.130. If the verification loss after refit is 0.121, then 0.121 is less than the rollback threshold of 0.130, the edge node retains the refitted structure hint vector, and increments the hint parameter version number by 1. If the verification loss after refit is 0.134, then 0.134 reaches the rollback threshold of 0.130, the edge node restores the previous version's structure hint vector, and records the rollback time, rollback reason, and verification loss before and after the update.

[0207] See Figure 10 In one device-based implementation, the edge-side rapid adaptation program may include an acquisition and caching unit, a window construction unit, a structure decomposition unit, a prompt generation unit, a dimension determination unit, a model training and version management unit, a prompt update unit, and a prediction maintenance output unit. These functional units can be executed by the same edge processor, or they can be executed collaboratively by a microcontroller, industrial computer, graphics processor, or neural network acceleration unit.

[0208] The acquisition and caching unit receives multi-source runtime sequence data, equipment operating condition information, and resource constraint information from industrial equipment and writes them to the edge cache queue. The window construction unit performs alignment, marking, padding, and segmentation on the multi-source runtime sequence data in the edge cache queue and outputs the current time-series sample window. The structure decomposition unit performs statistical, trend, periodic, and time-dependent structure decomposition on the current time-series sample window and outputs a set of structural features.

[0209] The prompt generation unit determines the current operating condition stage and structural prompt template based on equipment operating condition information, writes the structural feature set into the structural prompt template, and outputs the initial structural prompt vector. The dimension determination unit filters the verification windows consistent with the current operating condition stage, determines the prompt dimension based on structural contribution and resource constraint information, and outputs the structural prompt vector to be updated.

[0210] The model training and version management unit manages historical time-series sample windows, training labels, pre-trained time-series model parameters, and model version information, and saves the pre-trained time-series model parameters after training is complete. The cue update unit freezes the pre-trained time-series model parameters, updates the trainable cue parameters in the structure cue vector to be updated, and outputs the updated structure cue vector. The prediction and maintenance output unit constructs the input for the augmented model and outputs device status predictions, anomaly scores, and prediction and maintenance cue results.

[0211] In a specific application example, the industrial equipment is a semiconductor manufacturing device, and the edge node is deployed on the device's local industrial computer. The edge node receives temperature, current, voltage, cavity pressure, and vibration channel data every second. The window length is set to 128 sampling points, and the window step size is set to 16 sampling points. After performing structural decomposition on the current time-series sample window, the edge node generates statistical structural features, trend structural features, periodic structural features, and time-dependent structural features, respectively.

[0212] During the pre-training phase, the offline training server collects 30 consecutive days of historical operating data from the same type of semiconductor manufacturing equipment and forms a historical time-series sample window using the same data processing method as in step S101. The actual sampled values ​​of cavity pressure, temperature, and vibration within 30 minutes after the historical time-series sample window are used as prediction labels, and anomaly labels are generated using process alarm records and maintenance records. The offline training server trains the pre-trained time-series model based on the above training inputs and training labels, and saves the main parameters of the pre-trained time-series model after meeting the convergence conditions. After the edge node loads the main parameters of the pre-trained time-series model, it executes the edge-side fast adaptation process in this embodiment.

[0213] The edge node determines the current operating stage as the steady-state etching stage based on the equipment operating condition information and selects the corresponding structural hint template. Subsequently, the edge node writes the set of structural features into the structural hint template, generating an initial structural hint vector. The edge node then selects a verification window consistent with the steady-state etching stage and determines the hint dimension based on the structural contribution and resource constraint information on the verification window, forming the structural hint vector to be updated according to the hint dimension.

[0214] The edge nodes freeze the main parameters of the pre-trained temporal model and only update the trainable cue parameters in the structure cue vector to be updated, resulting in the updated structure cue vector. After the update, the edge nodes convert the current temporal sample window into a temporal input vector and concatenate the updated structure cue vector to the front of the temporal input vector to form the augmented model input. The pre-trained temporal model outputs the predicted values ​​of key sensors, equipment health, anomaly scores, and predicted maintenance cue results for the future target time period based on the augmented model input.

[0215] As another application example, this method can also be used for rotating machinery industrial equipment. For example, in the case of an industrial pump set, edge nodes collect data on pump body vibration, bearing temperature, motor current, outlet pressure, and flow channels. The edge nodes set a uniform time granularity of 2 seconds, a window length of 180 sampling points, and a future target duration of 20 minutes based on the equipment control cycle. If the vibration channel shows an increase in periodic amplitude within the current window, the bearing temperature channel shows an increase in trend slope, and the cross-channel correlation coefficient between the motor current channel and the outlet pressure channel decreases, then the edge nodes generate periodic structural features, trend structural features, and time-dependent structural features, respectively.

[0216] During the determination of the cue dimension, if the periodic structure cue brings the largest reduction in verification loss on the verification window, then the edge node prioritizes allocating a unit dimension increment to the initial structure cue vector corresponding to the periodic structure, and forms the structure cue vector to be updated according to the updated cue dimension. After completing the rapid adaptation, if the anomaly score reaches the second-level warning threshold, and the vibration channel residual contribution is ranked first, the edge node outputs the predicted maintenance cue result to indicate whether the bearing is worn, the rotor is unbalanced, or the pump body is loose.

[0217] The above application examples illustrate that, under the condition of limited resources at edge nodes, this application can achieve rapid adaptation of pre-trained temporal models through a continuous processing chain of structural feature sets, initial structural cue vectors, structural cue vectors to be updated, and updated structural cue vectors, and improve the real-time performance of prediction and maintenance results.

[0218] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims.

Claims

1. A fast adaptation method based on edge-side temporal models, characterized in that, Applications include edge nodes on the industrial equipment side, including: Step S101: Obtain multi-source runtime sequence data, equipment operating condition information and resource constraint information of industrial equipment in the current maintenance cycle; align, mark, complete and segment the multi-source runtime sequence data to obtain the current time series sample window. Step S102: Perform statistical, trend, periodic, and time-dependent structural decomposition on the current time series sample window to obtain a set of structural features; Step S103: Determine the current operating stage and structural prompt template based on the equipment operating condition information, write the structural feature set into the structural prompt template, and generate the initial structural prompt vector; Step S104: Filter the verification window that is consistent with the current operating stage according to the equipment operating condition information, and determine the prompt dimension according to the structural contribution and resource constraint information of the initial structural prompt vector on the verification window, and form the structural prompt vector to be updated according to the prompt dimension. Step S105: Freeze the main parameters of the pre-trained temporal model, update the trainable cue parameters in the structure cue vector to be updated, and generate the updated structure cue vector. Step S106: Convert the current time series sample window into a time series input vector, concatenate the updated structure prompt vector to the front of the time series input vector to form an enhanced model input and input it into the pre-trained time series model to obtain the equipment status prediction, anomaly score and predicted maintenance prompt results.

2. The fast adaptation method based on edge-side temporal model as described in claim 1, characterized in that, Step S101 involves acquiring multi-source runtime sequence data, equipment condition information, and resource constraint information of the industrial equipment during the current maintenance cycle. This includes: reading temperature, current, voltage, pressure, vibration, speed, and flow channel data through the equipment controller, sensor acquisition interface, or industrial communication bus, and recording each channel data as a data item including channel identifier, sampling timestamp, sampling value, and communication status; reading the equipment number, process stage, start / stop status, and load level from the equipment controller to generate equipment condition information; reading available memory, available computing power, maximum adaptation latency, and cache capacity from the edge node runtime monitoring program to generate resource constraint information; and associating the multi-source runtime sequence data, equipment condition information, and resource constraint information to form an edge cache record, which is used for window generation in step S101 and further called in steps S103 to S105.

3. The fast adaptation method based on edge-side temporal model as described in claim 2, characterized in that, Step S101 involves aligning, marking, padding, and segmenting the multi-source runtime sequence data, including: The edge nodes determine a unified time granularity based on the device control cycle and the effective sampling cycle of each channel in the edge cache records, and generate a unified time axis based on the unified time granularity; The data from each channel is mapped to a unified time axis according to the sampling timestamp, and quality markers are generated based on whether the sampled value is empty, whether it exceeds the channel range, whether the communication status is interrupted, and the start / stop status. When the quality marker indicates a missing value and the duration of the missing value does not exceed the upper limit of missing value completion, the missing sample value is completed by forward hold or the nearest mean. Then, based on the cache capacity and maximum adaptation latency in the resource constraint information, determine the window length and window step size, and truncate continuous windows according to the window length and window step size to generate the current time series sample window.

4. The fast adaptation method based on edge-side temporal model according to claim 1, characterized in that, Methods for obtaining a set of structural features include: The edge node reads the sampled values ​​and quality markers corresponding to each sensor channel in the current time-series sample window, first removes the stop marker positions and unusable fill positions, and then calculates the mean, variance, skewness and kurtosis of the remaining sampled values ​​to form statistical structure features; Perform moving average, linear regression, or first-order differencing on the retained sample values ​​to form trend structure characteristics; After removing the trend component from each channel window sequence, frequency domain transformation or periodic peak search is performed on the residual sequence to form periodic structure features; The autocorrelation coefficient is calculated according to the preset lag order and combined with the cross-channel correlation coefficient to form the time-dependent structural characteristics. Each structural feature is combined with the structural type, channel identifier, window number, and feature value to form a set of structural features, including statistical structural features, trend structural features, and periodic structural features.

5. The fast adaptation method based on edge-side temporal model according to claim 1, characterized in that, Step S103 determines the current operating stage and structural prompt template based on the equipment operating condition information, including: Read the process stage, start / stop status and load level from the equipment operating condition information, and determine the current operating condition stage based on the process stage; Select the structural prompt template that corresponds to the current working condition stage and structural type from the structural prompt template table; Following the order of data items in the structure prompt template, write the structure type, channel identifier, feature name and feature value in the structure feature set into the corresponding positions, and add the equipment number, current operating condition stage, load level and prediction task identifier into the structure prompt template to obtain the structure prompt record; Perform a hint encoding mapping on the structure hint record to generate an initial structure hint vector.

6. The fast adaptation method based on edge-side temporal model according to claim 1, characterized in that, Step S104 involves filtering verification windows that match the current operating condition stage based on equipment operating condition information, including: Read the historical sample window within the current maintenance cycle and compare the operating condition stage corresponding to the historical sample window with the current operating condition stage; The validation window is obtained by retaining historical sample windows that are consistent with the operating conditions and have a window missing ratio less than the missing ratio threshold, a noise amplitude less than the noise amplitude threshold, and a number of key channels not less than the number of key channels threshold. Generate status labels or weak labels for the verification window based on equipment maintenance records, process alarm records, or predicted residual thresholds; The verification window and its status label or weak label are used to calculate the structural contribution corresponding to the initial structural cue vector.

7. The fast adaptation method based on edge-side temporal model according to claim 1, characterized in that, Step S104 involves determining the cue dimension and forming the structure cue vector to be updated according to the cue dimension, including: The upper limit of the total dimension of the prompt is determined based on the input dimension of the pre-trained temporal model, available memory, available computing power, and maximum adaptation latency; Assign initial cue dimensions to the initial structure cue vectors corresponding to statistical structure, trend structure, periodic structure, and time-dependent structure; Each initial structural cue vector is appended with a unit dimension increment to form a candidate cue dimension scheme; Calculate the candidate validation loss corresponding to the candidate cue dimension scheme on the validation window, and take the difference between the baseline validation loss and the candidate validation loss as the corresponding structural contribution. The cue dimension of each initial structural cue vector is determined based on the structural contribution, and the initial structural cue vectors are truncated, linearly compressed, or linearly expanded according to the cue dimension to form the structural cue vector to be updated.

8. The fast adaptation method based on edge-side temporal model according to claim 1, characterized in that, Step S105 involves freezing the main parameters of the pre-trained temporal model and updating the trainable cue parameters in the structure cue vector to be updated, including: Edge nodes set the corresponding parameters of the input embedding layer, temporal coding layer, and prediction output layer of the pre-trained temporal model to an untrainable state; The parameters in the structure cue vector to be updated that are limited by the cue dimension are determined as trainable cue parameters; the cue learning rate and structure importance coefficient are read from the parameter configuration table according to the structure type; Calculate the adaptation loss based on at least one of the prediction loss, anomaly recognition loss, or cue parameter regularization term on the verification window; The trainable cue parameters are updated based on the adaptation loss, cue learning rate, and structural importance coefficient, generating an updated structural cue vector while keeping the main parameters of the pre-trained temporal model unchanged.

9. The fast adaptation method based on edge-side temporal model according to claim 1, characterized in that, Step S106 involves forming the input to the enhanced model and inputting it into the pre-trained temporal model, including: The sampled values ​​of each channel in the current time series sample window are normalized and vectorized to generate a time series input vector; The updated structure hint vectors are arranged in a fixed order of statistical structure, trend structure, periodic structure, and time-dependent structure. The arranged structural cue vector is concatenated to the beginning of the temporal input vector to obtain the enhanced model input; The augmented model input is used as the input data for the pre-trained time series model. The pre-trained time series model outputs at least one of the following: sensor prediction values, equipment health, or failure probability within the future target time period. Based on the prediction residual between the sensor prediction values ​​and the actual observation values, the equipment health, or the failure probability, the equipment status prediction and anomaly score are generated.

10. The fast adaptation method based on edge-side temporal model according to claim 1, characterized in that, After receiving the predicted maintenance prompts, the edge nodes continuously receive new multi-source runtime sequence data and write the new window prediction residuals, the current operating stage, and the verification loss into the adaptation log. When the current working condition changes, the prediction residuals of multiple consecutive windows reach the refit threshold, or the validation loss increases continuously, the edge node uses the structural cue vector and cue dimension obtained from the previous fit as the initial result and re-executes steps S102 to S105. When the verification loss after re-adaptation is less than the rollback threshold, the re-adapted structure hint vector is retained; when the verification loss after re-adaptation reaches the rollback threshold, the previous version's structure hint vector is restored, and the triggering reason, rollback time, and verification loss before and after the update are recorded in the adaptation log.