A speed reducer fault prediction method and system based on an industrial internet
Patent Information
- Application Number
- CN202611003395.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-10-02
AI Technical Summary
[0005]本申请提出一种基于工业互联网的减速机故障预测方法及系统,旨在解决在工业互联网边缘节点定时接收在线修正后的减速机故障预测模型时,因新模型激活瞬间丢失旧模型维持的时序状态导致输出预测值序列产生突变,使得闭环修正评估机制误将本次有效模型更新判定为性能恶化,从而错误触发模型回滚或过度修正,干扰了模型闭环演化过程的稳定性的技术问题
本申请通过同步采集旧模型的时序状态并构建上下文状态字典,为新模型提供了切换前的历史语境,避免了因新模型激活瞬间丢失旧模型维持的时序状态导致输出预测值序列产生突变;通过精准识别冷启动导致的时序状态断裂区间并将其标记为无效样本加以剔除,避免了闭环修正评估机制将瞬态偏差误判为性能恶化,从而防止了错误触发模型回滚或过度修正;通过时序上下文的强制对齐与承接补偿,提升了新模型在切换初期的预测稳定性,减少了因冷启动效应导致的系统干扰。本方法维护了工业互联网环境下模型闭环演化过程的稳定性,提高了减速机故障预测模型的更新效率和可靠性。
Smart Images

Figure CN122862928A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fault prediction technology, and in particular to a method and system for predicting gearbox faults based on the Industrial Internet. Background Technology
[0002] In continuous production within the high-end equipment manufacturing sector, gearbox failure prediction typically relies on an industrial internet architecture. The cloud platform uses recently collected operational data to continuously revise the prediction model online and distributes the new model to the edge computing gateway. While maintaining uninterrupted gearbox failure prediction, edge nodes receive and activate the new model, replacing the old model to output new health status predictions. A closed-loop correction and evaluation mechanism continuously monitors the model update effectiveness.
[0003] However, in scenarios involving model deployment and switching after conventional timing model corrections, when the reducer on the production line is continuously running and sensor data streams are constantly flowing in, the activation transition moment when the edge node loads the new model and replaces the old model's inference channel is often difficult to achieve in absolute silence. The new model typically needs to initialize its internal state at activation, such as the hidden state of a recurrent neural network or a sliding window for statistical features. At this moment, the temporal state maintained by the old model and synchronized with the current physical process is instantly lost. This state loss causes the new model to produce abrupt changes in its output prediction sequence when processing the first batch of data packets after the switch, clearly inconsistent with the previous continuous trend of the old model.
[0004] Existing closed-loop correction evaluation mechanisms typically rely on real-time captured prediction errors for judgment. This makes it difficult to identify sudden output jumps during switching as transient deviations caused by model initialization, instead interpreting them directly as deterioration in the predicted performance of the corrected model. This evaluation method is prone to errors when industrial internet edge nodes periodically receive online-corrected gearbox fault prediction models. The loss of the temporal state maintained by the old model upon activation of the new model causes abrupt changes in the output prediction value sequence. This leads the closed-loop correction evaluation mechanism to mistakenly classify the effective model update as performance degradation, incorrectly triggering model rollback or over-correction, thus interfering with the stability of the model's closed-loop evolution process. Summary of the Invention
[0005] This application proposes a method and system for predicting gearbox failures based on the Industrial Internet. It aims to solve the technical problem that when the edge node of the Industrial Internet receives the online-corrected gearbox failure prediction model at regular intervals, the output prediction value sequence will change abruptly due to the loss of the time sequence state maintained by the old model at the moment the new model is activated. This causes the closed-loop correction evaluation mechanism to mistakenly judge the current effective model update as performance degradation, thereby erroneously triggering model rollback or over-correction, which interferes with the stability of the model's closed-loop evolution process.
[0006] Firstly, this application provides a method for predicting gearbox failures based on the Industrial Internet, applied to an edge node deployed on the gearbox side. The edge node is connected to a cloud platform via the Industrial Internet to collect real-time gearbox operating data and run a prediction model. The method includes the following steps: The system synchronously collects the output prediction values, input data sequences, and time-series state vectors of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and constructs a context state dictionary to describe the current operating context of the reducer; wherein, the old model is the prediction model that is currently running and used to predict reducer faults, and the new model is the prediction model that is downloaded from the cloud and used to replace the old model; The system receives a new model from the cloud platform, constructs a shadow inference channel that runs parallel to the old model, mirrors the real-time sensor data stream and simultaneously sends it into both the old and new models, and obtains the output prediction sequence of each model. Based on the output prediction sequence, under the condition that the volatility of the input data sequence is lower than the preset physical mutation threshold and the output prediction of the old model is stable, if the error between the output prediction of the new model and the old model first exceeds the preset transient tolerance limit within the preset observation period and then shows a monotonically decreasing trend, then the time series state break interval caused by cold start is determined. Based on the temporal state vector stored in the context state dictionary, the new model is subjected to forced alignment and continuity compensation of the temporal context: if the new model has the same structure as the old model, the temporal state vector of the old model is directly injected into the new model; otherwise, the previous historical input data is switched and replayed to the new model at a rate higher than real-time for warm-up, and within the temporal state break interval, an output smoothing fusion mechanism is adopted to gradually transition from the output prediction value of the old model to the output prediction value of the new model. The identified time-series state break intervals are marked as invalid samples and removed using a time mask. The actual performance is verified only based on the steady-state output to obtain the verification results. Based on the verification results, if the new model outperforms the old model, the output channel for gearbox fault prediction will be switched to the new model and the old model will be shut down; otherwise, an abnormal rollback mechanism will be initiated to terminate the inference channel of the new model and generate a diagnostic report to be uploaded to the cloud platform.
[0007] According to some embodiments of this application, the step of synchronously acquiring the output prediction values, input data sequences, and time-series state vectors of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and constructing a context state dictionary to describe the current operating context of the reducer includes: Within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, a circular buffer queue of preset length is opened in the memory of the edge node; wherein, the preset length is set according to the temporal receptive field of the old model or the maximum memory depth of the recurrent neural network, and is used to cache the input data sequence within the preset historical time period, so as to provide the historical input data required for high-speed playback warm-up of the new model in subsequent steps. When the old model outputs a predicted value, the predicted value of the old model at the current moment, the input data sequence, and the temporal state vector are bound together, and a context state dictionary covering all inference cycles within the entire preset historical duration is constructed.
[0008] According to some embodiments of this application, the steps of receiving the new model from the cloud platform, constructing a shadow inference channel parallel to the old model, mirroring the real-time sensor data stream and simultaneously sending it into both the old and new models to obtain the output prediction value sequences of the old and new models respectively include: Receive the weight and network structure description file of the new model issued by the cloud platform through a security protocol; Based on the weights and network structure description file, a new model is loaded into the independent memory space and computing resource isolation area of the edge node to construct a shadow inference channel that runs parallel to the old model. The real-time sensor data stream is copied into two copies. One copy is sent to the old model to maintain the prediction output, and the other copy is sent to the shadow inference channel of the new model simultaneously, so as to obtain the output prediction value sequence of the old model and the new model respectively.
[0009] According to some embodiments of this application, the step of determining the time series state break interval caused by cold start, based on the output predicted value sequence, under the condition that the volatility of the input data sequence is lower than a preset physical mutation threshold and the output predicted value of the old model is stable, if the error between the output predicted values of the new model and the old model first exceeds a preset transient tolerance upper limit within a preset observation period and then shows a monotonically decreasing trend, includes: Calculate the volatility of one of the following: mean, variance, and first difference of the input data sequence over the most recent preset number of sampling periods; If the volatility is lower than a preset physical mutation threshold, the sliding window standard deviation or the absolute value of the first difference of the old model's output prediction value within the most recent preset number of sampling periods is calculated as the smoothness of the old model's output prediction value. If the smoothness is greater than the preset stability threshold, then the absolute value of the difference between the output prediction values of the new model and the old model at the same input time is calculated as the absolute error between the output prediction values of the new model and the old model. If the absolute error exceeds the preset transient tolerance limit within the preset observation period after the new model starts receiving real-time data streams, and the absolute error shows a monotonically decreasing trend over time after exceeding the preset transient tolerance limit and before the end of the preset observation period, then the moment when the absolute error first exceeds the preset transient tolerance limit is taken as the starting point and the moment when the preset observation period ends is taken as the ending point, which is taken as the time-series state break interval under the stable operating condition caused by cold start.
[0010] According to some embodiments of this application, the step of determining the time series state break interval caused by cold start, based on the output predicted value sequence, under the condition that the volatility of the input data sequence is lower than a preset physical mutation threshold and the output predicted value of the old model is stable, if the error between the output predicted values of the new model and the old model first exceeds a preset transient tolerance upper limit within a preset observation period and then shows a monotonically decreasing trend, further includes: If the volatility is higher than or equal to the preset physical mutation threshold, then determine whether the output prediction value of the old model triggers the preset alarm threshold. If not triggered, the absolute difference between the first-order difference of the predicted values of the new model and the old model is calculated as the relative logical deviation. If the relative logical deviation exceeds the preset transient tolerance limit within the preset observation period after the new model starts receiving real-time data streams, and the relative logical deviation shows a monotonically decreasing trend over time after exceeding the preset transient tolerance limit and before the end of the preset observation period, then the moment when the relative logical deviation first exceeds the preset transient tolerance limit is taken as the starting point and the moment when the preset observation period ends is taken as the ending point, which is taken as the time-series state break interval under the perturbation condition caused by cold start. If it has been triggered, it is determined to be a real device anomaly. The timing state break interval is not marked, and the old model output prediction value is kept as the current prediction result.
[0011] According to some embodiments of this application, in the step of warming up the new model by replaying historical input data from a period prior to the switch at a rate higher than real-time: If a time-series state break interval caused by cold start is identified, all historical input data sequences within the preset historical time before switching will be continuously input into the new model at a playback speed higher than the real-time data stream rate for warm-up. If a time-series state break interval under the perturbation condition caused by cold start is identified, the historical input data sequence containing the input data corresponding to the time-series state break interval under the perturbation condition is continuously input into the new model at a playback speed higher than the real-time data stream rate for warm-up.
[0012] According to some embodiments of this application, the output smoothing fusion mechanism includes: Within the time-series state break interval, weights that vary with time are constructed; The sum of the weights of the output predictions of the new model and the old model is always 1. At the initial moment, the weight of the output prediction of the new model is a preset small weight, and the weight of the output prediction of the old model is 1 minus the preset small weight. As time goes by, the weight of the output prediction of the new model increases linearly or exponentially, and the weight of the output prediction of the old model decreases accordingly, until the weight of the output prediction of the new model reaches its maximum value at the end of the time-series state break interval.
[0013] According to some embodiments of this application, the steps of marking the identified time-series state break intervals as invalid samples, removing them using a time mask, and performing real performance verification based solely on steady-state output to obtain the verification results include: The identified temporal state break intervals are marked as invalid samples; When performing real-world performance verification on the new model, a time mask is applied to remove invalid samples from the performance verification calculation. Based solely on the output prediction value sequence after the end of the time-series state break interval, the prediction accuracy, false alarm rate, and fault warning advance of the new and old models within the same preset verification time after the end of the time-series state break interval are calculated as performance verification parameters. If the new model has a higher prediction accuracy than the old model, a lower false alarm rate, and a greater early warning time for faults, then the new model is considered to outperform the old model; otherwise, the new model is considered to not outperform the old model.
[0014] According to some embodiments of this application, the steps of switching the output channel of the reducer fault prediction to the new model and shutting down the old model based on the verification results, and otherwise initiating an abnormal rollback mechanism to terminate the inference channel of the new model and generate a diagnostic report for uploading to the cloud platform, include: If the verification result shows that the new model performs better than the old model, then perform an atomic switching operation on the output channel of the reducer fault prediction to switch from the old model to the new model. After switching to the new model, if the standard deviation of the sliding window of the output predicted value sequence within the preset confirmation time is lower than the preset stability threshold, it is determined that the new model takeover is stable, the inference process of the old model is closed, and the memory and computing resources occupied by the old model are released; otherwise, it is determined that the new model takeover is abnormal, the inference channel of the new model is terminated, the output of the old model is restored, and a diagnostic report containing the reason for the takeover failure is generated and uploaded to the cloud platform. If the verification result shows that the performance of the new model is not better than that of the old model, then the shadow inference channel of the new model is terminated, the output of the old model is maintained, and a diagnostic report containing the reason for the switching failure, transient deviation characteristics and steady-state error comparison is generated and uploaded to the cloud platform.
[0015] Secondly, this application also provides a speed reducer fault prediction system based on the Industrial Internet, applied to an edge node deployed on the speed reducer side. The edge node is connected to a cloud platform via the Industrial Internet to collect speed reducer operating data in real time and to run a prediction model. The system includes: The data acquisition module is used to synchronously collect the output prediction values, input data sequences, and time-series state vectors of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and to construct a context state dictionary to describe the current operating context of the reducer; wherein, the old model is the prediction model that is currently running and used to predict reducer faults, and the new model is the prediction model that is downloaded from the cloud and used to replace the old model; The parallel inference module is used to receive new models from the cloud platform, build a shadow inference channel that runs parallel to the old model, mirror the real-time sensor data stream and send it into both the old and new models simultaneously, and obtain the output prediction value sequences of the old and new models respectively. The interval identification module is used to determine the time series state break interval caused by cold start based on the output prediction value sequence, under the condition that the volatility of the input data sequence is lower than the preset physical mutation threshold and the output prediction value of the old model is stable. If the error between the output prediction values of the new model and the old model first exceeds the preset transient tolerance upper limit within the preset observation time and then shows a monotonically decreasing trend. The context acceptance processing module is used to perform forced alignment and acceptance compensation of the temporal context on the new model based on the temporal state vector stored in the context state dictionary: if the new model has the same structure as the old model, the temporal state vector of the old model is directly injected into the new model; otherwise, the previous historical input data is switched and replayed to the new model at a rate higher than real time for warm-up, and within the temporal state break interval, an output smoothing fusion mechanism is adopted to gradually transition from the output prediction value of the old model to the output prediction value of the new model. The performance parameter acquisition module is used to mark the identified time-series state break intervals as invalid samples and remove them using a time mask. It performs real performance verification based only on steady-state output to obtain the verification results. The model switching module is used to switch the output channel of the reducer fault prediction to the new model and shut down the old model if the new model performs better than the old model based on the verification results; otherwise, it initiates an abnormal rollback mechanism, terminates the inference channel of the new model, and generates a diagnostic report to upload to the cloud platform.
[0016] The technical solution according to the embodiments of this application has at least the following beneficial effects: This application provides the new model with a historical context prior to the switch by synchronously collecting the temporal state of the old model and constructing a contextual state dictionary. This avoids abrupt changes in the output prediction sequence caused by the loss of the temporal state maintained by the old model upon the activation of the new model. By accurately identifying and removing temporal state breaks caused by cold starts, the closed-loop correction evaluation mechanism is prevented from misjudging transient deviations as performance degradation, thus preventing erroneous model rollback or over-correction. Through forced alignment and continuation compensation of the temporal context, the prediction stability of the new model in the initial stage of the switch is improved, and system interference caused by the cold start effect is reduced. This method maintains the stability of the model's closed-loop evolution process in the industrial internet environment and improves the update efficiency and reliability of the reducer fault prediction model.
[0017] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0018] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.
[0019] Figure 1 This is a flowchart illustrating a method for predicting gearbox faults based on the Industrial Internet, provided in an embodiment of this application.
[0020] Figure 2 This is a flowchart illustrating S120 in an embodiment of this application.
[0021] Figure 3 This is a flowchart illustrating S150 in an embodiment of this application.
[0022] Figure 4 This is a schematic diagram of the architecture of a speed reducer fault prediction system based on the Industrial Internet, provided for an embodiment of this application. Detailed Implementation
[0023] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0024] Traditional methods for predicting gearbox failures suffer from a significant loss of temporal state synchronization between the old and current physical processes during model updates and switching. This loss occurs because the new model typically initializes its internal state upon activation, causing a sudden change in the output prediction sequence when processing the initial data packets after the switch. This result in a clear inconsistency with the previous continuous trend of the old model. Existing closed-loop correction and evaluation mechanisms struggle to recognize this sudden output jump as a transient deviation caused by model initialization. Instead, they interpret it directly as a deterioration in the predicted performance of the corrected model, leading to erroneous model rollback or over-correction, severely disrupting the stability of the model's closed-loop evolution process.
[0025] In this regard, such as Figure 1 As shown, this application discloses a speed reducer fault prediction method based on the Industrial Internet, applied to an edge node deployed on the speed reducer side. The edge node is connected to a cloud platform via the Industrial Internet to collect speed reducer operating data in real time and to run a prediction model. The method includes the following steps: S110, synchronously collect the output prediction value, input data sequence and time-series state vector of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and construct a context state dictionary to describe the current operating context of the reducer; wherein, the old model is the prediction model that is currently running and used to predict reducer faults, and the new model is the prediction model that is downloaded from the cloud and used to replace the old model. S120 receives the new model from the cloud platform, constructs a shadow inference channel that runs parallel to the old model, mirrors the real-time sensor data stream and sends it into both the old and new models simultaneously, and obtains the output prediction value sequences of the old and new models respectively. S130, based on the output prediction value sequence, under the condition that the volatility of the input data sequence is lower than the preset physical mutation threshold and the output prediction value of the old model is stable, if the error between the output prediction values of the new model and the old model first exceeds the preset transient tolerance upper limit within the preset observation period and then shows a monotonically decreasing trend, then the time series state break interval caused by cold start is determined. S140, based on the temporal state vector stored in the context state dictionary, perform forced alignment and continuity compensation of the temporal context on the new model: if the new model has the same structure as the old model, the temporal state vector of the old model is directly injected into the new model; otherwise, the previous historical input data is switched and replayed to the new model at a rate higher than real-time for warm-up, and within the temporal state break interval, an output smoothing fusion mechanism is adopted to gradually transition from the output prediction value of the old model to the output prediction value of the new model. S150 marks the identified time-series state break intervals as invalid samples and applies a time mask to remove them. The actual performance is verified only based on the steady-state output to obtain the verification results. S160, based on the verification results, if the performance of the new model is better than that of the old model, the output channel of the reducer fault prediction is switched to the new model and the old model is turned off; otherwise, the abnormal rollback mechanism is started, the inference channel of the new model is terminated and a diagnostic report is generated and uploaded to the cloud platform.
[0026] Edge nodes, as a crucial component of the Industrial Internet architecture, are responsible for data collection, processing, and model inference close to the data source, thereby reducing network latency and improving response speed. The Industrial Internet provides a secure and reliable data transmission and command delivery channel between edge nodes and the cloud platform.
[0027] Before switching prediction models, edge nodes can continuously record the prediction results of the old model over a period of time (e.g., 1 hour), raw sensor input data (such as vibration signals, temperature signals, torque signals, and acoustic emission signals collected from the reducer via vibration sensors, temperature sensors, torque sensors, and acoustic emission sensors), and the temporal state of the model (such as the hidden layer state of a recurrent neural network). This data is organized into a dictionary structure, where each time step's data is an entry containing the corresponding output prediction value (e.g., the reducer's health status assessment value, remaining service life prediction value, and potential fault alarm information), the input data sequence, and the temporal state vector. This contextual state dictionary aims to capture the reducer's operating mode and model state before the switch, providing necessary historical information for the cold start of the new model.
[0028] Once the cloud platform completes the training or optimization of the new model, it will distribute the new model to the edge nodes. Upon receiving the new model, the edge nodes will not immediately replace the old model. Instead, they will create an independent runtime environment and load the new model. At this time, the real-time data stream collected from the reducer sensors will be copied. One copy will be fed into the old model for prediction, while the other will be fed into the new model for parallel prediction. This allows for simultaneous acquisition of the prediction outputs of both the old and new models under the same real-time input, facilitating comparison and evaluation.
[0029] By comparing the output prediction sequences of the old and new models during parallel inference, the error between them can be calculated. If the input data sequence (such as vibration signals, temperature signals, etc.) shows little change, and the predictions of the old model are relatively stable, but there is a large initial error between the predictions of the new model and the old model, and this error gradually decreases over a period of time, this usually indicates that the new model is gradually adapting to the current operating conditions from a cold start state. In this case, a time interval can be determined as the time-series state break interval caused by the cold start.
[0030] If both the old and new models are recurrent neural networks of the same type and have the same hidden layer dimension, then the last temporal state vector of the old model before the switch can be directly copied and set as the initial temporal state of the new model. If the old and new models have different structures and cannot be directly injected with temporal states, then previously collected historical input data can be used to input into the new model at a faster rate than real-time data streams, allowing the new model to quickly "warm up" and establish a temporal state relevant to the current operating conditions. For example, within the temporal state break interval, a weight function can be designed so that the weight of the new model's output prediction value gradually increases from zero, while the weight of the old model's output prediction value gradually decreases from one, ultimately achieving a smooth transition.
[0031] During the model performance evaluation phase, to avoid transient errors caused by cold starts interfering with the true performance evaluation, data points within the previously identified time-series state break intervals are marked as invalid. When calculating performance metrics such as prediction accuracy and false alarm rate of the new model, these invalid samples will be excluded, and only steady-state output data after the break interval ends will be used for calculation. This allows for a more accurate evaluation of the true performance of the new model under stable operating conditions.
[0032] If performance verification results show that the new model outperforms the old model in terms of prediction accuracy, false alarm rate, and fault warning lead time, the edge node will perform an atomic switch operation, seamlessly switching the output channel of the reducer fault prediction from the old model to the new model, and shutting down the inference process of the old model to release its occupied resources. Conversely, if the new model's performance is not superior to the old model, the shadow inference channel of the new model will be terminated, the output of the old model will be maintained, and a detailed diagnostic report will be generated and uploaded to the cloud platform. The report will include information such as the reason for the switch failure, transient deviation characteristics, and steady-state error comparison.
[0033] The working principle of this application is to distinguish between cold start mutations and actual performance degradation from a time-series perspective by using a triple progressive judgment of "input volatility + old model stability + new model error morphology (instantaneous rise followed by decay)"; the combination of "state injection / historical replay warm-up + output smooth fusion" actively repairs state breaks rather than passively waiting for the new model to stabilize on its own, thus eliminating output mutations; "time mask removal" directly applies the identification results of the broken intervals to the evaluation link, forming a complete closed loop of "identification → repair → evaluation", avoiding misjudgment and rollback.
[0034] In summary, this application provides the new model with a historical context prior to the switch by synchronously collecting the temporal state of the old model and constructing a contextual state dictionary. This avoids abrupt changes in the output prediction sequence caused by the loss of the temporal state maintained by the old model upon activation of the new model. By accurately identifying and removing temporal state breaks caused by cold starts as invalid samples, the closed-loop correction evaluation mechanism avoids misjudging transient deviations as performance degradation, thus preventing erroneous model rollback or over-correction. Through forced alignment and continuation compensation of the temporal context, the prediction stability of the new model in the initial stage of the switch is improved, and system interference caused by the cold start effect is reduced. This method maintains the stability of the model's closed-loop evolution process in the industrial internet environment and improves the update efficiency and reliability of the reducer fault prediction model.
[0035] In an embodiment of this application, the step of synchronously acquiring the output prediction values, input data sequences, and time-series state vectors of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and constructing a context state dictionary to describe the current operating context of the reducer, preferably includes: Within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, a circular buffer queue of preset length is opened in the memory of the edge node; wherein, the preset length is set according to the temporal receptive field of the old model or the maximum memory depth of the recurrent neural network, and is used to cache the input data sequence within the preset historical time period, so as to provide the historical input data required for high-speed playback warm-up of the new model in subsequent steps. When the old model outputs a predicted value, the predicted value of the old model at the current moment, the input data sequence, and the temporal state vector are bound together, and a context state dictionary covering all inference cycles within the entire preset historical duration is constructed.
[0036] The preset history duration refers to the time span of historical data that needs to be collected before model switching. Its purpose is to provide sufficient information to construct the context of the reducer's operation and to provide historical input for the warm-up of the new model. The preset history duration should at least cover the length of historical data that the old model needs to backtrack during a single complete inference. For example, for prediction models based on recurrent neural networks, the preset history duration should not be less than the time span that its hidden states can effectively remember; for models based on temporal convolutional networks, it should not be less than the time window length corresponding to its receptive field. In practical applications, the preset history duration can be adjusted according to the typical cycle of reducer operating conditions, for example, set to 1 hour, to ensure sufficient contextual information is provided in most operating condition change scenarios.
[0037] A pre-defined circular buffer queue is an efficient data structure characterized by new data overwriting the oldest data as it reaches the end of the queue, thus achieving fixed-length historical data storage. The temporal receptive field refers to the time range of historical data that a model can "see" when making predictions, while the maximum memory depth of a recurrent neural network determines the amount of historical information it can effectively utilize. By aligning the pre-defined length with these model characteristics, it ensures that the input data sequence buffered in the circular buffer queue can fully meet the historical data requirements for high-speed playback warm-up of new models.
[0038] The context state dictionary aims to cover data from all inference cycles within the entire preset historical duration, thereby comprehensively describing the current operating context of the reducer. This dictionary not only includes the model's inputs and outputs but also the model's internal temporal states, which is crucial for the forced alignment and continuation compensation of the temporal context for subsequent new models.
[0039] This application ensures that before model switching, the system can efficiently and systematically collect and store the historical data and internal states of the old model. This pre-preparation mechanism not only provides the new model with sufficient and high-quality historical input data for rapid warm-up, shortening the cold start time of the new model, but also enables the new model to more accurately understand and inherit the operating context of the old model by constructing a detailed context state dictionary. This reduces the decrease in prediction accuracy and system instability caused by temporal state breaks during model switching, and improves the smoothness and reliability of model switching.
[0040] In some embodiments of this application, such as Figure 2 As shown, the preferred steps of receiving the new model from the cloud platform, constructing a shadow inference channel parallel to the old model, mirroring the real-time sensor data stream and simultaneously sending it into both the old and new models to obtain the output prediction sequences of the old and new models include: S121, Receive the weight and network structure description file of the new model issued by the cloud platform through a security protocol; S122, Based on the weights and network structure description file, load the new model in the independent memory space and computing resource isolation area of the edge node, and construct a shadow inference channel that runs parallel to the old model; S123, the real-time sensor data stream is copied into two copies, one of which is sent to the old model to maintain the prediction output, and the other copy is sent to the shadow inference channel of the new model simultaneously, so as to obtain the output prediction value sequence of the old model and the new model respectively.
[0041] A secure communication link is established between the edge nodes and the cloud platform, using encryption protocols such as TLS / SSL, to ensure the integrity and confidentiality of model weights and network structure description files during transmission. The weights of the new model are the connection strength parameters of neurons in each layer of the deep learning model, while the network structure description file defines the topological information of the model, such as the number of layers, the type of each layer, the activation function, and the input and output dimensions. The purpose is to ensure that the new model can be securely and accurately deployed from the cloud platform to the edge nodes.
[0042] Independent memory space and isolated computing resources can be understood as dedicated resource areas allocated for new models on edge nodes, such as through containerization or virtual machine technology. This ensures that the operation of the new model does not interfere with the normal operation of the old model, while avoiding resource contention. Shadow inference channels refer to a mechanism where the new model receives real-time data and performs inference in parallel without affecting the output of the old model. The purpose is to perform real-time verification and performance evaluation of the new model without interrupting existing services.
[0043] Real-time sensor data streams refer to the operational data collected in real time from various sensors on the reducer (such as vibration sensors, temperature sensors, torque sensors, and acoustic emission sensors). Copying this data twice and feeding one copy into the old model and the other into the new model ensures that both models infer under identical input conditions, resulting in comparable sequences of output predictions. The old model continues to provide current fault predictions, while the new model runs in the background as a "shadow," its output prediction sequence used for subsequent performance validation and switching decisions. The aim is to provide fair and consistent input data for the performance validation of the new model and to ensure that the old model continues to provide services during validation.
[0044] The proposed solution receives model files through a secure protocol, ensuring the integrity and confidentiality of the model data. By loading the new model through an independent memory space and isolated computing resource area, resource conflicts and mutual interference between the old and new models are avoided, ensuring system stability. Through real-time data stream mirroring, the new model can be run and its performance evaluated in a real production environment without affecting the currently running old model. This provides a reliable basis for subsequent model switching and effectively improves the robustness and reliability of the reducer fault prediction system.
[0045] In the above embodiments of this application, the method for determining the time series state break interval caused by cold start based on the output predicted value sequence, under the condition that the volatility of the input data sequence is lower than a preset physical mutation threshold and the output predicted value of the old model is stable, if the error between the output predicted values of the new model and the old model first exceeds a preset transient tolerance upper limit within a preset observation period and then shows a monotonically decreasing trend, specifically includes the following steps: Calculate the volatility of one of the following: mean, variance, and first difference of the input data sequence over the most recent preset number of sampling periods; If the volatility is lower than a preset physical mutation threshold, the sliding window standard deviation or the absolute value of the first difference of the old model's output prediction value within the most recent preset number of sampling periods is calculated as the smoothness of the old model's output prediction value. If the smoothness is greater than the preset stability threshold, then the absolute value of the difference between the output prediction values of the new model and the old model at the same input time is calculated as the absolute error between the output prediction values of the new model and the old model. If the absolute error exceeds the preset transient tolerance limit within the preset observation period after the new model starts receiving real-time data streams, and the absolute error shows a monotonically decreasing trend over time after exceeding the preset transient tolerance limit and before the end of the preset observation period, then the moment when the absolute error first exceeds the preset transient tolerance limit is taken as the starting point and the moment when the preset observation period ends is taken as the ending point, which is taken as the time-series state break interval under the stable operating condition caused by cold start.
[0046] Volatility is designed to quantify the stability of the input data stream. For example, it can be characterized by calculating the standard deviation or mean absolute difference of sensor data over a recent period. The preset number of sampling periods can be configured based on the actual application scenario and data sampling frequency to ensure effective capture of input data fluctuations. The preset number is set comprehensively based on the sampling frequency and the typical cycle of gearbox operating conditions to ensure effective capture of the fluctuation characteristics of the input data stream. Specifically, the preset number should be large enough to filter out high-frequency interference such as sensor noise, while being small enough to reflect short-term trends in operating conditions. For example, if the sensor sampling frequency is 100Hz and the typical cycle of gearbox operating conditions is 1 second, then the preset number can be 100 sampling periods, corresponding to a 1-second time window; if the sampling frequency is 10Hz and the operating condition change cycle is 10 seconds, then the preset number can be 100 sampling periods, corresponding to a 10-second time window.
[0047] If the volatility is lower than the preset physical mutation threshold, it indicates that the current operating condition of the reducer is relatively stable, and there are no significant physical mutations. Based on this, the smoothness of the old model's output prediction is calculated to quantify the degree of volatility of the old model's output prediction; a smaller value generally indicates a more stable output. The preset stability threshold is used to define the upper limit for considering the old model's output prediction as stable. If the smoothness is greater than the preset stability threshold, it indicates that the old model's output prediction has a certain degree of volatility, but it is still under a stable operating condition where the input data volatility is lower than the preset physical mutation threshold. Under this condition, the absolute error between the output predictions of the new model and the old model is calculated to reflect the prediction differences between the new and old models under the same input. The preset physical mutation threshold can be set according to the statistical distribution of the reducer's volatility under historical stable operating conditions, for example, by taking a number of standard deviations of the mean or an empirical percentage value. The preset stability threshold can be set according to the statistical distribution of the smoothness of the old model's output predictions under historical stable operating conditions.
[0048] The preset transient tolerance upper limit is the maximum allowable transient error for a new model during the initial cold start phase. A monotonically decreasing trend indicates that the new model is gradually adapting and converging, rather than continuously diverging or exhibiting new anomalies. The preset transient tolerance upper limit can be set based on the normal fluctuation range of the old model's output predictions under historical stable conditions. For example, it can be a multiple of the standard deviation of the old model's output predictions (e.g., 3 times), or the upper limit of the old model's prediction error confidence interval. When the new model's output error exceeds this upper limit, it indicates that the transient bias caused by the cold start has exceeded the acceptable range, requiring context compensation or marking as a break in the interval. The preset observation duration limits the observation window for this transient behavior, aiming to ensure sufficient time to determine whether the error shows a decaying trend. The preset observation duration can be set based on the maximum time required for the new model's internal state to stabilize from cold start. For example, it can be a multiple of the new model's temporal receptive field or the memory depth of a recurrent neural network (e.g., 2 times), or an empirical value for the time required for historical data playback warm-up.
[0049] This application's solution, by comprehensively considering the volatility of input data, the stability of the old model's output, and the specific dynamic trends of the errors between the old and new models, can effectively distinguish between the cold start effect of the new model and real equipment anomalies or external interference, avoiding misjudgments. This allows for more accurate identification of the time window requiring context alignment and compensation during model switching, thus laying a solid foundation for subsequent smooth model transition and performance verification, ultimately improving the reliability and intelligence level of the gearbox fault prediction system.
[0050] The following is a specific example to illustrate this.
[0051] Suppose that the old model is running stably on the edge node of the reducer, and a new prediction model is deployed from the cloud platform. After receiving the new model and building a shadow inference channel, the edge node begins to run the old and new models in parallel.
[0052] The system first continuously monitors the volatility of the speed reducer sensor data stream. For example, it calculates the average of the absolute values of the first-order differences of key parameters such as temperature and vibration over the most recent 10 sampling periods. If this volatility remains below a preset physical abrupt change threshold (e.g., 0.5%), the input data environment is considered stable.
[0053] Next, the system calculates the standard deviation of the sliding window of the old model's output prediction value (e.g., the failure probability value) over the most recent 20 sampling periods, as the smoothness of the old model's output. If this smoothness is greater than a preset stability threshold (e.g., 0.01), it indicates that the old model's output has slight fluctuations but is still in an overall stable operating condition.
[0054] Under these conditions, the system begins to calculate the absolute error between the output predictions of the new model and the old model. Assuming that within a preset observation period (e.g., 30 minutes), the absolute error between the outputs of the new model and the old model first exceeds the preset transient tolerance limit (e.g., 0.2) at the 5th minute, and then between the 5th and 30th minutes, the absolute error value shows a monotonically decreasing trend, gradually converging towards zero.
[0055] Based on the above assessment, the system identifies the period from the 5th to the 30th minute as the temporal state disruption interval caused by a cold start under stable operating conditions. Within this interval, the new model exhibits transient instability due to a lack of historical context, but its predictive ability is gradually recovering. By identifying this interval, subsequent contextualization and output smoothing mechanisms can intervene in a targeted manner to ensure the smoothness of model switching.
[0056] In a further embodiment of this application, the step of determining the time series state break interval caused by cold start, based on the output predicted value sequence, under the condition that the volatility of the input data sequence is lower than a preset physical mutation threshold and the output predicted value of the old model is stable, if the error between the output predicted values of the new model and the old model first exceeds a preset transient tolerance upper limit within a preset observation period and then shows a monotonically decreasing trend, further preferably includes: If the volatility is higher than or equal to the preset physical mutation threshold, then determine whether the output prediction value of the old model triggers the preset alarm threshold. If not triggered, the absolute difference between the first-order difference of the predicted values of the new model and the old model is calculated as the relative logical deviation. If the relative logical deviation exceeds the preset transient tolerance limit within the preset observation period after the new model starts receiving real-time data streams, and the relative logical deviation shows a monotonically decreasing trend over time after exceeding the preset transient tolerance limit and before the end of the preset observation period, then the moment when the relative logical deviation first exceeds the preset transient tolerance limit is taken as the starting point and the moment when the preset observation period ends is taken as the ending point, which is taken as the time-series state break interval under the perturbation condition caused by cold start. If it has been triggered, it is determined to be a real device anomaly. The timing state break interval is not marked, and the old model output prediction value is kept as the current prediction result.
[0057] When the volatility of the input data sequence is higher than or equal to a preset physical mutation threshold, it indicates that the reducer may be in a non-stationary operating condition, such as load changes or slight vibrations. In this case, it is first necessary to determine whether the output prediction value of the old model has triggered the preset alarm threshold. The preset alarm threshold is usually set based on historical operating data and expert experience to indicate possible serious faults or abnormal states of the equipment. If the output prediction value of the old model has not triggered the preset alarm threshold, it indicates that although there is fluctuation in the input data, the old model has not detected a serious fault risk. In this case, the possible deviation between the new model and the old model is more likely to be caused by a cold start. To this end, the absolute difference of the first-order difference between the output prediction values of the new model and the old model is calculated and used as the relative logical deviation. The first-order difference can reflect the trend of the predicted value, and its absolute difference can measure the consistency of the two models in predicting the trend of change. If the relative logical deviation exceeds the preset transient tolerance limit within the preset observation period after the new model starts receiving real-time data streams, and then shows a monotonically decreasing trend after exceeding the limit and before the end of the preset observation period, it can be determined that this is a time-series state break interval under a perturbation condition caused by a cold start. This decreasing trend indicates that the new model is gradually adapting to the current operating conditions and converging with the predictions of the old model. The starting point of this break interval is defined as the moment when the relative logic deviation first exceeds the preset transient tolerance limit, and the ending point is the moment when the preset observation period ends. In practical applications, if the output prediction value of the old model has triggered the preset alarm threshold, it means that the equipment may have experienced a real anomaly or malfunction. In this case, it should not be attributed to the cold start effect of the new model, but should be judged as a real equipment anomaly.
[0058] This application's solution effectively distinguishes the causes of deviations between the old and new models under different operating conditions by introducing a judgment on the volatility of the input data sequence and combining it with the logic of whether the output prediction value of the old model triggers an alarm threshold. When the input data fluctuates significantly, it no longer simply relies on absolute error, but instead evaluates the adaptability of the new model to dynamic changes by calculating the relative logical deviation. This method avoids misjudging model deviations as cold start effects when equipment malfunctions, thus ensuring the accuracy and safety of model switching decisions in complex and ever-changing industrial environments. By identifying the temporal state break intervals caused by cold starts under perturbation conditions, targeted context alignment and continuation compensation can be performed, enabling the new model to take over the prediction task faster and more smoothly, reducing prediction interruptions or inaccuracies caused by model switching.
[0059] In a further embodiment of this application, in the step of warming up the new model by replaying historical input data from the period prior to the switch at a rate higher than real-time: If a time-series state break interval caused by cold start is identified, all historical input data sequences within the preset historical time before switching will be continuously input into the new model at a playback speed higher than the real-time data stream rate for warm-up. If a time-series state break interval under the perturbation condition caused by cold start is identified, the historical input data sequence containing the input data corresponding to the time-series state break interval under the perturbation condition is continuously input into the new model at a playback speed higher than the real-time data stream rate for warm-up.
[0060] The time-series state breakpoint caused by cold start under stable operating conditions refers to the interval where, under conditions of low volatility in the input data sequence and stable output predictions from the old model, the error between the output predictions of the new model and the old model first exceeds the preset transient tolerance limit within a preset observation period, and then shows a monotonically decreasing trend. Under this stable operating condition, the system is in a stable operating state, and the new model needs to fully learn and inherit the system's long-term stable behavior. Therefore, all historical input data sequences within the preset historical period before the switchover are continuously input into the new model at a playback speed higher than the real-time data stream rate for warm-up. The purpose is to provide the new model with a complete and comprehensive historical context, enabling it to fully initialize its internal time-series state and thus better adapt to subsequent stable operation.
[0061] The time-series state break interval under perturbation conditions caused by cold start refers to the interval where, under the condition that the volatility of the input data sequence is high but no real abnormal alarm is triggered, the absolute difference (relative logical deviation) of the first-order difference between the predicted values of the new model and the old model first exceeds the preset transient tolerance limit within a preset observation period and then shows a monotonically decreasing trend. Under this perturbation condition, the system may have experienced brief fluctuations or disturbances, and the new model needs to quickly capture and adapt to these transient changes. Therefore, historical input data sequences containing the input data corresponding to the time-series state break interval under the perturbation condition are continuously input into the new model at a playback speed higher than the real-time data stream rate for warm-up. The purpose is to enable the new model to focus on key historical data related to the perturbation, quickly adjust its internal state to cope with specific disturbance patterns, and avoid having its sensitivity to perturbations diluted by irrelevant stable data.
[0062] Playback speed exceeding the real-time data stream rate refers to inputting historical input data into the new model at a faster rate than actual data acquisition and processing. The aim is to rapidly update the internal temporal state of the new model within a short time, bringing it to a state that matches the current operating context as quickly as possible, thereby shortening the cold start impact time. The playback speed is determined comprehensively based on the idle computing power of the edge nodes and the warm-up time requirements. The upper limit of the playback speed is limited by the idle CPU / GPU computing power that the edge nodes can utilize under the current load; excessive speed may lead to exhaustion of computing resources or system response delays. The lower limit of the playback speed must meet the goal of completing the warm-up within the expected time, that is, inputting all historical data into the new model to update its internal temporal state in the shortest possible time.
[0063] Through the above technical solution, when the new model takes over the prediction task, its internal temporal state can be more accurately initialized and inherited according to the actual working conditions. This not only effectively shortens the duration of the cold start effect and reduces the prediction error during model switching, but also improves the adaptability and stability of the new model under different operating conditions. Compared with a single preheating strategy, the solution proposed in this application can more effectively handle complex and ever-changing industrial field environments, ensuring the continuity and high accuracy of reducer fault prediction.
[0064] The following is a specific example to illustrate this.
[0065] Suppose that a speed reducer needs to have its predictive model upgraded during normal operation.
[0066] In the first scenario, if the system identifies a break in the timing sequence under stable operating conditions caused by a cold start, it means that the reducer had been operating stably before the model switch, with minimal fluctuations in input data and stable output from the old model. In this case, to allow the new model to fully understand this stable operating condition, the system will continuously input all historical input data sequences (including sensor data such as vibration, temperature, and current) from the past 24 hours (e.g., before the switch) into the new model at a rate higher than the real-time data stream rate (e.g., 2 or 3 times the real-time rate) for preheating. This allows the new model to fully learn the characteristics of the reducer under normal stable operation and adjust its internal timing sequence to a state highly compatible with the current stable operating condition.
[0067] In the second scenario, if the system identifies a time-series state breakpoint under a perturbation condition caused by a cold start, this means that the reducer may have experienced a brief, non-faulty, minor perturbation (e.g., slight vibration fluctuations due to instantaneous changes in external load) before the model switch. In this case, to enable the new model to quickly adapt and accurately handle such perturbations, the system will only continuously input historical input data sequences (e.g., only data from a specific time period within the past hour where the perturbation occurred) into the new model at a rate higher than the real-time data stream rate for warm-up. This targeted warm-up method allows the new model to quickly capture and learn the characteristics of the perturbation, avoiding the dilution of sensitivity to perturbations due to processing a large amount of irrelevant stationary data, thus enabling more accurate identification and prediction of such perturbation events after the model switch.
[0068] It should be noted that the output smoothing fusion mechanism includes: Within the time-series state break interval, weights that vary with time are constructed; The sum of the weights of the output predictions of the new model and the old model is always 1. At the initial moment, the weight of the output prediction of the new model is a preset small weight, and the weight of the output prediction of the old model is 1 minus the preset small weight. As time goes by, the weight of the output prediction of the new model increases linearly or exponentially, and the weight of the output prediction of the old model decreases accordingly, until the weight of the output prediction of the new model reaches its maximum value at the end of the time-series state break interval.
[0069] Constructing time-varying weights refers to the system dynamically generating a series of weight values based on the duration of a time-series state break interval after it has been determined. These weights are used to perform a weighted average of the output predictions of the old and new models to form the final fused output. Initially, the weight of the new model's output prediction is set to a preset small weight, such as a value close to zero but not zero (e.g., between 0.01 and 0.1), while the weight of the old model's output prediction is 1 minus the preset small weight. This setting aims to ensure that the old model's output still dominates at the beginning of the transition, thereby avoiding the new model having an excessive influence on the output before it is fully stable or aligned. Over time, the weight of the new model's output prediction can increase linearly or exponentially, while the weight of the old model's output prediction decreases accordingly. This increasing or decreasing method can be a pre-defined function curve, such as a straight line or an exponential curve, the choice of which can be adjusted according to the smoothness requirements and computational resource limitations of the actual application scenario. Finally, until the end of the time-series state break interval, the weight of the new model's output prediction reaches its maximum value, usually 1. This means that the new model will completely take over the prediction output, while the weight of the old model drops to its minimum value, usually 0.
[0070] This application clarifies the specific strategies for weight construction, initial allocation, and time-varying changes. In particular, the new model weights start from preset small values and increase linearly or exponentially until they completely take over the output, enhancing the smoothness and reliability during the transition period. This avoids sudden changes or drastic fluctuations in output predictions caused by model switching, ensuring the continuity and accuracy of gearbox fault prediction, and thus reducing potential risks caused by improper model switching.
[0071] Based on the above implementation methods, such as Figure 3 As shown, the steps of marking the identified time-series state break intervals as invalid samples, removing them using a time mask, and verifying the actual performance based solely on the steady-state output to obtain the verification results include: S151, mark the identified temporal state break intervals as invalid samples; S152, When performing real performance verification on the new model, a time mask is applied to remove the invalid samples from the performance verification calculation. S153, based solely on the output prediction value sequence after the end of the time-series state break interval, calculate the prediction accuracy, false alarm rate, and fault warning advance of the new model and the old model within the same preset verification time after the end of the time-series state break interval, and use them as performance verification parameters. S154. If the prediction accuracy of the new model is higher than that of the old model, the false alarm rate is lower than that of the old model, and the early warning of faults is greater than that of the old model, then the verification result is that the performance of the new model is better than that of the old model; otherwise, the verification result is that the performance of the new model is not better than that of the old model.
[0072] Identified temporal state break intervals are marked as invalid samples because, during model switching, due to factors such as the cold start of the new model or temporal context mismatch, the output prediction values may exhibit significant deviations or instability over a period of time. These data points should not be used to evaluate the model's true performance. Therefore, these data points are explicitly marked as invalid to exclude them from subsequent performance evaluations.
[0073] When validating the real performance of a new model, a time mask is applied to remove invalid samples from the performance validation calculation. This can be understood as a mechanism that ignores or skips data points marked as invalid samples when calculating performance metrics. For example, a Boolean mask array with the same length as the output predicted value sequence can be set, with positions corresponding to invalid samples marked as false and positions corresponding to valid samples marked as true. During calculation, only data with a true mask are processed. The purpose is to ensure the objectivity and accuracy of the performance evaluation and avoid misjudging the model's true performance due to transient errors during the cold start period.
[0074] Prediction accuracy refers to the proportion of times a model correctly predicts a fault or normal state; false alarm rate refers to the proportion of times a model incorrectly predicts a normal state as a fault; fault warning lead time refers to how far in advance the model issues a warning before the actual fault occurs. These parameters are key indicators for measuring the performance of fault prediction models, comprehensively reflecting the model's predictive ability, reliability, and practicality. The aim is to provide a quantitative and comprehensive evaluation standard to objectively compare the performance of new and old models.
[0075] The method in this application requires the new model to show advantages in multiple key indicators before it can be considered as having superior performance. The purpose is to ensure that model switching is only carried out when performance is actually improved, thereby ensuring the accuracy and reliability of gearbox fault prediction.
[0076] Through the above technical solution, this application ensures that the performance evaluation of the new model is conducted under stable operating conditions, thereby avoiding evaluation bias caused by transient effects such as cold start or temporal context breaks. This performance verification mechanism improves the reliability of model switching decisions, preventing incorrect switching to a new model with poor performance due to inaccurate evaluation, or missing out on a new model with better performance. Furthermore, by comprehensively considering prediction accuracy, false alarm rate, and fault warning lead time, the performance evaluation becomes more comprehensive and scientific.
[0077] The following is a specific example to illustrate this.
[0078] Suppose an old model is deployed on an edge node to predict gearbox failure. A new model is deployed from the cloud platform, requiring a switch. During the parallel inference phase, the system identifies that within the first 30 minutes after receiving the real-time data stream, due to the cold start effect, the new model's output predictions have a large and monotonically decreasing error compared to the old model. Therefore, this 30-minute period is defined as the time-series state break interval.
[0079] According to the scheme in this application, the output predictions of the new and old models during these 30 minutes are marked as invalid samples. During real-world performance verification, the system will apply a time mask to ensure that the data from these 30 minutes is not included in the calculation of performance metrics. Subsequently, the system will calculate the prediction accuracy, false alarm rate, and fault warning advance based solely on the output prediction sequences of the new and old models within the next 2 hours (the preset verification duration) after the end of the time-series state break interval.
[0080] Specifically, if, within the 2-hour verification period, the new model achieves a prediction accuracy of 95%, a false alarm rate of 2%, and a fault warning lead time of 10 hours, while the old model achieves a prediction accuracy of 92%, a false alarm rate of 3%, and a fault warning lead time of 8 hours, then according to the preset judgment criteria, the new model's prediction accuracy (95% > 92%) is higher than the old model, its false alarm rate (2% < 3%) is lower, and its fault warning lead time (10 hours > 8 hours) is greater. Therefore, the system will obtain the verification result that the new model performs better than the old model, thereby triggering a model switching operation and switching the output channel of the reducer fault prediction to the new model. This method ensures that the model switching is based on the real performance advantages of the new model under stable operating conditions, rather than an evaluation result affected by transient fluctuations.
[0081] In some embodiments of this application, the step of switching the output channel of the reducer fault prediction to the new model and shutting down the old model if the new model outperforms the old model based on the verification results; otherwise, initiating an abnormal rollback mechanism, terminating the inference channel of the new model, and generating a diagnostic report for uploading to the cloud platform preferably includes: If the verification result shows that the new model performs better than the old model, then perform an atomic switching operation on the output channel of the reducer fault prediction to switch from the old model to the new model. After switching to the new model, if the standard deviation of the sliding window of the output predicted value sequence within the preset confirmation time is lower than the preset stability threshold, it is determined that the new model takeover is stable, the inference process of the old model is closed, and the memory and computing resources occupied by the old model are released; otherwise, it is determined that the new model takeover is abnormal, the inference channel of the new model is terminated, the output of the old model is restored, and a diagnostic report containing the reason for the takeover failure is generated and uploaded to the cloud platform. If the verification result shows that the performance of the new model is not better than that of the old model, then the shadow inference channel of the new model is terminated, the output of the old model is maintained, and a diagnostic report containing the reason for the switching failure, transient bias characteristics and steady-state error comparison is generated and uploaded to the cloud platform.
[0082] Atomic switching refers to seamlessly and uninterruptedly switching the source of prediction output from the old model to the new model within a very short time, ensuring that there is no data loss, prediction interruption, or output corruption during the switching process. This switching is typically achieved by updating pointers or configurations pointing to the currently active model to guarantee the immediacy and integrity of the switch.
[0083] After switching the output channel, the system continuously monitors the predicted value sequence output by the new model to further confirm whether it can stably take over the prediction task. The sliding window standard deviation is a statistic that measures data volatility. It is calculated by moving a fixed-size window across consecutive data points to reflect the stationarity of local data. If the sliding window standard deviation is lower than the preset stability threshold, it indicates that the new model's output predicted values have maintained good stability after the switch, without drastic fluctuations or abnormal jumps. At this point, the new model can be considered to have taken over stably. Once stable takeover is confirmed, the inference process of the old model will be shut down, and its occupied memory and computing resources will be released to optimize the resource utilization of edge nodes. The preset confirmation time is set based on the empirical value of the time required for the new model to go from cold start to output stability, and is usually several times the preset observation time. The preset stability threshold is set based on the statistical distribution of the sliding window standard deviation of the old model's output predicted values under historical stable conditions.
[0084] However, if the sliding window standard deviation of the new model's output prediction sequence within the preset confirmation period is higher than or equal to the preset stability threshold, it indicates that the new model exhibits instability or anomalies after takeover, and this will be judged as an abnormal takeover. Simultaneously, a diagnostic report containing the reasons for the takeover failure will be generated, indicating specific reasons such as excessive output fluctuations or abnormal deviations in predicted values, and this report will be uploaded to the cloud platform for further analysis and processing.
[0085] Furthermore, if the initial validation results show that the new model does not outperform the old model, the system will not perform the switch. In this case, the shadow inference channel of the new model will be terminated, and the old model will continue to serve as the primary prediction model. Simultaneously, a detailed diagnostic report will be generated, including the specific reasons for the switch failure, the transient bias characteristics of the new and old models during the switch process (e.g., the trend of error curve changes), and a comparison of steady-state errors (e.g., performance differences during stable operation). This report will be uploaded to the cloud platform, providing valuable data support for model optimization and iteration.
[0086] Through the above technical solutions, this application improves the reliability and security of online updates for gearbox fault prediction models. Atomic switching operations ensure a smooth and seamless model switching process, avoiding service interruptions. The takeover stability confirmation mechanism effectively identifies potential instability factors that may arise in the actual operation of the new model, thereby avoiding the deployment of poorly performing or unstable models to production and reducing deployment risks. The introduction of an anomaly rollback mechanism enables the system to quickly recover to a stable state in the event of model switching failure or takeover anomalies, ensuring the continuity of prediction services. Furthermore, detailed diagnostic reports provide rich fault information, facilitating model developers to analyze, optimize, and iterate model performance, thereby continuously improving the accuracy of the prediction model and providing solid technical support for the operation and maintenance of industrial equipment.
[0087] like Figure 4 As shown, this application also discloses a speed reducer fault prediction system based on the Industrial Internet, applied to an edge node deployed on the speed reducer side. The edge node is connected to a cloud platform via the Industrial Internet to collect speed reducer operating data in real time and to run a prediction model. The system includes: The data acquisition module 210 is used to synchronously collect the output prediction values, input data sequences, and time-series state vectors of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and to construct a context state dictionary to describe the current operating context of the reducer; wherein, the old model is the prediction model that is currently running and used to predict reducer faults, and the new model is the prediction model that is distributed from the cloud and used to replace the old model. Parallel inference module 220 is used to receive new models from the cloud platform, build a shadow inference channel that runs parallel to the old model, mirror the real-time sensor data stream and send it into the old model and the new model at the same time, and obtain the output prediction value sequence of the old model and the new model respectively. The interval identification module 230 is used to determine the time series state break interval caused by cold start based on the output prediction value sequence, under the condition that the volatility of the input data sequence is lower than the preset physical mutation threshold and the output prediction value of the old model is stable. If the error between the output prediction values of the new model and the old model first exceeds the preset transient tolerance upper limit within the preset observation time and then shows a monotonically decreasing trend. The context acceptance processing module 240 is used to perform forced alignment and acceptance compensation of the temporal context on the new model based on the temporal state vector stored in the context state dictionary: if the new model has the same structure as the old model, the temporal state vector of the old model is directly injected into the new model; otherwise, the previous historical input data is switched and replayed to the new model at a rate higher than real time for warm-up, and within the temporal state break interval, an output smoothing fusion mechanism is adopted to gradually transition from the output prediction value of the old model to the output prediction value of the new model. The performance parameter acquisition module 250 is used to mark the identified time-series state break intervals as invalid samples and apply a time mask to remove them. The actual performance is verified only based on the steady-state output to obtain the verification results. The model switching module 260 is used to switch the output channel of the reducer fault prediction to the new model and shut down the old model if the performance of the new model is better than that of the old model based on the verification results; otherwise, it initiates the abnormal rollback mechanism, terminates the inference channel of the new model and generates a diagnostic report to upload to the cloud platform.
[0088] The data acquisition module 210 can be implemented as a resident service process in the edge node's operating system, continuously monitoring the sensor data interface and the inference output interface of the legacy model, and storing the acquired data in the local storage medium of the edge node, such as a cache or non-volatile memory. Alternatively, the data acquisition module 210 can also be implemented as a plugin or callback function embedded in the legacy model's inference framework, automatically capturing and encapsulating the required data each time the legacy model completes inference.
[0089] The parallel inference module 220 can be implemented as a standalone containerized application on an edge node, such as a Docker container, which loads the new model and runs it in a separate computing resource and memory space from the old model to avoid resource conflicts. Alternatively, the parallel inference module 220 can also be implemented as a lightweight virtual machine in an edge node virtualization environment, dedicated to loading and inference of new models, ensuring the independence and security of its runtime environment.
[0090] The interval identification module 230 can be implemented as a data analysis service on an edge node. This service continuously monitors the predicted value sequence output by the parallel inference module and calculates in real time the volatility of the input data, the stability of the old model output, and the error between the old and new models. The data analysis service can use a rule-based engine or a lightweight machine learning model to determine whether the identification conditions for time-series state break intervals caused by cold start are met.
[0091] The context handling module 240 can be implemented as an initialization component during the new model loading process. Before the model is activated, it determines and executes the corresponding temporal state injection or historical data replay warm-up logic based on the structure of the new and old models. For example, the context handling module 240 can be a dynamic link library or a shared object, which is called when a new model is loaded. It is responsible for reading the context state dictionary and initializing the state of the new model according to a preset strategy.
[0092] The performance parameter acquisition module 250 can be implemented as a performance evaluation agent on an edge node. After parallel inference is completed, the performance evaluation agent obtains the output prediction sequence of the old and new models from the parallel inference module, and combines it with the temporal state break interval information provided by the interval identification module to clean and filter the data, and then calculates various performance indicators. The performance evaluation agent can be a standalone script or an evaluation service integrated into the edge computing framework.
[0093] The model switching module 260 can be implemented as a model lifecycle management service on the edge node. This service receives the verification results from the performance parameter acquisition module 250 and executes atomic model switching operations according to a preset strategy, including updating output routes, stopping the old model process, and releasing resources. In case of anomalies, the service also triggers a rollback operation and generates a diagnostic report. The service can interact with the resource scheduler and network configurator of the edge node to ensure a smooth switching process and efficient resource utilization.
[0094] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0095] The preferred embodiments of this application have been described in detail above, but this application is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application.
Claims
1. A method for predicting gearbox failures based on the Industrial Internet, applied to edge nodes deployed on the gearbox side, wherein the edge nodes are connected to a cloud platform via the Industrial Internet for real-time collection of gearbox operating data and for running a prediction model, characterized in that... The method includes the following steps: The system synchronously collects the output prediction values, input data sequences, and time-series state vectors of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and constructs a context state dictionary to describe the current operating context of the reducer; wherein, the old model is the prediction model that is currently running and used to predict reducer faults, and the new model is the prediction model that is downloaded from the cloud and used to replace the old model; The system receives a new model from the cloud platform, constructs a shadow inference channel that runs parallel to the old model, mirrors the real-time sensor data stream and simultaneously sends it into both the old and new models, and obtains the output prediction sequence of each model. Based on the output prediction sequence, under the condition that the volatility of the input data sequence is lower than the preset physical mutation threshold and the output prediction of the old model is stable, if the error between the output prediction of the new model and the old model first exceeds the preset transient tolerance limit within the preset observation period and then shows a monotonically decreasing trend, then the time series state break interval caused by cold start is determined. Based on the temporal state vector stored in the context state dictionary, the new model is subjected to forced alignment and continuity compensation of the temporal context: if the new model has the same structure as the old model, the temporal state vector of the old model is directly injected into the new model; otherwise, the previous historical input data is switched and replayed to the new model at a rate higher than real-time for warm-up, and within the temporal state break interval, an output smoothing fusion mechanism is adopted to gradually transition from the output prediction value of the old model to the output prediction value of the new model. The identified time-series state break intervals are marked as invalid samples and removed using a time mask. The actual performance is verified only based on the steady-state output to obtain the verification results. Based on the verification results, if the new model outperforms the old model, the output channel for gearbox fault prediction will be switched to the new model and the old model will be shut down; otherwise, an abnormal rollback mechanism will be initiated to terminate the inference channel of the new model and generate a diagnostic report to be uploaded to the cloud platform.
2. The method for predicting gearbox faults based on the Industrial Internet according to claim 1, characterized in that, The steps of synchronously acquiring the output prediction values, input data sequences, and time-series state vectors of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and constructing a context state dictionary to describe the current operating context of the reducer, include: Within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, a circular buffer queue of preset length is opened in the memory of the edge node; wherein, the preset length is set according to the temporal receptive field of the old model or the maximum memory depth of the recurrent neural network, and is used to cache the input data sequence within the preset historical time period, so as to provide the historical input data required for high-speed playback warm-up of the new model in subsequent steps. When the old model outputs a predicted value, the predicted value of the old model at the current moment, the input data sequence, and the temporal state vector are bound together, and a context state dictionary covering all inference cycles within the entire preset historical duration is constructed.
3. The method for predicting gearbox faults based on the Industrial Internet according to claim 1, characterized in that, The steps of receiving the new model from the cloud platform, constructing a shadow inference channel parallel to the old model, mirroring the real-time sensor data stream and simultaneously sending it into both the old and new models to obtain the output prediction sequences of the old and new models respectively include: Receive the weight and network structure description file of the new model issued by the cloud platform through a security protocol; Based on the weights and network structure description file, a new model is loaded into the independent memory space and computing resource isolation area of the edge node to construct a shadow inference channel that runs parallel to the old model. The real-time sensor data stream is copied into two copies. One copy is sent to the old model to maintain the prediction output, and the other copy is sent to the shadow inference channel of the new model simultaneously, so as to obtain the output prediction value sequence of the old model and the new model respectively.
4. The method for predicting gearbox faults based on the Industrial Internet according to claim 1, characterized in that, Based on the output predicted value sequence, under the condition that the volatility of the input data sequence is lower than a preset physical mutation threshold and the output predicted value of the old model is stable, if the error between the output predicted values of the new model and the old model first exceeds a preset transient tolerance upper limit within a preset observation period and then shows a monotonically decreasing trend, the step of determining the time series state break interval caused by the cold start includes: Calculate the volatility of one of the following: mean, variance, and first difference of the input data sequence over the most recent preset number of sampling periods; If the volatility is lower than a preset physical mutation threshold, the sliding window standard deviation or the absolute value of the first difference of the old model's output prediction value within the most recent preset number of sampling periods is calculated as the smoothness of the old model's output prediction value. If the smoothness is greater than the preset stability threshold, then the absolute value of the difference between the output prediction values of the new model and the old model at the same input time is calculated as the absolute error between the output prediction values of the new model and the old model. If the absolute error exceeds the preset transient tolerance limit within the preset observation period after the new model starts receiving real-time data streams, and the absolute error shows a monotonically decreasing trend over time after exceeding the preset transient tolerance limit and before the end of the preset observation period, then the moment when the absolute error first exceeds the preset transient tolerance limit is taken as the starting point and the moment when the preset observation period ends is taken as the ending point, which is taken as the time-series state break interval under the stable operating condition caused by cold start.
5. The method for predicting gearbox faults based on the Industrial Internet according to claim 4, characterized in that, Based on the output predicted value sequence, under the condition that the volatility of the input data sequence is lower than a preset physical mutation threshold and the output predicted value of the old model is stable, if the error between the output predicted values of the new model and the old model first exceeds a preset transient tolerance upper limit within a preset observation period and then shows a monotonically decreasing trend, the step of determining the time series state break interval caused by the cold start further includes: If the volatility is higher than or equal to the preset physical mutation threshold, then determine whether the output prediction value of the old model triggers the preset alarm threshold. If not triggered, the absolute difference between the first-order difference of the predicted values of the new model and the old model is calculated as the relative logical deviation. If the relative logical deviation exceeds the preset transient tolerance limit within the preset observation period after the new model starts receiving real-time data streams, and the relative logical deviation shows a monotonically decreasing trend over time after exceeding the preset transient tolerance limit and before the end of the preset observation period, then the moment when the relative logical deviation first exceeds the preset transient tolerance limit is taken as the starting point and the moment when the preset observation period ends is taken as the ending point, which is taken as the time-series state break interval under the perturbation condition caused by cold start. If it has been triggered, it is determined to be a real device anomaly. The timing state break interval is not marked, and the old model output prediction value is kept as the current prediction result.
6. The method for predicting gearbox faults based on the Industrial Internet according to claim 5, characterized in that, In the step of warming up the new model by replaying historical input data from the period prior to the switch at a rate higher than real-time: If a time-series state break interval caused by cold start is identified, all historical input data sequences within the preset historical time before switching will be continuously input into the new model at a playback speed higher than the real-time data stream rate for warm-up. If a time-series state break interval under the perturbation condition caused by cold start is identified, the historical input data sequence containing the input data corresponding to the time-series state break interval under the perturbation condition is continuously input into the new model at a playback speed higher than the real-time data stream rate for warm-up.
7. The method for predicting gearbox faults based on the Industrial Internet according to claim 1, characterized in that, The output smoothing fusion mechanism includes: Within the time-series state break interval, weights that vary with time are constructed; The sum of the weights of the output predictions of the new model and the old model is always 1. At the initial moment, the weight of the output prediction of the new model is a preset small weight, and the weight of the output prediction of the old model is 1 minus the preset small weight. As time goes by, the weight of the output prediction of the new model increases linearly or exponentially, and the weight of the output prediction of the old model decreases accordingly, until the weight of the output prediction of the new model reaches its maximum value at the end of the time-series state break interval.
8. The method for predicting gearbox faults based on the Industrial Internet according to claim 1, characterized in that, The steps of marking the identified time-series state break intervals as invalid samples, removing them using a time mask, and verifying the actual performance based solely on the steady-state output to obtain the verification results include: The identified temporal state break intervals are marked as invalid samples; When performing real-world performance verification on the new model, a time mask is applied to remove invalid samples from the performance verification calculation. Based solely on the output prediction value sequence after the end of the time-series state break interval, the prediction accuracy, false alarm rate, and fault warning advance of the new and old models within the same preset verification time after the end of the time-series state break interval are calculated as performance verification parameters. If the new model has a higher prediction accuracy than the old model, a lower false alarm rate, and a greater early warning time for faults, then the new model is considered to outperform the old model; otherwise, the new model is considered to not outperform the old model.
9. The method for predicting gearbox faults based on the Industrial Internet according to claim 1, characterized in that, Based on the verification results, if the new model outperforms the old model, the output channel for the reducer fault prediction is switched to the new model and the old model is shut down; otherwise, the abnormal rollback mechanism is activated, the inference channel of the new model is terminated, and a diagnostic report is generated and uploaded to the cloud platform. The steps include: If the verification result shows that the new model performs better than the old model, then perform an atomic switching operation on the output channel of the reducer fault prediction to switch from the old model to the new model. After switching to the new model, if the standard deviation of the sliding window of the output predicted value sequence within the preset confirmation time is lower than the preset stability threshold, it is determined that the new model takeover is stable, the inference process of the old model is closed, and the memory and computing resources occupied by the old model are released; otherwise, it is determined that the new model takeover is abnormal, the inference channel of the new model is terminated, the output of the old model is restored, and a diagnostic report containing the reason for the takeover failure is generated and uploaded to the cloud platform. If the verification result shows that the performance of the new model is not better than that of the old model, then the shadow inference channel of the new model is terminated, the output of the old model is maintained, and a diagnostic report containing the reason for the switching failure, transient bias characteristics and steady-state error comparison is generated and uploaded to the cloud platform.
10. A speed reducer fault prediction system based on the Industrial Internet, applied to an edge node deployed on the speed reducer side, wherein the edge node is connected to a cloud platform via the Industrial Internet for real-time collection of speed reducer operating data and for running a prediction model, characterized in that... The system includes: The data acquisition module is used to synchronously collect the output prediction values, input data sequences, and time-series state vectors of the old model within a preset historical time period before switching the output channel of the reducer fault prediction to the new model, and to construct a context state dictionary to describe the current operating context of the reducer; wherein, the old model is the prediction model that is currently running and used to predict reducer faults, and the new model is the prediction model that is downloaded from the cloud and used to replace the old model; The parallel inference module is used to receive new models from the cloud platform, build a shadow inference channel that runs parallel to the old model, mirror the real-time sensor data stream and send it into both the old and new models simultaneously, and obtain the output prediction value sequences of the old and new models respectively. The interval identification module is used to determine the time series state break interval caused by cold start based on the output prediction value sequence, under the condition that the volatility of the input data sequence is lower than the preset physical mutation threshold and the output prediction value of the old model is stable. If the error between the output prediction values of the new model and the old model first exceeds the preset transient tolerance upper limit within the preset observation time and then shows a monotonically decreasing trend. The context acceptance processing module is used to perform forced alignment and acceptance compensation of the temporal context on the new model based on the temporal state vector stored in the context state dictionary: if the new model has the same structure as the old model, the temporal state vector of the old model is directly injected into the new model; otherwise, the previous historical input data is switched and replayed to the new model at a rate higher than real time for warm-up, and within the temporal state break interval, an output smoothing fusion mechanism is adopted to gradually transition from the output prediction value of the old model to the output prediction value of the new model. The performance parameter acquisition module is used to mark the identified time-series state break intervals as invalid samples and remove them using a time mask. It performs real performance verification based only on steady-state output to obtain the verification results. The model switching module is used to switch the output channel of the reducer fault prediction to the new model and shut down the old model if the new model performs better than the old model based on the verification results; otherwise, it initiates an abnormal rollback mechanism, terminates the inference channel of the new model, and generates a diagnostic report to be uploaded to the cloud platform.