Multifunctional edge computing device and method based on artificial intelligence

By deploying the main and secondary models in the wind power edge node and setting the disturbance detection mechanism, a credibility scoring mechanism is established and the model status is automatically adjusted, the problem of model accuracy in the wind power system is solved, and the self-diagnosis and recovery of the edge AI system is realized, and the stability and robustness of the system are improved.

CN120295791AActive Publication Date: 2025-07-11CHINA DATANG GRP DIGITAL TECH CO LTD +1

Patent Information

Application Number
CN202510418184.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

Existing wind power edge AI systems are susceptible to memory bit flips in high humidity and high salt spray environments, resulting in a decrease in model inference output accuracy and lack of real-time verification and monitoring mechanisms, which may cause false alarms, misidentification of status or false triggering of key control logic, affecting the operating efficiency and safety of wind farms.

Method used

Deploy the main model and the secondary model in the edge node, set the disturbance detection mechanism, detect the model response offset by the disturbance sample, establish a confidence scoring mechanism, and automatically select the main model parameters to reload or switch to the secondary model to realize self-diagnosis and recovery of the model state.

Benefits of technology

Effectively identify the potential precision degradation behavior of the model, prevent silent failure, improve the stability and robustness of the long-term operation of the model, and ensure the safety and efficiency of the wind farm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295791A_ABST
    Figure CN120295791A_ABST
Patent Text Reader

Abstract

The invention discloses a multifunctional edge computing device and method based on artificial intelligence, and particularly relates to the technical field of wind power operation management, and the method comprises the following steps: deploying a main model and an auxiliary model at an edge node, enabling the main model to execute real-time reasoning, and enabling the auxiliary model to be used for credibility evaluation; setting a disturbance detection mechanism, injecting a disturbance sample with preset characteristics into the main model, comparing the disturbance sample with the output of the auxiliary model, and calculating a response offset degree; fusing the offset degree and the reasoning difference to generate a credibility score; when the score is lower than a threshold value, automatically selecting to execute parameter reloading or switch an auxiliary model, and reconstructing a reasoning session to guarantee stable operation of edge nodes; according to the method, real-time perception and structural offset judgment of the operation state of the edge computing device are achieved through disturbance detection, main and auxiliary model collaborative verification and a credibility scoring mechanism, parameter reloading or model switching is dynamically executed based on a strategy tree, and the stability, self-diagnosis capability and fault tolerance level of the model in an extreme environment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wind power operation management. More specifically, the present invention relates to a multifunctional edge computing device and method based on artificial intelligence. Background Art

[0002] With the in-depth application of artificial intelligence technology in industrial scenarios, the wind power generation system is gradually introducing edge computing and AI models to achieve local perception of the operating state of wind turbines, predictive maintenance, and intelligent control. Especially in offshore or isolated island-type wind farms far from the main power grid, edge intelligent nodes play a key role in operation judgment and decision support.

[0003] Existing edge AI models are usually deployed in embedded hardware devices, running on resource-constrained processing platforms and requiring long-term stable online operation. However, in the unique high-humidity, high-salt mist, and unmanned operation environment of the wind power system, such embedded devices are extremely vulnerable to hardware micro-errors induced by environmental factors. In particular, volatile memory units such as RAM may experience non-fatal bit flips under humid conditions, which, although not immediately causing device crashes, are sufficient to cause slight drift in the internal parameters of the AI model, and thus produce imperceptible but continuously accumulating accuracy deviations in the inference output.

[0004] Currently, most wind power edge AI systems lack a mechanism for real-time verification and structural monitoring of model inference states, and have not established a behavior evaluation logic based on perturbation response or replica comparison. This "silent operation" assumed environment allows the model to continue normal output even after the inference output has experienced accuracy degradation or structural anomalies after long-term operation, which may ultimately lead to false alarms of wind turbines, misidentification of states, or mis-triggering of key control logics, and seriously affect the operation efficiency and safety of wind farms when severe. Therefore, the present invention proposes a multifunctional edge computing device and method based on artificial intelligence to solve the above problems. Summary of the Invention

[0005] To achieve the above objectives, the present invention provides the following technical solutions:

[0006] A multifunctional edge computing method based on artificial intelligence, comprising the following steps:

[0007] Deploy a main model and a sub-model in the edge node respectively. The main model is used to perform real-time inference tasks, and the sub-model is configured as a static reference model and only participates in inference during the credibility evaluation process;

[0008] Set a perturbation detection mechanism, regularly inject perturbation samples with preset characteristics into the main model, and compare the response results of the main model to the perturbation samples with the output results of the sub-model to calculate the perturbation response offset.

[0009] Fuse the perturbation response offset and the inference differences between the primary and secondary models for the same real-time input data, establish a credibility scoring mechanism, and judge the stability and accuracy of the current model running state;

[0010] When the credibility score is lower than the threshold, automatically select to execute the primary model parameter reloading or switch to the secondary model based on the offset features and difference trends, and reconstruct the inference session to maintain the continuous and stable operation of the edge intelligent edge node.

[0011] In a preferred embodiment, the perturbation samples include two types, namely boundary samples and stability samples, which are respectively used to stimulate the sensitivity of the primary model to the input boundary and the stability of the intermediate layer inference structure. The boundary samples are used to test the prediction stability under extreme input conditions, and the stability samples are used to observe the response consistency to small input perturbations. The perturbation response offset is calculated by comparing any one of the cosine similarity, Euclidean distance, or Mahalanobis distance between the intermediate layer feature vectors before and after the perturbation, and is used to reflect the response consistency of the internal structure of the primary model to the perturbation.

[0012] In a preferred embodiment, the credibility scoring for fusing the perturbation response offset and the output differences between the primary and secondary models adopts a weighted scoring strategy. The weight of the perturbation term is dynamically adjusted according to the stability performance of the perturbation samples in the historical period. The output differences between the primary and secondary models are quantified by the root mean square error. The scoring function is constructed by one of the following methods: linear weighted combination, exponentially weighted average based on time series decay, or adaptive weight strategy based on information entropy distribution.

[0013] In a preferred embodiment, when the primary model triggers parameter reloading, it adopts a difference localization and local parameter recovery strategy. It locates the potential drift region through the credibility score, corresponding to specific sub-layers or parameter branches in the model structure, and then extracts the corresponding weights from the pre-stored parameter backups to replace a part of the parameter set in the current operation to achieve local reloading.

[0014] In a preferred embodiment, the judgment of the credibility score lower than the threshold adopts a threshold trend evaluation mechanism with a sliding window. Define an observation period. If the score falls into the preset warning interval multiple times within multiple consecutive observation periods, it is judged that the stability of the primary model degrades and the recovery behavior is triggered.

[0015] In a preferred embodiment, the secondary model is designed as a lightweight version optimized by knowledge distillation, with lower inference latency and resource overhead compared to the primary model. During the training phase, the secondary model achieves the retention of discriminative ability and the reduction of running load by aligning the output distribution, imitating intermediate features, and compressing the structure with the primary model. The input and output spaces are kept consistent, enabling seamless access during the verification process and switching scenarios.

[0016] In a preferred embodiment, the inference session reconstruction process after model overloading or switching includes releasing the current cache state, reallocating the execution context, performing an empty data inference operation once to activate the internal computational graph, and then generating event markers with timestamps and trigger information.

[0017] In a preferred embodiment, the model switching and overloading behaviors are determined by a policy tree structure containing multiple triggering factors. The evaluation conditions include the response trend of perturbation samples, the difference rate between the outputs of the primary and secondary models, the deviation degree between the model prediction and the observed data, and the load status of the edge nodes. Different response behaviors are executed according to the policy tree path selection.

[0018] In a preferred embodiment, the credibility score result is synchronously used to adjust the task scheduling policy of the edge nodes.

[0019] In a preferred embodiment, a multifunctional edge computing device based on artificial intelligence includes:

[0020] A model deployment module for respectively deploying a primary model and a secondary model in edge nodes. The primary model is used to perform real-time inference tasks, and the secondary model is configured as a static reference model and only participates in inference during the credibility evaluation process.

[0021] A perturbation detection module for setting a perturbation detection mechanism, regularly injecting perturbation samples with preset features into the primary model, and comparing the response result of the primary model to the perturbation samples with the output result of the secondary model to calculate the perturbation response offset.

[0022] A score calculation module for fusing the perturbation response offset and the inference difference between the primary and secondary models for the same real-time input data, establishing a credibility scoring mechanism, and judging the stability and accuracy of the current model running state.

[0023] A behavior switching module for automatically selecting to perform primary model parameter overloading or switching to the secondary model based on the offset feature and difference trend when the credibility score is lower than the threshold, and reconstructing the inference session to maintain the continuous and stable operation of the edge intelligent edge nodes.

[0024] The technical effects and advantages of the present invention:

[0025] By constructing a perturbation detection mechanism and a model output consistency evaluation system, the edge AI device can dynamically perceive the structural deviation risk of the model during the inference process in an environment without an external network and unattended for a long time. Especially in extreme working conditions such as high humidity and prone to memory bit flips in offshore wind farms, the device can timely capture abnormal perturbation responses and inference deviation trends, and effectively identify potential model accuracy degradation behaviors. Through the parameter local reloading mechanism triggered by the credibility score, the system can automatically repair the drifted sub-layer weights of the model, prevent the model from silently failing, and significantly improve the stability and robustness of the model during long-term operation.

[0026] The main and auxiliary model collaborative structure and policy tree behavior decision-making mechanism proposed by the present invention endow the edge AI device with the dynamic closed-loop ability of "operation self-check - health assessment - behavior adaptation". Through the credibility score system established by integrating the perturbation response deviation degree and the prediction difference between the main and auxiliary models, the device can not only evaluate the current state of the model in real time, but also select the most appropriate response behavior, including delayed reloading, auxiliary model switching or continuous observation, in combination with the model state trend and system operating conditions. This fault tolerance strategy avoids the resource waste and mis-shutdown risks caused by the "one-size-fits-all" restart mechanism of traditional systems, and provides a more flexible and intelligent model health management solution for critical mission scenarios such as wind farms. Brief Description of the Drawings

[0027] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings;

[0028] Figure 1 It is a schematic diagram of a multifunctional edge computing method based on artificial intelligence in the present invention.

[0029] Figure 2 It is a schematic diagram of a multifunctional edge computing device based on artificial intelligence in the present invention. Detailed Embodiments

[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] Refer to Figure 1 - Figure 2 The following embodiments are obtained:

[0032] Embodiment 1: The present invention mainly proposes a solution to the problem of maintaining model reliability in an extreme industrial operating environment for edge computing systems, and is particularly applicable to the offshore wind power island scenario deployed in an environment far from the main network. In such an application background, high-power electrical equipment such as wind turbines needs to continuously operate under harsh conditions such as high humidity, salt spray, drastic changes in wind pressure, and unstable power supply. Edge AI inference nodes must undertake functions such as local perception, prediction and judgment, and control and scheduling of equipment status, and are required to have the ability of long-term unattended operation, self-recovery, and self-decision-making.

[0033] However, there are a series of problems that cannot be avoided by existing technologies in such an environment, including implicit accuracy drift during model operation (such as parameter deviation caused by memory bit-flip), unidentifiable abnormal disturbance response, lack of local reload mechanism in the system, etc. More seriously, due to the inability to access the remote cloud model management system, the traditional model state evaluation and version update mechanism relying on cloud monitoring is difficult to implement in this environment. For example, the problem description: The edge node devices are eroded by high humidity all year round in the island environment. Some commercial embedded devices have non-fatal RAM bit errors (such as Flip) under humid conditions. Although the devices do not crash, there are slight deviations in the AI model inference results. Most existing technologies do not set up a model inference self-calibration mechanism, resulting in "silent" judgment deviations in the model after long-term operation. The model output is no longer accurate, the equipment status prediction error accumulates, and finally the wind turbine is miscontrolled or the protection is mis-triggered.

[0034] To meet the requirements of long-term stable operation of edge AI models in extreme operating environments such as wind farms, the present invention provides an edge computing device and method integrating multiple functions. While performing the conventional inference task of the main model, it integrates a secondary model collaborative verification mechanism, a disturbance response evaluation mechanism, a model credibility judgment mechanism based on scoring, a local parameter repair mechanism, and a model behavior scheduling mechanism driven by a policy tree, and has the ability of task-level resource adaptive adjustment, constituting an intelligent edge computing framework that can self-diagnose, self-recover, and self-select processing paths when the model state is abnormal.

[0035] Specifically applied to the operation management of AI models in edge devices, including the following steps:

[0036] Step 1: Deploy the main model and the secondary model in the edge node respectively. The main model is used to perform real-time inference tasks, and the secondary model is configured as a static reference model and only participates in inference during the credibility evaluation process;

[0037] Step 2: Set up a disturbance detection mechanism, regularly inject disturbance samples with preset characteristics into the main model, and compare the response results of the main model to the disturbance samples with the output results of the secondary model to calculate the disturbance response offset;

[0038] Step 3: Integrate the perturbation response offset and the inference differences between the primary and secondary models for the same real-time input data, establish a credibility scoring mechanism, and judge the stability and accuracy of the current model running state;

[0039] Step 4: When the credibility score is lower than the threshold, automatically select to execute the primary model parameter reloading or switch to the secondary model based on the offset features and difference trends, and reconstruct the inference session to maintain the continuous and stable operation of the edge intelligent edge node.

[0040] The perturbation samples include at least two types, which are respectively used to stimulate the sensitivity of the primary model to the input boundary and the stability of the intermediate layer inference structure. The boundary samples are used to test the prediction stability under extreme input conditions, and the stability samples are used to observe the response consistency to slight input perturbations. The perturbation response offset is calculated by comparing any one of the cosine similarity, Euclidean distance, or Mahalanobis distance between the intermediate layer feature vectors before and after the perturbation, and is used to reflect the response consistency of the internal structure of the primary model to the perturbation. The intermediate layer output is extracted by a predefined inference stability reference layer, and this output serves as the calculation basis for the response reliability.

[0041] The boundary samples are used to stimulate the sensitivity of the primary model to the input boundary, specifically referring to pushing the input data to a position close to the edge of the model training interval or beyond the normal distribution range to simulate extreme operating states or unconventional input signals, such as when the wind speed reaches an extremely high value or the voltage is at a critical state. The response of the primary model to such samples is used to test the stability performance of its prediction ability under extreme input conditions, that is, to judge whether its output still maintains continuity, reasonableness, and whether there are catastrophic predictions (such as output divergence, abnormal categories, etc.).

[0042] The stability samples are used to observe whether the model has consistent inference results under slightly perturbed inputs near the normal distribution. Such perturbations can be constructed by adding low-amplitude high-frequency noise, subtle signal jitters, or lightweight data rotations, aiming to simulate the effects of sensor errors or slight on-site interferences on the model, so as to detect the micro-response changes of the model under high-frequency operation.

[0043] To quantify the response differences of the primary model to different types of perturbations, the present invention calculates the perturbation response offset, that is, evaluates the response consistency of the model at the structural level by comparing the similarity degree between the intermediate layer feature vectors before and after the perturbation. The intermediate layer feature vector refers to the tensor representation output by the model at one or several specific layers during the inference process, usually containing high-dimensional semantic information and multi-channel activation states.

[0044] The calculation method of this offset includes any one of the following three methods:

[0045] Cosine similarity: used to measure the consistency of the directions of two vectors, and its value range is usually between [-1, 1]. The closer it is to 1, the more consistent the response directions are;

[0046] Euclidean distance: reflects the geometric distance between two vectors, and the smaller it is, the closer the outputs before and after perturbation are;

[0047] Mahalanobis distance: introduces the covariance matrix, used to consider the correlation and scale effects between different features, and is applicable to evaluating the anomaly degree of samples in multi-dimensional space.

[0048] In the present invention, these measurement methods are used to quantify the intermediate layer responses of the main model, for constructing the perturbation response deviation index. This index can directly reflect the consistency of the internal structure of the main model in dealing with perturbations, and serves as the key input for the subsequent credibility scoring mechanism. In addition, the feature vectors used for output comparison are extracted from a set of predefined inference-stable reference layers, which are usually selected from the middle part of the model, such as the intermediate convolutional blocks in a convolutional neural network, the intermediate attention layers in a Transformer architecture, etc. By continuously monitoring the output of this layer and sampling the perturbation responses, the change characteristics of the model's structural behavior under different input perturbations can be more accurately reflected.

[0049] The credibility scoring for fusing the perturbation response deviation and the output difference between the main and auxiliary models adopts a weighted scoring strategy. The weight of the perturbation term is dynamically adjusted according to the stability performance of the perturbation samples in the historical period, and the output difference between the main and auxiliary models is quantified by the root mean square error. The scoring function can be constructed in one of the following ways: linear weighted combination; exponential weighted average based on time series decay; or adaptive weight strategy based on information entropy distribution. This fusion mechanism takes into account both instantaneous deviation and long-term stability trends, which helps to improve the judgment accuracy of the scoring and the rationality of behavior triggering.

[0050] Periodically obtain the measurement indicators from two independent sources:

[0051] Perturbation response deviation: reflects the structural response consistency of the main model under the input of perturbation samples;

[0052] Output difference between the main and auxiliary models: reflects the degree of consistency of the prediction results of the two under the same real-time input.

[0053] To achieve the effective integration of two-dimensional metrics, the edge node adopts a weighted scoring strategy, where each metric is assigned a dynamically varying weight value. In particular, the weight of the perturbation term is dynamically adjusted according to the stability performance of the perturbation samples in the historical period. For example, if a certain type of perturbation sample has shown high consistency in multiple past periods (i.e., the model performance is stable), then once its current offset value fluctuates, it will be regarded as a stronger abnormal signal, so the weight corresponding to this sample will be increased; otherwise, the influence of its scoring will be decreased. This mechanism can be regarded as dynamic trust reconstruction based on the historical distribution of perturbation behaviors.

[0054] The output difference between the primary and secondary models is quantified by the root mean square error (RMSE), which reflects the average deviation degree of the two models in numerical prediction. RMSE is more sensitive to larger errors and can effectively amplify the inconsistency brought by prediction imbalance, and has high representativeness in model stability evaluation.

[0055] The credibility scoring function is used to fuse the above two measurement metrics into a unified scoring result, which can be constructed in one of the following three ways:

[0056] Linear weighted combination: Set fixed or dynamic proportional coefficients to linearly superimpose the two metrics, in the form of: Score = α·Δ1 + β·Δ2;

[0057] Exponentially weighted moving average (EWMA) based on time series decay: Introduce a time decay factor in consecutive periods to make the recent metrics have a greater impact on the scoring, which is suitable for responding to rapid offsets; the core idea of this strategy is to construct a weighted time series sliding value for each type of scoring metric before fusing the perturbation response offset degree or the output difference between the primary and secondary models, so that the data in the latest period has a higher impact weight, and the contribution of historical data decreases exponentially with time, thus forming a scoring basis that is more sensitive to current state changes. Assign a time weight to each consecutive period, and this weight shows a decreasing relationship over time. The newer metric data will participate in the scoring value update with a larger weight. This processing method has a stronger immediate response ability to behaviors such as sudden changes in the model state (such as unstable inference, drastic changes in perturbation response). Through this mechanism, even if the scoring result does not exceed the edge node threshold, if the recent fluctuation amplitude suddenly increases, it can also attract the attention of the edge node or trigger the evaluation logic in advance. In addition, EWMA does not need to store all historical data, only needs to retain the last sliding value, has simple calculation and low resource occupancy, and is especially suitable for being deployed on edge devices with limited computing power. This strategy is suitable for capturing abnormal trends in credibility within short periods and enhancing the response ability to "early signals" without causing excessive misjudgments.

[0058] Adaptive Weight Strategy Constructed Based on Information Entropy Distribution: Conduct information entropy analysis on the degree of change or error distribution of the input, dynamically adjust the importance of each indicator in scoring, and highlight the dimensions of prediction stability or structural robustness. The basic idea of this strategy is to use information entropy to measure the change distribution characteristics of scoring indicators, so as to judge whether the indicator has sufficient "information volume" or "recognition ability" at the current time period. When a certain type of indicator remains highly concentrated and has a narrow fluctuation range in multiple cycles, its information entropy is low, indicating that the contribution of this indicator to state discrimination is limited, and the edge node can appropriately reduce its weight in scoring; on the contrary, if the indicator shows high uncertainty and discrete change, it means that it may be in a stage sensitive to the state change of the edge node, and a higher weight should be assigned. This strategy dynamically identifies the characteristic dimensions with "high discrimination ability" in the current scoring signal by periodically calculating the entropy value of the historical values of each indicator, ensuring that the scoring mechanism focuses on the indicators with "high information density and strong change sensitivity", thus avoiding the interference of redundant indicators on judgment and improving the overall judgment efficiency and reliability of scoring. This adaptive strategy is particularly suitable for dealing with scenarios with diverse indicator dimensions and complex state evolution, and has good scalability and edge node self-learning ability. By introducing an entropy value dynamic feedback mechanism, the edge node can achieve structural dynamic optimization of the scoring weight and strengthen the driving ability of key abnormal signals on the model behavior.

[0059] The above scoring strategy design has strong configurability and adaptability, allowing for flexible selection according to the device operating environment, task type, and prediction period. The entire fusion mechanism takes into account both instantaneous deviation and long-term stability trends, not only being able to promptly detect short-term anomalies in the inference results but also capturing the impacts brought about by the long-term evolution of the internal state of the model. The finally generated credibility score is used to drive the behavior response of the edge node, including but not limited to model reloading, structure replacement, or inference priority adjustment, thereby improving the judgment accuracy of the score and the rationality of behavior triggering, and enhancing the self-regulation ability and reliable operation level of the edge intelligent edge node.

[0060] Assigning dynamically changing weight values means adjusting their influence in the final score at different time periods, operating states, or sample behaviors. The dynamic change of the weight value can be adjusted based on the following factors:

[0061] The change in the weight of the disturbance term driven by historical response stability. For the index of disturbance response deviation, its corresponding weight value not only reflects the deviation degree of the current disturbance result, but also reflects the response stability of this disturbance sample type in the past consecutive operation cycles. Specifically: when a certain disturbance type maintains its deviation within a stable range (such as within a low fluctuation range) in the recent N cycles, it is considered a low-noise signal source when the "model state is stable". Once this disturbance type has a large deviation in the current cycle, the signal meaning should be highly regarded, so the edge node will increase its weight; on the contrary, if this disturbance type shows high-frequency fluctuations or unpredictable behavior (such as frequent violent fluctuations) in the historical cycles, then even if its deviation jumps in the current cycle, the edge node still considers its indication reliability to be low, and the contribution to the scoring result should be reduced, and the weight will be adjusted downward accordingly. This change can be achieved by setting the historical stability score of the disturbance sample and mapping it to the current weight adjustment factor of the disturbance term.

[0062] The adjustment of the error weight guided by the current model output consistency. The output difference between the main and auxiliary models is used as the second scoring index, and the error is mainly calculated by comparing their prediction results for the same input. The weight of this error can be adjusted based on the following two dynamic factors: if the current edge node operation is in a critical control cycle (such as the edge device is in an extreme working condition or a high-speed operating state), the consistency of the outputs of the main and auxiliary models becomes more important, and the weight of the error term will be increased; if the edge node is in a non-critical cycle or a fault-tolerant window period (such as a maintenance state or a buffer period), the error weight will be appropriately reduced to avoid overreacting to short-term fluctuations. This adjustment process can be combined with the current task load, operating state label or control priority parameter of the edge device, and rely on existing technologies for adaptive automatic adjustment.

[0063] When the main model triggers parameter reloading, it adopts a strategy of difference localization and local parameter recovery. The edge node locates the potential drift area through credibility scoring, which corresponds to specific sub-layers or parameter branches in the model structure, and then extracts the corresponding weights from the pre-stored parameter backup to replace part of the parameter set in the current operation, achieving local reloading. This process uses the parameter hash mapping relationship for rapid localization and verification, thus avoiding the operation interruption caused by full-model reloading and restoring inference consistency with the premise of minimizing computing resources.

[0064] First, the running state of the main model is evaluated through a credibility scoring mechanism. When the scoring result is lower than the set threshold, the edge node further analyzes the scoring details and identifies some areas in the model structure that show signs of drift. These areas usually correspond to sub-layers or parameter branches in the model structure, such as the weight matrix of a certain convolutional layer, normalization layer, attention sub-module, fully connected layer, etc. After identification, the edge node will perform a parameter replacement operation on this specific sub-region. This operation is completed by extracting the corresponding weights from the pre-stored parameter backup.

[0065] The meaning of "extracting the corresponding weights from the pre - stored parameter backup" is as follows: After the initial deployment or training of the model, the edge node has completely saved the original parameter structure of the model in a non - volatile storage unit (such as local Flash, eMMC, or solid - state drive). This backup file not only contains the overall model structure information but also stores the parameters in blocks in a regional or hierarchical manner. Each parameter subset has an independent index for individual access and extraction. This backup represents the "healthy state" of the model during the training phase or the initial operation, and is not affected by potential disturbances or storage anomalies during operation.

[0066] To achieve fast access and verification, the edge node establishes a parameter hash mapping relationship for these parameter blocks, that is, calculates the unique verification identifier (such as CRC check code, SHA hash value) for each parameter partition, and establishes a mapping index with the corresponding layer in the model structure. During the actual operation, if a certain area is marked as "drift risk", the edge node will use this mapping to locate the corresponding weight block in the backup file, verify its integrity, and then load it into the current model running memory to replace the part of the parameter set being used.

[0067] This operation is called local reloading. Its core advantage is to avoid the time delay, memory jitter, and inference interruption risks brought by full - model reloading. After the reloading is completed, the edge node will automatically reconstruct the inference path, clear the old parameter cache, and continue to complete the current inference task. Through the above - mentioned differential positioning and local parameter recovery strategy, the present invention realizes an intelligent recovery mechanism with low resource consumption under edge devices while ensuring the overall stability of the model, significantly improving the self - maintenance ability and operation continuity of the model in complex field environments.

[0068] Among them, the potential drift area refers to a local sub - area in the main model structure that is presumed to have risks such as abnormal prediction behavior, decreased computational stability, or weakened feature expression ability. This area may correspond to a specific network layer (such as convolutional layer, fully - connected layer, normalization layer), a computational sub - unit (such as attention head, activation mapping group), or a parameter sub - structure (such as bias vector, scale parameter, etc.). Judging whether there is drift in this area depends on the joint analysis of multiple dimensions such as the edge node's response to intermediate - layer features, response to perturbations, and the stability of the model output. The specific judgment process includes the following aspects:

[0069] Intermediate layer response anomaly detection: In each round of inference tasks, the edge node inputs some perturbed samples into the main model and extracts the feature outputs of the intermediate layer simultaneously. If there is an obvious deviation in the feature distribution output by this layer from the historical distribution (such as mean, variance, activation map distribution), or the cosine similarity drops significantly, the edge node marks the structure where this intermediate layer is located as a potential drift area. Obvious deviation and significant change can be compared by setting a division threshold or a normal change range. For example, exceeding the normal change range can be considered as having an obvious deviation. Similar comparison methods can also be used in other parts of the present invention, which will not be repeated here. For example, the definition of "rapid change" in the perturbation response inconsistency index.

[0070] Perturbation response inconsistency index: When the model faces input samples of the same perturbation type and its response results to this perturbation change rapidly in a short period, the edge node will locate the feature extraction path or attention distribution path involved in these abnormal perturbation responses and further map them to its structural layer as candidate drift areas.

[0071] Analysis of the structural output difference between the main and auxiliary models: When the main model and the auxiliary model produce highly different outputs under the same input conditions, the edge node will trace back the calculation path of the difference in reverse. For example, if the output vector difference is mainly concentrated in certain output dimensions, combined with the model structure map, the key parameter path causing this error is deduced, and thus mapped to a specific sub-layer.

[0072] Comprehensive ranking of credibility score and impact sensitivity: All the candidate areas obtained from the above analyses will be uniformly entered into the credibility score and impact sensitivity evaluation process. Combining the impact factor ranking (such as the current score decline trend, parameter position sensitivity, fault history record, etc.), finally select the top several areas as the "potential drift areas" that need to perform local parameter recovery in this round. The ranking method can be selected from any method in the prior art to measure the overall impact degree. And once the potential drift area is confirmed, through the parameter hash mapping table, the parameter weight block corresponding to this area can be retrieved from the pre-stored parameter backup, and a replacement operation is performed to form a "local reload".

[0073] The judgment of the credibility score below the threshold does not depend on a single instantaneous value, but adopts a threshold trend evaluation mechanism with a sliding window. The edge node defines an observation period. If the score falls into the warning interval multiple times in consecutive periods, it is judged that the stability of the main model degenerates and a recovery behavior is triggered. This strategy introduces a combined judgment of abnormal count accumulation, time-weighted score, and interval confidence, effectively avoiding false triggering caused by short-term interference, and enhancing the tolerance of the edge node to fluctuation signals and the robustness of decision-making.

[0074] The sub-model is designed as a lightweight version optimized by knowledge distillation, with lower inference latency and resource overhead compared to the main model. During the training phase, the sub-model achieves the preservation of discriminative ability and the reduction of operating load by aligning the output distribution, imitating intermediate features, and compressing the structure with the main model. The input-output space remains consistent, enabling seamless access during the validation process and switching scenarios. This design allows the sub-model to operate stably under frequent calls while reducing the energy consumption accumulation of edge nodes under high load conditions.

[0075] Knowledge distillation is a model compression and transfer learning technique, whose basic idea is to transfer the knowledge learned in a large model (teacher model, main model) to a smaller-structured model (student model, i.e., sub-model). This process is usually completed during the training phase and specifically includes the following three key strategies:

[0076] Output distribution alignment: By minimizing the difference (such as Kullback-Leibler divergence) between the output probability distributions of the main model and the sub-model under the same input, the sub-model is made to mimic the judgment logic of the main model as much as possible in the final prediction results.

[0077] Intermediate feature imitation: During the training process, using the partial activation states of the intermediate layers of the main model (such as convolutional feature maps or attention weights) as a guide, forcing the sub-model to be structurally closer to the expression mode of the main model, thereby enhancing its representation ability.

[0078] Structure compression: At the model design level, the model is slimmed down by pruning the number of channels, reducing the parameter dimension, and replacing complex operations with efficient operators (such as replacing standard convolution with depthwise separable convolution), etc., to ensure fast response under tight computing resources.

[0079] Through the above training strategies, the sub-model achieves a high degree of fitting to the discriminative logic of the main model while significantly reducing the parameter scale and computational complexity. To ensure that the sub-model can be seamlessly integrated into the validation process and model switching scenarios, its input-output space remains consistent, that is, the input data dimension, preprocessing process, output vector structure are completely aligned with the main model. This consistency enables seamless task continuation during runtime without the need for additional data transformation, preprocessing reconstruction, or inference interface conversion when the edge node needs to call the sub-model to verify the results of the main model or perform a hot switch after the credibility fails.

[0080] Due to the optimized structure of the sub-model itself, it exhibits lower inference latency and resource overhead during actual operation, and can effectively handle frequent verification calls and high-frequency state evaluation tasks. At the same time, this structure greatly reduces the energy consumption accumulation of edge nodes under high-load conditions, reduces the chip temperature rise, power supply pressure and the risk of computational bottlenecks, and improves the energy efficiency ratio and reliability of edge nodes during long-term operation.

[0081] The distillation design of the sub-model not only ensures that edge nodes have structural complementarity and functional redundancy at the model level, but also achieves a balance between performance and energy consumption in the edge computing environment through optimizing the structure and behavior path, reflecting the innovative edge node advantages of the present invention in model design, edge node scheduling and task switching.

[0082] The inference session reconstruction process after model reloading or switching includes releasing the current cache state, reallocating the execution context, performing an empty data inference operation once to activate the internal computational graph, and then generating event tags with timestamps and trigger information for upper-layer edge node identification. The event tags contain the switching trigger reason, execution mode, and model identity, and are used for log tracking and subsequent analysis. This process ensures that the response time is controllable after inference switching and the edge node state is clear, which helps to maintain the continuous stability of the overall operation.

[0083] After model reloading or switching is triggered, the edge node will immediately perform the operation of releasing the current cache state. This step includes clearing the intermediate tensor data, gradient cache (if any), preprocessing cache, inference graph execution context, etc. in the running memory to avoid interference from historical states on the new model inference. This operation helps to release resources, eliminate state pollution, and prepare for subsequent context reconstruction. Then, reallocate the execution context, that is, allocate a new running session environment for the new model (main model reloading or sub-model switching) to be loaded currently, including model graph structure loading, weight parameter mapping, inference path construction, and necessary hardware binding processes (such as tensor RT binding, AI acceleration unit interface initialization, etc.), to ensure that the model can run correctly under the current edge node resources.

[0084] To activate the inference execution path of the new model and trigger the initial compilation or warm-up of its internal structures (such as dynamic computation graphs, control flows, cache mechanisms, etc.), the edge node will perform an inference operation with empty data. The so-called "empty data" refers to input samples with legal structures but invalid numerical values (such as all-zero tensors, pseudo-inputs, etc.). This operation is not used to obtain results but to initialize the kernel structure, cache construction, and operator optimization path of the model, ensuring that the response time of subsequent formal inference processes is controllable and the startup delay is minimized. After completing the model loading and inference path warm-up, the edge node will generate a set of event tags with timestamps and trigger information and send them to the upper-layer control edge node or logging module to identify this model switching event. This event tag is a structured data entity that includes at least the following three types of information:

[0085] Reason for switching trigger: such as "abnormal perturbation score", "output drift", "model freeze", etc.;

[0086] Execution mode: indicating whether this is "reloading the original model" or "switching to the secondary model";

[0087] Model identity information: indicating the version number, structural summary, verification hash, etc. of the currently active model.

[0088] The event tag is used by the edge node for log tracking and subsequent analysis to achieve model life cycle management, stability statistics, and behavior review. At the same time, if the edge node is designed with a remote monitoring or OTA upgrade mechanism, this event can also be used as the input basis for external control. This process ensures that the response time is controllable after the inference switch, and the status of the edge node is clear, which helps to maintain the continuous stability of the overall operation. Through this design, the present invention realizes the dynamic, transparent, and refined management of models in the edge computing environment, improves the controllability of the edge node over the model evolution process and the observability of the running behavior, and enhances the robustness of intelligent edge nodes under long-term deployment.

[0089] The credibility score result is synchronously used to adjust the task scheduling strategy of the edge computing device. In the state of low model score, the edge node reduces the execution frequency of non-critical tasks, releases resources to give priority to ensuring the main inference process, and at the same time increases the core sensor sampling rate to improve the environmental perception accuracy. After the score returns to normal, the original task weights are gradually restored. By dynamically associating the model credibility with the edge node resource allocation, the linkage and self-adaptation between the model state and the running strategy are realized, and the stability and energy efficiency of the edge node are improved.

[0090] Based on the model state judgment, the present invention introduces a task-level resource linkage mechanism, enabling the credibility scoring result to affect the task scheduling behavior of edge devices. When the device detects a decrease in the credibility of the main model, it can automatically reduce the execution frequency of non-critical tasks, release more resources to ensure the inference performance of the main model, and at the same time increase the data sampling frequency of key sensors to enhance the model input quality. This mechanism realizes the real-time linkage between the running state of the AI model and the system resource management strategy, thereby maximizing the model running efficiency under the condition of limited edge resources, prolonging the device service life, and improving the overall operation and maintenance efficiency.

[0091] In the state of low model scoring, for example, below a preset state threshold, that is, when the edge node identifies that there are uncertainties, possible drifts, or risks of performance degradation in the current main model inference state, the edge node will execute a series of resource allocation strategies to ensure the stable operation of the core inference process: First, identify all the computing tasks currently running, and divide the tasks into two categories, "critical tasks" and "non-critical tasks", according to the preset task priority table. Critical tasks include main model inference, fault detection, device control, etc., while non-critical tasks include local log archiving, remote heartbeat uploading, predictive analysis, visualization update, etc.

[0092] In the state of low scoring, the edge node reduces the execution frequency of non-critical tasks, such as converting the data visualization originally running once a minute to running once every 10 minutes, or pausing non-core task threads, thereby releasing edge node resources such as CPU, memory, and IO, giving priority to ensuring that the main model inference process is not interfered with, and improving the robustness and response speed of the edge node. At the same time, the edge node will increase the sampling rate of core sensors, such as signal sources highly related to model input, such as vibration, temperature, voltage, wind speed, etc., and upgrade the data originally collected at a normal frequency to a high-frequency sampling mode (such as from 10Hz to 50Hz). This helps to enhance the input data quality by increasing the perception accuracy when the model credibility decreases, and alleviates the prediction instability of the model externally.

[0093] Once the model scoring returns to the normal level, the edge node will gradually restore the original task weights, and reallocate the originally restricted resources in a "gentle regression" manner, restoring the execution frequency and task concurrency of non-critical tasks, and avoiding operation jitters or response mutations caused by instantaneous resource fluctuations. This mechanism forms a task scheduling strategy chain driven by credibility scoring, realizing an automatic response closed-loop of "prediction ability decline → edge node resource offset → perception intensity increase". By dynamically associating the model credibility with the edge node resource allocation, the present invention realizes the linkage self-adaptation between the model state and the operation strategy, that is, the uncertainty change of the model itself can trigger fine-tuning of the overall operation mode of the edge node.

[0094] The model switching and reloading behavior is determined by a policy tree structure that includes multiple triggering factors. The evaluation conditions include the response trend of perturbed samples, the difference rate between the outputs of the primary and secondary models, the deviation degree between the model prediction and the observed data, the load status of edge nodes, etc. Edge nodes execute different response behaviors according to the policy tree path, such as delaying reloading, directly switching, or extending the observation. The policy tree is composed of preset rules and supports online parameter fine-tuning to adapt to the model health management and task fault tolerance requirements in different deployment scenarios.

[0095] Each node in the policy tree structure represents an evaluation condition judgment point, each branch represents the state range and judgment result of a type of triggering factor, and the leaf nodes correspond to different response behavior strategies. During operation, edge nodes make condition judgments according to the tree structure based on the current model state input and finally fall into specific behavior nodes.

[0096] The evaluation conditions in the policy tree include but are not limited to the following types of factors. Each type can independently form a judgment node or serve as the intermediate decision logic for multi-condition combinations:

[0097] Response trend of perturbed samples: Edge nodes record the response offset curve of perturbed samples in consecutive cycles and analyze its change trend (such as whether it continues to grow, whether there are periodic anomalies) to judge whether the robustness of the model to structural perturbations has decreased. If it is observed that the perturbed response offset shows an accelerating upward trend, this factor is marked as a high-risk signal, and the policy tree branches downward into the "serious response offset" path.

[0098] Difference rate between the outputs of the primary and secondary models: By continuously comparing the output results of the primary model and the secondary model under the same input conditions, edge nodes can calculate their error growth rate. If the error increases rapidly within a unit time, it indicates that the behavior of the primary model drifts rapidly and there may be structural anomalies. This rate is used to evaluate the rate level of the degradation of the primary model's inference ability and is an important judgment index in the policy tree.

[0099] Deviation degree between the model prediction and the observed data: Edge nodes can compare the predicted output of the primary model with the actual sensor observation values to measure the deviation value. For example, the difference between the predicted value and the actual measured value of the fan vibration will be considered that the model no longer has real-world adaptability after exceeding a certain proportion. The deviation degree judgment node can be further superimposed with the historical deviation trend to judge whether it is a persistent deviation.

[0100] Edge node load status: including the CPU occupancy rate, memory usage rate, I / O pressure, etc. of edge devices. If the current running load of the edge node is at a high level, the policy tree can delay some non-critical operations according to the rules, such as delaying the execution of the reload behavior; conversely, model self-recovery operations are preferentially performed under low load. The load status can be used in combination with the model status judgment logic as an auxiliary condition to achieve self-regulation of edge node resources.

[0101] Behavior selection and response path: The edge node selects and executes different response behaviors according to the policy tree path, mainly including but not limited to the following response modes:

[0102] Delayed reload: In the case where abnormal risks exist in the model status but the current load of the edge node is high or the task criticality is strong, parameter reload is not immediately performed, but is executed in the next scheduling cycle or idle period.

[0103] Direct switch: When the model status is clearly abnormal, the response deviation is severe, and the output difference rate is high, the edge node directly switches the inference task from the main model to the secondary model to ensure real-time performance.

[0104] Observation extension: If the abnormal signal of the model status is not yet significant, the edge node enters the observation window and extends for a period of time to continuously observe the trend of key parameters, and then triggers the next action after the abnormal trend is established.

[0105] Embodiment 2: A multifunctional edge computing device based on artificial intelligence, including:

[0106] A model deployment module for respectively deploying a main model and a secondary model in the edge node. The main model is used to execute real-time inference tasks, and the secondary model is configured as a static reference model and only participates in inference during the credibility evaluation process.

[0107] A perturbation detection module for setting a perturbation detection mechanism, regularly injecting perturbation samples with preset characteristics into the main model, and comparing the response results of the main model to the perturbation samples with the output results of the secondary model to calculate the perturbation response deviation.

[0108] A scoring calculation module for fusing the perturbation response deviation and the inference difference between the main and secondary models for the same real-time input data, establishing a credibility scoring mechanism, and judging the stability and accuracy of the current model running state.

[0109] A behavior switching module for automatically selecting to perform main model parameter reload or switch to the secondary model based on the deviation characteristics and difference trends when the credibility score is lower than the threshold, and reconstructing the inference session to maintain the continuous and stable operation of the edge intelligent edge node.

[0110] The above formulas are all dimensionless and only take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0111] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution is prior or posterior. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0112] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0113] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described edge nodes, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0114] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-functional edge computing method based on artificial intelligence, characterized in that, It includes the following steps: Deploy the main model and the auxiliary model in the edge nodes respectively. The main model is used to perform real-time inference tasks, and the auxiliary model is configured as a static reference model and only participates in the inference during the credibility evaluation process; Set up a perturbation detection mechanism, regularly inject perturbation samples with preset features into the main model, and compare the response results of the main model to the perturbation samples with the output results of the auxiliary model to calculate the perturbation response offset; Fuse the perturbation response offset and the inference difference between the main and auxiliary models for the same real-time input data, establish a credibility scoring mechanism, and judge the stability and accuracy of the current model running state; When the credibility score is lower than the threshold, automatically select to perform main model parameter reloading or switch to the auxiliary model based on the offset features and difference trends, and reconstruct the inference session to maintain the continuous and stable operation of the edge intelligent edge node.

2. The multifunctional edge computing method based on artificial intelligence according to claim 1, characterized in that, The perturbation samples include two types, namely boundary samples and stability samples, which are used to stimulate the sensitivity of the main model to the input boundary and the stability of the intermediate layer inference structure respectively. The boundary samples are used to test the prediction stability under extreme input conditions, and the stability samples are used to observe the response consistency to minor input perturbations. The perturbation response offset is calculated by comparing any one of the cosine similarity, Euclidean distance, or Mahalanobis distance between the intermediate layer feature vectors before and after the perturbation, and is used to reflect the response consistency of the internal structure of the main model to the perturbation.

3. The multifunctional edge computing method based on artificial intelligence according to claim 2, characterized in that, The credibility scoring for fusing the perturbation response offset and the output difference between the main and auxiliary models adopts a weighted scoring strategy. The weight of the perturbation term is dynamically adjusted according to the stability performance of the perturbation samples in the historical period. The output difference between the main and auxiliary models is quantified by the root mean square error. The scoring function is constructed by one of the following methods: linear weighted combination, exponential weighted average based on time series decay, or adaptive weight strategy based on information entropy distribution.

4. A multifunctional edge computing method based on artificial intelligence according to claim 3, characterized in that When the main model triggers parameter reloading, it adopts a difference localization and local parameter recovery strategy. It locates the potential drift area through the credibility score, corresponding to specific sub-layers or parameter branches in the model structure, and then extracts the corresponding weights from the pre-stored parameter backup to replace part of the parameter set in the current operation to achieve local reloading.

5. The multifunctional edge computing method based on artificial intelligence according to claim 4, characterized in that, The judgment of the credibility score lower than the threshold adopts a threshold trend evaluation mechanism with a sliding window. Define an observation period. If the score falls into the preset warning interval multiple times within multiple consecutive observation periods, it is judged that the stability of the main model degenerates and the recovery behavior is triggered.

6. A multifunctional edge computing method based on artificial intelligence according to claim 5, characterized in that, The auxiliary model is designed as a lightweight version optimized by knowledge distillation. It has lower inference latency and resource overhead compared to the main model. During the training stage, the auxiliary model achieves the retention of discriminative ability and the reduction of running load by aligning the output distribution, imitating the intermediate features, and compressing the structure with the main model. The input and output spaces are kept consistent, enabling seamless access during the verification process and switching scenarios.

7. A multifunctional edge computing method based on artificial intelligence according to claim 6, characterized in that, The process of reconstructing the inference session after model reloading or switching includes releasing the current cache state, reallocating the execution context, performing an empty data inference operation to activate the internal computation graph, and then generating an event marker with a timestamp and trigger information.

8. A multifunctional edge computing method based on artificial intelligence according to claim 7, characterized in that The model switching and reloading behaviors are determined by a policy tree structure that includes multiple triggering factors. The evaluation conditions include the response trend of perturbed samples, the difference rate between the outputs of the primary and secondary models, the deviation degree between the model prediction and the observed data, and the load status of the edge nodes. Different response behaviors are executed according to the policy tree path selection.

9. A multifunctional edge computing method based on artificial intelligence according to claim 8, characterized in that, The credibility scoring results are synchronously used to adjust the task scheduling policy of the edge nodes.

10. A multi-functional edge computing device based on artificial intelligence, which is used to implement a multi-functional edge computing method based on artificial intelligence according to any one of claims 1-9, characterized in that, Including: A model deployment module for separately deploying the primary model and the secondary model in the edge nodes. The primary model is used to execute real-time inference tasks, and the secondary model is configured as a static reference model and only participates in the inference during the credibility evaluation process; A perturbation detection module for setting a perturbation detection mechanism, regularly injecting perturbation samples with preset features into the primary model, and comparing the response results of the primary model to the perturbation samples with the output results of the secondary model to calculate the perturbation response offset; A scoring calculation module for fusing the perturbation response offset and the inference difference between the primary and secondary models for the same real-time input data, establishing a credibility scoring mechanism, and judging the stability and accuracy of the current model running state; A behavior switching module for automatically selecting to execute the primary model parameter reloading or switching to the secondary model based on the offset feature and the difference trend when the credibility score is lower than the threshold, and reconstructing the inference session to maintain the continuous and stable operation of the edge intelligent edge nodes.

Citation Information

Patent Citations

  • Intelligent edge device and operation method and system thereof

    CN119578555A

  • Adaptive learning intelligent scheduling unified computing framework and system for industrial personalized customized production

    WO2022099596A1

  • Sample processing method and device and computer readable storage medium

    WO2023083176A1

  • Image processing method and apparatus, model training method and apparatus, and terminal device

    WO2024092590A1

Cited By

  • Method for detecting abnormal operation of intelligent metering box

    CN120508841A

  • Energy efficiency optimization control method for heterogeneous computing power cluster

    CN120993743A