Edge model evolution method and system for air gap isolation environments
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本申请提供了一种用于气隙隔离环境的边缘模型进化方法和系统,实现了边缘模型在气隙隔离环境下的持续进化,解决了物理隔离条件下边缘模型无法随业务变化而自主更新的技术难题
[0009]根据本申请的第五方面,本申请实施例提供了一种计算机程序产品,包括计算机程序,所述计算机程序在被处理器执行时实现如本申请实施例所述的用于气隙隔离环境的边缘模型进化方法。
Smart Images

Figure CN122569971A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of artificial intelligence, edge computing, object detection, network security and data isolation, and in particular to an edge model evolution method and system for air gap isolation environments. Background Technology
[0002] In high-security scenarios with mandatory data security requirements, such as chemical industrial parks, critical infrastructure, and military restricted areas, edge AI systems are widely deployed to perform tasks such as target detection and security monitoring. These scenarios typically require that "data not leave the park." Therefore, edge AI systems are usually deployed in air-gap isolated environments to physically disconnect from the network and prevent the risk of data leakage.
[0003] Edge AI systems in air-gap isolated environments primarily employ two implementation methods: One is a pure edge model solution, which deploys a pre-trained lightweight model on edge devices to perform object detection tasks. This solution processes data entirely locally, offering high security, but the model's capabilities are fixed at deployment time, making it unable to cope with performance degradation caused by environmental changes or task modifications. The other is a cloud-collaborative solution, which uploads data collected at the edge to the cloud, utilizing a large model deployed in the cloud for inference or retraining. This solution continuously receives cloud-based intelligent support and possesses a certain degree of evolutionary capability, but it relies on public network connections, compromising physical isolation and posing data leakage and compliance risks. It also fails to meet the mandatory requirement for data not to leave the premises in high-security scenarios. Summary of the Invention
[0004] This application provides an edge model evolution method and system for air gap isolation environments, which realizes the continuous evolution of edge models in air gap isolation environments and solves the technical problem that edge models cannot be updated autonomously with changes in business under physical isolation conditions.
[0005] According to a first aspect of this application, an edge model evolution method for air-gap isolation environments is provided, the method comprising: In the edge device domain, by deploying an agent based on the input data features of the initial edge model in the edge device and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task, the decay type of the initial edge model and its corresponding distillation strategy parameters are determined. In the server domain, the initial edge model is updated by fine-tuning the distillation strategy parameters of the agent based on the decay type, as well as the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, to obtain the evolutionary edge model. In the edge device domain, the evolutionary edge model is deployed to the edge device by deploying an agent; wherein, there is an air gap isolation between the edge device domain and the server domain.
[0006] According to a second aspect of this application, an edge model evolution system for an air-gap isolated environment is provided, the system comprising: a deployment agent, an edge device, a large visual model, and a fine-tuning agent, wherein the deployment agent and the edge device belong to an end-device domain, and the large visual model and the fine-tuning agent belong to a server domain; an air-gap isolation exists between the end-device domain and the server domain; In the edge device domain, by deploying an agent based on the input data features of the initial edge model in the edge device and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task, the decay type of the initial edge model and its corresponding distillation strategy parameters are determined. In the server domain, the initial edge model is updated by fine-tuning the distillation strategy parameters of the agent based on the decay type, as well as the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, to obtain the evolutionary edge model. In the edge device domain, the evolutionary edge model is deployed to the edge device by deploying an agent.
[0007] According to a third aspect of the present invention, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the edge model evolution method for an air gap isolation environment as described in embodiments of this application.
[0008] According to a fourth aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the edge model evolution method for an air gap isolation environment as described in the embodiments of the present application.
[0009] According to a fifth aspect of this application, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the edge model evolution method for air gap isolation environments as described in embodiments of this application.
[0010] This application's technical solution, under an air-gap isolation architecture between the edge device domain and the server domain, enables edge models to continuously adapt to business changes in a completely offline environment through the collaborative work of deployed agents, fine-tuned agents, and a large visual model. This solves the problem of edge models being unable to continuously evolve under physical isolation conditions. In the edge device domain, the deployed agent automatically detects model performance degradation and determines the degradation type and its corresponding distillation strategy parameters based on the input data features, confidence distribution, and prediction consistency of the initial edge model. This achieves autonomous drift detection and diagnosis without manual annotation or external network dependence. In the server domain, the fine-tuned agent updates the initial edge model using the classification vectors and intermediate layer feature maps output by pseudo-labels in the training pool, based on the distillation strategy parameters corresponding to the degradation type. This allows the model update strategy to adaptively adjust according to the cause of degradation, improving the retraining's relevance and efficiency. Air-gap isolation ensures that data does not leave the campus, meeting security and compliance requirements.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of the edge model evolution method for air gap isolation environments provided in Embodiment 1; Figure 2 This is a flowchart of the edge model evolution method for air gap isolation environment provided in Embodiment 2; Figure 3 This is a flowchart of the edge model evolution method for air gap isolation environments provided in Embodiment 3; Figure 4 This is a schematic diagram of the structure of an edge model evolution system for an air gap isolation environment provided in Embodiment 4 of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of this application. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0015] It should be noted that the terms "first," "second," "target," and "candidate," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Example 1 Figure 1 This is a flowchart of an edge model evolution method for air-gap isolation environments provided in Embodiment 1. This embodiment is applicable to scenarios where edge models are used for target detection in air-gap isolation environments, and the detection task needs to be adjusted according to changes in business operations, such as new violation categories being added in chemical industrial parks, new foreign objects appearing on rail transit, or new instruments being introduced into operating rooms. This method can be executed using an edge model evolution system for air-gap isolation environments, which can be integrated into the electronic device running this system.
[0017] like Figure 1 As shown, the method includes: S110. In the edge device domain, by deploying an agent, the decay type of the initial edge model and its corresponding distillation strategy parameters are determined based on the input data features of the initial edge model in the edge device and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task.
[0018] S120. In the server domain, by fine-tuning the distillation strategy parameters of the agent based on the decay type, and the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, the initial edge model is updated to obtain the evolutionary edge model.
[0019] S130. In the end device domain, the evolutionary edge model is deployed to the edge device by deploying an agent; wherein, there is an air gap isolation between the end device domain and the server domain.
[0020] Air gap isolation physically isolates the network between the end device domain and the server domain from other publicly accessible networks, ensuring data cannot be transmitted externally via any public network. In this solution, the end device domain and server domain communicate via unidirectional fiber optic cable, employing certificate fixing and transport layer security encryption. This enables necessary cross-domain data interaction while ensuring data remains within the campus environment. Unidirectional fiber optic cable refers to a physical transmission medium that allows optical signals to be transmitted unidirectionally from the end device domain to the server domain; the reverse optical path is physically removed, eliminating the possibility of reverse attacks at the hardware level. Certificate fixing involves pre-hardcoding the public key fingerprint of the server domain's receiving gateway's digital certificate into the edge device, verifying the server certificate during the TLS handshake phase to prevent man-in-the-middle attacks. Transport layer security encryption uses the TLS protocol for encryption and integrity protection on the communication link between the two domains.
[0021] The edge device domain refers to the physical area where multimodal sensors and edge devices are deployed. An edge agent cluster runs within the edge device domain, including at least deployed agents. The server domain refers to the physical area where a large visual model and a system agent cluster are deployed. The server domain runs at least a fine-tuning agent. The large visual model runs completely offline, and the system agent cluster supports the model's autonomous evolution.
[0022] The deployed intelligent agent refers to an autonomous software unit deployed in the edge device domain. It is a member of the edge intelligent agent cluster and is responsible for continuously monitoring the online edge models and triggering the model evolution process. The initial edge model refers to the lightweight inference model currently running on the edge device, used to perform object detection tasks. Input data features are obtained by extracting statistical features and numerical features from multimodal sensors from the input images of the initial edge model. Optionally, input data features include brightness histograms and contrast distributions. Input data features are used to compare with the statistical feature baseline of the model training process to detect data distribution shifts. Confidence distribution is obtained by obtaining the confidence values assigned to the prediction results of the initial edge model within a sliding window and analyzing the statistical distribution of these confidence values. Under normal circumstances, high confidence values dominate predictions. When the proportion of low confidence values increases abnormally, it indicates a significant increase in the uncertainty of the initial edge model regarding the current scene. Prediction consistency refers to the stability of the predicted label for the same target in consecutive keyframes to be identified. When the predicted category for the same target frequently jumps within a preset number of frames, it indicates that the inference of the initial edge model is disordered. The deployment agent determines whether the initial marginal model has experienced performance degradation based on three indicators: input data characteristics, confidence distribution, and prediction consistency. Concept drift is identified when the proportion of low confidence values increases abnormally and continuously, the predicted label for the same target changes frequently in a short period, or the distribution distance between the input data characteristics and the statistical baseline increases significantly. Concept drift refers to the phenomenon where the initial marginal model's prediction behavior and data input patterns inevitably undergo observable changes when facing environmental changes, its own capability degradation, or changes in business requirements, leading to a continuous decline in model prediction performance.
[0023] After determining that the initial edge model has experienced concept drift, further analysis is conducted on the spatiotemporal scope of the concept drift and the range of confidence changes to determine the type of decay. The type of decay refers to the classification of the reasons for the performance degradation of the edge model. Optionally, the decay type includes at least model aging, environmental changes, and task requirement changes. Model aging refers to the global decline in predictive ability of the initial edge model after long-term operation due to weight solidification, manifested as a synchronous decrease in confidence across all scenarios and categories. Environmental changes refer to local performance degradation caused by changes in external conditions such as lighting, weather, and background, manifested as feature shifts concentrated only in specific time periods or specific areas. Task requirement changes refer to the need for the target detection task to add the identification of new categories not supported by the current initial edge model, such as a security officer requiring the detection of newly added types of violations. Distillation strategy parameters refer to the training hyperparameter adjustment instructions determined by a preset mapping relationship based on the decay type. Optionally, when the decay type is model aging, the distillation temperature is adjusted; when the decay type is environmental change, the distillation loss weights are adjusted; and when the decay type is task requirement change, new category pseudo-label generation and new training code generation are triggered.
[0024] When concept drift occurs, the predictive behavior and data input patterns of the initial edge model will inevitably undergo observable changes. The deployed agent can sensitively capture and distinguish the causes of decline through multidimensional ground-value-free indicators without relying on externally labeled truth values, providing a basis for decision-making for subsequent targeted retraining.
[0025] A fine-tuning agent is an autonomous software unit deployed in the server domain, belonging to the system's agent cluster. It is responsible for receiving the degradation type and corresponding distillation strategy parameters obtained from the performance diagnosis of the initial marginal model by the deployed agent, and performing distillation training to update the initial marginal model, resulting in an evolved marginal model. Distillation training refers to the process of using a large visual model as the teacher model and the initial marginal model as the student model, providing supervision signals through the classification vectors output by the teacher model and intermediate layer feature maps, updating the weights of the student model, and obtaining the evolved marginal model.
[0026] A large-scale visual model refers to a large-scale visual foundational model deployed in a server domain and running completely offline. It possesses zero-shot segmentation and open-vocabulary detection capabilities, serving as a teacher model for distillation training. The training pool refers to the pseudo-labels stored in the server domain for object detection tasks. Pseudo-labels are training labels generated by a contrast agent deployed in the edge device domain. These labels are generated by cross-validating the predicted labels output by the large-scale visual model across at least two dimensions and adaptively calibrating the prediction confidence by invoking synchronous physical validation data from the acquisition agent. Pseudo-labels can replace manual annotation.
[0027] The classification vector refers to the class probability distribution output by the large visual model after processing the input image. It serves as the teacher signal for the distillation loss during distillation training. The intermediate layer feature map refers to the spatial feature representation output by the hidden layers of the large visual model. It also serves as the teacher signal for the feature distillation loss during distillation training.
[0028] Distillation training can compress the model into a lightweight model suitable for deployment on edge devices while preserving the generalization ability of the teacher model. Adaptive adjustment of distillation parameters according to the decay type makes retraining targeted and avoids the redundant overhead of full retraining and improves the scene adaptation efficiency of the new model.
[0029] Evolutionary edge model refers to a new version of edge model with improved performance obtained after distillation training. The deployment agent deploys the evolutionary edge model to the edge device and makes it run in place of the initial edge model.
[0030] Optionally, the agent can be deployed with a built-in hardware detection module. After receiving the evolutionary edge model, it can automatically identify the chip model of the edge device and dynamically select the corresponding compilation chain and quantization strategy according to the chip model. For example, when the chip is Rockchip RK3588, RKNN-Toolkit2 is selected for INT8 asymmetric quantization; when the chip is HiSilicon Hi3559A, NNIE is selected for INT16 or FP16 hybrid quantization; when the chip is NVIDIA Jetson, TensorRT is selected for FP16 and INT8 hybrid quantization; and when the chip is general ARM, ONNX Runtime is selected for INT8 dynamic quantization, thereby generating a model format adapted to the target chip.
[0031] After deployment, the deployed agent continues to monitor the input data features, confidence distribution, and prediction consistency of the evolutionary edge model. When concept drift occurs, it triggers the model evolution process again, forming a lifelong learning closed loop from perceptual decay to safe evolution.
[0032] This application's technical solution, under an air-gap isolation architecture between the edge device domain and the server domain, enables edge models to continuously adapt to business changes in a completely offline environment through the collaborative work of deployed agents, fine-tuned agents, and a large visual model. This solves the problem of edge models being unable to continuously evolve under physical isolation conditions. In the edge device domain, the deployed agent automatically detects model performance degradation and determines the degradation type and its corresponding distillation strategy parameters based on the input data features, confidence distribution, and prediction consistency of the initial edge model. This achieves autonomous drift detection and diagnosis without manual annotation or external network dependence. In the server domain, the fine-tuned agent updates the initial edge model using the classification vectors and intermediate layer feature maps output by pseudo-labels in the training pool, based on the distillation strategy parameters corresponding to the degradation type. This allows the model update strategy to adaptively adjust according to the cause of degradation, improving the retraining's relevance and efficiency. Air-gap isolation ensures that data does not leave the campus, meeting security and compliance requirements.
[0033] In an optional embodiment, the step of determining the decay type of the initial edge model and its corresponding distillation strategy parameters by deploying an agent based on the input data features of the initial edge model and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task includes: the deploying agent determining whether an input feature shift has occurred based on the proportion of frames with prediction confidence below a confidence threshold within a sliding window, the number of times the predicted label of the same target changes in consecutive frames, and the distribution distance between the input data features and the statistical feature baseline; if an input feature shift occurs, the deploying agent determines the spatiotemporal range of the feature shift and the range of change of the prediction confidence value, and determines the decay type of the initial edge model; the deploying agent determines the corresponding distillation strategy parameters for the decay type according to the preset mapping relationship between the decay type and the distillation strategy parameters.
[0034] Among them, the sliding window refers to the continuous frame sampling interval set by the deployed intelligent agent for statistical indicators during continuous monitoring. The window slides and updates frame by frame with the video stream to ensure the timeliness of the statistical results.
[0035] The deployed intelligent agent continuously monitors the initial edge model during runtime in the edge device domain, collecting three types of metrics data through a sliding window approach. The first metric is the percentage of frames with prediction confidence below a confidence threshold. Prediction confidence is the degree of certainty the initial edge model assigns to its prediction results for each inference, typically ranging from 0 to 1; a higher value indicates stronger confidence in the current prediction. The confidence threshold is a preset critical value used to distinguish between high-confidence and low-confidence predictions. When the prediction confidence is below the threshold, it is considered low confidence, indicating a significant increase in uncertainty about the current input by the initial edge model. When the percentage of low-confidence frames within the sliding window continues to rise abnormally, it indicates that the initial edge model generally lacks confidence in the current scene. The second metric is the number of times the predicted label for the same target changes in consecutive frames. The same target is a detected object associated with the same target identifier across consecutive frames through a target tracking algorithm. The predicted label is the category attribute judgment result output by the initial edge model after inference for that target. When the predicted label for the same target identifier changes frequently within a short period, it indicates that the initial edge model's inference is disordered. The third type of metric is the distribution distance between the input data features and the statistical feature baseline. The statistical feature baseline is the benchmark for the statistical distribution of input features in the training data used during the model training phase. The distribution distance is a measure of the statistical difference between the current input and the baseline. When the distribution distance increases significantly, it indicates that the current input data has deviated from the statistical pattern of the training data, that is, an input feature shift has occurred.
[0036] The deployed agent uses the above three types of indicators to determine whether input feature shift has occurred. If input feature shift is confirmed, the deployed agent further analyzes the spatiotemporal range of the shift and the range of changes in the prediction confidence value to determine the type of decay. The spatiotemporal range describes whether the shift is concentrated in a specific time period or region, or whether it is distributed throughout the entire time period or globally. The range of changes in the prediction confidence value reflects the fluctuations and overall trends in the prediction behavior of the initial marginal model within the shift interval.
[0037] If the offset is concentrated only in a specific time period or from a specific camera viewpoint, and differs significantly from data patterns outside this spatiotemporal range, it is diagnosed as an environmental change. If the offset is global and the confidence of all categories decreases synchronously and continuously, it is diagnosed as model aging. If the prediction confidence of the initial marginal model for a specific category drops sharply, and the input data features for that category have insufficient sample size in the training baseline or belong to a new category that did not appear during training, it is diagnosed as a change in task requirements.
[0038] After determining the decay type, the deployed agent automatically matches the corresponding distillation strategy parameters to the diagnosed decay type based on a preset mapping relationship. Optionally, model aging corresponds to increasing the distillation temperature, environmental changes correspond to increasing the feature distillation loss weight, and changes in business requirements correspond to generating new category pseudo-labels and new training code.
[0039] The aforementioned technical solution, by deploying an intelligent agent on the edge device domain to continuously monitor the initial edge model during runtime, achieves autonomous perception of input feature shifts without external annotation or human intervention. This is based on three ground-value-free indicators: the percentage of frames with prediction confidence below a confidence threshold within a sliding window, the number of times the predicted label for the same target changes in consecutive frames, and the distribution distance between input data features and the statistical feature baseline. Once an input feature shift is confirmed, the deployed agent further distinguishes the degradation type—environmental change, model aging, or task requirement change—by combining the spatiotemporal range of the feature shift and the range of changes in prediction confidence values, thus achieving accurate diagnosis of the causes of model performance degradation. Furthermore, based on the preset mapping relationship between degradation type and distillation strategy parameters, the deployed agent automatically matches the corresponding distillation strategy parameters to the diagnosed degradation type. This ensures that subsequent model update strategies accurately correspond to the causes of degradation, improving the retraining's relevance and efficiency, and avoiding the redundant overhead of full retraining.
[0040] In an optional embodiment, the step of updating the initial edge model by fine-tuning the agent based on the distillation strategy parameters corresponding to the decay type, and the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, to obtain an evolved edge model, includes: if the decay type is model aging or environmental change, then in the server domain, the fine-tuning agent loads the current training code of the initial edge model, and extracts the target distillation temperature and target loss weight from the distillation strategy parameters respectively; the fine-tuning agent adjusts the distillation temperature in the current training code based on the target distillation temperature, or adjusts the weight of the distillation loss in the total loss of the current training code based on the target loss weight to obtain a code adjustment result; the fine-tuning agent performs distillation training on the initial edge model based on the code adjustment result, and the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels corresponding to the target detection task in the training pool, to obtain an evolved edge model.
[0041] The fine-tuning agent first loads the current training code of the initial edge model, which is a complete executable script previously generated and used to train the current initial edge model. The fine-tuning agent extracts the target distillation temperature and target loss weights from the distillation policy parameters. The target distillation temperature is a preset temperature value used to replace the distillation temperature parameter in the current training code when the decay type is model aging. The target loss weights are preset weight values used to adjust the proportion of distillation loss in the total loss when the decay type is environmental change.
[0042] When the decay type is model aging, the fine-tuning agent adjusts the distillation temperature in the current training code based on the target distillation temperature, raising the distillation temperature from the default value to the target distillation temperature. This allows the student model to learn the inter-category similarity relationship implied in the classification vector output by the teacher model in a smoother way, thereby alleviating the weight solidification problem that occurs after long-term operation of the model.
[0043] When the decay type is environmental change, the fine-tuning agent adjusts the weight of distillation loss in the total loss of the current training code based on the target loss weight. The weight of feature distillation loss is increased from the default value to the target loss weight, which strengthens the student model's imitation of the intermediate layer feature map output by the teacher model, and enables the model to quickly adapt to changes in data distribution caused by changes in external conditions such as lighting, weather, and background.
[0044] After the above adjustments are completed, the code adjustment results are obtained. Based on the code adjustment results, as well as the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels corresponding to the object detection task in the training pool, the fine-tuning agent performs distillation training on the initial edge model to obtain the evolved edge model.
[0045] The above technical solution distinguishes between two types of degradation: model aging and environmental change. In the server domain, a fine-tuning agent directly loads existing training code and only adjusts the distillation temperature or loss weights, eliminating the need to regenerate training code. This achieves rapid evolution of the edge model at minimal cost. Increasing the distillation temperature to address model aging softens the distribution of the teacher signal and restores the generalization ability of the student model. Increasing the feature distillation loss weights to address environmental changes strengthens the alignment of the student model with the teacher's intermediate layer features, accelerating adaptation to new scenarios. Thus, under the constraint of air gap isolation, the edge model's continuous adaptability to business changes is ensured in a low-overhead and highly targeted manner.
[0046] In an optional embodiment, deploying the evolutionary edge model to the edge device via a deployment agent includes: the deployment agent loading the evolutionary edge model into a spare model slot and performing inference in parallel with the initial edge model running in the currently used model slot; within a preset monitoring period, the deployment agent comparing the operating metrics of the evolutionary edge model and the initial edge model; when the operating metrics of the evolutionary edge model are better than those of the initial edge model, the deployment agent switches the inference service to the spare model slot; when the operating metrics of the initial edge model are better than those of the evolutionary edge model, the deployment agent maintains the operation of the currently used model slot.
[0047] When deploying agents to execute the deployment of evolutionary edge models in the edge device domain, a dual-slot loading mechanism is adopted. The spare model slot refers to an independent storage area reserved in the edge device's memory for receiving new version models; the current model slot refers to the memory area currently running the initial edge model and providing responses for inference services.
[0048] The deploying agent loads the evolutionary edge model into the backup model slot. At this time, the initial edge model in the current model slot continues to provide inference services, and the two model slots do not affect each other. After loading, the evolutionary edge model and the initial edge model perform inference in parallel, that is, they independently perform forward propagation on the same input and produce prediction results. However, the output of the evolutionary edge model is only used for index recording and not for actual decision-making. This is the shadow mode operation.
[0049] Within a preset monitoring period, the deployed agent continuously compares the performance metrics of the evolving edge model and the initial edge model. These performance metrics include at least one of inference accuracy, inference latency, and resource consumption. When the performance metrics of the evolving edge model are better than those of the initial edge model, the deployed agent atomically switches the inference service from the current model slot to the backup model slot, with the evolving edge model taking over the inference service without interruption. When the performance metrics of the evolving edge model are worse than those of the initial edge model, the deployed agent maintains the operation of the current model slot, i.e., no switch is performed, and the evolving edge model is discarded or marked as invalid.
[0050] The above technical solution utilizes a dual-slot loading mechanism of backup and current model slots to enable the evolving edge model and the initial edge model to perform inference in parallel on the real data stream. Online metric comparison is conducted in shadow mode, allowing for the verification of the new model's performance without interrupting the inference service. The inference service is switched only when the evolving edge model's performance metrics are superior to the initial edge model; otherwise, the current model slot continues to operate. This provides an automated security check and rollback defense for model evolution deployment under air-gap isolation conditions, ensuring business continuity.
[0051] In an optional embodiment, the method further includes: the fine-tuning agent updating the initial edge model in a sandbox, and after obtaining the evolved edge model, digitally signing the evolved edge model and binding it to the hardware identifier of the edge device; before deploying the evolved edge model, the deployment agent verifying the digital signature and hardware identifier of the evolved edge model through the hardware root of trust of the security chip built into the edge device, and refusing to load the evolved edge model if the verification fails.
[0052] Within the server domain, the fine-tuning agent updates the initial edge model within a sandbox. A sandbox is a restricted execution environment isolated from other system resources within the server domain. The fine-tuning agent executes distilled training code within this environment. All file read / write and computation operations during training are confined within the sandbox boundaries, preventing potential defects or malicious logic in the training code from affecting other services and data within the server domain.
[0053] After the model update is complete, the fine-tuning agent digitally signs the evolutionary edge model and binds it to the hardware identifier of the edge device. The digital signature involves using a private key to calculate a hash value for the weight file of the evolutionary edge model and encrypting it to generate signature data. This signature data can be subsequently verified using the public key to confirm whether the model file has been tampered with. The hardware identifier is the unique serial number or hardware fingerprint of the security chip within the edge device. Binding the hardware identifier means that the evolutionary edge model is only authorized to run on the specified edge device. The model file is a complete set of model data saved after training, typically containing two parts: the model structure definition and the weight parameters. The model structure definition describes the architectural information of the neural network, such as the number of layers, the type of each layer, and the connection methods. The weight parameters are the parameter values learned by the model through backpropagation during training, including the convolutional kernel weights and biases of convolutional layers, and the weight matrix and bias vector of fully connected layers. These weight parameters determine the specific mapping method of the model to the input data and are the core carrier of the model's capabilities.
[0054] In the edge device domain, before deploying the evolutionary edge model, the deploying agent verifies the digital signature and hardware identifier of the evolutionary edge model through the hardware root of trust in the security chip built into the edge device. The hardware root of trust refers to the immutable key and verification logic embedded within the security chip. The deploying agent invokes the hardware root of trust to verify the digital signature, confirming that the model file has not been tampered with since its self-signing, and simultaneously compares the hardware identifier bound to the model with the hardware identifier of the current edge device. If the verification fails, the evolutionary edge model is refused to be loaded, the model update process is aborted, and the initial edge model in the current model slot continues to provide services.
[0055] The above technical solution updates the initial edge model within a sandbox, isolating the training process from server domain system resources and preventing the spread of security risks at the source. The fine-tuning agent digitally signs the evolving edge model and binds it to a hardware identifier, ensuring the model carries integrity verification credentials and anti-copying protection from its inception. The deployment agent verifies the digital signature and hardware identifier using the hardware root of trust of the secure chip; if verification fails, it refuses to load the evolving edge model. This constructs a chain of trust transfer from sandbox training and signature binding to hardware verification, preventing model tampering or unauthorized copying.
[0056] Example 2 Figure 2 This is a flowchart of the edge model evolution method for air gap isolation environments provided in Embodiment 2. This embodiment is a further optimization based on the above embodiments.
[0057] like Figure 2 As shown, the method includes: S210. In the edge device domain, by deploying an agent, the decay type of the initial edge model and its corresponding distillation strategy parameters are determined based on the input data features of the initial edge model in the edge device and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task.
[0058] S220. When the decay type is a change in task requirements, the fine-tuning agent in the server domain uses the visual big model to infer the keyframes to be identified corresponding to the change in task requirements, obtains new category pseudo-labels, and incorporates them into the training pool.
[0059] When the degradation type determined by the deploying agent is a change in task requirements, it indicates that the set of detection categories for the object detection task has changed, such as the addition of violation categories or object types that were not trained in the initial edge model.
[0060] In the server domain, the fine-tuning agent uses a large visual model to infer keyframes corresponding to changes in task requirements. Keyframes to be identified refer to image frames that the acquisition agent selects from the video stream in the end-device domain based on the new task description, potentially containing new categories of targets. The large visual model, leveraging its open-vocabulary detection capabilities, performs zero-shot inference on these keyframes, outputting predicted labels and their prediction confidence for the new category. During this process, the large visual model does not rely on pre-labeled samples of the new category but directly identifies them through semantic information in the task description. The fine-tuning agent uses the new category predicted labels obtained from the large visual model as pseudo-labels for the new category and incorporates them into the training pool, thereby rapidly expanding the training data for the new category. The training pool is used to aggregate and manage pseudo-labels; the addition of new category pseudo-labels expands the category range covered by the training pool.
[0061] S230, The fine-tuning agent calls the code-generating agent, which generates new training code based on the semantic context of the target detection task, the data structure of the new category pseudo-label in the training pool, and the hardware specifications of the edge device.
[0062] The code generation agent is an autonomous software unit deployed in a system agent cluster within a server domain. It is responsible for understanding the data structure of pseudo-labels in the training pool and generating training code containing lightweight model architecture definitions and distillation loss based on semantic context, data structure, and hardware specifications. The fine-tuning agent invokes the code generation agent, which generates new training code based on the semantic context of the object detection task, the data structure of new category pseudo-labels in the training pool, and the hardware specifications of the edge device. Semantic context refers to semantic information related to the object detection task, such as task descriptions and target historical states, used to guide the code generation agent in understanding the specific requirements of the current task. Data structure refers to the organization format and field definitions of new category pseudo-labels in the training pool. Hardware specifications refer to constraints such as computing power, memory, and inference latency of the target edge device.
[0063] The code-generating agent analyzes the data structure to determine the task type and updated category distribution, selects a suitable lightweight model architecture based on hardware specifications, and includes distillation loss definition, quantization scripts adapted to the target chip, and model signature code in the generated training code. This new training code is adapted to the updated number of categories and hardware constraints.
[0064] S240. The fine-tuned agent, based on the new training code and the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, performs distillation training on the initial edge model to obtain the evolved edge model.
[0065] The fine-tuned agent, based on new training code and using a large visual model to distill the initial marginal model with classification vectors and intermediate layer feature maps output from pseudo-labels in the training pool, yields an evolved marginal model. The classification vectors and intermediate layer feature maps serve as teacher-supervised signals to guide the student model's learning.
[0066] Distillation training enables evolutionary edge models to retain the ability of large visual models to recognize new categories while meeting the real-time and resource constraints of edge devices. S250. In the end device domain, the evolutionary edge model is deployed to the edge device by deploying an intelligent agent; wherein, there is an air gap isolation between the end device domain and the server domain.
[0067] In the edge device domain, the deployed agent will deploy the evolved edge model to the edge device, replacing the initial edge model and completing the expansion and update of the model's capabilities.
[0068] In this application's technical solution, when the decay type is a change in task requirements, the fine-tuning agent utilizes the open vocabulary capability of the large visual model to infer keyframes to be identified, generating new category pseudo-labels and incorporating them into the training pool, thus expanding the training data without manual annotation. The fine-tuning agent calls the code-generating agent, which automatically generates new training code based on semantic context, the data structure of the new category pseudo-labels, and the hardware specifications of the edge device, adapting to the updated number of categories and hardware constraints. Based on the new training code and the classification vectors and intermediate layer feature maps output by the large visual model, the initial edge model is distilled and trained to obtain an evolved edge model, enabling the edge model to detect new category targets under air gap isolation conditions, achieving rapid response to changes in task requirements.
[0069] In an optional embodiment, the code generation agent generates new training code based on the semantic context of the object detection task, the data structure of the new category pseudo-labels in the training pool, and the hardware specifications of the edge device. This includes: the code generation agent determining the task type and category distribution of the object detection task based on the semantic context and the data structure; the code generation agent selecting an appropriate lightweight model architecture according to the hardware specifications and generating model training code containing the definition of the lightweight model architecture; the code generation agent generating distillation loss in the model training code and generating a quantization script and model signature code adapted to the target chip of the edge device to obtain the new training code; wherein the distillation loss includes feature alignment loss between the intermediate layer feature map output by the large visual model and the intermediate layer features of the lightweight model architecture, and response matching loss between the classification vector output by the large visual model and the classification output of the lightweight model architecture.
[0070] The code generation agent, an autonomous software unit deployed in the server domain, is responsible for receiving semantic context, data structure, and hardware specifications, assembling them into prompt information according to a structured prompt template, inputting it into the large code generation model, and driving it to complete reasoning and code generation according to a preset thought chain. The large code generation model is a large language model with code generation capabilities; it is the underlying inference engine invoked by the code generation agent. First, the code generation agent determines the task type and category distribution of the object detection task based on semantic context and data structure. Semantic context refers to semantic information such as task descriptions and target historical states related to the object detection task, enabling the large-scale code generation model to understand the specific requirements of the current task. Data structure refers to the organization format and field definitions of new category pseudo-labels in the training pool, including specifications such as image paths, bounding box coordinates, and category labels. The large-scale code generation model performs insight analysis on the data structure, identifying whether the task type is object detection, object segmentation, or object tracking; it counts the updated number of categories and the sample distribution of each category; it infers the size and normalization method of the input image; and it confirms data quality indicators such as field missing rates, thereby completing the analysis of data distribution and task type.
[0071] Next, the code-generating agent selects a suitable lightweight model architecture based on the hardware specifications and generates model training code containing the definition of that lightweight model architecture. Hardware specifications refer to constraints such as computing power, memory, and inference latency of the target edge device. The code-generating large model selects an architecture based on the hardware specifications within the thought process and explains the reasons for the selection. For example, when computing power is below a preset threshold and latency requirements are strict, a lightweight detection architecture is selected, generating the corresponding model structure definition code.
[0072] Subsequently, the code-generating agent generates distillation loss within the model training code. This distillation loss comprises two parts: feature alignment loss and response matching loss. Feature alignment loss is calculated based on the difference between the intermediate layer feature maps output by the large visual model and the intermediate layer features of the lightweight model architecture, used to transfer the feature representation capabilities of the teacher model to the student model. Response matching loss is calculated based on the difference between the classification vector output by the large visual model and the classification output of the lightweight model architecture, using the classification vector as a soft label to guide the student model's learning by softening the differences in probability distribution.
[0073] Finally, the code generation agent generates quantization scripts and model signature code adapted to the target chip of the edge device, resulting in new training code. The quantization script is used to perform post-training quantization on the model after distillation training, converting it to a format supported by the target chip. The model signature code is used to digitally sign the trained model file, ensuring model integrity.
[0074] The above technical solution, through a code generation agent, drives a large-scale code model to generate training code according to a pre-defined thought chain based on structured prompt templates. This fully automates data structure analysis, lightweight architecture selection, distillation loss generation, quantization scripts, and model signature code, lowering the technical barrier to model updates under conditions of air-gap isolation and without the need for on-site professional AI engineers. The generated training code includes feature alignment loss and response matching loss, ensuring that distillation training effectively transfers the feature representations and classification knowledge of the large-scale visual model to the lightweight edge model. Simultaneously, the quantization scripts and signature code adapt the model to the target chip and embed integrity protection during the generation stage.
[0075] Example 3 Figure 3 This is a flowchart of the edge model evolution method for air-gap isolation environments provided in Embodiment 3. This embodiment further optimizes the above embodiments. Specifically, it defines the method for obtaining pseudo-labels in the training pool.
[0076] like Figure 3 As shown, the method includes: S310. By collecting multimodal sensor data from the target detection task in the end device domain of the intelligent agent, the data is fused and filtered to extract the key frame to be identified corresponding to the target detection task, the semantic context associated with the key frame to be identified, and the physical verification data, and the key frame to be identified and its associated semantic context are uploaded to the server domain.
[0077] In this context, the data acquisition agent refers to an autonomous software unit deployed in an edge agent cluster within the end-device domain. This unit is responsible for fusing and pre-screening multimodal sensor data, generating keyframes to be identified corresponding to the target detection task, and simultaneously acquiring the semantic context and physical verification data associated with the keyframes. Multimodal sensor data refers to a time-stamped collection of sensor data synchronously acquired by different types of sensors, including visible light images, infrared temperature data, and vibration amplitude.
[0078] In the edge device domain, the acquisition agent fuses and filters multimodal sensor data from cameras (video streams), infrared thermal imagers (temperature readings), and vibration sensors for target detection tasks. Based on the task description, the acquisition agent selects candidate image frames from the multimodal sensor data that may contain the target of interest, and scores them comprehensively. Frames with scores above a preset threshold are identified as keyframes to be identified. Simultaneously, the acquisition agent extracts the semantic context and physical verification data associated with the keyframes to be identified. Semantic context refers to semantic information such as the task description and target historical state associated with the keyframes to be identified, used to assist the large-scale visual model in understanding the focus and temporal background of the image. Physical verification data refers to numerical observation data synchronously acquired by the sensors and aligned with the keyframes to be identified in time, including infrared temperature readings, target bounding box coordinates synchronously obtained by the acquisition agent when extracting the keyframes, historical predicted labels of the same target in previous frames, and scene identifiers or region attributes corresponding to the keyframes to be identified. The acquisition agent uploads the keyframes to be identified and their associated semantic context to the server domain, while the physical verification data remains in the edge device domain.
[0079] S320. In the server domain, the visual big model outputs a predicted label and its corresponding prediction confidence based on the key frame to be identified and its associated semantic context, and then sends the predicted label and its corresponding prediction confidence back to the end device domain.
[0080] In the server domain, the visual big model receives keyframes to be identified and their associated semantic context. Based on its zero-shot inference capabilities, the visual big model analyzes the targets in the keyframes under the guidance of the semantic context, outputting predicted labels and their corresponding prediction confidence scores. The visual big model then sends the predicted labels and their corresponding prediction confidence scores back to the end device domain.
[0081] S330. In the terminal device domain, by comparing the predicted labels output by the intelligent agent based on the physical verification data in at least two dimensions to obtain the verification result, and calibrating the prediction confidence based on the verification result to obtain the calibration confidence.
[0082] Among them, the contrast agent refers to the autonomous software unit deployed in the edge agent cluster of the end device domain. It is responsible for performing multi-dimensional verification of the predicted labels output by the large visual model based on physical verification data, calibrating the prediction confidence, generating pseudo labels and storing them in the training pool.
[0083] In the end-device domain, the comparison agent receives the returned predicted labels and prediction confidence scores, and uses the retained physical verification data to perform cross-validation on the predicted labels in at least two dimensions. The comparison agent treats the predicted labels as hypotheses to be verified, and uses the physical verification data to examine whether the predicted labels are consistent with the local observation data in at least two dimensions, including geometric consistency, physical rationality, temporal continuity, and semantic self-consistency, to obtain the verification results for each dimension.
[0084] Subsequently, the comparative agent calibrates the prediction confidence based on the verification results to obtain the calibration confidence. That is, the original confidence is adaptively adjusted according to the pass rate of the verification dimensions. If all verifications pass, the confidence is maintained or increased, and if a certain dimension fails, a penalty discount is applied.
[0085] S340. The comparison agent determines the pseudo-label corresponding to the target detection task based on the predicted label and its corresponding calibration confidence, and stores the pseudo-label in the training pool of the server domain.
[0086] The comparative agent compares the calibrated confidence level corresponding to the predicted label with the preset confidence level, and stores the pseudo-labels with calibrated confidence levels higher than the preset confidence levels into the training pool of the server domain.
[0087] The technical solution of this application involves a smart agent fusing and filtering multimodal sensor data, extracting keyframes to be identified, and separating their semantic context and physical verification data. Only the semantic context is uploaded to the server domain for the large visual model to infer predicted labels and prediction confidence, while the physical verification data is stored in the end device domain. The smart agent then performs multi-dimensional verification of the predicted labels based on the physical verification data and calibrates the prediction confidence, obtaining high-confidence pseudo-labels which are stored in the training pool. Thus, under air-gap isolation conditions, pseudo-labels meeting training quality requirements can be generated without manual annotation, providing a reliable data foundation for model evolution.
[0088] In an optional embodiment, the comparison agent verifies the predicted label output by the large visual model in at least two dimensions based on the physical verification data to obtain a verification result, and calibrates the prediction confidence based on the verification result to obtain a calibration confidence. This includes: the comparison agent parses the physical verification data and the predicted label to obtain a reference bounding box and a predicted bounding box; wherein the reference bounding box is obtained synchronously when extracting the keyframe to be identified; the comparison agent verifies whether the geometric proportions of the predicted bounding box conform to preset structural constraints based on the reference bounding box to obtain geometric consistency; and the comparison agent verifies whether the current state represented by the predicted label conforms to the sensor values in the physical verification data. The system verifies the physical rationality by adhering to preset physical constraints. The comparison agent verifies the temporal continuity by checking whether the evolution of the predicted label between adjacent keyframes satisfies preset temporal constraints based on historical predicted labels of the same target in the physical verification data. The comparison agent also verifies the semantic consistency by checking whether the target behavior or current state represented by the predicted label conforms to preset association rules based on the scene identifier or region attribute corresponding to the keyframe in the physical verification data. Finally, the comparison agent calibrates the prediction confidence based on at least two dimensions of geometric consistency, physical rationality, temporal continuity, and semantic consistency.
[0089] In the end-device domain, the comparative agent performs localized verification and calibration of the predicted labels and their corresponding prediction confidence scores returned by the large visual model. First, the comparative agent parses the physical verification data and predicted labels to obtain reference bounding boxes and predicted bounding boxes. The reference bounding box refers to the position coordinate information synchronously obtained by the acquisition agent during the pre-screening process when extracting the keyframe to be identified, and is stored in the physical verification data. The predicted bounding box refers to the target position coordinates output by the large visual model after inference on the same keyframe to be identified. Based on the reference bounding box, the comparative agent verifies whether the geometric proportions of the predicted bounding box conform to preset structural constraints. For example, it checks whether the aspect ratio of the predicted bounding box is consistent with the geometry of the reference bounding box, and whether the target is located within a reasonable spatial range, obtaining a geometric consistency verification result.
[0090] Subsequently, the agent compares the sensor values in the physical verification data to verify whether the current state represented by the predicted label conforms to preset physical constraints. Sensor values refer to readings of physical quantities such as infrared temperature and vibration amplitude that are aligned with the time of the key frame to be identified. Preset physical constraints refer to the physical laws that the sensor values should satisfy under the target state. For example, when the predicted label determines that the target is a person, it checks whether the infrared temperature reading is within the range of human body temperature, thereby obtaining the physical rationality verification result.
[0091] The comparative agent also verifies whether the evolution of the current predicted label between adjacent frames satisfies preset temporal constraints based on the historical predicted labels of the same target in the physical verification data between adjacent frames. The same target is tracked by association through target identifiers, and the historical predicted labels refer to the label records that the target identifier was assigned and verified by the visual big data model in previous frames. When the predicted label of the same target changes between adjacent frames, it is checked whether the evolution conforms to a reasonable temporal pattern, thereby obtaining the temporal continuity verification result.
[0092] Furthermore, the comparative agent verifies whether the target behavior or current state represented by the predicted label conforms to preset association rules based on the scene identifier or region attribute corresponding to the key frame to be identified in the physical verification data. The scene identifier refers to the scene category label corresponding to the frame when it was collected, and the region attribute refers to the predefined spatial region type corresponding to the frame. When the predicted label indicates that the target is in a certain state or behavior, it checks whether the state or behavior is semantically consistent with the scene or region in which it is located, thereby obtaining the semantic self-consistency verification result.
[0093] Finally, by comparing the verification results of at least two dimensions among the agent's comprehensive geometric consistency, physical rationality, temporal continuity, and semantic self-consistency, the prediction confidence is adaptively calibrated to obtain the calibration confidence. Optional calibration methods are: if all dimensions involved in the verification pass, maintain or increase the prediction confidence; if a dimension fails, apply a corresponding penalty discount to the prediction confidence based on the importance of the failed dimension.
[0094] The above technical solution, under air gap isolation conditions, uploads the semantic context to the server domain for inference by the large visual model, and stores the physical verification data in the terminal device domain for the comparison agent to perform multi-dimensional verification of the predicted labels and calibrate the prediction confidence. High-confidence pseudo-labels can be generated and stored in the training pool without manual annotation, providing a reliable data foundation for model evolution.
[0095] Example 4 Embodiment 4 of this application provides an edge model evolution system for air gap isolation environments. This embodiment can be applied to scenarios in air gap isolation environments where edge models are used for target detection and the detection task needs to be adjusted with changes in business, such as new violation categories in chemical industrial parks, new foreign objects appearing in rail transit, or new instruments being introduced into operating rooms. The system can be integrated into electronic devices such as smart terminals.
[0096] The system may include: a deployment agent, an edge device, a large visual model, and a fine-tuning agent. The deployment agent and the edge device belong to the end device domain, while the large visual model and the fine-tuning agent belong to the server domain. There is an air gap between the end device domain and the server domain. In the edge device domain, by deploying an agent based on the input data features of the initial edge model in the edge device and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task, the decay type of the initial edge model and its corresponding distillation strategy parameters are determined. In the server domain, the initial edge model is updated by fine-tuning the distillation strategy parameters of the agent based on the decay type, as well as the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, to obtain the evolutionary edge model. In the edge device domain, the evolutionary edge model is deployed to the edge device by deploying an agent.
[0097] Figure 4 This is a schematic diagram of the structure of an edge model evolution system for an air gap isolation environment according to Embodiment 4 of this application. See also... Figure 4 This system adopts a three-layer, two-domain air-gap isolation architecture. The two domains are the end device domain and the private server domain, which are connected by a physically isolated local area network. This isolated network uses unidirectional fiber optic or cable with the WiFi 6 proprietary protocol, allowing only whitelisted protocols to pass through and communicating with certificate fixation and transport layer security encryption to ensure that data does not leave the campus.
[0098] In the edge device domain, camera devices such as Camera A and Camera B, which acquire visible light video streams with differentiated parameters, and sensor devices such as Sensor C, which acquire multimodal sensing data such as infrared temperature and vibration, are deployed. The video streams and sensor data acquired by the above devices are connected to the edge devices, which run on ARM or TPU chips. Lightweight models are deployed on these edge devices to perform real-time inference, and a data cache is provided to temporarily store data to be processed. The edge device domain also deploys an edge agent cluster, consisting of three types of agents: acquisition agents, comparison agents, and deployment agents. The acquisition agents are responsible for fusing and filtering multimodal sensor data, automatically extracting coordinates, time series, quality signals, and ROI regions, i.e., extracting the keyframes to be identified corresponding to the target detection task, and simultaneously acquiring the semantic context and physical verification data associated with the keyframes to be identified. The comparison agents are responsible for performing multi-dimensional cross-validation between the predicted labels and predicted confidence scores returned by the large visual model and the physical verification data retained by the acquisition agents, generating pseudo-labels and calibrating confidence scores. The deployment agent is responsible for model quantization, compilation, and hot updates. During the model deployment phase, it loads the evolutionary edge model onto the edge device and performs dual-slot switching.
[0099] In a private server domain, a large visual model and a system agent cluster are deployed. The system agent cluster supports the autonomous evolution of the model. The large visual model is a large-scale visual foundational model such as SAM3 or Qwen2.5-VL, possessing zero-shot segmentation capability and open-vocabulary detection capability. It operates completely offline. During the pseudo-label generation stage, it infers the keyframes to be identified uploaded from the end device domain and outputs predicted labels and prediction confidence. During the model update stage, it provides classification vectors and intermediate layer feature maps as supervision signals for distillation training. The system agent cluster consists of three agents: a code generation agent, a fine-tuning agent, and a version management agent. The code generation agent is responsible for understanding the data structure of pseudo-labels in the training pool and driving the code generation large model to generate complete training code, including lightweight model architecture definition, distillation loss, quantization script, and model signature code. The fine-tuning agent is responsible for using the large visual model as the teacher model and the initial edge model as the student model, performing distillation training based on the training code generated by the code generation agent to generate an evolving edge model. Version management agents are autonomous software units deployed in a cluster of system agents in a server domain. They are responsible for model version control and A / B testing effect tracking, and perform version archiving and lineage tracing for the evolutionary edge models produced by each model update.
[0100] During system operation, the acquisition agent in the edge device domain extracts the keyframes to be identified, along with their associated semantic context and physical verification data, from multimodal sensor data. It then uploads the keyframes and semantic context to the server domain, while the physical verification data remains in the edge device domain. The large-scale visual model in the private server domain infers the keyframes to be identified based on the semantic context, outputting predicted labels and their corresponding prediction confidence scores, which are then sent back to the edge device domain. The comparison agent in the edge device domain verifies the predicted labels based on the retained physical verification data across at least two dimensions: geometric consistency, physical plausibility, temporal continuity, and semantic self-consistency. It also calibrates the prediction confidence scores, obtaining high-confidence pseudo-labels which are stored in the training pool of the server domain. The deployment agent in the edge device domain continuously monitors the initial edge model running on the edge devices, determining the decay type and distillation strategy parameters based on input data features, confidence distribution, and prediction consistency, triggering the model evolution process. The fine-tuning agent in the server domain, based on the distillation strategy parameters, calls the code generation agent to generate or reuse model training code. It then uses the classification vectors output by the large visual model and intermediate layer feature maps to perform distillation training on the initial edge model, obtaining an evolved edge model. After the version management agent archives the evolved edge model, the deployment agent deploys it to edge devices, completing the autonomous evolution loop.
[0101] This application's technical solution, under an air-gap isolation architecture between the edge device domain and the server domain, enables edge models to continuously adapt to business changes in a completely offline environment through the collaborative work of deployed agents, fine-tuned agents, and a large visual model. This solves the problem of edge models being unable to continuously evolve under physical isolation conditions. In the edge device domain, the deployed agent automatically detects model performance degradation and determines the degradation type and its corresponding distillation strategy parameters based on the input data features, confidence distribution, and prediction consistency of the initial edge model. This achieves autonomous drift detection and diagnosis without manual annotation or external network dependence. In the server domain, the fine-tuned agent updates the initial edge model using the classification vectors and intermediate layer feature maps output by pseudo-labels in the training pool, based on the distillation strategy parameters corresponding to the degradation type. This allows the model update strategy to adaptively adjust according to the cause of degradation, improving the retraining's relevance and efficiency. Air-gap isolation ensures that data does not leave the campus, meeting security and compliance requirements.
[0102] Example 5 According to embodiments of this application, this application also provides an electronic device, a readable storage medium, and a computer program product.
[0103] Figure 5A schematic diagram of an electronic device 510, which can be implemented using an embodiment, is shown. The electronic device 510 includes at least one processor 511 and a memory, such as a read-only memory (ROM) 512, a random access memory (RAM) 513, etc., communicatively connected to the at least one processor 511. The memory stores computer programs executable by the at least one processor. The processor 511 can perform various appropriate actions and processes based on the computer program stored in the ROM 512 or loaded from storage unit 518 into the RAM 513. The RAM 513 may also store various programs and data required for the operation of the electronic device 510. The processor 511, ROM 512, and RAM 513 are interconnected via a bus 514. An input / output (I / O) interface 515 is also connected to the bus 514.
[0104] Multiple components in electronic device 510 are connected to I / O interface 515, including: input unit 516, such as keyboard, mouse, etc.; output unit 517, such as various types of displays, speakers, etc.; storage unit 518, such as disk, optical disk, etc.; and communication unit 519, such as network card, modem, wireless transceiver, etc. Communication unit 519 allows electronic device 510 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0105] Processor 511 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 511 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 511 performs the various methods and processes described above, such as edge model evolution methods for air-gap isolated environments.
[0106] In some embodiments, the edge model evolution method for an air-gap isolated environment can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 518. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 510 via ROM 512 and / or communication unit 519. When the computer program is loaded into RAM 513 and executed by processor 511, one or more steps of the edge model evolution method for an air-gap isolated environment described above can be performed. Alternatively, in other embodiments, processor 511 can be configured to execute the edge model evolution method for an air-gap isolated environment by any other suitable means (e.g., by means of firmware).
[0107] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0108] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable edge model evolution device for an air-gap isolation environment, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0110] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0111] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as an edge model evolution server for an air-gap isolated environment), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0112] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0113] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the edge model evolution method for air-gap isolation environments provided in any embodiment of this application. This program product shares the same inventive concept as the edge model evolution method for air-gap isolation environments disclosed in the embodiments of this application, and therefore will not be described further here.
[0114] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0115] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. An edge model evolution method for air-gap isolated environments, characterized in that, The method includes: In the edge device domain, by deploying an agent based on the input data features of the initial edge model in the edge device and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task, the decay type of the initial edge model and its corresponding distillation strategy parameters are determined. In the server domain, the initial edge model is updated by fine-tuning the distillation strategy parameters of the agent based on the decay type, as well as the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, to obtain the evolutionary edge model. In the edge device domain, the evolutionary edge model is deployed to the edge device by deploying an agent; wherein, there is an air gap isolation between the edge device domain and the server domain.
2. The method according to claim 1, characterized in that, The process of determining the decay type of the initial edge model and its corresponding distillation strategy parameters by deploying an agent based on the input data features of the initial edge model and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task includes: The deployed agent determines whether input feature shift has occurred based on the proportion of frames with prediction confidence below the confidence threshold within the sliding window, the number of times the predicted label of the same target changes in consecutive frames, and the distribution distance between the input data features and the statistical feature baseline. If an input feature shift occurs, the deploying agent determines the spatiotemporal range of the feature shift and the range of change in the predicted confidence value, and determines the decay type of the initial edge model; The deployment agent determines the corresponding distillation strategy parameters for the decay type based on a preset mapping relationship between the decay type and the distillation strategy parameters.
3. The method according to claim 1, characterized in that, In the server domain, the initial edge model is updated by fine-tuning the distillation strategy parameters of the agent based on the decay type, and the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, to obtain an evolutionary edge model, including: If the decay type is model aging or environmental change, then the fine-tuning agent in the server domain loads the current training code of the initial edge model and extracts the target distillation temperature and target loss weight from the distillation strategy parameters respectively. The fine-tuning agent adjusts the distillation temperature in the current training code based on the target distillation temperature, or adjusts the weight of the distillation loss in the total loss of the current training code based on the target loss weight to obtain the code adjustment result; The fine-tuning agent performs distillation training on the initial edge model based on the code adjustment results, as well as the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels corresponding to the target detection task in the training pool, to obtain the evolved edge model.
4. The method according to claim 1, characterized in that, In the server domain, the process of updating the initial edge model by fine-tuning the distillation strategy parameters of the agent based on the decay type, and the classification vector and intermediate layer feature map output by the large visual model for the pseudo-labels in the training pool, to obtain an evolutionary edge model, further includes: When the decay type is a change in task requirements, the fine-tuning agent in the server domain uses the visual big model to reason about the key frames to be identified corresponding to the change in task requirements, obtains new category pseudo-labels, and incorporates them into the training pool. The fine-tuning agent calls the code-generating agent, which generates new training code based on the semantic context of the target detection task, the data structure of the new category pseudo-labels in the training pool, and the hardware specifications of the edge device. The fine-tuned agent, based on the new training code and the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, distills the initial edge model to obtain the evolved edge model.
5. The method according to claim 4, characterized in that, The code-generating agent generates new training code based on the semantic context of the object detection task, the data structure of the new category pseudo-labels in the training pool, and the hardware specifications of the edge device, including: The code generation agent determines the task type and category distribution of the target detection task based on the semantic context and the data structure. The code generation agent selects an appropriate lightweight model architecture based on the hardware specifications and generates model training code containing the definition of the lightweight model architecture. The code generation agent generates distillation loss in the model training code and generates quantization scripts and model signature code adapted to the target chip of the edge device to obtain the new training code; The distillation loss includes feature alignment loss between the intermediate layer feature map output by the large visual model and the intermediate layer features of the lightweight model architecture, and response matching loss between the classification vector output by the large visual model and the classification output of the lightweight model architecture.
6. The method according to claim 1, characterized in that, The pseudo-labels in the training pool are obtained in the following way: By collecting and filtering multimodal sensor data from the target detection task in the end device domain, the agent extracts the key frame to be identified, the semantic context associated with the key frame to be identified, and the physical verification data, and uploads the key frame to be identified and its associated semantic context to the server domain. In the server domain, the visual big model outputs predicted labels and their corresponding prediction confidence based on the key frame to be identified and its associated semantic context, and then sends the predicted labels and their corresponding prediction confidence back to the end device domain. In the terminal device domain, the predicted label output by the visual large model is verified in at least two dimensions by comparing the physical verification data to obtain the verification result, and the prediction confidence is calibrated based on the verification result to obtain the calibration confidence. The comparative agent determines the pseudo-label corresponding to the target detection task based on the predicted label and its corresponding calibration confidence, and stores the pseudo-label in the training pool of the server domain.
7. The method according to claim 6, characterized in that, The comparative agent verifies the predicted labels output by the large visual model in at least two dimensions based on the physical verification data to obtain verification results, and calibrates the prediction confidence based on the verification results to obtain calibration confidence, including: The comparative agent parses the physical verification data and the predicted label to obtain a reference bounding box and a predicted bounding box; wherein the reference bounding box is obtained synchronously when extracting the keyframe to be identified. The comparative agent verifies whether the geometric proportions of the predicted bounding box conform to the preset structural constraints based on the reference bounding box, thereby obtaining geometric consistency. The comparative agent verifies whether the current state represented by the predicted label conforms to preset physical constraints based on the sensor values in the physical verification data, thereby obtaining physical rationality. The comparative agent verifies whether the evolution of the predicted label between adjacent key frames to be identified satisfies the preset temporal constraints based on the historical predicted label of the same target in the physical verification data, thereby obtaining the temporal continuity. The comparative agent, based on the scene identifier or region attribute corresponding to the key frame to be identified in the physical verification data, determines whether the target behavior or current state represented by the predicted label conforms to the scene identifier or region attribute according to a preset association rule, thereby obtaining semantic self-consistency. The comparative agent calibrates the prediction confidence based on at least two of the dimensions of geometric consistency, physical rationality, temporal continuity and semantic self-consistency to obtain a calibration confidence.
8. The method according to claim 1, characterized in that, The step of deploying the evolutionary edge model to the edge device by deploying an intelligent agent includes: The deploying agent loads the evolutionary edge model into the spare model slot and performs inference in parallel with the initial edge model running in the current model slot; Within a preset monitoring period, the deployed intelligent agent compares the operating metrics of the evolved edge model with those of the initial edge model; When the performance metrics of the evolutionary edge model are better than those of the initial edge model, the deploying agent will switch the inference service to the backup model slot; When the performance metrics of the initial edge model are better than those of the evolved edge model, the deploying agent maintains the operation of the existing model slot.
9. The method according to claim 1, characterized in that, The method further includes: The fine-tuning agent updates the initial edge model in the sandbox, and after obtaining the evolved edge model, digitally signs the evolved edge model and binds it to the hardware identifier of the edge device. Before deploying the evolutionary edge model, the deployment agent verifies the digital signature and hardware identifier of the evolutionary edge model through the hardware root of trust of the security chip built into the edge device. If the verification fails, the agent refuses to load the evolutionary edge model.
10. An edge model evolution system for air-gap isolated environments, characterized in that, The system includes: a deployment agent, an edge device, a large visual model, and a fine-tuning agent. The deployment agent and the edge device belong to the end device domain, while the large visual model and the fine-tuning agent belong to the server domain. There is an air gap isolation between the end device domain and the server domain. In the edge device domain, by deploying an agent based on the input data features of the initial edge model in the edge device and the confidence distribution and prediction consistency generated by the initial edge model when performing the target detection task, the decay type of the initial edge model and its corresponding distillation strategy parameters are determined. In the server domain, the initial edge model is updated by fine-tuning the distillation strategy parameters of the agent based on the decay type, as well as the classification vector and intermediate layer feature map output by the visual large model for the pseudo-labels in the training pool, to obtain the evolutionary edge model. In the edge device domain, the evolutionary edge model is deployed to the edge device by deploying an agent.