Collaborative learning method and system based on edge sample intelligent grading and cloud intelligent decision

By adopting a collaborative learning approach of intelligent classification of edge samples and intelligent decision-making in the cloud, the problems of high scenario adaptability and high operation and maintenance costs of cloud-edge collaborative AI models in specific scenarios are solved. This enables the autonomous evolution and efficient updating of AI models, and improves the robustness and automation level of the system.

CN121170549AActive Publication Date: 2025-12-19QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Patent Information

Application Number
CN202511706026.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2025-12-19
Estimated Expiration
2045-11-20

AI Technical Summary

Technical Problem

Existing cloud-edge collaborative AI models suffer from poor adaptability to specific scenarios, performance degradation, and high maintenance costs. In particular, when facing complex real-world scenarios, the models have low recognition accuracy and low update efficiency, leading to system performance bottlenecks and resource waste.

Method used

By employing a collaborative learning approach that combines intelligent classification of edge samples with intelligent decision-making in the cloud, challenging samples are screened and evaluated in real time. The cloud then makes intelligent training decisions based on the global state, enabling online updates and optimization of the model.

Benefits of technology

It enables AI models to evolve autonomously and update efficiently in specific scenarios, reduces data transmission and computing resource costs, improves system robustness and rapid response capabilities, and enhances the level of automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170549A_ABST
    Figure CN121170549A_ABST
Patent Text Reader

Abstract

The invention relates to a collaborative learning method and system based on edge sample intelligent grading and cloud intelligent decision, and belongs to the technical field of edge intelligence and big data processing, and the method comprises the steps: (1) AI model real-time reasoning and difficult sample preliminary screening; (2) intelligently grading edge end samples; (3) uploading data with priority labels; (4) cloud server state aggregation; (5) running a cloud intelligent training decision algorithm; (6) carrying out on-demand training and version management on the model; (7) online updating and closed-loop feedback of the model; according to the method, scene self-adaption and autonomous evolution of the AI model are realized, the efficiency and the economical efficiency of model iteration are remarkably improved, the robustness and the quick response capability of the system are enhanced, and the automation and intelligence levels of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of edge intelligence and big data processing, and in particular to a collaborative learning method and system based on edge sample intelligent grading and cloud intelligent decision-making. BACKGROUND

[0002] With the rapid development of artificial intelligence technology and the wide application of deep learning models, intelligent systems represented by real-time video analysis have played a crucial role in fields such as security monitoring, smart cities, and industrial automation. In order to meet the low-latency processing needs of massive video data, cloud-edge collaboration has become the mainstream computing architecture. Under this architecture, AI inference tasks are usually deployed on edge computing nodes close to data sources to achieve fast response, while model training and optimization tasks are undertaken by cloud servers with more resources.

[0003] However, in actual cloud-edge collaboration application scenarios, the full life cycle management of AI models (such as YOLO target detection models) deployed on the edge faces serious challenges. Typically, these models are pre-trained on general, standardized datasets, and although they have some generalization ability, their performance often declines significantly due to "scenario drift" after being deployed to specific, complex real-world scenarios. This approach is difficult to adapt to the dynamics and complexity of real-world scenarios, leading to the following problems:

[0004] (1) Poor scene adaptability and performance bottlenecks: General models lack targeted learning of subtle features of specific scenarios (such as indoor special lighting, partially occluded targets, and atypical poses), resulting in low recognition accuracy, frequent missed and false detections, and performance bottlenecks in actual applications.

[0005] (2) Static degradation of model capabilities: Once an AI model is deployed, its capabilities are fixed. As real-world scenarios continue to change (such as changes in environmental layout and the emergence of new activity patterns), the model's performance will inevitably degrade, and its matching degree with the real-world scenario will become increasingly low.

[0006] (3) Inefficient and costly operation mode: To address the above problems, the current mainstream solution is a passive, manual operation mode: model recognition errors are discovered by humans, and then data is manually collected and labeled, and model retraining is performed by algorithm engineers in an offline environment. Finally, the model on the edge is updated through downtime or complex release processes. This mode not only has a long response cycle and low efficiency, but also requires a large amount of human and computing resource costs, making it difficult to scale.

[0007] Although some cloud-edge collaborative closed-loop training schemes have appeared in the prior art, such as uploading edge-end data to the cloud for model fusion or ensemble learning. But these schemes still mostly stay at the level of "data carrier", lack of intelligent management of data and training process, and have the following deep-seated defects:

[0008] (1) Blindness of data flow: The edge usually "blindly" uploads all poorly recognized samples or all data to the cloud, which contains a large number of "garbage" samples with limited value for model improvement (such as poor image quality, repeated scene, etc.), which greatly wastes valuable network bandwidth and cloud storage resources.

[0009] (2) Extensive training decision: The retraining of the cloud usually relies on artificial experience or fixed time period to trigger, lacks intelligent, cost-aware judgment of training timing and training intensity, often leads to ineffective training or insufficient training, causing serious waste of computing resources or lag of model update.

[0010] Therefore, an intelligent system capable of realizing AI model scene adaptation and autonomous evolution is urgently needed. The system not only needs to build a data and model collaborative closed loop, but also needs to introduce intelligent decision-making capability in this closed loop to solve the core pain points of high cost, low efficiency and insufficient automation of model iteration in the prior art, and provide a new solution for long-term, stable and efficient operation of the cloud-edge collaborative AI system. SUMMARY

[0011] In view of the deficiencies of the prior art, the present application proposes a collaborative learning method and system based on edge sample intelligent grading and cloud intelligent decision-making, aiming to solve the problems of poor scene adaptability, low evolution efficiency and high cost of AI model in cloud-edge collaborative scenarios.

[0012] In the present method, the present application creatively proposes a double-loop collaborative mechanism including edge sample intelligent grading and cloud intelligent decision-making. By evaluating and grading the intrinsic value of the samples to be learned in real time at the edge, and uploading high-quality "intelligence" to the cloud; the cloud decision-making agent intelligently plans and executes the optimal model training strategy according to the received intelligence information combined with cost-benefit analysis, and finally downgrades the optimized model capability to the edge. The system realizes the autonomous evolution of AI model in a specific scene with low cost and high efficiency.

[0013] The application is suitable for real-time video analysis scenarios that require AI models to continuously learn and optimize. More specifically, the application proposes a collaborative learning system that combines edge sample intelligent grading and cloud intelligent decision-making, aiming to address the core challenge of performance degradation of general AI models in specific deployment scenarios, and to achieve automated and intelligent management of the model's entire life cycle, demonstrating the great potential of the new generation of intelligent systems in scenario adaptation and continuous evolution.

[0014] The technical solutions of the application are: A collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making, comprising: (1) AI model real-time inference and difficult sample preliminary screening; Deploy and run the main AI model on the edge computing node to process and analyze real-time video data streams; the real-time video data stream refers to a series of continuous and independent image frames formed after real-time collection by a video acquisition device and online decoding and frame extraction; the main AI model performs a target detection task for each input image; and samples with poor AI model output results are identified as difficult samples according to a preset preliminary screening rule; (2) Edge sample intelligent grading; Intelligently grade each difficult sample screened out; by extracting and analyzing the multi-dimensional feature vector of the difficult sample, the potential learning value of the difficult sample is quantitatively evaluated and a priority level label is assigned; (3) Data upload with priority label; Upload the difficult sample after intelligent grading together with its corresponding priority level label to the cloud server for storage and aggregation; (4) Cloud server state aggregation; The cloud server continuously receives and aggregates data from one or more edge computing nodes; the background service periodically calculates and updates the global state vector of the system, which includes the cumulative number of samples of each priority level, the time since the last successful training, the current model performance, the sample diversity, and the cloud resource load; (5) Running cloud intelligent training decision algorithm; In the cloud server, a training decision agent selects a currently optimal training action from a predefined action space according to the current latest global state vector of the system and a preset optimal decision strategy; (6) Model on-demand training and version management; According to the optimal training action output by the decision agent, the cloud server automatically executes the corresponding model training task; if the decision is "not training", step (6) is skipped; if the decision is "training", the training module is called, the existing model is fine-tuned or retrained using the collected high-value samples, and a new model with a version number is generated; (7) Online model updating and closed-loop feedback; The new model generated by training is distributed to the edge computing node to realize online hot updating of the main AI model, so that the edge computing node obtains stronger recognition capability without interrupting service.

[0015] Further preferably, in step (1), the main AI model is a YOLOv8 model.

[0016] Further preferably, in step (1), the real-time video data stream is processed and analyzed, including: The input image frame is preprocessed, including size normalization and numerical normalization; The preprocessed image frame is sent to the YOLOv8s model for a complete forward calculation; The YOLOv8s model finally outputs a tensor containing multiple prediction results; By post-processing the tensor, one or more detection results are finally obtained, each including the position information of the target, the target category, and a confidence score representing the prediction confidence.

[0017] According to the present application, preferably in step (2), the edge samples are intelligently graded, including: (2-1) Multi-dimensional feature vector construction: for each difficult sample preliminarily screened, a multi-dimensional feature vector F is extracted, which includes: Model confidence : confidence score of the target detection frame output by the main AI model , ; Image clarity : the blurring degree of the image is quantified by calculating the normalized Laplacian operator variance, as shown below: ; Wherein, is the normalized image, is the variance calculation; Laplacian is the Laplacian operator; target saliency : is the ratio of the target detection frame area to the total image area, which is used to measure the size of the target, ; (2-2) Sample learning value function: through the sample learning value function , a comprehensive value score is obtained by weighted aggregation of the multi-dimensional feature vector F; (2-3) Priority classification strategy: according to the output value of the sample learning value function , the priority of the sample is classified.

[0018] According to the application, the priority classification strategy is realized by setting different value thresholds (V , ) as follows: ; Among them, is the high priority, is the medium priority, is the low priority; P refers to the final output result of the priority level assigned to each evaluated difficult sample.

[0019] Further preferably, the value function is as follows: ; Among them, , , are the weight coefficients of the model confidence , image clarity and target saliency , respectively; is an activation function used to map the image clarity to the interval [0, 1], wherein and are adjustment parameters; According to the application, in step (3), aggregation refers to the process of automatic sorting, routing and collection of difficult samples with different priority labels from one or more edge nodes by a structured data stream processing and storage mechanism.

[0020] According to the application, in step (5), the cloud intelligent training decision algorithm is run, including: The goal of the training decision agent is to learn an optimal strategy to maximize the long-term cumulative reward; the decision logic of the training decision agent is modeled by the Q-Learning algorithm; (5-1) State space and action space: State is defined as a vector wherein, , , are the cumulative number of high, medium, and low priority samples, respectively, is the time since the last training; action space is defined as a discrete set, , , , represent waiting, fast training, and full training, respectively; (5-2) Reward function: a reward function is designed to evaluate the immediate reward obtained by transitioning to a new state after performing an action in a state ; the reward function is shown below: ; wherein, represents the improvement in AI model performance mAP after performing action , and mAP refers to the mean average precision; if is , then ; represents the computational resource cost consumed by performing action ; if is , then ; represents the time consumed by performing action ; (5-3) Action value function and policy: the decision-making agent is trained to estimate the long-term cumulative reward of performing an action in a state by learning an action value function ; the update of the Q function follows the Bellman equation, as shown below: ; wherein, is the learning rate, is the discount factor, refers to the immediate reward, refers to the maximum expected future return in the next state , is a placeholder variable used in the max operation; The optimal policy of the agent is: in any state , choose the action that maximizes Actions that maximize value As follows: .

[0021] According to the application, in step (7), the model is updated online and closed-loop feedback; comprising: (7-1) Model distribution and loading: the cloud server maintains a model repository to store all versions of the trained model; when a new and better model is generated, the cloud server distributes an update instruction to the edge computing node through a dedicated API interface, and the update instruction includes the download address and version number of the new model; (7-2) Edge hot update: after receiving the update instruction, the edge computing node downloads the new model file from the specified download address to the local; through an atomic model instance replacement operation, the main AI model object in the memory providing service is safely replaced by the newly loaded model instance; (7-3) Closed-loop verification: after the model is updated, the edge computing node continues to identify the video stream.

[0022] A computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the above-mentioned collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making when executing the computer program.

[0023] A computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the above-mentioned collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making.

[0024] A collaborative learning system based on edge sample intelligent grading and cloud intelligent decision-making, comprising: A preliminary screening module configured to perform real-time inference of an AI model and preliminary screening of difficult samples; specifically including: deploying and running a main AI model on an edge computing node to process and analyze real-time video data streams; the real-time video data stream refers to a series of continuous and independent image frames formed after real-time acquisition by a video acquisition device and online decoding and frame extraction; the main AI model performs a target detection task for each input image; and according to a preset preliminary screening rule, samples with poor AI model output results are identified as difficult samples; An intelligent grading module configured to perform intelligent grading of edge samples; specifically including: intelligent grading of each difficult sample preliminarily screened; quantitatively evaluating the potential learning value of the difficult sample by extracting and analyzing the multi-dimensional feature vector of the difficult sample, and dividing a priority level label; The data uploading module is configured to upload data with priority labels, specifically including uploading the intelligent-classified difficult samples together with corresponding priority level labels to a cloud server for storage and aggregation. The state aggregation module is configured to aggregate states of the cloud server, specifically including that the cloud server continuously receives and aggregates data from one or more edge computing nodes, and a background service periodically calculates and updates a global state vector of the system, the global state vector including cumulative number of each priority sample, time since last successful training, current model performance, sample diversity and cloud resource load. The cloud intelligent training decision algorithm running module is configured to run a cloud intelligent training decision algorithm, specifically including that a training decision agent in the cloud server selects a currently optimal training action from a predefined action space according to a currently latest global state vector of the system and according to a preset optimal decision strategy. The model on-demand training and version management module is configured to automatically execute a corresponding model training task according to the optimal training action output by the decision agent, and if the decision is "no training", step (6) is skipped, and if the decision is "training", a training module is called to fine-tune or retrain an existing model using collected high-value samples, and a new model with a version number is generated. The online updating and closed-loop feedback module is configured to update models online and provide closed-loop feedback, specifically including that the new model generated by training is distributed to edge computing nodes to realize online hot updating of the main AI model, so that the edge computing nodes obtain stronger recognition capability without interrupting service.

[0025] Compared with the prior art, the above technical solutions of the present application can achieve the following beneficial effects: 1. Scene adaptation and autonomous evolution of AI models are realized: the present application builds a collaborative learning closed loop of edge analysis, cloud decision and model updating, and gives the AI system the ability of scene adaptation and autonomous evolution. It can continuously learn and adapt to the unique characteristics of a specific deployment environment (such as a specific classroom), dynamically and automatically convert a general AI model into a highly customized special model, and continuously improve and maintain its high recognition performance in a specific scene without human intervention, thereby fundamentally solving the performance degradation problem of traditional static models caused by "scene drift".

[0026] 2. Significantly improve the efficiency and economy of model iteration: the present application introduces intelligent sample grading at the edge and intelligent training decision strategy at the cloud, significantly improving the efficiency and economy of AI model iteration. It filters out a large number of low-value samples by evaluating the value at the data source, greatly reducing the data transmission and storage cost between the cloud and the edge; and through cost-benefit trade-off in the cloud, unnecessary and inefficient computing resource consumption is avoided, so that the model can achieve the most efficient performance iteration at the lowest cost, solving the technical pain points of high cost and serious resource waste in the traditional model updating process.

[0027] 3. Enhance the robustness and rapid response capability of the system: the system architecture of the present application, especially the linkage mechanism of edge intelligent grading and cloud strategy decision, significantly enhances the robustness and rapid response capability of the AI system to environmental changes. The system can sensitively perceive the sudden or unknown situation in the scene (reflected as high priority P1 sample), and trigger the cloud decision system for rapid and high-intensity retraining, so that the system can adapt to the scene mutation in a short time, quickly recover and improve its recognition ability, and ensure the long-term availability and reliability of the system in the real and variable environment.

[0028] 4. Improve the automation and intelligence level of the system: the present application upgrades the traditional model updating process which relies on manual discovery and manual triggering to a double-loop linkage intelligent system containing data and decision loops. The analysis result of the edge directly drives the strategic decision of the cloud, realizing the full-process automation from "finding problems" to "solving problems", greatly reducing the dependence on manual operation, and providing a feasible technical path for large-scale, low-cost deployment of high-availability AI system. BRIEF DESCRIPTION OF DRAWINGS

[0029] Figure 1 the overall flowchart of a collaborative learning method based on edge intelligent grading and cloud strategy optimization proposed by the present application; Figure 2 the architecture and flowchart of the edge AI sample intelligent grading module in the present application; Figure 3 the architecture and flowchart of the cloud intelligent training decision module in the present application; Figure 4 the relationship diagram of sample learning value function and priority grading in the present application; Figure 5 the strategy representation diagram of the cloud decision agent in the present application; Figure 6 the performance index (mAP) curve diagram of the new model on the validation set in the embodiment of the present application; Figure 7Figure of loss function (loss) change curve in training process in the embodiment of the application. DETAILED DESCRIPTION

[0030] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0031] In order to facilitate those skilled in the art to understand the present application, the key terms appearing in the specification and claims will now be explained: 1. Edge intelligence grading: refers to the process of using one or more models / algorithms to perform real-time, multi-dimensional value assessment on AI inference results or raw data generated locally on edge computing nodes close to data sources without relying on the cloud, and dividing the priority level of the data.

[0032] 2. Sample learning value: refers to the potential value that a specific data sample (such as an image) can contribute to improving the performance of an AI model. The present application believes that the value depends not only on the recognition confidence of the model for the sample, but also on factors such as the data quality (such as clarity) of the sample itself.

[0033] 3. Cloud intelligent training decision: refers to the intelligent strategy of dynamically and non-fixedly deciding whether to trigger and how to trigger (such as training intensity, round) the model training task on the cloud server according to the real-time state of the data queue with different characteristics (such as priority) gathered from the edge, as well as the historical state of the system (such as the time since the last training).

[0034] 4. Decision agent: refers to the logical entity in the cloud module of the present application for executing intelligent training decisions. In a preferred embodiment, the agent is modeled by the idea of reinforcement learning (such as Q-Learning), which aims to seek the optimal balance between model performance benefits and computing resource costs.

[0035] 5. Model online hot update: refers to the process of loading new model weights trained by the cloud into memory and safely replacing the running old model instance without interrupting the edge AI inference service, aiming to achieve a non-susceptible and smooth upgrade of the model.

[0036] 6. Laplacian Variance: A computational method used in digital image processing to measure the focus and sharpness of an image. The principle is to apply the Laplacian operator (a second-order differential operator) to the image to highlight the edges and details in the image, and then calculate the gray variance of the resulting image. Generally, the larger the variance value, the richer the edges and details of the image, and the clearer the image; on the contrary, the smaller the variance value, the more blurred the image. In this invention, this index is used as a key parameter to quantify the multi-dimensional feature of "image sharpness".

[0037] 7. YOLOv8 Model: An advanced, real-time, open-source object detection AI model. YOLOv8 model is known for its excellent balance between speed and accuracy, it can directly predict multiple bounding boxes and target class probabilities from a complete image in a single forward pass. In this invention, yolov8s.pt is used as its official pre-trained small size version, deployed on edge computing nodes as the main AI model for real-time pedestrian detection tasks.

[0038] Example 1 A collaborative learning method based on edge intelligent grading and cloud strategy optimization, as shown in Figure 1 , comprising: (1) AI model real-time inference and difficult sample preliminary screening; Deploy and run the main AI model on the edge computing node to process and analyze real-time video data streams; Real-time video data streams refer to a series of continuous and independent image frames formed by real-time acquisition by front-end monitoring cameras and other video acquisition devices, and online decoding and frame extraction by video processing tools such as FFmpeg. These image frames are continuously fed into the data pipeline in sequence for subsequent processing. For each input image, the main AI model performs an object detection task; according to the pre-set preliminary screening rules (for example, below the confidence threshold), samples with poor AI model output results are identified as difficult samples; as a candidate for subsequent intelligent grading. Poor samples refer to cases where the main AI model detects the target (for example, identifies a "person"), but the output confidence score is lower than a pre-set screening threshold (for example, the threshold can be set to 0.4). The identification process is: the system compares the confidence score of each detection result output by the model with the pre-set threshold to identify poor samples. Any image frame with a confidence score less than the threshold is considered a sample that the model is "not confident" about its own prediction, and is automatically identified as a "difficult sample" by the system.

[0039] (2) Edge sample intelligent grading; Referring to Figure 2For each hard case sample, a sample value assessment model deployed on the edge is used to intelligently rank the sample; the sample value assessment model extracts and analyzes the multi-dimensional feature vector (at least including model confidence, image clarity, target saliency, etc.) of the hard case sample, quantitatively evaluates the potential learning value of the hard case sample, and divides a priority level label; for example, high, medium, and low priority.

[0040] The module first extracts a multi-dimensional feature vector , which at least includes model confidence , image clarity , and target saliency . Then, as shown in Figure 4 , a sample learning value function is used to calculate the feature vector to obtain a quantitative value score. Finally, the score is compared with the preset threshold and to divide a priority level P (P1, P2, or P3) for the sample.

[0041] (3) Data upload with priority labels; The hard case samples after intelligent ranking are uploaded to the cloud server together with their corresponding priority level labels for storage and aggregation; in a preferred embodiment, to facilitate subsequent training and manual verification, the system uploads the original picture (for AI training) and the labeled picture with the detection box (for manual analysis) together, and distinguishes them by adding suffixes such as --hardcase-P1-original.jpg or --hardcase-P2-annotated.jpg to the file names.

[0042] (4) Cloud server state aggregation; Referring to Figure 3 , the cloud server (such as a data flow engine such as NiFi) continuously receives and aggregates data from one or more edge computing nodes; the background service periodically calculates and updates the global state vector of the system, which refers to the "intelligent training decision module" or "decision agent" deployed in the cloud. It is a background service or a script that is executed at a fixed time, and its only responsibility is to serve as the "brain" of the entire collaborative learning system, continuously observing state indicators reflecting the global situation, and making optimal training decisions based on these indicators. The global state vector includes: the cumulative number of samples of each priority, the time since the last successful training, the current model performance, the sample diversity, and the cloud resource load.

[0043] Current model performance: the performance score (mAP) of the current model deployed on the edge, on a fixed, small validation set in the cloud. This metric can reflect whether the model performance has started to "degrade".

[0044] Sample diversity: measures the difference between newly collected difficult samples and existing samples. For example, the average distance between the features of new sample images and historical sample features can be calculated to quantify it. This metric can help the system determine whether new samples bring "new knowledge" or not.

[0045] Cloud resource load: refers to the load of the GPU, CPU and other computing resources of the current cloud training server. This metric can prevent the start of new, non-urgent training tasks when the server itself is already highly loaded, thus avoiding resource conflicts.

[0046] The above features are preferred technical features designed to make the system more intelligent.

[0047] The update process of the state vector is automatically and periodically (e.g., every 10 minutes) performed by the cloud backend service. Although not detailed above, the specific update mechanism is as follows:

[0048] Update of the cumulative number of samples of each priority ( , , ) : These sample images have been automatically sorted by NiFi and stored in the specified directories of the cloud server. The background service will periodically scan these directories and update the state values of , , etc. by counting the number of files in each directory. When a training task is triggered and completed, the sample images used for this training can be moved to an archive directory, so that the count is reset or reduced, and a new accumulation begins.

[0049] Update of the time since the last successful training ( ) : The background service records a timestamp in a persistent storage (such as a simple state file or database entry) to mark the time point of the last successful training task. In each decision cycle, the background service subtracts the recorded timestamp from the current time to dynamically calculate the value of . When a new training task is successfully triggered, the service immediately replaces the old timestamp with the current time, completing the reset.

[0050] (5) Run the cloud intelligent training decision algorithm; At the cloud server, a training decision agent selects a current optimal training action from a predefined action space (e.g., no training, fast training, full training) according to the current latest global state vector of the system and a preset optimal decision strategy, which aims to achieve a long-term balance between model performance improvement (benefit) and computing resource consumption (cost).

[0051] (6) Model on-demand training and version management; According to the optimal training action output by the decision agent, the cloud server automatically executes the corresponding model training task; if the decision is "no training", step (6) is skipped; if the decision is "training", the training module is called to fine-tune or retrain the existing model using the collected high-value samples, and after training, a better model file (best.pt) with a new version number is generated. As shown in Figure 6 and Figure 7 As shown, during the training process, the performance indicator (mAP) of the model is significantly improved, and the loss function (loss) is steadily decreased, proving the effectiveness of the training.

[0052] In step (6), when the intelligent decision agent of the cloud outputs the "training" action, the system automatically executes a clear training process that includes the following core steps: 1) Load data and base model: After the training task is started, the system first locates and loads the high-value samples (mainly P1 and P2 levels) that have been collected and labeled according to the preset configuration file. At the same time, the system loads the official pre-trained yolov8s.pt model currently deployed on the edge into the GPU memory as the starting point for this fine-tuning.

[0053] 2) Perform model fine-tuning: The system uses the loaded high-value sample data to iteratively train the base model. In this process, the model learns new features in these "difficult samples" and adapts to specific scenarios (such as classroom monitoring). According to the instructions of the decision agent (see Figure 3 , fast training or full training), the system automatically executes a preset number of training rounds.

[0054] 3) Performance evaluation and optimal model saving: After each round of training ends, the system automatically uses the validation set to evaluate the performance of the current model (calculates mAP and other indicators). The system continuously tracks the model state with the best performance on the validation set and automatically saves it as a best.pt (self-defined name) file.

[0055] 4) Versioned archiving: After the completion of the entire training, the final best.pt new model, along with the logs, performance charts and configuration files of this training, will be saved completely in a dedicated directory with a unique experiment name. In this way, the system not only generates a new model with better performance, but also automatically creates a new version for it that can be traced and managed.

[0056] (7) Online model updating and closed-loop feedback; The newly trained model is stored in the cloud model repository. The cloud issues an update instruction to the edge node through the API. After receiving the instruction, the edge node downloads the new model and safely replaces the running AI model instance in memory with the new model instance without interrupting existing services, completing the online upgrade of capabilities.

[0057] After the model is updated, the edge node immediately uses the more powerful new model for real-time inference. For the same scene that was previously poorly recognized, the optimized model can significantly improve the accuracy and recall rate of recognition, successfully identifying previously missed targets. This improvement in performance directly verifies the success of the entire collaborative learning closed loop at the business level, and the system thus completes a complete, low-cost, and automated self-evolution.

[0058] Embodiment 2 A collaborative learning method based on edge intelligent grading and cloud strategy optimization according to Embodiment 1, the difference lies in: In step (1), the main AI model is a YOLOv8 model.

[0059] In step (1), real-time video data streams are processed and analyzed; including: The input image frame is preprocessed, including: size normalization (e.g., scaling to 640x640 pixels), and numerical normalization (converting pixel values from the 0-255 range to the 0-1 range); to meet the model input requirements.

[0060] The preprocessed image frame is sent to the YOLOv8s model for a complete forward calculation (recognition); The YOLOv8s model finally outputs a tensor containing multiple prediction results; Through decoding and non-maximum suppression post-processing operations on the tensor, one or more detection results are finally obtained, each including the location information (bounding box coordinates) of the target, the target class (e.g., "person"), and a confidence score representing the prediction confidence.

[0061] In step (2), edge sample intelligent grading; refer to Figure 2 andFigure 4 , comprising: (2-1) Multi-dimensional feature vector construction: for each difficult sample preliminarily screened out, a multi-dimensional feature vector F is extracted, which aims to quantify the learning value of the sample from multiple dimensions. The multi-dimensional feature vector F includes:

[0062] Model confidence : the confidence score of the target detection box output by the main AI model , ; After the main AI model (such as YOLOv8s) completes processing and analysis on each frame of image, its output result is a set containing one or more detection instances. For each detection instance, its output information must include three core elements: target position, target category, and confidence score. 1. Target position: usually represented by the coordinates of the bounding box; 2. Target category: refers to what object the model recognizes (e.g., "person"); Confidence score: a floating-point value between 0 and 1, used to quantitatively represent the model's confidence in its own judgment (i.e., "this box is this category"). The higher the score, the more confident the model is.

[0063] Image sharpness : the blurring degree of the image is quantified by calculating the normalized Laplacian operator variance of the image, as shown below: ; where, is the normalized image, which is a common and necessary preprocessing step before the AI model processes the image. The normalized image ( ) refers to a series of standardized preprocessing that the original image undergoes before being input into the main AI model (such as YOLOv8) or used for feature calculation (such as sharpness calculation).

[0064] This process at least includes: 1. Gray scale conversion: in order to calculate the sharpness, the three-channel color image (RGB) needs to be converted into a single-channel gray scale image first. This can eliminate the interference of color information, so that the subsequent operator can focus more on the brightness, edge and texture changes of the image; 2. Numerical normalization: convert the pixel values of the image from the original integer range (usually 0-255) to a standard, smaller floating-point number range (e.g. 0-1). Therefore, is a standardized, mathematically suitable image data representation obtained after the original image has undergone gray scale conversion and numerical normalization. for variance calculation; The higher the value, the clearer the image. The Laplacian operator is a classic and important differential operator in digital image processing and computer vision; specifically, it is a second-order differential operator. In this embodiment of the invention, the Laplacian operator is applied to the normalized image. The process generates a new image where only edges and details are highlighted. The variance of this new image can then be calculated. There are two cases for variance: 1. If the original image is sharp, it has many edges and details, resulting in a new image with many highlights and strong contrast, leading to a large variance value; 2. If the original image is blurry, it has few edges and details, resulting in a very dark and smooth new image, leading to a small variance value. Therefore, by calculating the variance of the Laplacian operator, a quantifiable value can be obtained to accurately determine the sharpness of the original image, thus providing a reliable input feature for the invented "sample learning value evaluation model".

[0065] Target salience : It refers to the ratio of the area of ​​the target detection box to the total area of ​​the image, and is used to measure the size of the target. ; An image refers to an image frame that has undergone preprocessing (especially size normalization) and is about to be fed into the sample value assessment model for analysis. In this embodiment of the invention, to ensure consistency in calculation, all original image frames of varying sizes extracted from the video stream are first uniformly scaled and placed on a canvas of fixed resolution. Therefore, the "total image area" refers to the total pixel area of ​​this fixed-resolution canvas (e.g., 854 * 480). The "target detection box area" refers to the pixel area covered by the detection box on this canvas. By calculating the ratio of these two, a standardized "target saliency" feature value, unaffected by the original video resolution, can be obtained for subsequent value assessment.

[0066] (2-2) Learning the value function from samples: Learning the value function from samples We perform weighted aggregation on the multidimensional feature vector F to obtain a comprehensive value score. (2-3) Priority-based ranking strategy: learning the value function based on samples The output value is used to prioritize the samples.

[0067] Priority tiering strategies use different value thresholds ( , This can be achieved as follows: ; in, As a high priority, Medium priority Low priority; Value threshold and These are the key decision boundaries used in this invention to divide different priority intervals. These two values ​​are configurable hyperparameters; their specific values ​​are not fixed but can be empirically adjusted and optimized based on the needs of the actual deployment scenario, data distribution, and the setting of the weight coefficients w in the adopted value function V(F). For example:

[0068] Suppose that in V(F), weights are set Therefore, we can assume three scenarios: An ideal high-value sample ( Its characteristics may include: confidence level Extremely low (e.g., 0.1), image sharpness Very high ( Approximately 1 (counted as 1)), target size Moderate (e.g., 0.2). The V(F) value for this sample can be calculated and is approximately 1.38.

[0069] A typical medium-value sample ( Its characteristics may include: confidence level Lower (e.g., 0.3), image sharpness Very high ( Approximately 1), target size Moderate (e.g., 0.2). The V(F) for this sample can be calculated and is approximately equal to 1.18.

[0070] A low-value sample ( Its characteristics may include: confidence level Lower (e.g., 0.3), but the image is very blurry. (Close to 0). The V(F) value for this sample can be calculated and is approximately 0.68.

[0071] Based on the calculation example above, we can then... and Set a reasonable set of example values: It can be set to 1.25 to filter out the samples with the highest learning value. It can be set to 0.8 to distinguish between medium-value and low-value samples.

[0072] therefore, and The range of values ​​is related to the output range of the value function V(F), and together they constitute the core mechanism for quantifying and classifying the learning value of samples in this invention.

[0073] P refers to the final output result of the priority level assigned to each evaluated difficult sample. It is a classification label that identifies the value and urgency of the sample for cloud model training. In embodiments of the present invention, the value set of P is {P1, P2, P3}, representing high, medium, and low priority respectively.

[0074] This grading strategy converts complex rules into a continuous value-based, quantifiable, and more generalizable grading model.

[0075] The value function is as follows: ; where, , , are the weight coefficients of model confidence , image clarity , and target saliency respectively; Weight coefficients are key parameters in the present invention for adjusting and balancing the contribution of different features to the final "sample learning value". They are configurable hyperparameters whose value range can theoretically be any positive real number. In practical applications, their absolute values are not important, what matters more is their relative proportion. For example:

[0076] In the experimental scenario of the present invention, = 1.0 (model confidence weight), which means that confidence is considered the most important basic measurement indicator. A sample with very low confidence (c close to 0), its contribution to this item is close to 1.0, indicating that it has a high basic value.

[0077] = 0.5 (image clarity weight), clarity is an important bonus item, but its importance is slightly lower than confidence. Therefore, it can be given a medium-sized positive weight. A clear picture (c close to 1) can get a value bonus of about 0.5, while a fuzzy picture (c close to 0) does not have this bonus. This ensures that the system will prefer clear pictures for training.

[0078] ​​= 0.1 (Target Salience Weight), target salience (i.e., area) is used here as a slight penalty. It can be given a small negative weight. Because a very large target occupying most of the image area is usually easy to identify, its learning value is relatively limited. By introducing this small negative weight, the system slightly reduces the priority of such overly "simple" samples, thus focusing more on learning medium-sized or more difficult-to-identify targets. This is used to balance the contribution of different features to the final value score; This indicates that the lower the confidence level, the higher the value. It is an activation function (such as the Sigmoid function) used to adjust image sharpness. Mapped to the interval [0, 1] ,in, and To adjust the parameters; to adjust the parameters and These are the key hyperparameters used in this invention for finely controlling the shape of the sharpness scoring function σ(b). Together, they define how the system transforms a continuous, physically meaningful sharpness value (Laplace operator variance b) into a standardized value score representing "sharpness confidence" ( (The range is between 0 and 1).

[0079] for It defines the center point or activation threshold of the Sigmoid function. Technically, it represents the set critical point of "acceptable minimum sharpness." When the sharpness value b of an image is exactly equal to... hour, The output value is exactly 0.5; when Much larger hour, The output approaches 1 (indicating very clear); when much smaller hour, The output approaches 0 (indicating very blurry). The value range is positive real numbers, and its specific value should be determined by statistical analysis of the sharpness values ​​of a batch of samples, based on factors such as lighting conditions and camera quality in the actual scene. For example, suppose that in a scene, when the variance of the Laplacian operator is below 100, the image is usually too blurry for the human eye or the model, while it becomes clear and usable when it is above 100. In this case, a sample value can be set. = 100, defining the boundary between "clear" and "fuzzy" here.

[0080] for , The Sigmoid function curve was controlled at the center point. The steepness of the vicinity, i.e. the width of the transition zone. A larger value will make the function curve very steep, forming a "step-like" effect. This means that the system is very "harsh" and "black-and-white" in judging the sharpness, only pictures with sharpness significantly above will get a high score. A smaller value will make the function curve very flat, with a wide transition zone. This means that the system is more "lenient" in judging the sharpness, i.e. even if the sharpness value fluctuates around , the final score will change smoothly. The value range of k is positive real number. As an example, to achieve a smooth but effective transition, we can set an example value = 0.1. This value can ensure that when the picture sharpness fluctuates around the center point = 100, its value score will also correspondingly, non-suddenly, transition from around 0.5 to 0 or 1.

[0081] In summary, by jointly adjusting the two parameters and , the present application can flexibly and accurately define how to convert the physical image sharpness into a standardized learning value score that is meaningful for AI model training in a specific scene.

[0082] In step (3), aggregation is not a complex calculation process, but refers to: the cloud server, through a structured data stream processing and storage mechanism, automatically sorts, routes and collects the difficult samples with different priority labels from one or more edge nodes.

[0083] The aggregation process is: 1) Unified reception: deploy an HTTP listening endpoint (provided by the ListenHTTP processor of NiFi in the experiments of the present application) in the cloud as the only entrance for all edge nodes to report data. The edge nodes send difficult sample pictures with priority labels P1, P2, P3, etc. to this endpoint through POST requests.

[0084] 2) Attribute extraction and routing: in the experiments of the present application, when a sample picture enters NiFi, the system will first extract key attributes from its metadata, especially the file name. Then, a RouteOnAttribute processor (which will route based on attributes) will check the file name of each file according to the pre-set rules.

[0085] 3) Sorting and Aggregation: RouteOnAttribute processors will distribute data flows to different processing paths according to the inspection results, achieving automated sorting and aggregation. In the experiments of the present application:

[0086] If the file name contains --hardcase-P1-original, the file is routed to the "high-priority original sample" path.

[0087] If the file name contains --hardcase-P2-annotated, the file is routed to the "medium-priority annotated sample" path. And so on.

[0088] 4) Structured Storage: At the end of each processing path, there is a PutFile processor. This processor is responsible for storing the sorted sample pictures into a pre-planned, structured directory on the server. In the experiments of the present application:

[0089] High-priority original samples are uniformly stored (aggregated) in the / mnt / hardcase / classroom_dataset / images / p1 / folder.

[0090] Medium-priority original samples are uniformly stored (aggregated) in the / mnt / hardcase / classroom_dataset / images / p2 / folder. And so on.

[0091] Through the above pipeline process of receiving -> routing -> sorting -> storage, the present application can automatically and structurally aggregate scattered and difficult samples from different edges, different times, and with different priority labels to the designated storage location of the cloud server, providing a neat and orderly data foundation for subsequent online annotation, model training, and intelligent decision-making. For the convenience of subsequent analysis and manual verification, the original samples and annotated samples with model inference results can be selectively uploaded together.

[0092] In step (5), the cloud intelligent training decision algorithm is run; including: The core of this algorithm is a training decision agent based on reinforcement learning, aiming to solve the Markov Decision Process (MDP) problem. The goal of the training decision agent is to learn an optimal policy to maximize long-term cumulative rewards; in a preferred embodiment, the decision logic of the training decision agent is modeled through the Q-Learning algorithm;

[0093] (5-1) State space and action space: State is defined as a vector where, , , cumulative number of high, medium, and low priority samples, time since last training; action space is defined as a discrete set, , , , represent waiting, fast training, and full training, respectively; (5-2) Reward function: To guide the agent learning, a reward function is designed to evaluate the immediate reward obtained after performing action from state to new state ; the reward function is as follows: ; where, represents the performance mAP improvement of the AI model (specifically, the main AI model deployed on the edge node and continuously learned and optimized in the cloud. In the preferred embodiment of the present application, it is the YOLOv8s model.) after performing action ; mAP refers to the average precision mean; mAP (mean Average Precision) is the most authoritative and commonly used core quantitative indicator in the field of target detection, used to evaluate the performance of the model. It is not just a simple "accuracy", but a combination of the precision (i.e. how many of all the frames predicted by the model are really drawn correctly) and recall (i.e. how many of all the targets that should be found are successfully found by the model) dimensions. The mAP score is a value between 0 and 1, the higher the score, the better the overall performance of the model, i.e. both accurate and complete. In the present application, is the mAP score measured on a fixed cloud validation set after training the new model, minus the mAP score of the old model on the validation set before training, thus obtaining a quantitative "performance improvement value". If is , then ; represents the computational resource cost (such as GPU hours) consumed by performing action ; if is , then ; represents the time spent performing action ; the weight coefficient , , are the core of the reward function of the present application, which are configurable hyper-parameters to define and balance the relationship between "model performance gain" and "resource consumption cost", thus guiding the decision agent to learn an optimal policy that meets the expectation. Their value range is positive real number, and their relative proportion is more important than absolute value.

[0094] For example: The goal of the present application is to enable the "gain" of model performance to be numerically compared and weighed with the "cost" of resources.

[0095] = 100 (performance gain weight): The amount of mAP improvement is usually a small number (e.g., 0.1). By setting a large weight (e.g., 100), it is amplified into a significant reward score. For example, mAP improvement of 0.1 can bring a positive reward of 100 * 0.1 = +10.

[0096] = 5 (computation cost weight): Suppose that is defined as the GPU hours used for training, and the cost per GPU hour is estimated to be 5 units. This weight makes the execution of a 1-hour training bring a negative reward of -5 * 1 = -5.

[0097] = 0.1 (time cost weight): Time itself is also a cost. By giving it a small negative weight, it is used to encourage the agent to choose actions that take less time when other conditions are similar. For example, a 1-hour training will additionally bring a slight penalty of -0.1 * 1 = -0.1.

[0098] With the above value settings, the reward function can perform a cost-benefit analysis: if a training brings more than 0.05 mAP improvement (gain > +5), it can offset the 1-hour GPU cost (cost = -5), thus obtaining a positive total reward, encouraging the agent to make more similar "high-return investments" in the future. Conversely, if a training only brings a negligible performance improvement, the total reward will be negative, thus "punishing" this decision, so that the agent learns to avoid such "inefficient investments".

[0099] (5-3) Action value function and policy: training the decision agent by learning an action value function to estimate the long-term cumulative reward of executing action in state ; the update of the Q function follows the Bellman equation as follows: ; where, is the learning rate, is the discount factor, refers to the immediate reward, refers to the next state with the maximum expected future return, is a placeholder variable used in the max operation; First, this formula is the core of the Q-Learning algorithm, which is an iterative form of the Bellman equation. It describes how the decision agent corrects its own “past” knowledge (old values) according to the “reality” (immediate reward R + best estimate of the future) and thus constantly learns and improves.

[0100] For R: “R” here stands for the immediate reward. It is a specific, quantifiable numerical value that evaluates the “short-term gain or cost” that the decision agent obtains immediately after performing an action a in state s. For example, if the agent decides to perform “full training” (action a), and the model performance (AmAP) improves significantly after training, while the computation cost (C(a)) is not high, then the reward function will calculate a larger positive number as the R value (such as +15). Conversely, if the performance improvement is very small after training, the R value may be negative (such as -3).

[0101] For , it is the maximum expected future return in the next state . The meaning of is: according to the agent’s current experience, if the future arrives at a new state , and a certain action is chosen in the new state, how much long-term gain can be obtained from time to the end of the task; The meaning of is: in the new state , the agent will review all possible next actions (such as waiting, fast training, full training) and find the one that maximizes its long-term gain . In simple terms, it is the agent’s estimate of the long-term value of choosing the optimal action after entering the next state based on its existing knowledge. It means that the agent will not only look at the immediate reward R.

[0102] For , it is a placeholder variable used in the max operation. It represents all possible, alternative actions that the agent can choose when the system enters the next state s’. Specifically, the value range of The actions are exactly the same, i.e. wait, fast training, full training. Its role is, in the calculation , the program will internally make a simulation deduction: it will assume that, in the state, if the choice is to wait, what will happen? If the choice is fast training, what will happen? If the choice is full training, what will happen? Then, it will find the action with the largest Q value from these deduction results, and take this largest Q value as the result of the entire max expression.

[0103] Through the coordinated work of the three parts, the Q-learning algorithm ensures that each decision of the agent is an optimal judgment made after comprehensively considering the "actual return (R) at the moment" and "long-term planning for the future (max Q)".

[0104] The optimal strategy of the agent is: in any state , choose the action that maximizes the value , as follows: ; In one specific embodiment of the present application, a solidified strategy table is predefined, which is a concrete embodiment of the optimal strategy obtained by the Q-Learning algorithm after convergence under a specific reward function. The solidified strategy table refers to the optimal strategy obtained by the Q-Learning algorithm after sufficient training and learning , which is implemented in the form of a clear and directly queryable decision rule set. This strategy table defines the unique optimal action a that the decision agent should perform in any given system state S.

[0105] For ease of understanding, the query table is shown in Table 1: Table 1

[0106] Among them, T_p1 is the high-priority sample quantity threshold for triggering "full training", i.e. when the system accumulates at least 10 P1 level samples, it is considered that there is enough high-value new knowledge worthy of a deep learning; T_p2 is the medium-priority sample quantity threshold for triggering "fast training", i.e. when P1 samples are insufficient, but 50 regular P2 samples are accumulated, it is worth a consolidating fine-tuning training. T_time is the longest time interval threshold for triggering "maintenance training", which can be set to 24 (hours) to ensure that the model capability does not become "outdated" due to long-time non-updating.

[0107] The execution logic of the query table is shown in Figure 5 ; In step (7), the model is updated online and closed-loop feedback; including: (7-1) Model distribution and loading: the cloud server maintains a model repository to store all versions of the trained model; when a new and better model is generated, the cloud server distributes an update instruction to the edge computing node through a dedicated API interface, and the update instruction includes the download address and version number of the new model; (7-2) Edge hot update: after the edge computing node receives the update instruction, it downloads the new model file from the specified download address to the local; through an atomic model instance replacement operation, the main AI model object in the memory providing service is safely replaced by the newly loaded model instance; the whole process does not need to restart the service, which guarantees the continuity of the business. (The "atomic model instance replacement operation" is realized through a two-step mechanism of "background loading and instantaneous switching". 1) Background loading: when receiving the update instruction, the system does not immediately touch the old model that is providing service. It will completely load and initialize the new model file (best.pt) in a temporary and independent memory space in the background. This process is time-consuming, but completely isolated from the main service and does not affect any ongoing recognition task.

[0108] 2) Instantaneous switching: only when the new model is completely ready, the system will perform the most core replacement operation: the global variable (such as MODEL) pointing to the AI model is instantly switched from the memory address pointing to the old model to the memory address pointing to the new model. In Python, this "reference switching" action is atomic, that is, it cannot be interrupted. This means that all subsequent recognition requests will seamlessly start using the new model, while all requests before this switching instant will normally complete their tasks using the old model.

[0109] (7-3) Closed-loop verification: after the model is updated, the edge computing node continues to recognize the video stream. The new recognition result (for example, the previously missed target is now detected with high confidence) is returned or displayed through the regular business process, thereby verifying the effectiveness of the entire collaborative learning closed loop at the business level.

[0110] In the present application, Figure 5 The decision tree shown is a concrete and executable embodiment of the optimal strategy obtained after the Q-Learning algorithm converges under a specific reward function.

[0111] The present embodiment further illustrates the beneficial effects that can be achieved by the above technical solutions through specific experimental data.

[0112] Referring to Figure 6 and Figure 7The two figures show the detailed quantitative indicators recorded in a complete collaborative learning closed loop of the embodiment, which are automatically generated by the YOLOv8 training framework in the validation stage, for evaluating the performance of the optimized model.

[0113] Figure 6 The average precision mean (mAP) change curve of the model on the validation set is shown. It can be seen that with the increase of training rounds (epochs), the key performance indicators mAP50 and mAP50-95 are steadily improved from the initial lower level. This shows that the model has effectively and quantitatively improved its generalization ability and overall performance on unseen validation data by learning the high-value difficult samples selected by the application.

[0114] Figure 7 The loss function (Loss) change curve in the training process is shown. The figure shows key indicators such as positioning loss (box_loss) and classification loss (cls_loss). All loss curves show a steady downward trend and eventually converge, indicating that the model's learning process is stable and effective, and it has successfully learned from new training data how to more accurately locate and classify targets, correcting its original "knowledge blind spot".

[0115] Embodiment 3 A computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the collaborative learning method based on edge intelligent grading and cloud strategy optimization of embodiments 1 or 2 when executing the computer program.

[0116] Embodiment 4 A computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the collaborative learning method based on edge intelligent grading and cloud strategy optimization of embodiments 1 or 2.

[0117] Embodiment 5 A collaborative learning system based on edge intelligent grading and cloud strategy optimization, comprising: The preliminary screening module is configured to perform real-time inference of the AI model and preliminary screening of difficult samples. Specifically, it includes deploying and running the main AI model on the edge computing node to process and analyze real-time video data streams. Real-time video data streams are a series of continuous and independent image frames formed after real-time acquisition by a video acquisition device and online decoding and frame extraction. The main AI model performs a target detection task for each input image. According to the preset preliminary screening rules, samples with poor AI model output results are identified as difficult samples. The intelligent grading module is configured to perform edge sample intelligent grading, specifically including: performing intelligent grading on each difficult sample screened out initially; quantitatively evaluating the potential learning value of the difficult sample by extracting and analyzing the multi-dimensional feature vector of the difficult sample, and dividing a priority level label; The data uploading module is configured to upload data with priority labels, specifically including: uploading the difficult sample after intelligent grading together with the corresponding priority level label to the cloud server for storage and aggregation; The state aggregation module is configured to perform cloud server state aggregation, specifically including: the cloud server continuously receives and aggregates data from one or more edge computing nodes; the background service regularly calculates and updates the global state vector of the system, and the global state vector includes: the cumulative number of each priority sample, the time from the last successful training, the current model performance, the sample diversity and the cloud resource load; The cloud intelligent training decision algorithm running module is configured to run a cloud intelligent training decision algorithm, specifically including: in the cloud server, a training decision agent selects a currently optimal training action from a predefined action space according to the current latest global state vector of the system and according to a preset optimal decision strategy; The model on-demand training and version management module is configured to automatically execute the corresponding model training task according to the optimal training action output by the decision agent; if the decision is "no training", step (6) is skipped; if the decision is "training", the training module is called to fine-tune or retrain the existing model using the collected high-value samples, and a new model with a version number is generated; The online updating and closed-loop feedback module is configured to perform model online updating and closed-loop feedback, specifically including: the new model generated by training is distributed to the edge computing node to realize online hot updating of the main AI model, so that the edge computing node obtains stronger recognition capability without interrupting service.

[0118] Those skilled in the art will readily understand that the above description is only the preferred embodiment of the present application, and is not intended to limit the present application, and any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making, characterized in that, Comprise: (1) AI model real-time inference and difficult sample preliminary screening; Deploy and run the main AI model on the edge computing node to process and analyze real-time video data streams; Real-time video data streams refer to a series of continuous and independent image frames formed after real-time collection by a video acquisition device, online decoding and frame extraction; The main AI model performs a target detection task for each input image; According to the preset preliminary screening rules, samples with poor AI model output results are identified as difficult samples; (2) Edge sample intelligent grading; Intelligently grade each difficult sample screened out; By extracting and analyzing the multi-dimensional feature vector of the difficult sample, the potential learning value of the difficult sample is quantitatively evaluated, and a priority level label is divided; (3) Data upload with priority label; Upload the difficult sample after intelligent grading together with its corresponding priority level label to the cloud server for storage and aggregation; (4) Cloud server state aggregation; The cloud server continuously receives and aggregates data from one or more edge computing nodes; The background service regularly calculates and updates the global state vector of the system, which includes the cumulative number of samples of each priority, the time since the last successful training, the current model performance, sample diversity, and cloud resource load; (5) Run cloud intelligent training decision algorithm; In the cloud server, a training decision agent selects a currently optimal training action from a predefined action space according to the current latest global state vector of the system and a preset optimal decision strategy; (6) Model on-demand training and version management; According to the optimal training action output by the decision agent, the cloud server automatically performs the corresponding model training task; If the decision is "no training", step (6) is skipped; If the decision is "training", the training module is called to fine-tune or retrain the existing model using the collected high-value samples, and a new model with a version number is generated; (7) Model online update and closed-loop feedback; The newly trained model is sent to the edge computing node to realize online hot update of the main AI model, so that the edge computing node can obtain stronger recognition ability without interrupting service.

2. The collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making according to claim 1, characterized in that, In step (1), the main AI model is a YOLOv8 model; In step (1), the real-time video data stream is processed and analyzed, including: Preprocessing the input image frames, including size normalization and numerical normalization; The preprocessed image frames are input into the YOLOv8s model for a complete forward calculation; The YOLOv8s model finally outputs a tensor containing multiple prediction results; Through post-processing operations on the tensor, one or more detection results are obtained, each including the position information of the target, the target category, and a confidence score representing the prediction confidence. 3.The method of claim 1, wherein, In step (2), edge sample intelligent grading includes: (2-1) Multi-dimensional feature vector construction: for each difficult sample screened out initially, a multi-dimensional feature vector F is extracted, the multi-dimensional feature vector F includes: Model confidence : confidence score of the target detection box output by the main AI model , ; Image sharpness : The degree of blurring of an image is quantified by computing the normalized Laplacian variance of the image, as follows: ; wherein, is the normalized image, is the variance calculation; Laplacian is the Laplacian operator; target saliency : is the ratio of the target detection frame area to the total image area, which is used to measure the size of the target, ; (2-2) Sample learning value function: through the sample learning value function , a comprehensive value score is obtained by weighting and aggregating the multi-dimensional feature vector F. (2-3) Priority ranking strategy: prioritize samples according to the output value of the value function learned from the samples .

4. The collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making of claim 1, wherein, Priority tiering strategies use different value thresholds ( , This can be achieved as follows: ; wherein, high priority, medium priority, low priority; P refers to the final output result of the priority level assigned to each evaluated difficult sample; Value function As follows: ; wherein, , , are weight coefficients for model confidence , image sharpness , and target saliency , respectively. is an activation function that maps the image sharpness to the interval [0, 1], where, and are tuning parameters.

5. The collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making according to claim 1, characterized in that, In step (3), aggregation refers to: the cloud server sorts, routes and collects the difficult samples with different priority labels from one or more edge nodes through a structured data stream processing and storage mechanism.

6. The collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making of claim 1, wherein, In step (5), the cloud intelligent training decision algorithm is run; including: The goal of training the decision agent is to learn an optimal policy that maximizes long-term cumulative reward; the decision logic of the trained decision agent is modeled by a Q-Learning algorithm; (5-1) State space and action space: State is defined as a vector wherein, , , are the cumulative number of high, medium, and low priority samples, respectively, is the time since last training; Action space is defined as a discrete set, , , , represent waiting, fast training and full training, respectively; (5-2) Reward function: design a reward function for evaluating the immediate reward obtained after transitioning to a new state from performing an action in a state ; the reward function is as follows: ; wherein, represents performing an action the amount of performance improvement of the post-AI model mAP, mAP refers to the average precision mean; if is , then ; represents performing an action the computing resource cost consumed, if is , then ; represents the time consumed by performing an action ; (5-3) Action-value function and policy: training the decision agent by learning an action-value function to estimate the long-term cumulative reward of performing an action in state ; the update of the Q-function follows the Bellman equation, as follows: ; wherein, is the learning rate, is the discount factor, refers to the immediate reward, refers to the next state with the maximum expected future return, is a placeholder variable used in the max operation; The optimal strategy of an agent is to choose the action that maximizes the expected value of the reward, given the current state of the environment is to choose the action that maximizes the expected value of the reward, given the current state of the environment is to choose the action that maximizes the expected value of the reward, given the current state of the environment is to choose the action that maximizes the expected value of 。 7. The collaborative learning method based on edge sample intelligent grading and cloud intelligent decision-making according to any one of claims 1-6, characterized in that, In step (7), the model is updated online and fed back in a closed loop; including: (7-1) Model distribution and loading: the cloud server maintains a model repository to store all versions of the trained model; when a new and better model is generated, the cloud server distributes an update instruction to the edge computing node through a dedicated API interface, the update instruction including the download address and version number of the new model; (7-2) Edge hot update: after receiving the update instruction, the edge computing node downloads the new model file from the specified download address to the local; through an atomic model instance replacement operation, the main AI model object in the memory providing service is safely replaced by the newly loaded model instance; (7-3) Closed-loop verification: after the model is updated, the edge computing node continues to identify the video stream. 8.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-7. The processor executes the computer program to realize the steps of the collaborative learning method based on edge sample intelligent grading and cloud intelligent decision of any one of claims 1-7.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the collaborative learning method based on edge sample intelligent grading and cloud intelligent decision of any one of claims 1-7.

10. A collaborative learning system based on edge sample intelligent grading and cloud intelligent decision-making, characterized in that, Including: The preliminary screening module is configured to perform AI model real-time inference and difficult sample preliminary screening; specifically including: deploying and running the main AI model on the edge computing node to process and analyze real-time video data streams; the real-time video data stream refers to a series of continuous and independent image frames formed after real-time acquisition by a video acquisition device and online decoding and frame extraction; the main AI model performs a target detection task for each input image; according to the preset preliminary screening rule, samples with poor AI model output results are identified as difficult samples; The intelligent grading module is configured to perform edge sample intelligent grading; specifically including: intelligent grading of each difficult sample screened out initially; by extracting and analyzing the multi-dimensional feature vector of the difficult sample, the potential learning value of the difficult sample is quantitatively evaluated and a priority level label is assigned; The data upload module is configured to upload data with priority labels; specifically including: uploading the difficult sample after intelligent grading together with its corresponding priority level label to the cloud server for storage and aggregation; The state aggregation module is configured to aggregate the states of the cloud server, specifically including: the cloud server continuously receives and aggregates data from one or more edge computing nodes; a background service periodically calculates and updates the global state vector of the system, which includes: the cumulative number of each priority sample, the time since the last successful training, the current model performance, the sample diversity, and the cloud resource load; The cloud intelligent training decision algorithm running module is configured to run the cloud intelligent training decision algorithm, specifically including: at the cloud server, a training decision agent selects a currently optimal training action from a predefined action space according to the current latest global state vector of the system and according to a preset optimal decision strategy; The model on-demand training and version management module is configured to automatically execute a corresponding model training task according to the optimal training action output by the decision agent; if the decision is "no training", it is skipped; if the decision is "training", the training module is called to fine-tune or retrain the existing model using the collected high-value samples, and a new model with a version number is generated; The online updating and closed-loop feedback module is configured to update the model online and implement closed-loop feedback, specifically including: the new model generated by training is distributed to the edge computing nodes to realize online hot updating of the main AI model, so that the edge computing nodes obtain stronger recognition capability without interrupting the service.

Citation Information

Patent Citations

  • Transform large model training method based on cloud edge collaboration

    CN119294444A

  • AI robot intelligent target identification and tracking platform based on deep learning

    CN120070496A

  • Data-driven learning model combination optimization method based on edge cloud collaboration

    CN120128577A

  • Large model parameter adjustment method based on cloud edge collaboration and related equipment

    CN120851140A

  • Systems and methods for safe policy improvement for task oriented dialogues

    US20220036884A1

Cited By

  • Cloud side-end collaborative unmanned aerial vehicle cluster intelligent sensing and decision-making system

    CN121477980A