A multi-terminal oriented multi-model collaborative continual learning method

By adopting a multi-terminal, multi-model collaborative continuous learning method, the problem of excessive network bandwidth load in multi-robot parallel operations is solved, achieving efficient model deployment and evidence feedback, and improving the real-time performance and stability of the operation.

CN121724100BActive Publication Date: 2026-05-01XIAMEN YAYA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN YAYA INFORMATION TECH CO LTD
Filing Date
2026-02-11
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In multi-robot parallel operations, existing technologies cause excessive network bandwidth load and increased latency due to high-frequency image data backhaul, affecting the effectiveness and continuity of operation decisions. Furthermore, under bandwidth-constrained conditions, it is difficult to continuously backhaul key evidence, resulting in slower task scheduling response and inconsistent model updates.

Method used

The central terminal selects model combinations and their routing parameters based on scenario priors and terminal capabilities. The terminal performs local inference and sends back structured learning logs. The central terminal performs quality assessment and anomaly detection, triggers small-volume evidence back transmission, and performs group clustering and federated aggregation updates to realize personalized parameters and routing strategies.

Benefits of technology

It reduces uplink traffic, improves link utilization efficiency, enhances the real-time and continuity of operations, ensures the traceability of key evidence and timely location of anomalies under low bandwidth conditions, and strengthens the convergence stability and generalization ability of learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121724100B_ABST
    Figure CN121724100B_ABST
Patent Text Reader

Abstract

The application discloses a multi-terminal-oriented multi-model collaborative continuous learning method, which comprises the following steps: a center end determines an initial combined model and routing parameters for each terminal from a model pool based on scene priori and terminal capability, and issues base model parameters, incremental parameters and inference configuration to the terminal; the terminal performs inference on local multi-source perception data by using the issued model combination, generates a training signal and returns a structured learning log; the center end performs quality evaluation and abnormality detection according to the structured learning log, triggers an active learning evidence collection instruction, and the terminal returns small-volume evidence samples and context according to the instruction; the center end fuses the structured learning log, the evidence samples and the update amount, groups and clusters the terminals, updates the routing strategy, generates group-shared updates by using federal aggregation, and outputs personalized incremental parameters and the routing strategy.
Need to check novelty before this filing date? Find Prior Art

Description

A Multi-Model Collaborative Continuous Learning Method for Multiple Terminals Technical Field

[0001] This invention relates to the field of model collaborative continuous learning technology, and in particular to a multi-model collaborative continuous learning method for multiple terminals. Background Technology

[0002] Existing laser weeding equipment typically acquires images of the farmland operating environment using sensors such as vehicle-mounted RGB cameras and near-infrared / infrared cameras. At the edge or center, algorithms such as target detection and semantic segmentation are used to identify and locate weeds and crops in the images to generate a list of targets and provide a basis for laser operations. To improve recognition performance in complex lighting and occlusion environments, some solutions employ multi-source image acquisition and fusion processing. For example, this involves registering RGB and infrared images, enhancing and denoising multi-source images, or introducing multi-model collaborative inference mechanisms to output multiple recognition results for the same operating area and then fusing them to obtain more stable target area and boundary information.

[0003] In large-scale agricultural operations, several laser weeding robots and several fields awaiting treatment often coexist. To increase the area covered per unit time, existing systems typically employ a multi-robot parallel operation approach: multiple robots simultaneously enter different plots or different work sections of the same plot, establishing communication connections with edge stations / clouds via cellular networks, Wi-Fi, or ad hoc networks for tasks such as operation monitoring, task allocation, operation log aggregation, anomaly alarms, and model parameter distribution and updates. Because laser weeding operations are highly dependent on environmental perception, to ensure recognition quality and traceability, existing technologies often tend to transmit high-frequency image frames, short video clips, or multi-source image data, or perform unified inference and fusion at the central end, significantly increasing the data volume carried by the communication link and the computational collaboration requirements.

[0004] However, in the above-mentioned multi-robot parallel operation mode, the images and related metadata generated by a single robot during the operation process have a high data volume. When the number of robots increases, the uplink data demand is linearly superimposed and may even be further amplified due to multi-source fusion, repeated backhaul and convergence of multi-model results, which can easily exceed the available uplink bandwidth of the cellular network or local wireless network in the farmland, resulting in the communication link being in a state of high load.

[0005] In situations with limited bandwidth or multiple concurrent users, existing technologies that rely on central data aggregation mechanisms are prone to issues such as upload queue backlog, increased retransmissions, and decreased effective throughput. Meanwhile, image data is highly time-sensitive, and increased latency can cause the images received by the central terminal to be inconsistent with the robot's current position / posture, thereby reducing the effectiveness of central terminal analysis, monitoring, and scheduling decisions.

[0006] Uneven network coverage, signal obstruction, bandwidth fluctuations, latency jitter, and packet loss are common problems in farmland operations. Existing technologies employing centralized inference, unified fusion, or high-frequency synchronous update strategies are prone to issues such as missing data from some robots, intermittent data transmission, or asynchronous version updates between terminals when the link is unstable. This leads to discontinuous operation monitoring, slower task scheduling response, and difficulty in consistently implementing model updates, thus limiting the large-scale deployment of multi-robot parallel operations.

[0007] To alleviate bandwidth pressure, some solutions attempt to reduce image upload frequency or compress image quality. However, such methods often weaken the central end's visibility into the operation process and its ability to judge key scenarios. At the same time, when different plots of land and different crop stages require differentiated strategies, if large-scale data backhaul is still required to complete the analysis or update, communication and maintenance costs will continue to rise, making it difficult to meet the engineering requirements of long-term, large-scale, and multi-terminal deployment.

[0008] The purpose of this invention is to design a multi-terminal, multi-model collaborative continuous learning method to address the problems existing in the prior art. Summary of the Invention

[0009] In view of this, the purpose of this invention is to propose a multi-model collaborative continuous learning method for multiple terminals, which can solve the above-mentioned problems.

[0010] This invention provides a multi-model collaborative continuous learning method for multiple terminals, comprising:

[0011] Based on scenario priors and terminal capabilities, the S1 central terminal determines the initial combined model and its routing parameters for each terminal from the model pool, and sends the base model parameters, incremental parameters, and inference configuration to the terminal.

[0012] The S2 terminal uses the distributed model combination to infer local multi-source perception data, generates training signals, and sends back structured learning logs.

[0013] The S3 central terminal performs quality assessment and anomaly detection based on the structured learning logs, triggering active learning forensics instructions. The terminal then sends back a small sample of evidence and context according to the instructions.

[0014] The S4 central terminal integrates structured learning logs, evidence samples, and update volumes to group and cluster terminals and update routing policies. It uses federated aggregation to generate group-shared updates and outputs personalized incremental parameters and routing policies.

[0015] The beneficial effects of this invention are:

[0016] First, by using the central terminal to select model combination models and model routing parameters from the model pool based on scenario priors and terminal capabilities, and then distributing base parameters, expert incremental parameters and inference configurations, the problem of poor generalization and difficulty in adapting a single model under different plots / crop stages / light shading differences is solved. At the same time, the problem of models being undeployable or experiencing instability under different terminal computing power / storage / latency constraints is solved, realizing differentiated deployment of models according to scenarios and capabilities.

[0017] Secondly, by using multiple models distributed from the edge to complete local inference and compressing the learning signal into structured learning logs for back transmission, the high bandwidth consumption and concurrency caused by high-frequency image / video uplink, as well as the latency and pose asynchrony risks caused by unified inference at the central end, are resolved. This significantly reduces uplink communication volume and congestion probability, improves link utilization efficiency, and the edge-side closed loop enables more timely operation decisions, thereby improving the real-time performance, continuity, and anti-network jitter capability of robot operations.

[0018] Third, the central terminal performs quality assessment and anomaly detection based on structured learning logs, triggering active learning forensics. The terminal combines network status to hierarchically transmit small-volume evidence samples and context, solving the problem of not being able to continuously transmit large amounts of data under bandwidth-limited / link fluctuations but still needing traceable evidence, as well as the problems of difficulty in locating and attributing anomalies and lack of effective training samples. Even under low bandwidth conditions, key evidence can still be obtained for verification and attribution.

[0019] Fourth, by merging logs, evidence samples, and update volume, grouping and clustering are performed. Within each group, evidence-weighted federated aggregation is executed and expert routing strategies are updated. Group-shared updates and personalized incremental parameters are distributed to form group-based shared updates and personalized adaptations for terminals, thereby improving the convergence stability and generalization ability of continuous learning. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 is a flowchart of the method in this embodiment. Detailed Implementation

[0022] To facilitate understanding by those skilled in the art, the structure of the present invention will now be described in further detail with reference to the accompanying drawings. It should be understood that, unless otherwise specified, the order of the steps mentioned in this embodiment can be adjusted according to actual needs, and they can even be executed simultaneously or partially simultaneously.

[0023] As shown in Figure 1, this embodiment of the invention provides a multi-model collaborative continuous learning method for multiple terminals, including:

[0024] Based on scenario priors and terminal capabilities, the S1 central terminal determines the initial combined model and its routing parameters for each terminal from the model pool, and sends the base model parameters, incremental parameters, and inference configuration to the terminal.

[0025] In this step, instead of placing detection and decision-making directly on the central server, the identification and detection are performed on the edge as much as possible. First, a usable model is selected based on scene characteristics and terminal resources / network capabilities, and then a lightweight adaptation is performed. Subsequently, only the base model parameters, incremental model parameters, and inference configuration of the target model combination are distributed to the corresponding terminal. After receiving the model, the terminal performs inference locally, only needing to upload a small amount of result data instead of continuously uploading large amounts of raw data. Therefore, the dependence on communication bandwidth is significantly reduced, allowing for stable completion of identification and detection tasks even under low bandwidth conditions.

[0026] The S101 central terminal creates an operation project, binds the operation plot and the corresponding operation terminal, and generates operation scene features based on the operation plot information and prior data. The operation scene features include: crop type and growth period, surface reflectivity probability, weed density prediction, row shading level, operation time period, and soil moisture content. The operation scene features are encoded to generate a scene feature vector.

[0027] The S102 central terminal acquires the operating data of each working terminal, including: processor computing power, memory size, storage space, network bandwidth, and latency data;

[0028] S103 constructs a model pool by using several weed target detection models oriented towards different input modalities, several weed target segmentation models, several multimodal fusion models, and crop risk assessment models. It establishes metadata descriptions for each model and marks the applicable scenario feature range and corresponding terminal running data constraint parameters.

[0029] Based on scene feature vectors and terminal operation data, the S104 central terminal selects model combinations that meet the conditions from the model pool and maps them to generate parameters and inference configurations for model routing strategies.

[0030] The S1041 central terminal determines its deployment constraints and optimization objectives based on the terminal's operating data. The deployment constraints include at least: the maximum available storage limit on the terminal side, the maximum acceptable inference latency or minimum frame rate on the terminal side, and the network bandwidth and latency limits. The optimization objectives include at least: maximizing weed recognition accuracy, minimizing false positive rate, and maximizing operational continuity.

[0031] The S1042 central terminal performs similarity classification and capability grading on scene feature vectors and terminal operation data. Based on scene clusters and terminal clusters, it retrieves a set of candidate models from the model pool and performs input modality matching and resource constraint verification on the set of candidate models in combination with model metadata to obtain a subset of candidate models that meet the current scene and deployment constraints.

[0032] The S10421 central terminal clusters the scene feature vectors of historical operation scenarios to obtain several scene clusters. It also establishes a cluster center vector, intra-cluster distribution statistics, and cluster identifier for each scene cluster. The central terminal statistically analyzes the historical performance indicators and resource consumption indicators of different models under each scene cluster to form a correlation profile between scene clusters and models.

[0033] S10422 The central end calculates the distance between the current scene feature vector and the center vector of each scene cluster based on the current scene feature vector, determines the matching target scene cluster, and selects a set of candidate models from the scene cluster and model-related profile corresponding to the target scene cluster.

[0034] The S10423 central terminal clusters the terminal operation data to obtain several terminal capability clusters and establishes a resource constraint template for each terminal capability cluster. The central terminal performs capability constraint pre-screening on the candidate model set according to the terminal capability cluster to which the terminal belongs, and generates a subset of candidate models that meet the resource constraint template.

[0035] Based on the determined deployment constraints and network backhaul budget, the S10424 central terminal performs constraint verification on a subset of candidate models, eliminating models that do not meet the input modality requirements, peak memory constraints, storage upper limit constraints, inference latency constraints, and network backhaul budget constraints.

[0036] In this step, the central terminal clusters the scene feature vectors of historical operation scenarios to form scene clusters (S10421) and clusters the terminal operation data to form terminal capability clusters (S10423). Based on this, in the online decision-making stage, the current scene vector is matched to the target scene cluster (S10422), and the model is retrieved and pre-screened by combining the resource constraint template of the terminal capability cluster (S1042). This allows for the rapid and effective acquisition of model combinations that match the terminal.

[0037] The S1043 central terminal generates a candidate model combination set based on a subset of candidate models. The candidate model combination includes at least: single model combination, detection and segmentation cascade combination, and multimodal fusion combination. Among them, the cascade combination defines the pre-model and post-model and their triggering conditions, and the multimodal fusion combination defines the input dependency relationship and output fusion method of each modality.

[0038] S1044 generates parameters and inference configurations for the model routing strategy for each model mapping in the target model combination, and sends the base model parameters and model incremental parameters of each model in the target model combination to the corresponding terminal.

[0039] In this step, routing parameters refer to the decision parameters (such as gating rules / thresholds, cascading trigger thresholds, hysteresis parameters, fusion weights, and degradation strategies) that the edge device uses to determine "which model to run, when to switch, whether to trigger post-segmentation, and how to fuse the output" among multiple models / inference links. These parameters need to be distributed because the optimal calling strategy differs under different scenarios and terminal constraints. Only by using these parameters can the edge device obtain the target effect stably and with low latency.

[0040] The base model parameters refer to the backbone network weights (general feature extraction / general detection capabilities, etc.) that serve as the foundation for general capabilities. They need to be distributed because edge inference must have a basic weight file, and using a unified base facilitates reuse and version consistency.

[0041] Model incremental parameters refer to lightweight differential parameters (such as Adapter / LoRA / incremental weights / patches) adapted to specific scenarios / modalities / crop stages on top of the base model. They are used to quickly "adjust" the general model to the current scenario. They need to be distributed because transmitting only the incremental parameters can significantly reduce the distribution volume, speed up updates, and allow the same base model to adapt to multiple scenarios.

[0042] Inference configuration refers to the runtime configuration of "how to run the model" (input resolution, frame rate / sampling rate, quantization accuracy, backend CPU / GPU / NPU, parallel / batch processing, timeout and post-processing thresholds, etc.); it needs to be distributed because the same model needs to be used at different settings on different terminals to meet latency / memory / frame rate constraints and ensure continuous operation.

[0043] The S2 terminal uses the distributed model combination to infer local multi-source perception data, generates training signals, and sends back structured learning logs.

[0044] In this step, the terminal is a laser weeding robot, and the local multi-source sensing data may include, but is not limited to: images / video frames from the operating camera, pose / velocity information, IMU / odometry information, and the execution status of the laser and galvanometer. The model combination is used to output candidate weed targets, crop risk assessment results, and operation judgment results. To adapt to low communication conditions, this embodiment does not generate an upload payload of raw sensing data (raw image or video data), but instead organizes the minimum necessary information required for training into a structured learning log.

[0045] The structured learning log consists of compressed log blocks, each of which corresponds to at least one data record of a job segment. The specific steps are as follows:

[0046] After the S201 terminal side completes target recognition and task determination based on the distributed model combination on the local multi-source sensing data, it generates learning sample entries for each candidate target and forms a candidate learning sample sequence sorted by time.

[0047] In this step, the learning sample entries are used to carry structured fields of the training signal, including at least: timestamps or frame numbers, target geometry fields, weed confidence fields, crop risk fields, uncertainty fields, operation stage codes, laser execution status codes, and model / model version identifiers. By converting the continuous sensing stream into a sequence of target-level entries, subsequent segmentation, encoding, and compression processes have clearly defined data objects and field boundaries.

[0048] S202 segments the candidate learning sample sequence according to the semantic stages of laser weeding operations to obtain multiple operation segments. The semantic stages include at least: entry stable segment, stable weeding segment, dense weed segment, high reflectivity / obstruction segment, and turning non-operation segment.

[0049] In this step, the terminal side divides the candidate learning sample sequence generated during a continuous operation into different semantic stages according to preset logic. For example, in the stable entry stage, the robot enters the work row from an unstable movement state, and the crop row detection / localization is still converging; in the stable weeding stage, the intra-row localization is stable, the perception quality is normal, and the operation continues. This makes the data change patterns within each work segment more consistent, so that a higher compression ratio can be obtained in the subsequent compression process. Furthermore, different fidelity strategies can be set for different semantic stages to avoid compressing and masking key anomalies (such as the risk of mis-hitting due to reflection).

[0050] S203 performs field shaping and hierarchical description on the learning sample bars within each work segment, and encodes and marks the learning sample bars with field shaping and hierarchical description to obtain the relative coded segment;

[0051] S2031 quantizes the continuous value field of the learning sample bar into fixed-point integers;

[0052] S2032 maps the discrete fields of the learning sample bar to short codes;

[0053] S2033 uses a hierarchical representation for the geometric fields of the learning sample bars;

[0054] S2034 selects the first entry of each segment from the learning sample bars that are shaped and hierarchically divided into segments, encodes the remaining entries as relative quantities, and generates fidelity tags for the learning key fields.

[0055] In this step, continuous value fields can include target coordinates, dimensions, fractions, uncertainty, distance, etc. Converting continuous values ​​to integers using a preset scaling factor reduces field storage overhead and facilitates subsequent differential (relative quantity) encoding to generate more values ​​with small variations or zero variations.

[0056] Discrete fields can include model ID, job stage code, laser execution status code, interception reason code, etc. Mapping them to short codes using a dictionary table can reduce the representation overhead of discrete fields and create conditions for subsequent compression of repetitive state sequences;

[0057] Geometric fields are preferentially represented in a lightweight form (e.g., (x,y,w,h)). When the target is near a crop protection zone, or when there is a high risk or uncertainty in the crop, a more refined geometric description (e.g., simplified polygon boundaries) can be attached. Through hierarchical representation, most common targets occupy only a small volume of geometric fields, while retaining sufficient training evidence at key samples.

[0058] The first entry of a segment is retained as the baseline entry; the remaining entries are represented, at least in terms of the change relative to the previous entry or the baseline entry, for target geometry, fractional fields, pose summaries, etc. For key learning fields (including but not limited to crop risk grade changes, laser execution state changes, uncertainty mutations, etc.), a fidelity marker is generated to indicate whether the absolute value of the field must be recorded or whether it must be fully retained, so as to avoid the loss of key details of the training signal due to the compression process.

[0059] S204 performs repeating pattern compression on the short code sequence, zero-difference sequence, and repeating state sequence within each relative coding segment to form a compressed log block.

[0060] In this step, repetition pattern compression can be achieved by compressing consecutive identical values ​​as tuples. For example, consecutively occurring identical short codes or consecutively occurring zero differences can be represented in the form of (value, number of repetitions), thereby significantly reducing the occupation of redundant fields within the segment.

[0061] The compressed log block includes metadata for parsing and tracing, which includes at least: segment type identifier, segment start timestamp, segment baseline entry, quantization scale identifier, dictionary version identifier, and compressed differential payload.

[0062] The S3 central terminal performs quality assessment and anomaly detection based on the structured learning logs, triggering active learning forensics instructions. The terminal then sends back a small sample of evidence and context according to the instructions.

[0063] The S301 central end obtains intra-segment learning sample entries based on decoding compressed log blocks, and calculates several quality evaluation indicators for each semantic job segment;

[0064] The S3011 central terminal reads the confidence level of each learning sample entry in the decoded semantic task segment, and determines whether each sample belongs to the low-confidence sample according to the low-confidence threshold. The number of samples determined to be low-confidence is accumulated and divided by the total number of samples in the segment to obtain the low-confidence ratio of the semantic task segment, which is used as the first core quality indicator. The calculation formula is as follows:

[0065] ,

[0066] in, This represents the first core quality indicator, namely the low confidence percentage of this semantic task segment. This indicates the number of entries within the semantic segment. Indicates the entry index. This indicates an indicator function; the condition is 1 if true, and 0 otherwise. Indicates the first The confidence level of the target weeds This indicates the low confidence threshold.

[0067] The S3012 central terminal reads the uncertainty measure for each learning sample entry in the semantic task segment, and determines whether the entry is a high-uncertainty sample based on the high uncertainty threshold. The number of high-uncertainty samples is accumulated and divided by the total number of samples in the segment to obtain the high uncertainty ratio of that semantic task segment, which serves as the second core quality indicator. The calculation formula is as follows:

[0068] ,

[0069] in, This indicates the second core quality indicator, namely the proportion of semantic task segments with excessively high uncertainty. This represents the uncertainty measure of the i-th objective. This indicates the threshold for determining high uncertainty.

[0070] The S3013 central terminal reads the crop risk score for each learning sample entry in the semantic task segment, and counts the number of samples with risks exceeding the high-risk threshold based on the high-risk threshold. This number is then divided by the total number of samples to obtain the high-risk ratio, which serves as the third core quality indicator. The calculation formula is as follows:

[0071] ,

[0072] in, This indicates the third core quality indicator, namely the proportion of crops with excessively high risk within this semantic segment. This represents the crop risk score for the i-th objective. This indicates the threshold for determining high risk;

[0073] In this step, the weed confidence score is obtained by the weed target detection model issued by the model pool through inference on the collected operation images and / or multimodal data, and is used to characterize the confidence level that the target is a weed. The uncertainty measure is directly output by the weed target detection model during the inference process, or calculated by the terminal based on the category probability distribution output by the detection model, and is used to characterize the uncertainty of the detection result. The crop risk score is output by the crop risk assessment model issued by the model pool. The crop risk assessment model takes the identification result of the detection model and / or the spatial geometric information related to the crop as input, and outputs the crop accidental damage risk of the corresponding target, which is used to constrain the safety of laser strikes.

[0074] The S3014 central terminal reads the indicator variable indicating whether the laser was triggered and the execution success indicator variable for each learning sample entry in the semantic task segment, counts the number of triggers within the segment, and counts the number of times the laser was triggered and executed successfully.

[0075] When the number of triggers within a segment is greater than 0, the central terminal calculates the execution success rate and obtains the execution failure rate as the fourth core quality indicator. The calculation formula is as follows:

[0076] ,

[0077] ,

[0078] in, This represents the fourth core quality metric, namely the execution failure rate after triggering. This indicates the success rate of execution after departure. This indicates whether the i-th laser is triggered (1 for trigger, 0 otherwise). This indicates whether the i-th laser execution was successful (1 for success, 0 for failure; typically only when...). (Valid when =1) This represents a very small constant, used to avoid a denominator of 0;

[0079] In this step, whether to trigger the laser is decided by the terminal. The decision is made by comprehensively considering the weed confidence level, uncertainty measure, crop risk score, and safety interlock status to determine whether to issue a laser trigger command. The decision is based on a preset strategy. Successful execution information is returned by the terminal execution link. Specifically, it is the execution result status fed back by the laser controller or hardware driver after receiving the trigger command, which is used to characterize whether the laser command was successfully executed.

[0080] S302 obtains the historical baseline statistics of the same semantic type as the current semantic operation segment, calculates the standardized deviation of each quality assessment indicator relative to the historical baseline, and performs non-negative truncation on the deviation values ​​to obtain the abnormal scores of each indicator. The abnormal scores are then weighted and summed according to preset weights to obtain the comprehensive abnormal score.

[0081] In this step, for the semantic task segment to be evaluated, the central end first obtains the historical baseline data corresponding to the historical task segments with the same semantic type. The historical baseline data is used to characterize the normal fluctuation range of each quality evaluation indicator under the same semantic type and similar task conditions.

[0082] After obtaining the historical baseline, the standardized deviation of each quality assessment indicator relative to the historical baseline is calculated, and the direction of deviation is constrained so that it only scores deviations of "quality deterioration", thus obtaining the anomaly score corresponding to each quality assessment indicator. Further, the anomaly scores are weighted and summarized according to preset weights to obtain the comprehensive difference (comprehensive anomaly score) used to characterize the overall quality deviation of the semantic operation segment.

[0083] S303 compares the overall anomaly score with the anomaly threshold. If the overall anomaly score exceeds the anomaly threshold, the semantic operation segment is determined to be an anomaly segment and an anomaly flag is output. At the same time, the quality assessment index with the largest anomaly score is selected as the main source of anomaly, and corresponding evidence is obtained for evidence collection and attribution.

[0084] In this step, the overall difference obtained in step S302 is compared with a preset anomaly threshold. When the overall difference exceeds the anomaly threshold, the semantic work segment is determined to be an anomaly segment and an anomaly flag is output to trigger subsequent alarm, downgrade, review, or evidence collection processes. To achieve interpretable localization of the anomaly cause, the quality assessment index with the highest anomaly score in the semantic work segment is further determined as the primary anomaly source index, and corresponding evidence is obtained based on this primary anomaly source index.

[0085] When the main anomaly source indicators correspond to a decrease in the credibility of the perception results (e.g., abnormal confidence in weed identification) and an abnormal increase in uncertainty, the central terminal sends an evidence request to the terminal, or the terminal actively returns evidence data according to preset rules. The evidence data includes at least: key frame images and / or short video clips within the semantic operation segment, the corresponding detection / segmentation result overlay, the confidence and uncertainty distribution of the model output, and the timestamp and pose information synchronized with the key frame. The central terminal performs random checks or manual checks on the recognition results based on the returned image evidence to verify whether the anomaly is caused by perception failure, data quality degradation, or scene boundary violation.

[0086] After determining that the semantic operation segment is an abnormal segment and obtaining the corresponding evidence, S304 obtains the network status parameters of the terminal, classifies the evidence according to the network status parameters, and determines the return strategy corresponding to each level of evidence.

[0087] S3041 determines the network level based on network bandwidth and latency data, and the network level includes: low bandwidth and high latency level, general level and high bandwidth and low latency level;

[0088] If S3042 is classified as a low-bandwidth, high-latency class, then only the time point corresponding to the time point with the largest contribution of the abnormal score will be retained. Zhang keyframes are scaled down to a preset low resolution and encoded and compressed according to a preset compression ratio. The compressed keyframes and their corresponding segmentation results are sent back, along with the timestamp, semantic job segment identifier, and anomaly index summary corresponding to the keyframes.

[0089] If S3043 is a general level, then the acquired image will be extracted. Zhang keyframes are scaled to a preset medium resolution and encoded and compressed according to a preset medium compression ratio. The processed keyframes and their corresponding segmentation results are sent back, along with the timestamp, semantic job segment identifier, and anomaly index summary corresponding to the keyframes.

[0090] If S3044 is in the high bandwidth, low latency class, then the acquired image will be extracted. The system retrieves keyframes, transmits the segmentation result evidence corresponding to the keyframes, and transmits the timestamp, semantic job segment identifier, and anomaly index summary corresponding to the keyframes.

[0091] In this step, after a semantic task segment is identified as abnormal, the evidence is adaptively transmitted back based on the actual network conditions of the terminal to balance the timeliness of transmission and the verifiability of the evidence. The terminal first obtains the network status parameters obtained in step S102 and maps the network status to a preset network level. Under low bandwidth and high latency conditions, the minimum necessary evidence is prioritized, and only a small number of key frames contributing most to the anomaly are transmitted back, using low resolution and a high compression ratio to ensure priority and rapid arrival of key evidence. Under normal network conditions, the number of key frames is moderately increased, and a medium resolution and compression ratio are used to improve coverage while maintaining controllable latency. Under high bandwidth and low latency conditions, the number of key frames is further increased, and transmission is performed using the original or high resolution with a lower compression ratio to maximize detail retention and verification accuracy. The transmitted content at each level includes key frames and their corresponding segmentation result evidence (segmentation mask data and / or overlaid visualization), along with timestamps, semantic task segment identifiers, and anomaly index summaries for central-end alignment and localization, rapid verification, and traceability analysis. The number of key frames meets the following requirements. .

[0092] The S4 central terminal integrates structured learning logs, evidence samples, and update volumes to group and cluster terminals and update routing policies. It uses federated aggregation to generate group-shared updates and outputs personalized incremental parameters and routing policies.

[0093] In this step, under the scenario of multi-terminal parallel operation, the data distribution of each terminal exhibits significant heterogeneity due to differences in the plot of land where the terminal is located, the crop growth stage, light shading, weed density, and camera / sensor differences. If an indiscriminate global federated aggregation is used, the feature distribution of some terminals may dominate the global update and cause negative transfer to other terminals, resulting in decreased recognition accuracy or an increased risk of crop misidentification. At the same time, under bandwidth fluctuation conditions, the quality of evidence transmission and update volume is inconsistent. Without the introduction of quality measurement and weighting mechanisms, low-reliability updates are easily treated the same as high-value updates, reducing the overall convergence stability.

[0094] After receiving the data returned by each terminal, the S401 central terminal binds the evidence samples and update volume corresponding to the abnormal segment based on the semantic job segment identifier and timestamp to obtain the abnormal segment record. For each abnormal segment record, the evidence aggregation weight is calculated based on its comprehensive abnormal score.

[0095] In this step, a correspondence is established between the evidence samples and the update quantities on the same semantic segment, so that the central end can clearly identify "which anomalous segment evidence corresponds to a certain update quantity", which facilitates subsequent attribution and quality control. The comprehensive anomaly score is used as the weighting basis, so that the aggregation process has a greater impact on the update of anomalous segments that are "more likely to represent the model's weaknesses and have higher learning value", thereby reducing overfitting of normal or noisy segment updates.

[0096] S402 generates an evidence representation vector based on the evidence sample, and performs weighted fusion of the evidence representation vectors of multiple abnormal segment records of the same terminal to obtain the terminal representation vector corresponding to the terminal.

[0097] In this step, the original evidence samples are transformed into a unified vector representation that can be used for clustering and similarity calculation, enabling the central end to align terminal scenarios and problem types without relying on large-scale original data. By weighted fusion of multiple abnormal segments, the interference of occasional anomalies on the terminal profile is reduced, making the terminal representation more stable and better able to reflect the main difficult scenarios and error patterns of the terminal over a period of time.

[0098] S403 central terminal performs grouping and clustering of terminals based on the terminal representation vector to obtain at least one terminal group and the group identifier of each terminal, and generates a group prototype vector for each terminal group.

[0099] In this step, terminals with similar anomaly patterns and scenario characteristics are grouped into the same group, making updates within the group more consistent and transferable. The group prototype vector is used to characterize the typical scenario / error distribution of the group, providing a reference for subsequent routing policy updates, thereby realizing group-based learning and group-based adaptation.

[0100] For each terminal group, the S404 central terminal performs weighted federated aggregation based on the update volume of each terminal in the group and according to the evidence aggregation weight to obtain the group shared update corresponding to the terminal group. The weighted federated aggregation satisfies at least the following condition: the terminal with the larger the evidence weight, the higher its contribution to the group shared update.

[0101] In this step, compared to global aggregation, intra-group aggregation reduces cross-distribution conflicts, alleviates negative migration, and improves update stability. Compared to equal-weighted aggregation, evidence-weighted aggregation makes highly reliable terminal updates contribute more to group-shared updates, thereby improving the ability of aggregated updates to correct abnormal scenarios and reducing the perturbation of low-quality updates on the overall model.

[0102] The S405 central terminal performs group sharing updates based on the group prototype vector, determines the model routing strategy corresponding to the terminal group, and generates personalized incremental parameters corresponding to the group to which the terminal belongs for each terminal in the group.

[0103] In this step, by updating the routing strategy, the terminal is more likely to call the expert model that matches its group scene in the subsequent inference and learning stages, thereby improving the robustness of recognition and segmentation under complex lighting, occlusion or high reflectivity conditions; personalized incremental parameters enable the terminal to retain the necessary scene adaptation differences on the basis of shared updates, taking into account both the improvement of group sharing capabilities and the personalized needs of the terminal.

[0104] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0105] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0106] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0107] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0108] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The words first, second, and third, etc., do not indicate any order. These words can be interpreted as names.

[0109] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.

[0110] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

[0111] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0112] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

Claims

1. A multi-terminal, multi-model collaborative continuous learning method, characterized in that, include: S1, based on scenario priors and terminal capabilities, determines the initial combined model and its routing parameters for each terminal from the model pool, and sends the base model parameters, incremental parameters, and inference configuration to the terminals; S2, the terminal uses the sent model combination to infer local multi-source perception data, generates training signals, and sends back structured learning logs. The terminal is a laser weeding robot, comprising: S201, after the terminal side completes target recognition and task determination based on the sent model combination and local multi-source perception data, it generates learning sample entries for each candidate target and forms a candidate learning sample sequence sorted by time; S202, the candidate learning sample sequence is segmented according to the semantic stage of the laser weeding operation to obtain multiple... The operational segment, the semantic stage, includes at least: a stable entry segment, a stable weeding segment, a dense weeding segment, a high reflectivity / occlusion segment, and a non-operational turning segment; S203 performs field shaping and hierarchical description on the learning sample bars in each operational segment, and encodes and marks the learning sample bars with field shaping and hierarchical description to obtain a relative coding segment; S204 performs repetition pattern compression on the short code sequence, zero-difference sequence, and repetitive state sequence in each relative coding segment to form a compressed log block; S3 the central end performs quality assessment and anomaly detection based on the structured learning log, triggers an active learning evidence collection instruction, and the terminal returns a small amount of evidence samples and context according to the instruction; S4 the central end integrates the structured learning log, Evidence samples and update quantities are used to group and cluster terminals and update routing strategies. Federated aggregation is employed to generate group-shared updates, outputting personalized incremental parameters and routing strategies. This includes: S401: After receiving data from each terminal, the central terminal binds the evidence samples and update quantities corresponding to abnormal segments based on semantic job segment identifiers and timestamps to obtain abnormal segment records. For each abnormal segment record, an evidence aggregation weight is calculated based on its comprehensive abnormality score. S402: An evidence representation vector is generated based on the evidence samples, and the evidence representation vectors of multiple abnormal segment records from the same terminal are weighted and fused to obtain the terminal representation vector corresponding to that terminal. S403: The central terminal uses the terminal representation vector... The terminals are grouped and clustered to obtain at least one terminal group and the group identifier of each terminal, and a group prototype vector is generated for each terminal group; S404 For each terminal group, the central end performs weighted federated aggregation based on the update amount of each terminal in the group and according to the evidence aggregation weight to obtain the group shared update corresponding to the terminal group, wherein the weighted federated aggregation satisfies at least the following condition: the terminal with the larger the evidence weight, the higher its update amount contributes to the group shared update; S405 The central end performs group shared update based on the group prototype vector, determines the model routing strategy corresponding to the terminal group, and generates personalized incremental parameters corresponding to the group to which the terminal belongs for each terminal in the group.

2. The multi-terminal, multi-model collaborative continuous learning method according to claim 1, characterized in that, The central terminal, based on scenario priors and terminal capabilities, determines the initial combined model and its routing parameters for each terminal from the model pool, and sends base model parameters, incremental parameters, and inference configurations to the terminals, including: S101 The central terminal creates an operation project, binds the operation plot and the corresponding operation terminal, and generates operation scenario features based on the operation plot information and prior data. The operation scenario features include: crop type and growth period, surface reflectivity probability, weed density estimation, inter-row shading level, operation time period, and soil moisture content. The operation scenario features are encoded to generate a scene feature vector; S102 The central terminal obtains the operation data of each operation terminal. The operational data includes: processor computing power, memory size, storage space, network bandwidth, and latency data; S103 constructs a model pool through several weed target detection models oriented towards different input modalities, several weed target segmentation models, several multimodal fusion models, and crop risk assessment models, establishes metadata descriptions for each model, and marks its applicable scenario feature range and corresponding terminal operational data constraint parameters; S104 the central terminal selects model combinations that meet the conditions from the model pool based on scenario feature vectors and terminal operational data, and maps and generates parameters and inference configurations for model routing strategies.

3. The multi-terminal, multi-model collaborative continuous learning method according to claim 2, characterized in that, The central terminal, based on scene feature vectors and terminal operation data, selects model combinations that meet certain conditions from the model pool and maps them to generate parameters and inference configurations for model routing strategies. This includes: S1041 The central terminal determines deployment constraints and optimization objectives based on terminal operation data. The deployment constraints include at least: maximum available storage limit on the terminal side, maximum acceptable inference latency or minimum frame rate on the terminal side, and network bandwidth and latency limits. The optimization objectives include at least: maximizing weed recognition accuracy, minimizing false positive rate, and maximizing operational continuity. S1042 The central terminal performs similarity classification and capability grading on scene feature vectors and terminal operation data. Based on scene clusters and terminal clusters, it retrieves a set of candidate models from the model pool and combines them with... The model metadata is used to perform input modality matching and resource constraint verification on the candidate model set to obtain a subset of candidate models that meet the current scenario and deployment constraints. In S1043, the central terminal generates a set of candidate model combinations based on the candidate model subset. The candidate model combinations include at least: single model combinations, detection and segmentation cascade combinations, and multimodal fusion combinations. Among them, the cascade combination defines the pre-model and post-model and their triggering conditions, and the multimodal fusion combination defines the input dependency relationship and output fusion method of each modality. In S1044, parameters and inference configurations of the model routing strategy are generated for each model mapping in the target model combination, and the base model parameters and model incremental parameters of each model in the target model combination are sent to the corresponding terminal.

4. The multi-terminal, multi-model collaborative continuous learning method according to claim 3, characterized in that, The central terminal performs similarity classification and capability grading on scene feature vectors and terminal operation data. Based on scene clusters and terminal clusters, it retrieves a set of candidate models from the model pool and performs input modality matching and resource constraint verification on the candidate model set in conjunction with model metadata to obtain a subset of candidate models that meet the current scene and deployment constraints. This includes: S10421 The central terminal clusters the scene feature vectors of historical operation scenarios to obtain several scene clusters, and establishes a cluster center vector, intra-cluster distribution statistics, and cluster identifier for each scene cluster. The central terminal statistically analyzes the historical performance indicators and resource consumption indicators of different models under each scene cluster to form a correlation profile between scene clusters and models; S10422 The central terminal calculates its... The distance to the center vector of each scene cluster is used to determine the target scene cluster to match. The center selects a set of candidate models from the scene cluster and model association profile corresponding to the target scene cluster. In S10423, the center clusters the terminal running data to obtain several terminal capability clusters and establishes a resource constraint template for each terminal capability cluster. The center performs capability constraint pre-screening on the candidate model set according to the terminal capability cluster to which the terminal belongs, and generates a subset of candidate models that meet the resource constraint template. In S10424, the center performs constraint verification on the candidate model subset according to the determined deployment constraints and network backhaul budget, and eliminates models that do not meet the input modality requirements, peak memory constraints, storage upper limit constraints, inference latency constraints, and network backhaul budget constraints.

5. The multi-terminal, multi-model collaborative continuous learning method according to claim 1, characterized in that, The process of shaping and hierarchically describing the learning sample bars within each work segment, and encoding and marking the learning sample bars with field shaping and hierarchical description to obtain relative encoded segments includes: S2031 quantizing the continuous value fields of the learning sample bars into fixed-point integers; S2032 mapping the discrete fields of the learning sample bars into short codes; S2033 representing the geometric fields of the learning sample bars using hierarchical representation; S2034 selecting the first entry of each segment of shaped and hierarchical learning sample bars, encoding the remaining entries into relative quantities, and generating fidelity marks for key learning fields.

6. The multi-terminal, multi-model collaborative continuous learning method according to claim 1, characterized in that, The central terminal performs quality assessment and anomaly detection based on the structured learning logs, triggering an active learning forensics instruction. The terminal then transmits small-volume evidence samples and context according to the instruction, including: S301 The central terminal decodes the compressed log blocks to obtain learning sample entries within the segment, and calculates several quality assessment indicators for each semantic job segment; S302 The central terminal obtains the historical baseline statistics of the same semantic type as the current semantic job segment, calculates the standardized deviation of each quality assessment indicator relative to the historical baseline, and performs non-negative truncation on the deviation values ​​to obtain the anomaly score for each indicator. The anomaly scores are then weighted and summarized according to preset weights to obtain a comprehensive anomaly score; S303 The comprehensive anomaly score is compared with an anomaly threshold. If the comprehensive anomaly score exceeds the anomaly threshold, the semantic job segment is determined to be an anomaly segment and an anomaly flag is output. At the same time, the quality assessment indicator with the largest anomaly score is selected as the main source of anomaly, and corresponding evidence is obtained accordingly for evidence collection and attribution; S304 After determining that the semantic job segment is an anomaly segment and obtaining the corresponding evidence, the terminal's network status parameters are obtained, the evidence is classified according to the network status parameters, and the corresponding transmission strategy for each level of evidence is determined.

7. A multi-terminal, multi-model collaborative continuous learning method according to claim 6, characterized in that, The central endpoint obtains learning sample entries within a segment by decoding compressed log blocks. For each semantic task segment, it calculates several quality evaluation indicators, including: S3011 The central endpoint reads the confidence level of each learning sample entry in the decoded semantic task segment, and determines whether each sample belongs to a low-confidence sample based on a low-confidence threshold. The number of samples determined to be low-confidence is accumulated and divided by the total number of samples in the segment to obtain the low-confidence ratio of the semantic task segment, which is used as the first core quality indicator. The calculation formula is as follows: ,in, This represents the first core quality indicator, namely the low confidence percentage of this semantic task segment. This indicates the number of entries within the semantic segment. Indicates the entry index. This indicates an indicator function; the condition is 1 if true, and 0 otherwise. Indicates the first The confidence level of the target weeds This represents the low confidence threshold. The S3012 central terminal reads the uncertainty measure for each learning sample entry in the semantic task segment and determines whether the entry is a high uncertainty sample based on the high uncertainty threshold. The number of high uncertainty samples is accumulated and divided by the total number of samples in the segment to obtain the high uncertainty ratio of the semantic task segment, which serves as the second core quality indicator. The calculation formula is as follows: ,in, This indicates the second core quality indicator, namely the proportion of semantic task segments with excessively high uncertainty. This represents the uncertainty measure of the i-th objective. This represents the high uncertainty threshold. The S3013 central terminal reads the crop risk score for each learning sample entry in the semantic task segment, and counts the number of samples with risks exceeding the high risk threshold based on the high risk threshold. This number is then divided by the total number of samples to obtain the high risk ratio, which serves as the third core quality indicator. The calculation formula is as follows: ,in, This indicates the third core quality indicator, namely the proportion of crops with excessively high risk within this semantic segment. This represents the crop risk score for the i-th objective. This indicates the high-risk judgment threshold; S3014 The central end reads the indicator variable for whether the laser is triggered and the execution success indicator variable for each learning sample entry in the semantic task segment, counts the number of triggers within the segment, and counts the number of times the triggers are successful; S3015 When the number of triggers within the segment > 0, the central end calculates the execution success rate and obtains the execution failure rate, which is used as the fourth core quality indicator. The calculation formula is as follows: , ,in, This represents the fourth core quality metric, namely the execution failure rate after triggering. This indicates the success rate of execution after departure. Indicates whether the i-th step triggers the laser. This indicates whether the i-th laser execution was successful. This represents a very small constant, used to avoid a denominator of 0.

8. A multi-terminal, multi-model collaborative continuous learning method according to claim 6, characterized in that, After determining that the semantic operation segment is an abnormal segment and obtaining the corresponding evidence, the process of obtaining the terminal's network status parameters, classifying the evidence according to the network status parameters, and determining the corresponding backhaul strategy for each level of evidence includes: S3041 determining the network level based on network bandwidth and latency data, wherein the network levels include: low bandwidth high latency level, general level, and high bandwidth low latency level; S3042 if it is a low bandwidth high latency level, then only the time point corresponding to the time point with the largest contribution of the abnormal score will be retained. S3043 involves taking a keyframe, scaling it to a preset low resolution, and encoding and compressing it according to a preset compression ratio. The compressed keyframe and its corresponding segmentation result evidence are then transmitted back, along with the timestamp, semantic task segment identifier, and anomaly index summary corresponding to the keyframe. If the level is normal, the acquired image is extracted... S3044 involves taking a keyframe, scaling it to a preset medium resolution, encoding and compressing it according to a preset medium compression ratio, and transmitting back the processed keyframe and its corresponding segmentation result evidence, along with the timestamp, semantic task segment identifier, and anomaly index summary corresponding to the keyframe; S3044 If it is a high bandwidth low latency level, then extract the acquired image... The system retrieves keyframes, transmits the segmentation result evidence corresponding to the keyframes, and transmits the timestamp, semantic job segment identifier, and anomaly index summary corresponding to the keyframes.

Citation Information

Patent Citations

  • Centralized federal learning method and system based on multi-model collaborative fusion and application

    CN116562395A

  • System for secure federated learning

    US20200285980A1