Model updating method and device based on incremental data, equipment and storage medium
By combining comparative learning and knowledge distillation, an incremental data model update method is proposed to improve the model's perception and performance of incremental data based on the adaptability problem of cloud models and edge models when facing new categories of data.
Patent Information
- Application Number
- CN202510479061.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Cloud-based models and edge models are difficult to adapt quickly when facing new categories of data, resulting in the inability to effectively extract and utilize the characteristics of these data, affecting their perception and performance of incremental data.
A model update method based on incremental data is proposed. By obtaining the training sub-data set of each new category, comparative training of the category gap between the cloud model and the edge model, the target model and the corresponding category prototype are obtained. Send these prototypes to the edge and perform parameter adjustments to improve the adaptability and recognition capabilities of the model.
By combining comparative learning and knowledge distillation, the coordinated update of cloud models and edge models is achieved, and the model's perception of incremental data is improved.
Smart Images

Figure CN120215989A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model updating, and particularly to a model updating method, device, equipment and storage medium based on incremental data. Background Art
[0002] Cloud models are large machine learning or deep learning models that can run in a cloud computing environment. They can utilize the powerful computing capabilities and storage resources of the cloud platform to process and analyze large amounts of data. However, with the increase in Internet of Things devices, sending all data processing tasks to the cloud for processing will cause the cloud server to be overloaded and reduce the data processing efficiency.
[0003] In related technologies, key knowledge can be extracted from cloud models through technologies such as model compression, and an edge model can be deployed at the edge based on the key knowledge to process data, so as to improve the processing capacity and efficiency of the overall system. However, in the face of new categories of data, cloud models and edge models may be difficult to quickly adapt to these changes, making it impossible for the models to effectively extract and utilize the features of these data, thereby affecting the perception ability of cloud models and edge models for incremental data and reducing the performance of the models. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a model updating method, device, equipment and storage medium based on incremental data, which can improve the perception ability of the model for incremental data and the performance of the model.
[0005] To achieve the above object, a first aspect of the embodiments of this application proposes a model updating method based on incremental data, which is applied to a cloud server, and the method includes: Obtain a training sub-dataset corresponding to each new category; For each new category, according to a plurality of target training samples included in the corresponding training sub-dataset, compare and train a preset cloud model with a plurality of first category prototypes corresponding to a plurality of new categories to obtain a target cloud model and corresponding plurality of target first category prototypes; wherein, the first category prototype of each new category is obtained by the preset cloud model performing feature averaging according to the corresponding training sub-dataset; Send the plurality of target first category prototypes to the edge side, so that the edge side adjusts the parameters of the intermediate edge model based on the plurality of target first category prototypes to obtain a target edge model; wherein, the intermediate edge model is obtained by the edge side comparing and training a preset edge model with a plurality of second category prototypes corresponding to the plurality of target training samples and the plurality of new categories; the second category prototype of each new category is obtained by the preset edge model performing feature averaging according to the corresponding training sub-dataset.
[0006] Correspondingly, a second aspect of the embodiments of the present application proposes a model update device based on incremental data, which is applied to a cloud server. The device includes: An acquisition module, configured to acquire a training sub-dataset corresponding to each new category; A training module, configured to, for each new category, compare and train a preset cloud model for category gap according to a plurality of target training samples included in the corresponding training sub-dataset and a plurality of first category prototypes corresponding to the plurality of new categories, to obtain a target cloud model and corresponding plurality of target first category prototypes; wherein, the first category prototype of each new category is obtained by the preset cloud model performing feature averaging according to the corresponding training sub-dataset; A first sending module, configured to send the plurality of target first category prototypes to the edge side, so that the edge side adjusts parameters of an intermediate edge model based on the plurality of target first category prototypes to obtain a target edge model; wherein, the intermediate edge model is obtained by the edge side comparing and training a preset edge model for category gap based on the plurality of target training samples and a plurality of second category prototypes corresponding to the plurality of new categories; the second category prototype of each new category is obtained by the preset edge model performing feature averaging according to the corresponding training sub-dataset.
[0007] In some embodiments, the model update device based on incremental data further includes a second sending module, configured to: For each new category, acquire a plurality of classification fuzzy samples uploaded by the edge side, and compare and train a category gap according to the corresponding plurality of target training samples and the plurality of classification fuzzy samples, and a target first category prototype corresponding to each new category and a plurality of historical category prototypes obtained from any historical task, to obtain an updated cloud model and corresponding plurality of updated first category prototypes; Wherein, the plurality of classification fuzzy samples are determined by a plurality of uncertainty indices of each test sample in the plurality of new categories obtained by classifying and testing a plurality of test samples included in a test sub-dataset corresponding to each new category through the target edge model; Send the plurality of updated first category prototypes to the edge side, so that the edge side adjusts parameters of the target edge model based on the plurality of updated first category prototypes to obtain an updated edge model.
[0008] In some embodiments, the training module is further configured to: For each new category, perform feature averaging on a plurality of training samples included in the corresponding training sub-dataset by a preset cloud model to obtain a corresponding first category prototype; Obtain the first feature representation corresponding to each training sample through the preset cloud model, and determine the first distance between each first feature representation and the corresponding first category prototype; Obtain the second distance between each first feature representation and the other first category prototypes of other new categories; Determine the first category gap loss based on the difference between the first distance and the second distance; Adjust the parameters of the preset cloud model based on the first category gap loss to obtain a target cloud model and the corresponding multiple target first category prototypes.
[0009] To achieve the above object, a third aspect of the embodiments of the present application proposes a model update method based on incremental data, which is applied to the edge side. The method includes: Obtain the training sub-dataset corresponding to each new category; For each new category, according to the multiple target training samples included in the corresponding training sub-dataset, compare and train the preset edge model with the multiple second category prototypes corresponding to the multiple new categories to obtain an intermediate edge model and the corresponding multiple target second category prototypes; wherein, the second category prototype of each new category is obtained by the preset edge model through feature averaging according to the corresponding training sub-dataset; Receive the multiple target first category prototypes sent by the cloud server, and adjust the parameters of the intermediate edge model based on the multiple target first category prototypes to obtain a target edge model.
[0010] In some embodiments, the model update method based on incremental data applied to the edge side corresponds to a model update device based on incremental data applied to the edge side. The model update device based on incremental data includes a test module for: For each new category, through the target edge model, classify and test the multiple test samples included in the test sub-dataset corresponding to each new category to obtain the uncertainty index of each test sample under the multiple new categories; Determine multiple classification fuzzy samples according to the multiple uncertainty indexes corresponding to the multiple test samples, and send the multiple classification fuzzy samples to the cloud server, so that the cloud server compares and trains the target cloud model according to the multiple target training samples and the multiple classification fuzzy samples corresponding to each new category, and the target first category prototype and the multiple historical category prototypes obtained from any historical task to obtain multiple updated first category prototypes; Obtain the multiple updated first category prototypes sent by the cloud server, and adjust the parameters of the target edge model based on the multiple updated first category prototypes to obtain an updated edge model.
[0011] In some embodiments, the model updating device based on incremental data applied to the edge side further includes a calculation module, configured to: For each newly added category, through the preset edge model, calculate the Euclidean distance between each test sample included in the test sub-dataset corresponding to each newly added category and the multiple second category prototypes corresponding to the multiple newly added categories; For each test sample, calculate the logarithm value of the corresponding Euclidean distance, and based on the product of the Euclidean distance and the corresponding logarithm value, obtain the uncertainty measure of each test sample with respect to the corresponding second category prototype; Based on the multiple uncertainty measures corresponding to the multiple second category prototypes, obtain the uncertainty index of each test sample under the multiple newly added categories.
[0012] In some embodiments, the model updating device based on incremental data applied to the edge side further includes a determination module, configured to: For each newly added category, obtain the corresponding target second category prototype and target first category prototype, and determine the third distance between the target second category prototype and the target first category prototype; Obtain the other target first category prototypes of other newly added categories, and determine the fourth distance between the target second category prototype and the other target first category prototypes; Based on the difference between the third distance and the fourth distance, determine the distance distillation loss; Based on the distance distillation loss, adjust the parameters of the intermediate edge model to obtain the target edge model.
[0013] Correspondingly, a fourth aspect of the embodiments of the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method for updating a model based on incremental data according to any one of the first aspect embodiments or the third aspect embodiments of the present application.
[0014] Correspondingly, a fifth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the method for updating a model based on incremental data according to any one of the first aspect embodiments or the third aspect embodiments of the present application.
[0015] Embodiments of the present application obtain a training sub-dataset corresponding to each new category; for each new category, based on multiple target training samples included in the corresponding training sub-dataset, and multiple first-category prototypes corresponding to multiple new categories, a preset cloud model is subjected to contrast training of category differences to obtain a target cloud model and corresponding multiple target first-category prototypes; wherein, the first-category prototype of each new category is obtained by the preset cloud model through feature averaging according to the corresponding training sub-dataset; the multiple target first-category prototypes are sent to the edge side, so that the edge side adjusts the parameters of the intermediate edge model based on the multiple target first-category prototypes to obtain a target edge model; wherein, the intermediate edge model is obtained by the edge side through contrast training of category differences on a preset edge model based on multiple target training samples and multiple second-category prototypes corresponding to multiple new categories; the second-category prototype of each new category is obtained by the preset edge model through feature averaging according to the corresponding training sub-dataset. In this way, through the combination of contrast learning and knowledge distillation, the collaborative update of the cloud model and the edge model can be realized. Specifically, the present application fully trains the cloud model and the edge model respectively based on new samples. During the training process, the contrast learning technology is used to enhance the distinguishability of old and new categories in the feature space, and category prototypes are introduced to accurately represent the central features of new categories, providing a stable reference benchmark for model training. Furthermore, during the model training process, the parameter optimization mechanism guided by category prototypes calibrates the representation vectors of category prototypes in the parameter space, prompting the model weights to quickly converge to the optimization interval of the new category distribution, thereby effectively improving the adaptability and optimization efficiency of the model to new categories. Further, the present application can make full use of the computing power advantage of the cloud model, and transfer the knowledge of the cloud model to the edge model through distance distillation between category prototypes, so that the edge model can quickly and accurately identify new categories while maintaining the original recognition ability. In summary, the present application can improve the perception ability of the model to incremental data and improve the performance of the model. Description of the Drawings
[0016] Figure 1 FIG. is a schematic architecture diagram of a model update system based on incremental data provided by an embodiment of the present application; Figure 2 FIG. is a flowchart of a method for updating a model based on incremental data on a cloud server provided by an embodiment of the present application; Figure 3 FIG. is a flowchart of a method for updating a model based on incremental data on an edge side provided by an embodiment of the present application; Figure 4 FIG. is a general flowchart of a method for updating a model based on incremental data provided by an embodiment of the present application; Figure 5 FIG. is another general flowchart of a method for updating a model based on incremental data provided by an embodiment of the present application; Figure 6 It is a schematic diagram of the functional modules of the model update device based on incremental data provided by an embodiment of the present application; Figure 7 It is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0017] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0018] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0020] A cloud model is a large-scale machine learning or deep learning model that can run in a cloud computing environment. It can utilize the powerful computing power and storage resources of the cloud platform to process and analyze a large amount of data. However, with the increase in Internet of Things devices, sending all data processing tasks to the cloud for processing will cause the cloud server to be overloaded and reduce the data processing efficiency.
[0021] In the related art, key knowledge can be extracted from the cloud model through technologies such as model compression, and an edge model can be deployed at the edge based on the key knowledge to process data, so as to improve the processing ability and efficiency of the overall system. However, when facing new categories of data, the cloud model and the edge model may be difficult to quickly adapt to these changes, making it impossible for the model to effectively extract and utilize the features of these data, thereby affecting the perception ability of the cloud model and the edge model for incremental data and reducing the performance of the model.
[0022] Based on this, the embodiments of the present application provide a model update method, device, device and storage medium based on incremental data, which can improve the perception ability of the model for incremental data and the performance of the model.
[0023] The model update method, device, equipment, and storage medium based on incremental data provided by the embodiments of the present application will be specifically described through the following embodiments. First, the model update system based on incremental data in the embodiments of the present application will be described.
[0024] Please refer to Figure 1 , in some embodiments, the model update system includes a terminal 11, a cloud server 12, and an edge device 13.
[0025] Exemplarily, the terminal 11 can be an intelligent camera, an industrial sensor, a mobile device, etc. The terminal 11 can be used for data acquisition and annotation, user interaction and configuration, data preprocessing and uploading, etc.
[0026] Furthermore, the cloud server 12 can be a high-performance computing server, a distributed storage system, a network infrastructure, etc. The cloud server 12 can be used for large-scale data storage and management, training and updating of cloud models, and unified management of the device status of the edge device 13, etc.
[0027] Furthermore, the edge device 13 can be an embedded computing device, an industrial-grade edge server, etc. The edge device 13 can be used to deploy a lightweight small model inference engine and perform lightweight inference, local decision-making, etc.
[0028] Exemplarily, after receiving the training data uploaded by the terminal 11, the cloud server 12 and the edge device 13 can respectively perform model training and optimize the class prototypes through contrast learning techniques. Further, the cloud server 12 can perform further contrast training on the cloud model based on the uncertainty samples uploaded by the edge device 13, optimize the class prototypes, and transfer the knowledge of the cloud model to the edge model of the edge device 13 through knowledge distillation to achieve efficient data processing and model update. After the models of both the cloud server 12 and the edge device 13 are trained, the data collected by the terminal 11 can be classified in real time and accurately.
[0029] The model update method based on incremental data in the embodiments of the present application can be described through the following embodiments.
[0030] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0031] In the embodiments of the present application, a description will be made from the dimension of a model update device based on incremental data, and the model update device based on incremental data can be specifically integrated in a computer device. Refer to Figure 2 , Figure 2 FIG. is a flowchart of the steps of a model update method based on incremental data applied to a cloud server provided by an embodiment of the present application. In the embodiments of the present application, taking the model update device based on incremental data being specifically integrated in a computer device as an example, when the processor on the computer device executes the program instructions corresponding to the model update method based on incremental data, the specific process is as follows: Step 101, obtain a training sub-dataset corresponding to each new category.
[0032] In some embodiments, to ensure that the cloud model and the edge model can still execute efficiently and accurately in the face of changing application requirements, a training sub-dataset corresponding to each new category can be obtained to integrate the incremental data that the model needs to learn, so as to facilitate the model to learn.
[0033] Among them, the new category may be a category that did not appear in the historical training stage of the model, that is to say, the new category is a category that the model does not have the classification ability for, but there is a classification need in the future. For example, in an intelligent transportation system, the model was initially only trained to identify pedestrians and vehicles, but over time, it may be necessary to identify new categories such as bicycles and electric vehicles, and bicycles and electric vehicles belong to the new categories.
[0034] Among them, the training sub-dataset may be a sample set collected for each new category.
[0035] Specifically, in the current training task, it may include training tasks for multiple new categories, and each new category has corresponding training sub-data and test sub-data. The multiple training sub-data of the current training task form a training dataset, and the multiple test sub-data form a test dataset.
[0036] Exemplarily, for different scenarios, the newly added categories obtained are different. Taking the field of intelligent transportation as an example, the newly added categories can be electric vehicles, new traffic signs, and so on. The corresponding training sub-datasets can be collected by intelligent cameras or uploaded by users. Each training sub-dataset can contain multiple samples corresponding to the newly added category. For example, the newly added category of electric vehicles can contain 15 training samples of electric vehicles, and so on. It should be noted that this application is applicable to any scenario of data classification, such as intelligent security scenarios, industrial manufacturing scenarios, medical imaging scenarios, crop recognition scenarios in agriculture, etc. Therefore, specific scenarios and data categories are not limited.
[0037] By obtaining the training sub-datasets corresponding to each newly added category, it is convenient for the subsequent model to train the training sub-datasets to improve the generalization ability of the model.
[0038] Step 102, for each newly added category, according to the multiple target training samples included in the corresponding training sub-dataset, compare and train the preset cloud model with the multiple first-category prototypes corresponding to the multiple newly added categories to obtain the target cloud model and the corresponding multiple target first-category prototypes; among them, the first-category prototype of each newly added category is obtained by the preset cloud model through feature averaging according to the corresponding training sub-dataset.
[0039] In some embodiments, in order to enable the cloud model to more accurately distinguish between new and old categories, for each newly added category, the training sub-dataset of the newly added category can be used to perform contrast training on the preset cloud model, so that the training samples of the same category are more concentrated in the feature space and the training samples of different categories are more dispersed.
[0040] Among them, the target training sample can be a specific sample selected from the training sub-dataset of each newly added category.
[0041] Among them, the first-category prototype can be the feature average value calculated by the preset cloud model based on the corresponding training sub-dataset for each newly added category in the initial stage or the current state. The first-category prototype represents the central position of the newly added category in the feature space.
[0042] Among them, the preset cloud model can be a pre-trained large cloud model that already has the ability to classify existing categories. After receiving the training sub-dataset corresponding to the newly added category, the preset cloud model can adapt to the newly added category through contrast training.
[0043] Among them, the target first-class prototype can be the new-class prototype generated by the target cloud model after the contrast training is completed. The target first-class prototype can be the class center position recalculated by the cloud server for each new class based on the new target training samples during the process of adjusting the parameters of the preset cloud model, so as to more accurately represent the overall characteristics of the new class and improve the recognition ability of the cloud model for the new class.
[0044] In some embodiments, the first-class prototype of each new class is the geometric center obtained by extracting features of all target training samples of this class through the preset cloud model and performing feature averaging. Exemplarily, for the b-th training task containing multiple new classes, for any new class , the first-class prototype corresponding to the preset cloud model can be obtained by the following averaging method : ; Among them, represents the total number of sample of the target training samples belonging to the new class , ( ) is the target training sample in where the new class is
[0045] Furthermore, during the process of training the preset cloud model based on each new class, the calculated first-class prototype can be used to train the preset cloud model through contrast learning. The goal of contrast learning is to reduce the distance between each target training sample of the same new class and the corresponding first-class prototype, while increasing the distance between each target training sample and the first-class prototypes of different new classes. Specifically, for each target training sample of the new class p, calculate the similarity between its feature and the first-class prototype of the new class p (positive sample pair), and the similarity with other first-class prototypes (negative sample pair). By optimizing the contrast loss function, make the features of the same-class samples closer to the first-class prototype of the new class to which they belong, and farther away from the first-class prototypes of other new classes, so as to improve the discrimination ability of the cloud model for new classes. Exemplarily, according to the feature of any one target training sample calculate the following contrast loss : ; Among them, is the first-class prototype corresponding to the new class to which belongs (positive sample pair), The first-category prototypes (negative sample pairs) corresponding to different newly added categories.
[0046] Through the above method, not only the model's perception ability for newly added categories is enhanced, but also the discrimination ability of the model for each newly added category is improved, greatly improving the training efficiency.
[0047] In some embodiments, in order to enhance the model's perception ability for newly added categories, for each newly added category, the preset cloud model can be trained by comparison, and the model parameters can be gradually optimized to improve the overall robustness and generalization ability of the model, so that the model can remain efficient and accurate when facing continuously changing newly added data. For example, step 102 may include: (102.1) For each newly added category, the preset cloud model averages the features according to the multiple training samples included in the corresponding training sub-dataset to obtain the corresponding first-category prototype; (102.2) Obtain the first feature representation corresponding to each training sample through the preset cloud model, and determine the first distance between each first feature representation and the corresponding first-category prototype; (102.3) Obtain the second distance between each first feature representation and the other first-category prototypes of other newly added categories; (102.4) Determine the first-category gap loss based on the difference between the first distance and the second distance; (102.5) Adjust the parameters of the preset cloud model based on the first-category gap loss to obtain the target cloud model and the corresponding multiple target first-category prototypes.
[0048] Among them, the first feature representation may be a vector representation obtained by the preset cloud model after feature extraction of each training sample.
[0049] Among them, the first distance may be the distance between the first feature representation of each training sample and the first-category prototype of the newly added category to which it belongs. The first distance can be measured using the Euclidean distance or cosine similarity.
[0050] Among them, the other first-category prototypes may be, for each newly added category, except for the first-category prototype of its own category, the first-category prototypes of all other newly added categories. The other first-category prototypes represent the central positions of different newly added categories in the feature space.
[0051] Among them, the second distance may be the distance between the first feature representation of each training sample and the other first-category prototypes of other newly added categories.
[0052] Among them, the first-category gap loss may be a loss value calculated based on the difference between the first distance and the second distance.
[0053] Further, during the process of training the preset cloud model based on each newly added category, the calculated first-category prototype can be used to train the preset cloud model through contrastive learning. The goal of contrastive learning is to reduce the distance between each target training sample of the same newly added category and the corresponding first-category prototype, while increasing the distance between each target training sample and the first-category prototypes of different newly added categories. Specifically, for each target training sample of the newly added category p, calculate the similarity between its features and the first-category prototype of the newly added category p (positive sample pair), as well as the similarity with the first-category prototypes of other newly added categories (negative sample pair). By optimizing the contrastive loss function, the features of samples of the same category are made closer to the first-category prototype of the newly added category to which they belong, while being far from the first-category prototypes of other newly added categories, thereby enhancing the discrimination ability of the cloud model for newly added categories.
[0054] Exemplarily, according to any one target training sample 's first feature representation calculate the following first-category gap loss : ; where is the first-category prototype corresponding to the newly added category to which it belongs (positive sample pair), is the first-category prototype corresponding to a different newly added category from a hyperparameter set according to the actual situation, such as taking 0.8.
[0055] In some embodiments, the first distance and the second distance can be calculated by Euclidean distance or cosine distance. Taking the calculation by Euclidean distance as an example, the calculation formula of the first distance is as follows: ; where represents the first feature representation of the training sample and represents the first-category prototype corresponding to the newly added category to which it belongs.
[0056] Taking the vehicle recognition scenario as an example, if the newly added categories of the system are ambulance, fire truck, and truck, these three newly added categories can be used as a training task, and the preset cloud model is trained for each newly added category in turn. Taking the ambulance as an example, if there are 100 ambulance pictures in the training sub-dataset corresponding to the ambulance, these 100 pictures are input into the preset cloud model to extract features, and the 100 first feature representations obtained by extraction are averaged to obtain the first-category prototype corresponding to the ambulance. Similarly, the first-category prototype corresponding to the fire truck can be obtained The first category prototype corresponding to the truck .
[0057] Furthermore, for each ambulance picture, such as the first feature representation of the first ambulance picture , the first distance between it and the first category prototype corresponding to the ambulance can be calculated, and calculate respectively with the second distance between them, and the second distance between it and them, and based on the first category gap loss between the first distance and the second distances respectively with the fire truck and the truck, adjust the parameters of the preset cloud model to obtain the target cloud model and the corresponding multiple target first category prototypes.
[0058] By constructing the first category gap loss to update the model parameters, when the sum of the first distance and the hyperparameter is greater than the second distance, the first category gap loss can be made positive, triggering gradient update, forcing the preset cloud model to adjust the parameters, recalculate the first category prototypes of all new categories, making the same-class training samples gather towards the first category prototypes (during this period, the first category prototypes also belong to one of the parameters to be adjusted), and at the same time making the same-class training samples far away from the different-class first category prototypes.
[0059] In the above way, the intra-class compactness of the same new category can be improved, so that the sample features of the same new category are highly concentrated, the distances between the prototypes of different categories are enlarged, and the efficiency and accuracy of the cloud model for identifying new types are effectively improved.
[0060] Step 103: Send the multiple target first category prototypes to the edge side, so that the edge side adjusts the parameters of the intermediate edge model based on the multiple target first category prototypes to obtain the target edge model; wherein, the intermediate edge model is obtained by the edge side comparing and training the preset edge model for the category gap based on the multiple target training samples and the multiple second category prototypes corresponding to the multiple new categories; the second category prototype of each new category is obtained by the preset edge model averaging the features according to the corresponding training sub-dataset.
[0061] In some embodiments, in order to ensure that the edge model can obtain the latest category information in a timely manner and improve its ability to identify new categories, the knowledge of the target cloud model can be distilled into the edge model, so that the edge model can quickly adapt to new categories and maintain a high recognition accuracy even when the computing power is limited and the training effect is limited.
[0062] Among them, the edge device can be a device close to the data source or application scenario, such as a smart camera, a smartphone, etc. The edge device deploys a lightweight edge model, which has limited computing resources but needs to process data in real time and make classification decisions.
[0063] Among them, the intermediate edge model can be a small intermediate-state model running on the edge device, which has been adjusted through a round of contrast training for new categories but has not been fully optimized.
[0064] Among them, the target edge model can be the final optimized model obtained by sending multiple target first-category prototypes updated from the cloud large model to the edge device and further adjusting the parameters of the intermediate edge model. The target edge model has higher accuracy and robustness than the intermediate edge model and can better handle the recognition tasks of new and old categories.
[0065] Among them, the second-category prototype can be the average value of features calculated by a preset edge model based on the training sub-dataset of each new category at the edge device. The second-category prototype is the initial center point used by the edge device model when learning new categories.
[0066] In some embodiments, the second-category prototype of each new category is the geometric center obtained by extracting features of all target training samples of this category through a preset edge model and then averaging the features. Exemplarily, for the b-th training task including multiple new categories, for any new category , the second-category prototype corresponding to the preset edge model can be obtained by the following averaging method : ; Among them, represents the total number of target training samples belonging to the new category , ( ) is the target training sample in with the new category being
[0067] Further, during the training of the preset edge model based on each newly added category, the calculated second-category prototype can be used to train the preset edge model through contrastive learning. The goal of contrastive learning is to reduce the distance between each target training sample of the same newly added category and the corresponding second-category prototype, while increasing the distance between each target training sample and the second-category prototypes of different newly added categories. Specifically, for each target training sample of the newly added category p, calculate the similarity between its features and the second-category prototype of the newly added category p (positive sample pair), and the similarity with the second-category prototypes of other newly added categories (negative sample pair). By optimizing the contrastive loss function, the features of the same-class samples are made closer to the second-category prototype of the newly added category to which they belong, while being far from the second-category prototypes of other newly added categories, thereby improving the discrimination ability of the edge model for the newly added categories.
[0068] Exemplarily, according to any one target training sample the second feature representation calculate the following second-category gap loss : ; wherein, is the second-category prototype corresponding to the newly added category to which it belongs (positive sample pair), is the second-category prototype corresponding to a different newly added category from (negative sample pair), is a hyperparameter set according to the actual situation, such as taking 0.8.
[0069] In some embodiments, the distance between the second feature representation and the second-category prototype corresponding to the newly added category to which it belongs, and the distance between the second-category prototypes corresponding to other newly added categories can both be calculated by the Euclidean distance or the cosine distance. The specific calculation process will not be elaborated here.
[0070] By constructing the second-category gap loss to update the model parameters of the preset edge model, the same-class training samples can be aggregated towards the second-category prototype (during this period, the second-category prototype also belongs to one of the parameters to be adjusted), while making the same-class training samples far from the second-category prototypes of different classes, effectively improving the efficiency and accuracy of the edge model in recognizing newly added types.
[0071] Further, multiple target first-category prototypes can be distilled to the edge side, and the distance between the target first-category prototype and the target second-category prototype of the intermediate edge model can be reduced through backpropagation, so that the target first-category prototype and the target second-category prototype are aligned, making the distilled target edge model have classification performance equivalent to that of the target cloud model. In some embodiments, the distance distillation loss calculated during the distillation process It can be calculated in the following way: ; Wherein, represents the total number of new categories included in the current training task in the middle, is the target first category prototype corresponding to the new category of this training, is the target second category prototype corresponding to the new category of this training, is the target first category prototype of other new categories of the new category of this training, is a hyperparameter set according to the actual situation.
[0072] The distance distillation loss calculated by the above formula can encourage the target second category prototype of the intermediate edge model to be close to the target first category prototype of the same new category (such as a car) in the target cloud model and far from the target first category prototype of different new categories (such as a bicycle, a tractor) in the target cloud model. After the edge side receives multiple target first category prototypes corresponding to multiple new categories, it can use the multiple target first category prototypes to adjust the parameters of the intermediate edge model. Specifically, the intermediate edge model adjusts its parameters by minimizing the distance distillation loss function to continuously optimize its classification performance.
[0073] By sending the target first category prototype trained on the cloud server side to the edge side, the edge side can effectively adjust the parameters of the intermediate edge model based on these category prototypes, thereby obtaining an optimized target edge model. This process not only realizes the efficient transfer of the knowledge of the cloud model, but also significantly improves the recognition accuracy and robustness of the edge model for new categories in a resource-constrained environment, and further enhances the real-time response ability and deployment flexibility of the entire system.
[0074] In some embodiments, in order to further improve the classification accuracy of the cloud model and the edge model, samples that are difficult to classify can be screened out by the edge model to more quickly identify samples that need further analysis, and the cloud model can be retrained based on the fuzzy samples and then distilled again to effectively improve the classification accuracy of the cloud model and the edge model and the sensitivity to data. Exemplarily, after step 103, it may further include: (A.1) For each new category, obtain multiple classification fuzzy samples uploaded by the edge side, and compare and train the category gaps between the corresponding multiple target training samples and multiple classification fuzzy samples with the target first category prototype corresponding to each new category and multiple historical category prototypes obtained from any historical task to obtain an updated cloud model and corresponding multiple updated first category prototypes; Among them, multiple classification ambiguous samples are used by the target edge model to classify and test multiple test samples included in the test sub-dataset corresponding to each newly added category, and multiple uncertainty indices of each test sample under multiple newly added categories are determined. (A.2)Send multiple updated first-category prototypes to the edge side, so that the edge side adjusts the parameters of the target edge model based on the multiple updated first-category prototypes to obtain an updated edge model.
[0075] Among them, the classification ambiguous samples can be test samples that are identified as difficult to clearly classify after the target edge model classifies and tests the test sub-dataset of the newly added category. The classification ambiguous samples have relatively high uncertainty indices under multiple newly added categories, have a large overlap with the features of multiple newly added categories, and are difficult to distinguish.
[0076] Among them, the historical task can be a training task that has been completed. The historical task can include at least one historical newly added category that has been trained before retraining the preset cloud model and the preset edge model. For example, during the training phase corresponding to the historical task, the historical newly added categories of trucks and tractors have been trained, and in the current training task, the newly added categories of cars and bicycles need to be trained.
[0077] Among them, the historical category prototype can be the category prototype corresponding to the historical newly added category included in the historical task, and the calculation method can refer to the calculation process of the first-category prototype and the second-category prototype above.
[0078] Among them, the updated first-category prototype can be the new category prototype generated by the cloud large model after contrast training, and the updated first-category prototype is used to better represent the central category features of each newly added category.
[0079] Among them, the test sub-dataset can be a data set collected from the actual application scenario and is used to evaluate the performance of the model. Each newly added category corresponds to a test sub-dataset.
[0080] Among them, the test sample can be a specific sample instance in the test sub-dataset and is used to classify and test the model.
[0081] Among them, the uncertainty index can be used to characterize the classification confidence or uncertainty degree of each test sample under multiple newly added categories. A relatively high uncertainty index means that the test sample is difficult to be clearly classified into a specific category by the target edge model.
[0082] Among them, the updated edge model can be the target edge model after parameter adjustment.
[0083] In some embodiments, in order to quickly screen out fuzzy samples, multiple classification fuzzy samples can also be used to classify and test multiple test samples included in the test sub-dataset corresponding to each new category through a preset edge model or an intermediate edge model, and multiple uncertainty indices of each test sample under multiple new categories are determined.
[0084] In some embodiments, in order to effectively avoid the catastrophic forgetting problem during the training process, an updated dataset can be formed based on multiple target training samples and multiple classification fuzzy samples of each new category, and compared and trained with cross-task prototypes (i.e., prototypes stored in the prototype set after the historical task is trained). Exemplarily, for each sample data in the updated dataset corresponding to each new category, after inputting the sample data into the target cloud model, a corresponding feature representation can be obtained. , through contrastive training, the distance between the target first category prototype corresponding to the same new category can be reduced, and the distance between the multiple historical category prototypes stored in the historical task can be increased. Specifically, the category gap loss corresponding to this contrastive training is calculated as follows ( is a preset hyperparameter): ; ; Further, through the above category gap loss for backpropagation, the parameters of the target cloud model are updated until the model converges or reaches the preset number of training rounds, and the final updated cloud model can be obtained. In this way, new categories can be recognized more effectively while maintaining the recognition ability for existing categories.
[0085] Further, since the target edge model cannot accurately classify classification fuzzy samples, the updated first category prototype obtained from the updated cloud model can be distilled into the target edge model, so that the target edge model adjusts its parameters based on the updated first category prototype to achieve rapid alignment of category prototypes. Specifically, the distillation loss of this process is as follows: ; where represents the total number of categories of new samples in the current training task, represents the target third category prototype corresponding to the target edge model, represents the updated first category prototype corresponding to each new category in the current training (for example, the current training is for the new category of truck), Represents the updated first prototype of other new categories except for the newly added categories currently being trained (such as newly added categories like cars and bicycles).
[0086] The distilled loss calculated by the above formula can encourage the updated third-category prototype of the edge model to be close to the updated first-category prototype of the same new category (such as cars) in the updated cloud model, and be far from the updated first-category prototype of different new categories (such as bicycles and tractors) in the updated cloud model. After the edge side receives multiple updated first-category prototypes corresponding to multiple new categories, it can use these multiple updated first-category prototypes to adjust the parameters of the target edge model. Specifically, the edge model adjusts its parameters by minimizing the distance distilled loss function to continuously optimize its classification performance.
[0087] In some embodiments, after obtaining the updated cloud model and the updated edge model, the updated first-category prototype corresponding to the updated cloud model and the updated third-category prototype corresponding to the updated edge model can be stored in their respective prototype sets, so as to facilitate subsequent classification of the model and training of other new categories.
[0088] In some embodiments, the process of determining classification fuzzy samples is based on the results of the edge model classifying the samples in the new category test sub-dataset. The purpose of this process is to identify samples that the model is less certain about during classification, and these samples may belong to new categories or may be old category samples that are difficult to distinguish.
[0089] First, the target edge model (or the preset edge model, or the target edge model) can use its current knowledge to classify each test sample in the test sub-dataset. During the classification process, for some test samples, the target edge model may give predictions for multiple categories, and the confidence of each prediction is not high. Therefore, the model's classification of these test samples is uncertain. To quantify this uncertainty, the information entropy (i.e., the uncertainty index) of the confidence distribution of each test sample over all possible new categories can be calculated. The higher the information entropy, the more uncertain the target edge model is about the classification of the test sample. Therefore, all test samples can be sorted in descending order according to their uncertainty indices, and the top K% (the value of K can be determined according to the actual situation, such as 10%, 20%, etc.) of the test samples with the highest scores after sorting are selected as classification fuzzy samples, and these classification fuzzy samples are used for subsequent training to improve the overall performance and adaptability of the model.
[0090] By screening fuzzy samples of new categories and retraining the target cloud model, and distilling the trained updated first-category prototypes to the target edge model to obtain an updated edge model, the recognition ability of the model for new-category data can be significantly improved while maintaining the recognition performance for old-category data. This approach enables the updated edge model to operate efficiently on resource-constrained edge devices while leveraging the powerful computing capabilities of the cloud server to update and optimize the model.
[0091] In some embodiments, after training the current new category (such as a bicycle) to obtain a trained cloud model and edge model, the next new category (such as a car) of the current training task can be trained, and then the parameters of the cloud model and edge model can be adjusted until all new categories in the current training task are trained. In this way, the model can be efficiently trained within a small scope of the same task, enabling the model to have the ability to efficiently distinguish new categories, and also preventing the model from forgetting old knowledge and maintaining good performance.
[0092] In some embodiments, after training the cloud model and edge model, the prototypes corresponding to multiple new categories can be added to the prototype sets of the corresponding models, and the performance of the models can be tested. Taking the performance test of the cloud model as an example, for any test data X, the trained cloud model can be used to predict the category (category i = 1, 2,... ) to which the data X belongs, and the probability can be calculated as follows: ; where , is the feature representation extracted by the cloud model, is the total number of categories learned so far in the current training task, and represents the Euclidean distance function between two given vectors. The prediction result of the cloud model for X is the category i that maximizes .
[0093] The embodiments of the present application obtain a training sub-dataset corresponding to each new category; for each new category, based on a plurality of target training samples included in the corresponding training sub-dataset and a plurality of first category prototypes corresponding to the plurality of new categories, the preset cloud model is subjected to contrastive training of category gaps to obtain a target cloud model and the corresponding plurality of target first category prototypes; wherein, the first category prototype of each new category is obtained by the preset cloud model through feature averaging according to the corresponding training sub-dataset; the plurality of target first category prototypes are sent to the edge side, so that the edge side adjusts the parameters of the intermediate edge model based on the plurality of target first category prototypes to obtain a target edge model; wherein, the intermediate edge model is obtained by the edge side through contrastive training of category gaps on the preset edge model based on a plurality of target training samples and a plurality of second category prototypes corresponding to the plurality of new categories; the second category prototype of each new category is obtained by the preset edge model through feature averaging according to the corresponding training sub-dataset. In this way, through the combination of contrastive learning and knowledge distillation, the collaborative update of the cloud model and the edge model can be realized. Specifically, the present application fully trains the cloud model and the edge model based on the new samples. During the training process, the contrastive learning technology is used to enhance the distinguishability between the old and new categories in the feature space, and the category prototype is introduced to accurately represent the central features of the new category, providing a stable reference benchmark for model training. Furthermore, during the model training process, the parameter optimization mechanism guided by the category prototype calibrates the representation vector of the category prototype in the parameter space, prompting the model weights to quickly converge to the optimization interval of the new category distribution, thereby effectively improving the adaptability and optimization efficiency of the model to the new category. Further, the present application can make full use of the computing power advantage of the cloud model, and transfer the knowledge of the cloud model to the edge model through the distance distillation between the category prototypes, so that the edge model can quickly and accurately identify the new category while maintaining the original recognition ability. In summary, the present application can improve the perception ability of the model to incremental data and improve the performance of the model.
[0094] In some embodiments, to ensure that the edge model can dynamically adapt to the changing application requirements and still maintain a high recognition accuracy when facing new categories. The parameter adjustment can be performed by receiving a plurality of target first category prototypes sent by the cloud server, so as to improve the model performance without significantly increasing the burden on the edge side. Therefore, in a second aspect, the present application proposes a model update method based on incremental data, which is applied to the edge side. The method includes: Step 201, obtain a training sub-dataset corresponding to each new category; Step 202: For each newly added category, based on multiple target training samples included in the corresponding training sub-dataset, compare and train the preset edge model with multiple second-category prototypes corresponding to the multiple newly added categories to obtain an intermediate edge model and corresponding multiple target second-category prototypes; wherein, the second-category prototype of each newly added category is obtained by the preset edge model through feature averaging based on the corresponding training sub-dataset. Step 203: Receive multiple target first-category prototypes sent by the cloud server, and adjust the parameters of the intermediate edge model based on the multiple target first-category prototypes to obtain a target edge model.
[0095] Among them, the target second-category prototype can be the category center point obtained by the intermediate edge model through feature averaging based on the training sub-dataset of each newly added category. Specifically, during the process of comparing and training the preset edge model at the edge side, the original category prototypes will be continuously adjusted. After the model training is completed to obtain the intermediate edge model, the second target prototype can be obtained. And after receiving the multiple target first-category prototypes sent by the cloud server, parameter adjustment can be performed to obtain the target edge model and the target third-category prototype.
[0096] In some embodiments, the embodiment in which the edge side performs feature averaging based on the training sub-dataset corresponding to each newly added category to obtain the first-category prototype corresponding to the newly added category has been introduced above. Specifically, reference can be made to the above embodiments for obtaining the first-category prototype and the second-category prototype, and the embodiments of the present application will not elaborate on this.
[0097] In some embodiments, the process of comparing and training the preset edge model based on multiple target training samples and multiple second-category prototypes corresponding to multiple newly added categories to obtain the intermediate edge model has been introduced in detail above. Moreover, the above also introduces the process in which the cloud server compares and trains the preset cloud model based on multiple target training samples and multiple first-category prototypes corresponding to multiple newly added categories to obtain the target cloud model. Specifically, reference can be made to the above implementation process, and the embodiments of the present application will not elaborate on the process of training the preset edge model to obtain the intermediate edge model.
[0098] In some embodiments, the process of receiving multiple target first-category prototypes sent by the cloud server and adjusting the parameters of the intermediate edge model accordingly to obtain the target edge model has been introduced in detail above and will not be elaborated here.
[0099] By combining contrastive learning and knowledge distillation, the collaborative update of the cloud model and the edge model can be achieved. Specifically, in this application, the cloud model and the edge model are fully trained based on newly added samples. During the training process, the contrastive learning technique is used to enhance the distinguishability of new and old categories in the feature space, and the category prototype is introduced to accurately represent the central features of the newly added categories, providing a stable reference benchmark for model training. Furthermore, during the model training process, the parameter optimization mechanism guided by the category prototype calibrates the representation vector of the category prototype in the parameter space, prompting the model weights to quickly converge to the optimization interval of the newly added category distribution, thereby effectively improving the adaptability and optimization efficiency of the model to the newly added categories. Additionally, the computing power advantage of the cloud model can be fully utilized, and the knowledge of the cloud model is transferred to the edge model through the distance distillation between category prototypes, enabling the edge model to quickly and accurately identify newly added categories while maintaining its original recognition ability. In summary, this application can improve the model's perception ability of incremental data and enhance the model's performance.
[0100] In some embodiments, to ensure that the cloud model and the edge model have similar performance for newly added categories in the feature space, and at the same time ensure that samples of the same category are clustered and samples of different categories are dispersed, the knowledge of the target cloud model can be distilled into the intermediate edge model to enhance the robustness and generalization ability of the edge model, enabling it to better handle complex and changing application scenarios. For example, step 203 may include: (203.1) For each newly added category, obtain the corresponding target second category prototype and target first category prototype, and determine the third distance between the target second category prototype and the target first category prototype; (203.2) Obtain the other target first category prototypes of other newly added categories, and determine the fourth distance between the target second category prototype and the other target first category prototypes; (203.3) Based on the difference between the third distance and the fourth distance, determine the distance distillation loss; (203.4) Based on the distance distillation loss, adjust the parameters of the intermediate edge model to obtain the target edge model.
[0101] Among them, the third distance may be the distance between the target second category prototype of each newly added category and the target first category prototype corresponding to the same newly added category.
[0102] Among them, the other target first category prototypes may be, for each newly added category, except for the target first category prototype of the current newly added category, the target first category prototypes of all other newly added categories.
[0103] Among them, the fourth distance can be the distance between the target second-class prototype of each newly added category and the other target first-class prototypes of other newly added categories. The fourth distance can be used to distinguish different newly added categories, ensure the clustering of samples of the same class and the dispersion of samples of different classes, thereby improving the discrimination ability of the edge model.
[0104] Among them, the distance distillation loss can be a loss value calculated based on the difference between the third distance and the fourth distance. By minimizing the distance distillation loss, the target second-class prototype can be made as close as possible to the target first-class prototype while maintaining a sufficient distance from other target first-class prototypes.
[0105] In some embodiments, multiple target first-class prototypes can be received at the edge side, and the distance between the target first-class prototype and the target second-class prototype of the intermediate edge model can be reduced through backpropagation, so that the target first-class prototype and the target second-class prototype are aligned, and the distilled target edge model has classification performance comparable to that of the target cloud model. In some embodiments, the distance distillation loss calculated during the distillation process can be calculated in the following way: ; Among them, represents the total number of newly added categories included in the current training task in, is the target first-class prototype corresponding to the newly added category in this training, is the target second-class prototype corresponding to the newly added category in this training, is the target first-class prototype of other newly added categories of the newly added category in this training, is a hyperparameter set according to the actual situation.
[0106] The distance distillation loss calculated by the above formula can encourage the target second-class prototype of the intermediate edge model to be close to the target first-class prototype of the same newly added category (such as a car) in the target cloud model and far from the target first-class prototypes of different newly added categories (such as a bicycle, a tractor) in the target cloud model. After receiving multiple target first-class prototypes corresponding to multiple newly added categories at the edge side, the intermediate edge model can be parameter-adjusted using the multiple target first-class prototypes. Specifically, the intermediate edge model adjusts its parameters by minimizing the distance distillation loss function to continuously optimize its classification performance.
[0107] Receiving the target first-class prototypes trained by the cloud server at the edge can enable the edge to effectively adjust the parameters of the intermediate edge model based on these class prototypes, thereby obtaining an optimized target edge model. This process not only realizes the efficient transfer of cloud model knowledge but also significantly improves the recognition accuracy and robustness of the edge model for new classes in resource-constrained environments, thereby enhancing the real-time response ability and deployment flexibility of the entire system.
[0108] In some embodiments, to further improve the classification accuracy of the cloud model and the edge model, samples that are difficult to classify can be screened out by the edge model to more quickly identify samples that require further analysis, and the cloud model can be retrained based on the fuzzy samples and then distilled again to effectively improve the classification accuracy of the cloud model and the edge model and the sensitivity to data. Exemplarily, after step 203, the following may further be included: (B.1) For each new class, through the target edge model, classify and test multiple test samples included in the test sub-dataset corresponding to each new class to obtain the uncertainty index of each test sample under multiple new classes; (B.2) Determine multiple classification fuzzy samples according to the multiple uncertainty indices corresponding to the multiple test samples, and send the multiple classification fuzzy samples to the cloud server, so that the cloud server can compare and train the target cloud model for the class gap according to the multiple target training samples and the multiple classification fuzzy samples corresponding to each new class, and the multiple target first-class prototypes corresponding to each new class and the multiple historical class prototypes obtained from any historical task, to obtain multiple updated first-class prototypes; (B.3) Obtain the multiple updated first-class prototypes sent by the cloud server, and adjust the parameters of the target edge model based on the multiple updated first-class prototypes to obtain an updated edge model.
[0109] Among them, the uncertainty index can be an index used to measure the classification uncertainty degree of each test sample under multiple new classes, and is used to reflect the overall classification uncertainty of the edge model for a test sample under multiple new classes.
[0110] Among them, the updated edge model can be the target edge model after parameter adjustment.
[0111] In some embodiments, to quickly screen out classification fuzzy samples, the multiple classification fuzzy samples can also be determined by performing classification tests on multiple test samples included in the test sub-dataset corresponding to each new class through a preset edge model or an intermediate edge model, and obtaining the multiple uncertainty indices of each test sample under multiple new classes.
[0112] In some embodiments, to effectively avoid the catastrophic forgetting problem during training, on the cloud server side, an updated dataset can be formed based on multiple target training samples and multiple classification ambiguous samples of each newly added category, and compared and trained with cross-task prototypes (i.e., prototypes stored in the prototype set after the historical tasks are trained). Exemplarily, for each sample data in the updated dataset corresponding to each newly added category, after inputting the sample data into the target cloud model, the corresponding feature representation can be obtained , through comparative training, the distance from the target first-category prototype corresponding to the same newly added category can be reduced , and the distance from the multiple historical-category prototypes stored for historical tasks can be increased . Specifically, the formula for the category gap loss corresponding to this comparative training is as follows ( is a preset hyperparameter): ; Furthermore, the cloud server can perform backpropagation through the above category gap loss to update the parameters of the target cloud model until the model converges or reaches the preset number of training epochs, and then the final updated cloud model can be obtained. In this way, it is possible to more effectively identify newly added categories while maintaining the recognition ability for existing categories
[0113] Furthermore, since the target edge model cannot accurately classify classification ambiguous samples, the updated first-category prototype obtained from the updated cloud model can be distilled into the target edge model, so that the target edge model adjusts its parameters based on the updated first-category prototype to achieve rapid alignment of category prototypes. Specifically, the formula for the distillation loss in this process is as follows: ; wherein, represents the total number of categories of newly added samples in the current training task , represents the target third-category prototype corresponding to the target edge model, represents the updated first-category prototype corresponding to each newly added category in the current training (such as the newly added category of truck in the current training), represents the updated first prototypes of other newly added categories except the newly added category in the current training (such as newly added categories of car, bicycle, etc.)
[0114] The distilled loss calculated by the above formula can encourage the updated third-class prototypes of the edge model to approach the updated first-class prototypes of the same new category (such as cars) in the updated cloud model, and be far from the updated first-class prototypes of different new categories (such as bicycles and tractors) in the updated cloud model. After receiving multiple updated first-class prototypes corresponding to multiple new categories at the edge, the edge can use the multiple updated first-class prototypes to adjust the parameters of the target edge model. Specifically, the edge model adjusts its parameters by minimizing the distance distilled loss function to continuously optimize its classification performance.
[0115] In some embodiments, to improve the accuracy of the model in classifying new categories, by evaluating in detail the classification uncertainty of the edge model for each test sample, key samples that the model has difficulty classifying accurately can be found, enabling the entire system to continuously improve and optimize, ensuring that the model can still adjust and enhance its classification ability in a timely manner when facing new categories. For example, (B.1) may include: (B.1.1) For each new category, through the target edge model, calculate the Euclidean distance between each test sample included in the test sub-dataset corresponding to each new category and multiple second-class prototypes corresponding to multiple new categories; (B.1.2) For each test sample, calculate the logarithm of the corresponding Euclidean distance, and based on the product of the Euclidean distance and the corresponding logarithm value, obtain the uncertainty measure of each test sample for the corresponding second-class prototype; (B.1.3) Based on the multiple uncertainty measures corresponding to multiple second-class prototypes, obtain the uncertainty index of each test sample for multiple new categories.
[0116] Among them, the Euclidean distance can be the distance between each test sample and multiple second-class prototypes corresponding to multiple new categories.
[0117] Among them, the logarithm of the Euclidean distance can be the natural logarithm or the common logarithm of the calculated Euclidean distance.
[0118] Among them, the uncertainty measure can be an index calculated based on the Euclidean distance between each test sample and multiple second-class prototypes and its logarithm value, used to characterize the uncertainty degree of the target edge model in classifying this test sample.
[0119] In some embodiments, for any test sample X in the test set, the corresponding feature representation can be obtained through the target edge model , calculate the Euclidean distance between and the second-class prototypes corresponding to each new category , where Indicates the number of second-category prototypes corresponding to all newly added categories. Further, the uncertainty measure of each test sample with respect to the corresponding second-category prototype is .
[0120] Further, based on the multiple uncertainty measures corresponding to multiple second-category prototypes, the uncertainty index of each test sample under multiple newly added categories can be obtained : ; wherein, represents the Euclidean distance between each test sample and the corresponding second-category prototype.
[0121] Further, for each test sample in the test sub-dataset, the corresponding uncertainty index can be calculated through the above formula, and all the test samples can be sorted in descending order according to the corresponding uncertainty indices, and a preset proportion of test samples are selected as classification fuzzy samples. The preset proportion can be 15%, 20%, etc., and can be specifically set according to the actual situation. For example, if 100 test samples are sorted in descending order according to the uncertainty index, 20% of the test samples can be selected as classification fuzzy samples, that is, the first 20 test samples are selected as classification fuzzy samples.
[0122] By calculating the uncertainty index of each test sample to screen classification fuzzy samples, the uncertainty degree of the model for classifying each test sample can be effectively quantified, so as to ensure that the samples that are difficult for the model to classify are sent to the cloud for further analysis and learning, so as to facilitate subsequent optimization of the adaptability of the cloud model and the edge model to newly added categories, and improve the accuracy and robustness of the model in practical applications.
[0123] Please refer to Figure 4 and Figure 5 , in some embodiments, Figure 4 and Figure 5 are schematic diagrams of two general embodiments of this application. Now, in combination with Figure 4 and Figure 5, an overall embodiment of the present application is introduced. Exemplarily, for each newly added category in the current training stage, the model is mainly trained in two stages. Specifically, in the first stage, for each newly added category in the current training task, the corresponding training sub-dataset can be used to train the preset cloud model and the preset edge model respectively. On the cloud server side, the preset cloud model is optimized through contrastive learning, the first category prototype (center point of the feature space) of each newly added category is calculated, and the target cloud model and the optimized target first category prototype are generated; on the edge side, the preset edge model is trained by contrastive learning based on the same training sub-dataset to generate an intermediate edge model and the corresponding second category prototype, so as to endow the model with the ability to recognize newly added categories. Since the computing power of the cloud server side is relatively strong and the training of the model is relatively sufficient, after the target cloud model is trained on the cloud server side, the intermediate edge model can be distilled through the target first category prototype, and the parameters of the intermediate edge model can be adjusted through the distance distillation loss, so that the second category prototype at the edge is aligned with the target first category prototype at the cloud, and at the same time, the distance from other categories is increased to obtain the target edge model, thereby realizing the efficient transfer of high-order knowledge from the cloud to the edge model.
[0124] Furthermore, the target edge model can classify the test samples in the test sub-dataset, calculate the uncertainty index based on the Euclidean distance between the test samples and the second category prototype, and select the samples with the lowest classification confidence as classification fuzzy samples and upload them to the cloud server for retraining. Specifically, on the cloud server side, a new data set can be formed based on the uploaded classification fuzzy samples and the original training samples to perform cross-task contrastive training on the target cloud model. During the contrastive training process, by optimizing the category distance loss, the features of the newly added categories are decoupled from the features of the historical categories, an updated cloud model and an optimized updated first category prototype are generated, and the updated first category prototype of the updated cloud model is transferred to the target edge model through knowledge distillation, and the final updated edge model is obtained through parameter adjustment, thereby further improving the ability of the model to recognize new categories.
[0125] Furthermore, the updated cloud model and edge model can also be tested to evaluate the recognition accuracy of newly added categories and the retention ability of old categories to ensure that the model performance meets the requirements.
[0126] It can be understood that the training process for each newly added category is an iterative loop. Until all newly added categories of the current training task are trained, the model can possess the ability to accurately identify new and old categories. If there are still untrained newly added categories subsequently, the above two-stage process is looped until all newly added categories are trained. In this way, the model can continuously learn and adapt to new categories in a changing environment while maintaining the ability to identify old categories. This collaborative update technology of the model can be widely applied in fields such as intelligent transportation and intelligent security.
[0127] Please refer to Figure 6 , the embodiment of the present application also provides a model update device based on incremental data, which can implement the above-mentioned model update method based on incremental data. The model update device based on incremental data is applied to the cloud server side. The model update device based on incremental data includes: An acquisition module 61, configured to acquire a training sub-dataset corresponding to each newly added category; A training module 62, configured to, for each newly added category, compare and train a preset cloud model for category gap according to a plurality of target training samples included in the corresponding training sub-dataset and a plurality of first category prototypes corresponding to a plurality of newly added categories, to obtain a target cloud model and corresponding plurality of target first category prototypes; wherein, the first category prototype of each newly added category is obtained by the preset cloud model through feature averaging according to the corresponding training sub-dataset; A first sending module 63, configured to send the plurality of target first category prototypes to the edge side, so that the edge side adjusts the parameters of the intermediate edge model based on the plurality of target first category prototypes to obtain a target edge model; wherein, the intermediate edge model is obtained by the edge side comparing and training a preset edge model for category gap based on the plurality of target training samples and a plurality of second category prototypes corresponding to a plurality of newly added categories; the second category prototype of each newly added category is obtained by the preset edge model through feature averaging according to the corresponding training sub-dataset.
[0128] The specific implementation manner of the model update device based on incremental data is basically the same as the specific embodiment of the above-mentioned model update method based on incremental data, and will not be elaborated here. On the premise of meeting the requirements of the embodiment of the present application, other functional modules can also be set for the model update device based on incremental data to implement the model update method based on incremental data in the above embodiment.
[0129] The embodiment of the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned model update method based on incremental data. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0130] Please refer toFigure 7 , Figure 7 schematically shows the hardware structure of a computer device according to another embodiment. The computer device includes: a processor 71, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application; a memory 72, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 72 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 72 and are called by the processor 71 to execute the model update method based on incremental data according to the embodiments of the present application; an input / output interface 73, which is used to implement information input and output; a communication interface 74, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.); a bus 75, which transmits information between various components of the device (such as the processor 71, the memory 72, the input / output interface 73, and the communication interface 74); wherein the processor 71, the memory 72, the input / output interface 73, and the communication interface 74 are communicatively connected to each other inside the device through the bus 75.
[0131] The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned model update method based on incremental data is implemented.
[0132] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0133] The embodiments described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0134] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0135] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0136] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0137] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0138] It should be understood that in this application, "at least one (item)" and "several" mean one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0139] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0140] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0142] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0143] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall fall within the scope of the rights of the embodiments of this application.
Claims
1. A model updating method based on incremental data, characterized in that: Applied to a cloud server, the method includes: Get the training sub-dataset corresponding to each newly added category; For each newly added category, a preset cloud model is trained for category gap comparison based on a plurality of target training samples contained in a corresponding training sub-dataset and a plurality of first category prototypes corresponding to the plurality of newly added categories, so as to obtain a target cloud model and a plurality of corresponding target first category prototypes; wherein the first category prototype of each newly added category is obtained by averaging the features of the preset cloud model based on the corresponding training sub-dataset; The multiple target first category prototypes are sent to the edge end, so that the edge end adjusts the parameters of the intermediate edge model based on the multiple target first category prototypes to obtain the target edge model; wherein the intermediate edge model is obtained by the edge end performing comparative training on the category gap of the preset edge model based on the multiple target training samples and the multiple second category prototypes corresponding to the multiple newly added categories; the second category prototype of each newly added category is obtained by the preset edge model according to the corresponding training sub-dataset by averaging the features.
2. The model updating method based on incremental data according to claim 1, characterized in that: The method further includes sending the plurality of target first category prototypes corresponding to the target cloud model to the edge end so that the edge end adjusts the parameters of the intermediate edge model based on the plurality of target first category prototypes to obtain the target edge model: For each newly added category, multiple classification fuzzy samples uploaded by the edge are obtained, and based on the corresponding multiple target training samples and the multiple classification fuzzy samples, comparative training of category gaps is performed with the target first category prototype corresponding to each newly added category and multiple historical category prototypes obtained from any historical task to obtain an updated cloud model and the corresponding multiple updated first category prototypes; The multiple classified fuzzy samples are classified and tested by the target edge model on multiple test samples contained in the test sub-dataset corresponding to each newly added category, and multiple uncertainty indexes of each test sample under the multiple newly added categories are determined; The multiple updated first category prototypes are sent to the edge end, so that the edge end adjusts the parameters of the target edge model based on the multiple updated first category prototypes to obtain an updated edge model.
3. The model updating method based on incremental data according to claim 1, characterized in that: For each newly added category, a preset cloud model is trained for category gap comparison based on a plurality of target training samples included in the corresponding training sub-dataset and a plurality of first category prototypes corresponding to the plurality of newly added categories, to obtain a target cloud model and a plurality of corresponding target first category prototypes, including: For each newly added category, the preset cloud model averages the features based on multiple training samples included in the corresponding training sub-dataset to obtain the corresponding first category prototype; Obtaining a first feature representation corresponding to each training sample through the preset cloud model, and determining a first distance between each first feature representation and the corresponding first category prototype; Obtain a second distance between each first feature representation and other first category prototypes of other newly added categories; determining a first category gap loss based on a difference between the first distance and the second distance; Based on the first category gap loss, parameters of the preset cloud model are adjusted to obtain a target cloud model and a corresponding plurality of target first category prototypes.
4. A model updating method based on incremental data, characterized in that: Applied to the edge end, the method includes: Get the training sub-dataset corresponding to each newly added category; For each newly added category, a preset edge model is trained for category gap comparison based on a plurality of target training samples contained in a corresponding training sub-dataset and a plurality of second category prototypes corresponding to the plurality of newly added categories, to obtain an intermediate edge model and a plurality of corresponding target second category prototypes; wherein the second category prototype of each newly added category is obtained by averaging the features of the preset edge model based on the corresponding training sub-dataset; A plurality of target first category prototypes sent by the cloud service end are received, and parameters of the intermediate edge model are adjusted based on the plurality of target first category prototypes to obtain a target edge model.
5. The model updating method based on incremental data according to claim 4, characterized in that: The receiving of multiple target first category prototypes sent by the cloud service end, and adjusting parameters of the intermediate edge model based on the multiple target first category prototypes to obtain the target edge model further includes: For each of the newly added categories, a classification test is performed on a plurality of test samples included in a test sub-dataset corresponding to each of the newly added categories by using the target edge model to obtain an uncertainty index of each test sample under the plurality of newly added categories; Determine a plurality of classification fuzzy samples according to a plurality of uncertainty indexes corresponding to the plurality of test samples, and send the plurality of classification fuzzy samples to a cloud server, so that the cloud server performs comparative training on the target cloud model based on the plurality of target training samples corresponding to each newly added category and the plurality of classification fuzzy samples, the target first category prototype corresponding to each newly added category, and the plurality of historical category prototypes obtained from any historical task, to obtain a plurality of updated first category prototypes; A plurality of updated first category prototypes sent by the cloud service end are obtained, and parameters of the target edge model are adjusted based on the plurality of updated first category prototypes to obtain an updated edge model.
6. The model updating method based on incremental data according to claim 5, characterized in that: For each of the newly added categories, a classification test is performed on a plurality of test samples included in a test sub-dataset corresponding to each of the newly added categories by using the target edge model to obtain an uncertainty index of each test sample under the plurality of newly added categories, including: For each of the newly added categories, calculating, by using the target edge model, the Euclidean distance between each test sample included in the test sub-dataset corresponding to each of the newly added categories and a plurality of second category prototypes corresponding to the plurality of newly added categories; For each of the test samples, calculate the logarithmic value of each corresponding Euclidean distance, and obtain the uncertainty measure of each of the test samples in the corresponding second category prototype based on the product between the Euclidean distance and the corresponding logarithmic value; Based on the multiple uncertainty measures corresponding to the multiple second category prototypes, an uncertainty index of each test sample under the multiple newly added categories is obtained.
7. The model updating method based on incremental data according to claim 4, characterized in that: The receiving of multiple target first category prototypes sent by the cloud service end, and adjusting parameters of the intermediate edge model based on the multiple target first category prototypes to obtain the target edge model, includes: For each newly added category, obtain the corresponding target second category prototype and the target first category prototype, and determine a third distance between the target second category prototype and the target first category prototype; Obtaining other target first category prototypes of other newly added categories, and determining a fourth distance between the target second category prototype and the other target first category prototype; determining a distance distillation loss based on a difference between the third distance and the fourth distance; Parameters of the intermediate edge model are adjusted based on the distance distillation loss to obtain a target edge model.
8. A model updating device based on incremental data, characterized in that: Applied to a cloud server, the device comprises: An acquisition module is used to obtain a training sub-dataset corresponding to each newly added category; A training module is used to perform category gap comparison training on a preset cloud model for each newly added category based on a plurality of target training samples included in a corresponding training sub-dataset and a plurality of first category prototypes corresponding to the plurality of newly added categories, so as to obtain a target cloud model and a plurality of corresponding target first category prototypes; wherein the first category prototype of each newly added category is obtained by averaging the features of the preset cloud model based on the corresponding training sub-dataset; The first sending module is used to send the multiple target first category prototypes to the edge end, so that the edge end adjusts the parameters of the intermediate edge model based on the multiple target first category prototypes to obtain the target edge model; wherein the intermediate edge model is obtained by the edge end performing comparative training on the category gap of the preset edge model based on the multiple target training samples and the multiple second category prototypes corresponding to the multiple newly added categories; the second category prototype of each newly added category is obtained by the preset edge model according to the corresponding training sub-dataset by feature averaging.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the incremental data-based model updating method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the model updating method based on incremental data described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Incremental training method and device of classification model and computer equipment
CN115438755A
Incremental learning method and device based on cloud edge collaborative architecture, equipment and medium
CN116128036A
Comparative incremental learning-based model training method and malicious traffic classification method and system
CN116244645A
Model training method and device, image classification method and device and electronic equipment
CN117218399A
Edge cloud collaborative learning method and system based on cloud model decomposition
CN119398138A