Model training method and device and storage medium

By performing unsupervised pre-training of the training model in the edge computing scenario, using the training samples without labels and the characteristics of larger models, the problem of limited computing power of edge nodes is solved, and the performance and generalization ability of the model are improved.

CN120020821APending Publication Date: 2025-05-20HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311550937.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

Smart Images

  • Figure CN120020821A_ABST
    Figure CN120020821A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device and a storage medium. In the process of training the model, on one hand, the model can depend on a training sample which does not carry a label; and on the other hand, a model which has been pre-trained and has a larger scale can be introduced as a reference, training samples which do not carry labels are simultaneously input into the model to be trained and the introduced model, and output features and intermediate layer features are respectively extracted from the two models. And the unsupervised pre-training can be implemented by taking an output feature and an intermediate layer feature of the introduced model as optimization targets. Therefore, pre-training of the model according to an unsupervised mode can be supported, the generalization ability of the model can be remarkably improved in a scene with scarce labels, and a larger-scale model is introduced as a reference, so that the to-be-trained model can obtain baseline performance similar to that of the reference model, the box opening accuracy of the model is improved, and the user experience is improved. And the model performance can be effectively optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a model training method, device, and storage medium. Background Art

[0002] In scenarios such as urban governance, considering requirements such as network bandwidth and computing real-time performance, computing tasks are usually deployed on edge nodes. That is, edge computing technology is relied on to complete the computing tasks in the scenario.

[0003] However, due to the limited computing power of edge nodes, it is difficult to support various large models represented by Transformer, and it is necessary to introduce lighter small models on edge nodes to execute computing tasks.

[0004] Therefore, how to optimize the performance of these lightweight or medium-weight models has become an urgent problem to be solved. Summary of the Invention

[0005] Multiple aspects of this application provide a model training method, device, and storage medium to optimize model performance.

[0006] An embodiment of this application provides a model training method, including:

[0007] In response to a pre-training instruction for a first model, obtain a training sample set, where the training samples in the training sample set do not carry labels;

[0008] Input the training samples without labels into the first model and a second model respectively. The second model uses a model that has completed pre-training, and the scale of the second model is larger than the scale of the first model;

[0009] Extract output features and intermediate layer features from the first model and the second model respectively;

[0010] Use the output features and intermediate layer features of the second model as the optimization target for the first model, so as to perform unsupervised pre-training on the first model using the training samples without labels.

[0011] An embodiment of this application also provides a computing device, including a memory, a processor, and a communication component;

[0012] The memory is used to store one or more computer instructions;

[0013] The processor is coupled to the memory and the communication component, and is used to execute the one or more computer instructions to execute the foregoing model training method.

[0014] The embodiments of the present application further provide a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to execute the foregoing model training method.

[0015] In the embodiments of the present application, during the process of training a model, on the one hand, it is possible to completely rely on training samples without labels. In this way, it is possible to support unsupervised pre-training of the model with the help of a large amount of unlabeled data. On the other hand, it is also possible to introduce a pre-trained model that is larger in scale than the model to be trained as a reference, input the training samples without labels into both the model to be trained and the introduced larger-scale model at the same time, and extract the output features and intermediate layer features from the two models respectively. Based on this, the output features and intermediate layer features of the larger-scale model can be used as the optimization target to implement the foregoing unsupervised pre-training. In this way, it is possible to support pre-training the model in an unsupervised manner, which can effectively improve the knowledge richness of the model and significantly improve the generalization ability of the model in scenarios where labels are scarce. Moreover, by introducing a larger-scale model as a reference, the model to be trained can obtain a baseline performance similar to that of the reference model, improve the out-of-the-box accuracy of the model, and thus effectively optimize the model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0017] Figure 1 is a schematic flowchart of a model training method provided by an exemplary embodiment of the present application;

[0018] Figure 2 is a schematic logical diagram of a model training method provided by an exemplary embodiment of the present application;

[0019] Figure 3 is a schematic flowchart of another model training method provided by an exemplary embodiment of the present application;

[0020] Figure 4 is a schematic diagram of an application scenario provided by an exemplary embodiment of the present application;

[0021] Figure 5 is a schematic structural diagram of a computing device provided by another exemplary embodiment of the present application. DETAILED DESCRIPTION

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Apparently, the described embodiments are only a part of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0023] Before elaborating on the technical solutions provided in the embodiments of this application, several technical concepts involved in this application are briefly explained as follows.

[0024] Unsupervised learning, a machine learning / deep learning training method relying on unlabeled data.

[0025] Pre-training, a technique of pre-training a model using data far larger than that required for a specific scenario.

[0026] Unsupervised pre-training: In unsupervised pre-training, the model is pre-trained on a large scale of unlabeled data.

[0027] In the process of research, the inventors found that in some scenarios with the problem of scarce labels represented by government affairs scenarios, it is usually necessary to rely on large models represented by Transformer to carry out computing tasks. However, for edge computing, the computing power of edge nodes is limited, and the parameter scale and computational complexity of these large models are very high, far exceeding the carrying capacity of edge nodes. Therefore, the computing requirements in scenarios dominated by edge computing and with scarce labels cannot be met. For this reason, it is proposed in the field to introduce lighter small models on edge nodes to execute computing tasks.

[0028] Currently, the field usually adopts the method of supervised training to train these introduced small models. However, as mentioned above, in scenarios with scarce labels, it is difficult to obtain sufficient labeled data to support supervised training, which leads to overfitting problems and insufficient generalization ability of the introduced small models.

[0029] For this reason, in this embodiment, a model training method is proposed, which provides an improved training scheme for the small models introduced in the aforementioned scenarios dominated by edge computing and with scarce labels, so as to optimize the feature representation ability of the small models and ensure that the small models after training have excellent enough performance.

[0030] The following will elaborate on the technical solutions provided in the embodiments of this application in conjunction with the drawings.

[0031] Figure 1 It is a schematic flowchart of a model training method provided for an exemplary embodiment of this application. Figure 2A logic diagram of a model training method provided by an exemplary embodiment of the present application. Refer to Figure 1 , this method can be executed by a data processing device, which can be implemented as software, hardware, or a combination of software and hardware, and the data processing device can be integrated in a computing device. Refer to Figure 1 , this method may include:

[0032] Step 100, in response to a pre-training instruction for the first model, obtain a training sample set, and the training samples in the training sample set do not carry labels;

[0033] Step 101, input the training samples without labels into the first model and the second model respectively. The second model is a model that has completed pre-training, and the scale of the second model is larger than that of the first model;

[0034] Step 102, extract output features and intermediate layer features from the first model and the second model respectively;

[0035] Step 103, use the output features and intermediate layer features of the second model as the optimization target of the first model, so as to perform unsupervised pre-training on the first model by using the training samples without labels.

[0036] In this embodiment, the first model is a model to be trained. It can be understood that the first model can be a lightweight or medium-sized model that can support deployment on edge nodes. The scale of the first model is not specifically limited in this embodiment. In addition, in this embodiment, the model structure of the first model is not specifically limited either. Theoretically, a model including a backbone network and an output layer can be used as the first model in this embodiment.

[0037] Refer to Figure 1 , in step 100, in response to a pre-training instruction for the first model, start the pre-training process of the first model. It should be understood that at this time, the first model in this embodiment can be a model that has not been trained at all, and the initial parameters in the first model can take random values or other preset values. On this basis, the pre-training process of the first model can be started in this embodiment.

[0038] Refer to Figure 2, in this embodiment, a training sample set can be provided for the first model, and the training samples in the training sample set do not carry labels. Here, not carrying labels means that no labels are required to be used during the pre-training process in this embodiment, but the training samples are regarded as unlabeled data. In practical applications, some of the training samples in the training sample set may already be labeled, and in this embodiment, such training samples are added to the training sample set during operation, but the labels of such training samples are not used. In this way, the training sample set in this embodiment can contain a large number of training samples without being limited to labeled samples. Moreover, the training samples in the training sample set in this embodiment do not have to be limited to sample data related to a single downstream task, but can be extended to more sample data without being limited by the task, which enables the training sample set in this embodiment to contain a large number of training samples. Ideally, for any downstream task, relevant training samples can be adapted in the training sample set, which can effectively ensure the richness of knowledge provided by the training samples.

[0039] Continue to refer to Figure 2 , in this embodiment, a second model is also introduced. Among them, the second model can adopt a model that has completed pre-training and has a larger scale than the first model. Here, the scale can refer to the parameter scale or the layer scale, etc.

[0040] In this embodiment, the pre-training process of the second model is not specifically limited and will not be elaborated. In this embodiment, a large model can be used as the second model. Here, the large model can refer to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. The large model can also be called a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one billion parameters is produced. This kind of model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc. It should be noted that in actual applications, the pre-trained model can be fine-tuned with a small number of samples so that the large model can be applied to different tasks. For example, large models can be widely used in fields such as natural language processing (NLP), computer vision, and speech processing. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), and image generation, and can also be widely used in tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. Preferably, in this embodiment, a large model adapted to the basic function of the first model can be used. That is, preferably, the basic function of the large model adopted is similar or consistent with the basic function required by the first model.

[0041] Based on this, in this embodiment, the model training method provided in this embodiment can be executed separately for different basic functions, and models with different basic functions can be constructed. Among them, the basic functions in this embodiment can be general functions such as visual detection and natural language understanding. In this way, the first model generated according to the model training method provided in this embodiment can be used as a general model with a certain basic function. Furthermore, the first model can be used as a general model for downstream tasks. Before carrying the downstream tasks, the pre-trained first model can be processed such as model fine-tuning so that the model function of the first model focuses on the corresponding downstream tasks. The subsequent model fine-tuning and other links will be described in detail later.

[0042] It should be understood that the second model introduced in this embodiment already has excellent baseline performance.

[0043] On this basis, referring to Figure 1 and Figure 2, in step 101, the training samples without tags can be input into the first model and the second model respectively. That is to say, from the perspective of a certain training sample, this training sample can be input into the first model and the second model respectively. In this way, this training sample can forward propagate along the model structure in the first model and the second model respectively.

[0044] In this embodiment, in step 102, the output features and the intermediate layer features can be extracted from the first model and the second model respectively. Among them, the output features are the features extracted from the output result of the model. The intermediate layer features are the features extracted from the intermediate layer of the model. Still from the perspective of a certain training sample, after this training sample is input into the first model and the second model respectively, during the process of its forward propagation along the model structure of the first model, the intermediate layer features generated by the forward propagation can be extracted from the backbone of the first model; after the forward propagation is completed, the features of the specified type can be extracted from the output result of the first model as the output features. Similarly, during the process of this training sample forward propagating along the model structure of the second model, the intermediate layer features generated by the forward propagation can be extracted from the backbone of the second model; after the forward propagation is completed, the features of the specified type can be extracted from the output result of the second model as the output features.

[0045] In practical applications, the intermediate layer features can be extracted from the backbones of the first model and the second model. Among them, the backbone can contain one or more intermediate layers, and the backbone can also be divided into one or more stages, and each stage can contain one or more intermediate layers. Based on this, the features can be extracted from one or more stages in the backbone as the intermediate layer features. In this way, the intermediate layer features can contain the features extracted from one or more stages in the backbone. If it is set that the intermediate layer features only contain the features extracted from one stage in the backbone, it is preferably to extract the features from the last stage of the backbone, that is, the features output by the backbone are used as the intermediate layer features. And if it is set that the intermediate layer features contain the features extracted from multiple stages in the backbone, in addition to specifying the last stage of the aforementioned backbone, other stages in the backbone can also be specified as needed, and the features are extracted from these specified stages respectively to form the intermediate layer features.

[0046] During the research process, the inventor found that there may be differences in the feature space scales between the first model and the second model. Therefore, in this embodiment, it is proposed that if the feature space scales of the first model and the second model are different, the intermediate layer features generated from the first model are mapped to the feature space with the same scale as that of the second model as the intermediate layer features extracted from the first model. Among them, in this embodiment, a fully convolutional operation can be performed on the intermediate layer features generated from the first model, and by setting parameters such as the convolution kernel specifications in the fully convolutional operation, the intermediate layer features generated from the first model are mapped to the feature space with the same scale as that of the second model. In this way, it can be ensured that the feature scales of the intermediate layer features extracted from the first template are the same as those of the intermediate layer features extracted from the second model, so as to improve the training effect in the subsequent step 103.

[0047] In this way, for any training sample used, the output features and intermediate layer features corresponding to the training sample can be extracted from the first model and the second model respectively. Of course, since the model parameters of the first model are still being optimized, there are obvious deficiencies in the accuracy and representation ability of the output features and intermediate layer features extracted from the first model. However, the output features and intermediate layer features extracted from the second model are already relatively excellent features.

[0048] Reference Figure 1 and Figure 2 , in this embodiment, it is proposed in step 103 that the output features and intermediate layer features of the second model are used as the optimization objectives of the first model to perform unsupervised pre-training on the first model using unlabeled training samples.

[0049] It should be emphasized here that in this embodiment, the unsupervised pre-training of the first model is completely dependent on unlabeled training samples. That is, during the pre-training of the first model, no labeled tags are referred to. Instead, it is proposed that in terms of output, the output features of the first model are used as the optimization objective, so that the output features of the second model can be used as a constraint to narrow the difference in output between the first model and the second model to optimize the output performance of the first model.

[0050] For the intermediate layer of the first model, it is proposed that the intermediate layer features of the second model are used as the optimization objective, so that the intermediate layer features of the second model can be used as a constraint to narrow the difference in feature representation between the first model and the second model in the intermediate layer to optimize the feature representation ability of the first model in the intermediate layer.

[0051] In this way, by using the output features and intermediate layer features extracted from the second model as the optimization objective of the first model, during the unsupervised pre-training of the first model, the baseline capabilities of the second model can be continuously passed on to the first model. Thus, after the pre-training is completed, the first model not only learns rich and generalized knowledge from a large number of unlabeled training samples, but also inherits similar baseline performance from the second model.

[0052] In summary, in the embodiments of the present application, during the process of training the model, on the one hand, it is completely dependent on the unlabeled training samples. In this way, the unsupervised pre-training of the model can be supported by means of a large amount of unlabeled data. On the other hand, a larger-scale model that has completed pre-training can be introduced as a reference. The unlabeled training samples are input into the model to be trained and the introduced larger-scale model at the same time, and the output features and intermediate layer features are respectively extracted from the two models. Based on this, the output features and intermediate layer features of the larger-scale model are used as the optimization objective to implement the aforementioned unsupervised pre-training. In this way, the pre-training of the model can be supported in an unsupervised manner, which can effectively improve the knowledge richness of the model and significantly improve the generalization ability of the model in scenarios where labels are scarce. Moreover, by introducing a larger-scale model as a reference, the model to be trained can obtain baseline performance similar to that of the larger-scale model, improving the out-of-the-box accuracy of the model, and thus effectively optimizing the model performance.

[0053] As mentioned above, the first model after pre-training in this embodiment will be a general model with basic functions. In this embodiment, it is proposed that the pre-trained first model can be used to process downstream tasks. The types of downstream tasks are not limited in this embodiment. Several exemplary downstream tasks can be object detection tasks, image classification tasks, or semantic segmentation tasks, etc., and no more examples are given here. Among them, the labeled data related to the downstream task requirements in the scenario may be scarce. In view of this situation, this embodiment further proposes:

[0054] After the unsupervised pre-training of the first model is completed, in response to a tuning instruction for the first model, a sample set related to the downstream task requirements is obtained, and the samples in the sample set carry labels;

[0055] The pre-trained first model is supervised-trained based on these labeled samples so that the first model adapts to the downstream task requirements.

[0056] Since the pre-trained first model in this embodiment has excellent generalization performance and out-of-the-box accuracy, even if a small-scale labeled dataset is used for supervised training of the first model, it can very efficiently and accurately learn the knowledge corresponding to the downstream task in these labeled data. This is mainly because during the pre-training process of the first model, the first model has already learned the basic knowledge related to various downstream tasks. On this basis, with only a small amount of labeled data, the first model can quickly focus on the downstream task requirements. That is to say, the pre-trained first model in this embodiment also has excellent few-shot learning ability and will no longer suffer from overfitting problems. No matter which downstream task needs to be switched to, the pre-trained first model can be quickly adapted to the downstream task requirements through few-shot learning.

[0057] Furthermore, this embodiment also proposes:

[0058] After completing the unsupervised pre-training of the first model, in response to a transfer learning instruction for the first model, transfer the backbone network in the pre-trained first model to other models.

[0059] In practical applications, in this embodiment, the backbone network in the pre-trained first model can be solidified, so that the performance of the backbone network can be solidified. On this basis, the backbone network can be disassembled from the pre-trained first model and transferred to other models as the backbone network in other models. This enables other models to obtain the baseline capabilities of the backbone network in the first model without training. Through the transfer learning scheme provided in this embodiment, more high-performance models can be derived based on the first model, thus supporting more types of downstream tasks.

[0060] It should be noted that in this embodiment, the structural parts of the first model that can be transferred may include other structural parts such as the feature pyramid and output head in addition to the aforementioned backbone network, which is not limited herein.

[0061] Figure 3 It is a schematic flowchart of another model training method provided by an exemplary embodiment of the present application. Refer to Figure 3 , this method may include:

[0062] Step 300, in response to a pre-training instruction for the first model, obtain a training sample set, and the training samples in the training sample set do not carry labels;

[0063] Step 301, input the training samples without labels into the first model and the second model respectively. The second model uses a model that has completed pre-training, and the scale of the second model is larger than that of the first model;

[0064] Step 302: Extract the output features and intermediate layer features from the first model and the second model respectively;

[0065] Step 303: Evaluate the output feature loss and intermediate layer feature loss of the first model relative to the second model to construct a loss function for the first model;

[0066] Step 304: Based on the loss function, optimize the parameters of the first model to perform unsupervised pre-training on the first model using the training samples without labels.

[0067] Among them, Steps 300 - 302 can refer to the relevant descriptions in the foregoing embodiments and will not be repeated here. In this embodiment, an exemplary implementation manner of using the output features and intermediate layer features of the second model as the optimization target of the first model is provided based on Steps 303 and 304.

[0068] Reference Figure 3 , in Step 303, the output feature loss and intermediate layer feature loss of the first model relative to the second model can be evaluated to construct a loss function for the first model. The loss function in Step 303 can be characterized as:

[0069] L total = L emb + L match

[0070] Among them, L match represents the output feature loss of the first model relative to the second model. L emb characterizes the intermediate layer feature loss of the first model relative to the second model.

[0071] It can be understood that from this loss function, it can be more clearly perceived that in the process of pre-training the first model in this embodiment, it completely depends on the training samples without labels and does not need to rely on any labeled tags. The output features and intermediate layer features extracted from the second model are used as the optimization target of the first model.

[0072] In this embodiment, in Step 302, after the training samples without labels pass through the backbone network of the first model, the intermediate layer features can be extracted. After the same training samples pass through the backbone network or the encoder network of the second model, the intermediate layer features can also be extracted. Based on this, the foregoing intermediate layer feature loss can be characterized as:

[0073] L emb = λ e ·||Conv(F img1 ) - F img3 ||

[0074] Among them, λe Denote the intermediate layer feature loss weight. Conv() can represent mapping the intermediate layer features to the feature space of a specified scale, F img represents the intermediate layer features.

[0075] As mentioned above, in this embodiment, for different basic functions, the model training method provided in this embodiment can be respectively executed to train models with different basic functions. For models with different basic functions, the information items included in the output results may not be exactly the same. Therefore, in this embodiment, it is proposed that according to the basic function corresponding to the first model, the extraction rule for the output features in step 302 can be preset, and at least the type of the information items to be extracted from the output results is specified in the extraction rule. Based on this, in step 302 of this embodiment, the output features for the same training sample can be respectively extracted from the first model and the second model according to the extraction rule for the output features set for the first model. Correspondingly, in step 303, a loss function for the first model can be constructed based on the output features extracted in step 302.

[0076] Taking the visual detection model as an example, the output feature extraction process in step 302 will be exemplarily described below.

[0077] In this embodiment, if the first model is a visual detection model, then after any unlabeled training sample is respectively input into the first model and the second model, the object box vector and the foreground score can be respectively extracted from the output results of the first model and the second model; based on the object box vector and the foreground score, the output features are constructed.

[0078] The inventors found during the research process that the basic functions of visual detection models are usually foreground object detection, classifying the recognized objects, etc. The model structure of its model usually includes structural parts such as a backbone network, a feature pyramid, and a detection head. Among them, during the process of the image passing through the backbone network and the feature pyramid, high-dimensional features representing different scales can be generated, and these high-dimensional features can be extracted as the aforementioned intermediate layer features in this embodiment. The detection head can generate object boxes, and the typical output results after the detection head generates object boxes can include: foreground scores and object box vectors, etc. Among them, the foreground score can be used to represent the confidence that there is a foreground object in the object box. The object box vector is used to represent the position information of the object box, and the object box vector can be expressed as [x, y, w, h], where x is the abscissa of the center point of the object box, y is the ordinate of the center point of the object box, w is the width of the object box, and h is the height of the object box.

[0079] Based on this, in this embodiment, the object box vector and the foreground score can be extracted from the output result of the visual detection model to characterize the output features corresponding to the first model.

[0080] It should be understood that the foreground score and the object box vector mentioned above can also be provided in the output result of the second model. Therefore, similarly, in step 302 above, the foreground score and the object box vector can also be extracted from the output result of the second model to characterize the output features corresponding to the second model.

[0081] Continuing with the output features extracted in step 302, in step 303: If the first model is a visual detection model, then pair the object boxes in the output results of the second model and the first model for the same training sample to determine the object box pairs that meet the preset pairing requirements; evaluate the differences in the object box vectors and foreground scores corresponding to the two object boxes in the object box pair to characterize the output feature loss.

[0082] In this way, the output feature loss can be characterized as:

[0083] L match = λ cls ·L cls (S obj1 , S obj3 ) + λ box ·II obj ·L box (B obj1 , B obj3 )

[0084] Among them, λ cls represents the foreground loss weight; λ box represents the object box loss weight; L cls () represents the foreground score difference; L box () represents the object box vector difference; S obj represents the foreground score; B obj represents the object box vector; II obj represents a logical operation. For the object box identified as the foreground, this parameter is set to 1, otherwise it is set to 0.

[0085] The inventor found during the research process that there are usually more than one object box output by the visual detection model. For example, there may be multiple foreground objects in a training sample. In this case, after inputting the training sample into the first model, multiple object boxes can be generated, and the generated object boxes are described in the output result through the foreground score and the object box vector mentioned above. Based on this, after inputting the training sample into the second model, multiple object boxes will also be generated. After the same training sample is input into the first model and the second model respectively, due to the performance differences between the two models, the number and positions of the generated object boxes may be different.

[0086] To this end, in this embodiment, the foregoing object box pairing step is proposed. Each determined single object box pair contains two object boxes, one of which comes from the first model and the other comes from the second model. After the object box pairing step, the two object boxes in the same object box pair can be made to correspond to and point to the same object as much as possible, so that the object boxes generated by the first model can be more accurately constrained.

[0087] This embodiment provides an exemplary object box pairing solution: the Hungarian bipartite matching technique is used to pair the object boxes in the output results of the second model and the first model for the same training sample, so as to determine the object box pairs that meet the preset pairing requirements. The Hungarian bipartite matching technique will not be elaborated here. In this exemplary solution, the object boxes output by the first model and the second model for the same training sample are used as the matching objects, and by applying the Hungarian bipartite matching technique, the object box pairs can be reasonably determined. The aforementioned preset pairing requirements can be set according to actual needs. For example, a matching degree threshold can be set in the pairing requirements. This embodiment does not limit this.

[0088] It should be understood that in addition to the above Hungarian bipartite matching technique, other pairing solutions can also be used in this embodiment to determine the aforementioned object box pairs, as long as it is ensured that the two object boxes in the determined object box pair point to the same object as much as possible. No more examples will be given here.

[0089] After the object box pairs are determined, as mentioned above, the output results of the first model and the second model include foreground scores and object box vectors corresponding to the object boxes. In this way, the differences in object box vectors and foreground scores corresponding to the two object boxes in the object box pair can be evaluated, and the evaluation results obtained for each object box pair can be integrated to generate the aforementioned output feature loss. This can ensure the comprehensiveness and accuracy of the input feature loss.

[0090] In summary, in this embodiment, an exemplary implementation method is provided to use the output features and intermediate layer features of the second model as the optimization target of the first model. In this exemplary implementation method, by evaluating the output feature loss and intermediate layer feature loss of the first model relative to the first model, a loss function for the first model is constructed. In this way, on the one hand, the loss function does not need to rely on any labeled tags at all, so the pre-training process of the first model can completely rely on unlabeled data; on the other hand, the relevant features of the second model are used as constraints in the loss function, so that the first model can continuously learn the baseline capabilities in the second model.

[0091] Further, in this embodiment, in the aforementioned step 302, when the output result of the second model supports classification labels, classification labels can also be extracted from the output results of the first model and the second model respectively; in this way, output features can be constructed based on the object box vector, foreground score, and classification label.

[0092] Correspondingly, in the aforementioned step 303, the classification label difference between the first model and the second model can also be evaluated and used as part of the output feature loss. In this way, the output feature loss for the first model can include: foreground score difference, object box vector difference, and the classification label difference here.

[0093] It should be noted that although the loss evaluation of the classification label is proposed in this further improvement scheme, in this further improvement scheme, the classification label extracted from the second model is used as a constraint, and it does not rely on the labeled labels on the training samples. That is to say, in this further improvement scheme, it completely relies on unlabeled data to optimize the first model in terms of the output classification label and improve the baseline ability of the classification label output by the first model.

[0094] In addition, it should be understood that if the second model does not support the output of classification labels, then in this embodiment, the loss evaluation of the classification label mentioned in this further improvement scheme can be ignored, and the output feature loss can still be determined from the two aspects of foreground score and object box vector as described above.

[0095] Figure 4 The figure is a schematic diagram of an application scenario provided for an exemplary embodiment of the present application. In this application scenario, it is expected to improve the urban governance effect based on methods such as AI mobile patrol. AI mobile patrol can automatically detect illegal and irregular events through mobile vehicles, providing basic perception capabilities for the intelligent upgrade of urban governance. Moreover, AI mobile patrol can further improve the perception ability of the city through multi-dimensional integration of drones, patrol vehicles, fixed cameras, etc. Considering that no matter which perception method is adopted, visual detection based on pictures or videos is a basic function required in this scenario. Therefore, the model training method provided in this embodiment can be used to build a lightweight general visual detection model for this application scenario, hereinafter referred to as the "general visual detection model", which corresponds to the first model in the previous text.

[0096] Reference Figure 4 , in this application scenario, a pre-trained large visual model is introduced as the second model mentioned in this embodiment.

[0097] Based on this, in this application scenario, the model training can be completed according to the following several links to provide the "general visual detection model" for this application scenario.

[0098] Step 1: Design of Pre-training Tasks

[0099] Considering the objective limitations of unsupervised learning, in the pre-training stage of this solution, it is not required that the training samples have class labels. Instead, it is only required that the pre-training tasks can achieve the following outputs:

[0100] 1. Foreground object detection. Detect the foreground objects in the scene. Each object should include a foreground score S obj , which is normalized to [0, 1]. The larger the value, the more likely it is to be a foreground object; it also includes an object bounding box vector B obj , which is represented as [x, y, w, h]. Here, x is the abscissa of the center point of the target bounding box, y is the ordinate of the center point of the target bounding box, w is the width of the target bounding box, and h is the height of the target bounding box.

[0101] 2. Global feature extraction (embedding). Extract features from the training samples to obtain a high-dimensional feature vector F img .

[0102] Step 2: Design of Model Structure

[0103] Considering the generality of the model, the model structure of the required "general visual detection model" is not strictly defined in this solution. Theoretically, any model structure that includes a backbone network and object bounding box regression can be used. For the convenience of solution description and also considering the computational cost, this solution can mainly design the structure around lightweight visual detection models (including but not limited to YOLO, etc.).

[0104] Reference Figure 4 , a typical lightweight visual detection model generally includes three structural parts: a backbone network, a feature pyramid, and a detection head. After the image passes through the backbone network and the feature pyramid, high-dimensional features representing different scales will be generated. To adapt to the aforementioned pre-training task design, at least the following two aspects of processing can be carried out here:

[0105] 1. Feature space mapping. Perform a fully convolutional operation on the output of the backbone network to map it into a feature space of a specific scale. The dimension of this space should be exactly the same as the scale of the feature vector in the pseudo ground-truth mentioned in the next step, which can be denoted as CONV(Fimg).

[0106] 2. Generation of recognition bounding boxes. Still use the detection head of the original network to generate recognition bounding boxes. The typical output of the recognition bounding boxes can include the category (One-Hot), foreground score, and object bounding box vector. Among them, the foreground score S obj and the object bounding box Bobj It has been described in the previous section.

[0107] Step 3: Generation of Pre-training Pseudo Ground-Truth

[0108] After completing the design of the pre-training task, it is necessary to complete the design of the corresponding pre-training objectives. Similarly, considering the objective limitations of unsupervised learning, the training objectives should not contain any object category labels. To align with the above pre-training tasks, the pseudo ground-truth should also correspondingly include the following two parts of information:

[0109] Foreground object detection. Foreground object detection is performed through a pre-trained large vision model. Symmetrically, the model needs to output a foreground score and an object box vector. All basic mainstream pre-trained large vision models can meet the above requirements, including but not limited to Swin, Transformer, DETR, etc.

[0110] Global feature extraction. Global feature extraction is performed through a pre-trained large vision model. Whether it is a convolutional neural network or a Transformer, mainstream object detectors include a backbone part. Conventionally, this solution can directly use the output of the backbone network as the global feature.

[0111] The above two parts of information provided by the large vision model are constructed into pre-training pseudo ground-truth.

[0112] Step 4: Unsupervised Pre-training

[0113] After completing the design of the overall task, model structure, and pseudo ground-truth acquisition mentioned above, this solution can start unsupervised pre-training.

[0114] 1. After passing through the backbone network part of the lightweight vision detection model, the unlabeled training samples obtain high-dimensional feature vectors. The same training samples also generate a high-dimensional feature vector after passing through the backbone network or the encoder network of the pre-trained large vision model. Through a learnable convolution operation, the two are mapped into the same feature space, and a knowledge distillation loss function in the feature space is constructed, corresponding to the intermediate layer feature loss in the previous text.

[0115] 2. The training samples continue to propagate forward along the lightweight visual detection network. After passing through the feature pyramid, they are input into the multi-scale detection heads, respectively obtaining object box vectors and foreground scores. Without loss of generality, the object class labels output by the detection heads are not included in the pre-training loss function. The same training samples will also output object box vectors and foreground scores after passing through the decoder (Decoder) of the large visual model. This solution can construct the loss functions of the large and small models in foreground object detection through Hungarian bipartite matching, as shown in the following formula. Among them, Lcls is the foreground and background classification loss, represented by binary cross-entropy; Lbox is the foreground box regression loss, which can be represented by the IoU loss. IIobj represents a logical operation, setting the candidate boxes matched as foreground to 1 and vice versa to 0. λcls is the foreground loss weight, and λbox is the object box recognition weight.

[0116] L match =λ cls ·L cls (S obj1 ,S obj3 )+λ box ·II obj ·L box (B obj1 ,B obj3 )

[0117] The final loss function for the lightweight visual detection model is the sum of the aforementioned knowledge distillation loss and the foreground object detection loss.

[0118] Based on this loss function, backpropagation is performed on the lightweight visual detection model to continuously optimize the model parameters in the lightweight visual detection model. Finally, the "general visual detection model" required for this application scenario can be obtained. The obtained "general visual detection model" not only has excellent generalization performance and the ability of few-shot learning for downstream tasks, but also can obtain the baseline ability of the large visual model camera. The obtained "general visual detection model" can be used in real production scenarios such as edge computing or end intelligence. It can also be fine-tuned through the model to achieve precise perception of different types of objects in the aforementioned urban governance scenarios, such as illegal occupation of roads, outdoor advertisements, and overflowing trash cans. No more examples are given here.

[0119] It should be noted that in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different models, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0120] Figure 5 FIG. is a schematic structural diagram of a computing device provided by another exemplary embodiment of the present application. As Figure 5 shown, the computing device includes: a memory 50, a processor 51, and a communication component 52.

[0121] The processor 51 is coupled to the memory 50 and the communication component 52 and is configured to execute a computer program in the memory 50 for:

[0122] In response to a pre-training instruction for the first model, obtain a training sample set, and the training samples in the training sample set do not carry labels;

[0123] Input the training samples without labels into the first model and the second model respectively. The second model uses a pre-trained model, and the scale of the second model is larger than that of the first model;

[0124] Extract output features and intermediate layer features from the first model and the second model respectively;

[0125] Use the output features and intermediate layer features of the second model as the optimization target of the first model to perform unsupervised pre-training on the first model using the training samples without labels.

[0126] It should be understood that the pre-trained first model in this embodiment can be deployed on an edge node for downstream tasks. However, the model training method provided in this embodiment is not limited to being implemented on an edge node, but can be implemented on any suitable computing device. Preferably, in order to improve training efficiency, it can be implemented on a central node in a cloud environment, which is not limited herein.

[0127] In an optional embodiment, when the processor 51 uses the output features and intermediate layer features of the second model as the optimization target of the first model, it may specifically be used for:

[0128] Evaluate the output feature loss and intermediate layer feature loss of the first model relative to the first model to construct a loss function for the first model;

[0129] Based on the loss function, perform parameter tuning on the first model.

[0130] In an alternative embodiment, when the processor 51 extracts output features from the first model and the second model respectively, it may specifically be configured to:

[0131] If the first model is a visual detection model, extract object box vectors and foreground scores from the output results of the first model and the second model respectively,

[0132] Based on the object box vectors and the foreground scores, construct the output features;

[0133] Wherein, the foreground score is used to characterize the confidence that there is a foreground object within the object box.

[0134] In an alternative embodiment, when the processor 51 constructs the output features based on the object box vectors and the foreground scores, it may specifically be configured to:

[0135] When the output result of the second model supports classification labels, extract classification labels from the output results of the first model and the second model respectively;

[0136] Based on the object box vectors, the foreground scores, and the classification labels, construct the output features.

[0137] In an alternative embodiment, when the processor 51 evaluates the output feature loss and intermediate layer feature loss of the first model relative to the first model, it may specifically be configured to:

[0138] If the first model is a visual detection model, pair the object boxes in the output results of the second model and the first model for the same training sample to determine object box pairs that meet preset pairing requirements;

[0139] Evaluate the object box vector difference and foreground score difference corresponding to the two object boxes in the object box pair to characterize the output feature loss.

[0140] In an alternative embodiment, when the processor 51 pairs the object boxes in the output results of the second model and the first model for the same training sample to determine object box pairs that meet preset pairing requirements, it may specifically be configured to:

[0141] Using the Hungarian bipartite matching technique, pair the object bounding boxes in the output results generated by the second model and the first model for the same training sample to determine object bounding box pairs that meet the preset pairing requirements.

[0142] In an alternative embodiment, when the processor 51 extracts intermediate layer features from the first model and the second model respectively, it may specifically be used to:

[0143] Extract intermediate layer features from the backbone networks of the first model and the second model;

[0144] Wherein, the intermediate layer features include features extracted at one or more stages within the backbone network.

[0145] In an alternative embodiment, the processor 51 extracting intermediate layer features from the first model and the second model respectively includes:

[0146] If the feature space scales of the first model and the second model are different, map the intermediate layer features generated from the first model to a feature space consistent with the feature space scale of the second model as the intermediate layer features extracted from the first model.

[0147] In an alternative embodiment, the processor 51 may also be used to:

[0148] After completing the unsupervised pre-training of the first model, in response to a tuning instruction for the first model, obtain a sample set associated with the downstream task requirements, where the samples in the sample set carry labels;

[0149] Perform supervised training on the pre-trained first model based on the samples with labels so that the first model adapts to the downstream task requirements.

[0150] In an alternative embodiment, the downstream task includes one or more of an object detection task, an image classification task, or a semantic segmentation task.

[0151] In an alternative embodiment, the processor 51 may also be used to:

[0152] After completing the unsupervised pre-training of the first model, in response to a transfer learning instruction for the first model, transfer the backbone network in the pre-trained first model to other models.

[0153] In an alternative embodiment, the second model uses a large model that matches the basic functions required by the first model.

[0154] Further, as Figure 5 shown, the computing device further includes other components such as a power supply component 53.Figure 5 Only some components are schematically shown, which does not mean that the computing device only includes Figure 5 the components shown.

[0155] It should be noted that for the technical details in the foregoing embodiments of the computing device, reference may be made to the relevant descriptions in the foregoing method embodiments. To save space, they will not be elaborated herein, but this should not cause loss of the protection scope of this application.

[0156] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps executable by the computing device in the foregoing method embodiments.

[0157] The foregoing Figure 5 The memory is used to store a computer program and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method for operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0158] The foregoing Figure 5 The communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0159] The foregoing Figure 5 The power component provides power for various components of the device where the power component is located. The power component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power component is located.

[0160] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0161] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0162] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0164] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0165] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0166] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A model training method, characterized in that: include: In response to a pre-training instruction for the first model, a training sample set is obtained, wherein the training samples in the training sample set do not carry labels; Inputting the training samples without labels into the first model and the second model respectively, wherein the second model adopts a pre-trained model, and the scale of the second model is larger than that of the first model; Extracting output features and intermediate layer features from the first model and the second model respectively; The output features and intermediate layer features of the second model are used as optimization targets of the first model, so as to perform unsupervised pre-training on the first model using the training samples without labels.

2. The method according to claim 1, characterized in that Using the output features and intermediate layer features of the second model as optimization targets of the first model includes: Evaluating the output feature loss and the intermediate layer feature loss of the first model relative to the first model to construct a loss function for the first model; Based on the loss function, parameters of the first model are tuned.

3. The method according to any one of claims 1 or 2, characterized in that: Extracting output features from the first model and the second model respectively includes: If the first model is a visual detection model, the object frame vector and the foreground score are extracted from the output results of the first model and the second model respectively. constructing the output feature based on the object box vector and the foreground score; The foreground score is used to characterize the confidence of the foreground object existing in the object frame.

4. The method according to claim 3, characterized in that Constructing the output feature based on the object frame vector and the foreground score includes: When the output result of the second model supports classification labels, extracting classification labels from the output results of the first model and the second model respectively; The output feature is constructed based on the object box vector, the foreground score, and the classification label.

5. The method according to claim 2, characterized in that: Evaluating the output feature loss and the intermediate layer feature loss of the first model relative to the first model, including: If the first model is a visual detection model, pairing the object frames in the output results generated by the second model and the first model for the same training sample to determine an object frame pair that meets preset pairing requirements; The object box vector difference and the foreground score difference corresponding to the two object boxes in the object box pair are evaluated to characterize the output feature loss.

6. The method according to claim 5, characterized in that Pairing the object frames in the output results generated by the second model and the first model for the same training sample to determine an object frame pair that meets preset pairing requirements, including: The Hungarian bipartite matching technique is used to pair the object frames in the output results generated by the second model and the first model for the same training sample to determine an object frame pair that meets the preset pairing requirements.

7. The method according to claim 1, characterized in that Extracting intermediate layer features from the first model and the second model respectively includes: Extracting intermediate layer features from the backbone networks of the first model and the second model; The intermediate layer features include features extracted at one or more stages in the backbone network.

8. The method according to any one of claims 1 or 7, characterized in that: Extracting intermediate layer features from the first model and the second model respectively includes: If the feature space scales of the first model and the second model are different, the intermediate layer features generated from the first model are mapped to a feature space having the same feature space scale as the second model to serve as the intermediate layer features extracted from the first model.

9. The method according to claim 1, characterized in that: Also includes: After completing the unsupervised pre-training of the first model, in response to a tuning instruction for the first model, obtaining a sample set associated with downstream task requirements, wherein samples in the sample set carry labels; The pre-trained first model is supervisedly trained based on samples carrying labels so that the first model adapts to the downstream task requirements.

10. The method according to claim 9, characterized in that The downstream tasks include one or more of a target detection task, an image classification task, or a semantic segmentation task.

11. The method according to claim 1, characterized in that: Also includes: After completing the unsupervised pre-training of the first model, in response to a transfer learning instruction for the first model, the backbone network in the pre-trained first model is migrated to other models.

12. The method according to claim 1, characterized in that The second model uses a large model that matches the basic functions required by the first model.

13. A computing device, characterized in that: Includes memory, processor, and communication components; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component, and is used to execute the one or more computer instructions to execute the model training method described in any one of claims 1-12.

14. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the model training method described in any one of claims 1-12.