Model training method and device, electronic equipment and computer program product

Through the method of querying and evaluating candidate models from the preset algorithm library in the field of computer vision, the problem of low model training efficiency in the existing technology is solved, and efficient and accurate model training is achieved.

CN120105307APending Publication Date: 2025-06-06CHINA TELECOM CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510253455.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing technology has low model training efficiency in the field of computer vision, high application technology threshold, low resource utilization rate, and cannot scale horizontally, which has the problem of low model training efficiency.

Method used

By obtaining the training data set and task requirements in the model training task, query candidate models that meet the task requirements from the preset algorithm library, evaluate the performance of the candidate model using the representative data set, and filter the candidate models based on the model performance. Finally, simply train the selected candidate models using the training data set to obtain the target model.

Benefits of technology

It realizes the rapid determination of candidate models that meet the required model training tasks, improves the efficiency and accuracy of model training, reduces the complexity of model training, and solves the problem of low model training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105307A_ABST
    Figure CN120105307A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, electronic equipment and a computer program product. The method comprises the steps that a model training task is acquired, and the model training task at least comprises a training data set and a task requirement; querying at least one candidate model meeting the task requirement in a plurality of preset models recorded in a preset algorithm library; a representative data set in the training data set is used for evaluating the model performance of each candidate model, multiple representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set are used for training each candidate model, and the representative sample data in the pre-verification set are used for verifying the candidate model; the pre-verification set is used for performing performance evaluation on each trained candidate model to obtain model performance; and using the training data set to train the candidate model of which the model performance accords with a preset suitability standard to obtain a target model. The technical problem of low model training efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a model training method, device, electronic equipment and computer program product. Background Art

[0002] With the continuous development of artificial intelligence (AI), computer vision (CV) technology, as an important branch of the field of artificial intelligence, has shown great potential and value in many industries and application scenarios, including but not limited to autonomous driving, medical diagnosis, security monitoring, retail analysis, entertainment interaction, etc. However, the application of CV technology is not a simple process that can be achieved overnight. It involves complex theoretical foundations, algorithm design, data processing, and engineering implementation. This makes its application threshold relatively high and usually requires professional technicians and teams to complete.

[0003] Figure 1 It is a schematic diagram of a solution process for CV problems according to the prior art, such as Figure 1 As shown in the figure, for a specific CV scenario problem, the steps from "requirements definition and scenario analysis" -> "data collection and preprocessing" -> "model selection and training" -> "model evaluation and tuning" -> "deployment and maintenance" all require professional technicians or teams to write and execute codes. The application threshold is quite high and the problem cannot be solved quickly.

[0004] Since existing technologies need to rely on professional technicians or teams to solve CV problems, the current computer vision (CV) technology has the defects of high application technology threshold, low resource utilization, inability to horizontally expand needs, and low model training efficiency when training machine learning models.

[0005] Currently, no effective solution has been proposed to the problem of low training efficiency of the above-mentioned models. Summary of the invention

[0006] Embodiments of the present invention provide a model training method, device, electronic device and computer program product to at least solve the technical problem of low model training efficiency.

[0007] According to one aspect of an embodiment of the present invention, a model training method is provided, comprising: obtaining a model training task, wherein the model training task includes at least: a training data set and a task requirement; querying at least one candidate model that meets the task requirement among multiple preset models recorded in a preset algorithm library; using a representative data set in the training data set to evaluate the model performance of each candidate model, wherein the training data set includes multiple preset sample data, the representative data set includes multiple representative sample data, the representative sample data is the preset sample data that can represent the data characteristics of the training data set, the number of the representative sample data in the representative data set is less than the number of the preset sample data in the training data set, and the multiple representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain the model performance; using the training data set, training the candidate model whose model performance meets the preset adaptability standard to obtain a target model.

[0008] Optionally, among the multiple preset models recorded in the preset algorithm library, querying at least one candidate model that meets the task requirements includes: identifying the model requirement label indicated by the model training task, wherein the model requirement label is used to indicate the algorithm model required for the model training task; among the multiple preset models recorded in the preset algorithm library, querying at least one candidate model that matches the model requirement label, wherein the preset algorithm library is also used to record the model preset label corresponding to each preset model.

[0009] Optionally, among multiple preset models recorded in the preset algorithm library, querying at least one candidate model that meets the task requirements includes: analyzing the task requirements, decomposing the model training task into multiple model training subtasks, wherein the task requirements are at least used to describe multiple target sub-models constituting the target model, and each of the model training subtasks is used to indicate the selection of the target sub-model in the preset algorithm library; identifying the sub-model requirement label indicated by each of the model training subtasks, wherein the sub-model requirement label is at least used to indicate the target sub-model type required for the model training subtask; among the multiple preset sub-model types recorded in the preset algorithm library, querying the target sub-model type corresponding to each sub-model requirement label, wherein the multiple preset models of the preset algorithm library are pre-divided into multiple preset sub-model types, and each of the preset sub-model types has a corresponding sub-model preset label; combining the preset models in the multiple target sub-model types to obtain at least one candidate model.

[0010] Optionally, combining the preset models in multiple target sub-model types to obtain at least one candidate model includes: determining a combination order of multiple target sub-model types based on multiple model training sub-tasks decomposed from the model training task, wherein the combination order represents the structural order of the multiple model training sub-tasks in the model training task; and randomly selecting the preset models from at least one preset model corresponding to each target sub-model type according to the combination order to combine to obtain at least one candidate model.

[0011] Optionally, after querying at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library, the method also includes: identifying candidate training tasks performed to train each of the candidate models, wherein the candidate training tasks are at least used to describe the feature space, data statistical attributes and task target similarity used to train the candidate models; determining the task similarity between the model training task and the candidate training task, wherein the model training task is also used to describe the feature space, data statistical attributes and task target similarity used to train the target model, and the task similarity is based on the feature space similarity, data statistical similarity and task target similarity between the model training task and the candidate training task, the feature space similarity is determined based on the feature space of the model training task and the candidate training task, the data statistical similarity is determined based on the data statistical attributes of the model training task and the candidate training task, and the task target similarity is determined based on the task target similarity between the model training task and the candidate training task; and deleting the candidate models whose task similarity does not meet the preset task similarity threshold.

[0012] Optionally, using the training data set to train the candidate model whose model performance meets the preset adaptability standard, obtaining the target model includes: determining the candidate model whose model performance meets the preset adaptability standard as the model to be trained; using the representative data set in the training data set to perform transfer learning on the model to be trained to obtain the target model.

[0013] Optionally, before querying at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library, the method also includes: obtaining a historical model trained historically, wherein the historical model is a combination of multiple historical sub-models with different functions; decomposing the historical model to obtain multiple historical sub-models and dependency relationships corresponding to each historical sub-model; containerizing each historical sub-model and the dependency relationships corresponding to the historical sub-models to obtain a model container corresponding to each historical sub-model; and storing each model container in the preset algorithm library, wherein the historical sub-model encapsulated in each model container is the preset model in the preset algorithm library.

[0014] According to another aspect of an embodiment of the present invention, a model training device is also provided, comprising: an acquisition module, used to acquire a model training task, wherein the model training task includes at least: a training data set and a task requirement; a query module, used to query at least one candidate model that meets the task requirement among multiple preset models recorded in a preset algorithm library; an evaluation module, used to use a representative data set in the training data set to evaluate the model performance of each candidate model, wherein the training data set includes multiple preset sample data, the representative data set includes multiple representative sample data, the representative sample data is the preset sample data that can represent the data characteristics of the training data set, the number of the representative sample data in the representative data set is less than the number of the preset sample data in the training data set, and the multiple representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain the model performance; a training module, used to use the training data set to train the candidate model whose model performance meets the preset adaptability standard to obtain a target model.

[0015] According to another aspect of an embodiment of the present invention, there is further provided an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-mentioned model training method through the computer program.

[0016] According to another aspect of an embodiment of the present invention, a computer program product is also provided, comprising computer instructions, which implement the steps of the above-mentioned model training method when executed by a processor.

[0017] In an embodiment of the present invention, when it is necessary to perform model training based on a training data set and task requirements in a model training task, at least one candidate model can be screened from a plurality of preset models recorded in a preset algorithm library according to the task requirements in the model training task, and then a small number of samples that can represent the data features in the training data set are selected from the training data set as a representative data set, and then the representative data set is split into a pre-training set used to train the candidate models, and a pre-verification set for performance evaluation of the trained candidate models, and then based on the model performance of each candidate model, the candidate models can be screened according to the model performance, and then the selected candidate models are simply trained using the training data set to obtain the target model requested for training by the model training task, thereby using a representative data set with a small amount of data to screen the candidate models, and the candidate models that meet the requirements of the model training task can be quickly determined, and then training based on the candidate models can ensure training efficiency and accuracy, thereby achieving the technical effect of reducing the complexity of model training, and thus solving the technical problem of low model training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0019] Figure 1 It is a schematic diagram of a solution process for CV problems according to the prior art;

[0020] Figure 2 is a flow chart of a model training method according to an embodiment of the present invention;

[0021] Figure 3 is a schematic diagram of an automated full-link method for computer vision (CV) scene problems according to an embodiment of the present invention;

[0022] Figure 4 is a schematic diagram of an Efficient Neural Architecture Search (ENAS) according to an embodiment of the present invention;

[0023] Figure 5 is a schematic diagram of a Hyperband algorithm process according to an embodiment of the present invention;

[0024] Figure 6 is a schematic diagram of a Ring-AllReduce algorithm according to an embodiment of the present invention;

[0025] Figure 7 is a schematic diagram of an ONNX graph structure according to an embodiment of the present invention;

[0026] Figure 8 is a schematic diagram of a Kubernetes deployment process according to an embodiment of the present invention;

[0027] Fig. 9 is a schematic diagram of a monitoring system architecture according to an embodiment of the present invention;

[0028] Fig.10 is a schematic diagram of an early warning process according to an embodiment of the present invention;

[0029] Fig.11 is a schematic diagram of an Isolation Forest algorithm according to an embodiment of the present invention;

[0030] Fig.12 is a schematic diagram of an AutoEncoder architecture according to an embodiment of the present invention;

[0031] Fig.13 is a schematic diagram of an algorithm library structure according to an embodiment of the present invention;

[0032] Fig.14 is a schematic diagram of a version control structure according to an embodiment of the present invention;

[0033] Fig.15 is a core architecture schematic diagram of a rapid adaptability evaluation according to an embodiment of the present invention;

[0034] Fig.16 is a schematic diagram of transfer learning according to an embodiment of the present invention;

[0035] Fig.17 is a schematic diagram of few-sample learning according to an embodiment of the present invention;

[0036] Fig.18 is a schematic diagram of a model training device according to an embodiment of the present invention;

[0037] Fig.19 It is a structural block diagram of a computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0040] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:

[0041] AI (Artificial Intelligence) refers to the intelligent behaviors exhibited by systems or machines created by humans. These behaviors are usually similar to those of humans or other animals, including the abilities of learning, understanding, reasoning, planning, perception, problem solving, and language comprehension.

[0042] CV (Computer Vision): In the field of computer science and artificial intelligence, C refers to computer vision, which involves enabling computers to gain high-level understanding from images or videos, including tasks such as object detection, scene recognition, image segmentation, motion analysis, etc.

[0043] According to an embodiment of the present invention, a model training method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0044] Figure 2 is a flow chart of a model training method according to an embodiment of the present invention. Figure 2 As shown, the method comprises the following steps:

[0045] Step S202, obtaining a model training task, wherein the model training task at least includes: a training data set and a task requirement;

[0046] Step S204, searching for at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library;

[0047] Step S206, using a representative data set in the training data set to evaluate the model performance of each candidate model, wherein the training data set includes a plurality of preset sample data, the representative data set includes a plurality of representative sample data, the representative sample data is preset sample data that can represent the data features of the training data set, the number of representative sample data in the representative data set is less than the number of preset sample data in the training data set, the plurality of representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain model performance;

[0048] Step S208, using the training data set, training the candidate model whose model performance meets the preset adaptability standard to obtain the target model.

[0049] In an embodiment of the present invention, when it is necessary to perform model training based on a training data set and task requirements in a model training task, at least one candidate model can be screened from a plurality of preset models recorded in a preset algorithm library according to the task requirements in the model training task, and then a small number of samples that can represent the data features in the training data set are selected from the training data set as a representative data set, and then the representative data set is split into a pre-training set used to train the candidate models, and a pre-verification set for performance evaluation of the trained candidate models, and then based on the model performance of each candidate model, the candidate models can be screened according to the model performance, and then the selected candidate models are simply trained using the training data set to obtain the target model requested for training by the model training task, thereby using a representative data set with a small amount of data to screen the candidate models, and the candidate models that meet the requirements of the model training task can be quickly determined, and then training based on the candidate models can ensure training efficiency and accuracy, thereby achieving the technical effect of reducing the complexity of model training, and thus solving the technical problem of low model training efficiency.

[0050] In the above step S202, the model training task is at least used to instruct to train a target model that meets the task requirement indication according to preset sample data in the training data set.

[0051] In the above step S204, the preset model recorded in the preset algorithm library is determined in advance based on the historical model trained historically. The preset model can be a complete historical model trained historically, or a historical sub-model obtained by splitting the historical model trained historically.

[0052] In the above step S204, the candidate model is selected according to the task requirements, wherein the task requirements may indicate the model structure (or algorithm model) required for the model training task.

[0053] Optionally, the model structure (ie, the algorithm model) is used to represent the types of sub-models in the target model indicated by the composition model training task, as well as the combination relationship between the sub-models.

[0054] Optionally, when the preset model is a complete historical model that has been pre-selected and trained, each historical model can be pre-labeled with a corresponding model structure (ie, an algorithm model), and each candidate model is a historical model (ie, a preset model) whose model structure meets the task requirements.

[0055] Optionally, in the case where the preset model is a historical sub-model pre-selected from a complete historical model, each historical sub-model can be pre-labeled with the corresponding model type, and the candidate model can be a combination relationship between model types indicated by the model structure (that is, the algorithm model), and historical sub-models (that is, preset models) that meet each model type are selected in turn for combination.

[0056] For example, the combination relationship between the model types indicated by the model structure is, in sequence: model type A and model type B; and then the multiple preset models that conform to model type A can be: preset model a1 and preset model a2, and the multiple preset models that conform to model type B can be: preset model b1 and preset model b2; at least one candidate model obtained by combining multiple preset models (that is, historical sub-models) can be: a combination of preset model a1 and preset model b1, a combination of preset model a1 and preset model b2, a combination of preset model a2 and preset model b1, and a combination of preset model a2 and preset model b2.

[0057] In the above step S206, the training data set includes a plurality of preset sample data, and the representative sample data in the representative data set may be screened out from the plurality of representative sample data in the training data set.

[0058] In the above step S206, the multiple representative sample data in the representative data set can be divided into a pre-training set and a pre-verification set. The multiple representative sample data in the pre-training set are used to train each candidate model, and then the performance of the trained candidate model is evaluated using the multiple representative sample data in the pre-verification set.

[0059] In the above step S208, the model performance evaluated by each candidate model is compared with the preset adaptability standard in turn to determine the model performance that meets the preset adaptability standard and the candidate model corresponding to the model performance. The candidate model that best meets the model training task can be screened out, and then the candidate model is trained using the training data set in the model training task to obtain the most accurate target model.

[0060] In the above-mentioned embodiments of the present application, a representative data set with a small number of samples in the training data set is used to screen the candidate models, and then the target model is obtained by training based on the screened candidate models. Therefore, there is no need to use the complete training data set to train multiple models and then screen the trained models, thereby improving the model training efficiency while ensuring the accuracy of the training model.

[0061] As an optional embodiment, querying at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library includes: identifying the model requirement label indicated by the model training task, wherein the model requirement label is used to indicate the algorithm model required for the model training task; querying at least one candidate model that matches the model requirement label among multiple preset models recorded in the preset algorithm library, wherein the preset algorithm library is also used to record the model preset label corresponding to each preset model.

[0062] In the above-mentioned embodiments of the present application, the task requirements in the model training task can be represented by a model requirement label. Each preset model recorded in the corresponding preset algorithm library can be a complete historical model that is pre-trained. Each preset model (that is, the historical model) is pre-set with a corresponding model preset label. Then, when it is necessary to query the candidate model from the preset algorithm library, the model preset label corresponding to each preset model in the preset algorithm library can be matched with the model requirement label indicated by the model training task, and the preset model corresponding to the successfully matched model preset label can be used as the candidate model.

[0063] Optionally, when there is no at least one candidate model matching the model requirement label among the multiple preset models recorded in the preset algorithm library, the task requirements can be analyzed and the model training task can be decomposed into multiple model training subtasks.

[0064] As an optional embodiment, querying at least one candidate model that meets the task requirements from among multiple preset models recorded in a preset algorithm library includes: analyzing the task requirements, decomposing the model training task into multiple model training subtasks, wherein the task requirements are at least used to describe multiple target sub-models that constitute the target model, and each model training subtask is used to indicate the selection of a target sub-model in the preset algorithm library; identifying a sub-model requirement label indicated by each model training subtask, wherein the sub-model requirement label is at least used to indicate a target sub-model type of the algorithm sub-model required for the model training subtask; querying the target sub-model type corresponding to each sub-model requirement label from among multiple preset sub-model types recorded in the preset algorithm library, wherein the multiple preset models of the preset algorithm library are pre-divided into multiple preset sub-model types, and each preset sub-model type has a corresponding sub-model preset label; combining the preset models in the multiple target sub-model types to obtain at least one candidate model.

[0065] In the above-mentioned embodiments of the present application, the target model required for the model training task can be composed of multiple target sub-models. Therefore, the task requirement indicating the training of the target model can be decomposed into model training sub-tasks for training each target sub-model, and a combination of multiple model training sub-tasks, and each model training sub-task can use the corresponding sub-model requirement label to indicate the target sub-model type of the required target sub-model. Each preset model recorded in the corresponding preset algorithm library can be a historical sub-model split from a pre-trained complete historical model, and each preset model (i.e., historical sub-model) is pre-set with a corresponding sub-model preset label; and then, when it is necessary to query a candidate model from the preset algorithm library, at least one preset model (i.e., historical sub-model) corresponding to each target sub-model type can be queried based on the sub-model requirement label, and then according to the combination of the model training sub-tasks, at least one preset model (i.e., historical sub-model) queried by each model training sub-task is combined to obtain at least one candidate model.

[0066] As an optional embodiment, combining preset models in multiple target sub-model types to obtain at least one candidate model includes: determining a combination order of multiple target sub-model types based on multiple model training sub-tasks decomposed from the model training task, wherein the combination order represents the structural order of the multiple model training sub-tasks in the model training task; and randomly selecting preset models from at least one preset model corresponding to each target sub-model type according to the combination order to combine to obtain at least one candidate model.

[0067] In the above-mentioned embodiment of the present application, when multiple model training subtasks are decomposed from a model training task, a combination method of the multiple model training subtasks can be obtained, and based on the combination method of the multiple model training subtasks, the combination order of the target sub-model types queried by each model training subtask can be determined, and then when at least one preset model corresponding to each target sub-model type is queried, a preset model can be randomly selected from at least one preset model corresponding to each target sub-model type according to the above-mentioned combination order for combination to obtain at least one of the candidate models.

[0068] As an optional embodiment, after querying at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library, the method also includes: identifying candidate training tasks performed to train each candidate model, wherein the candidate training tasks are at least used to describe the feature space, data statistical attributes and task target similarity used to train the candidate model; determining the task similarity between the model training task and the candidate training task, wherein the model training task is also used to describe the feature space, data statistical attributes and task target similarity used by the training target model, and the task similarity is based on the feature space similarity, data statistical similarity and task target similarity of the model training task and the candidate training task, the feature space similarity is determined based on the feature space of the model training task and the candidate training task, the data statistical similarity is determined based on the data statistical attributes of the model training task and the candidate training task, and the task target similarity is determined based on the task target similarity of the model training task and the candidate training task; and deleting candidate models whose task similarity does not meet the preset task similarity threshold.

[0069] In the above-mentioned embodiment of the present application, the preset model recorded in the preset algorithm library is a historical model or a historical sub-model of historical training. Therefore, each preset model has a corresponding model training task in the historical training stage, which is a candidate training task. The model training task and the candidate training task are also used to describe the training indicators such as the feature space, data statistical attributes and task target similarity used to train the target model. Therefore, after querying the candidate model from the preset algorithm library, it is also necessary to calculate the task similarity between the candidate training task of the candidate model and the current model training task to further determine whether the candidate model meets the task requirements indicated by the current model training task, and then determine the candidate model whose task similarity meets the preset task similarity threshold as meeting the task requirements indicated by the current model training task, and determine the candidate model whose task similarity does not meet the preset task similarity threshold as not meeting the task requirements indicated by the current model training task, thereby deleting the candidate model to achieve preliminary screening of the candidate model.

[0070] Optionally, when the preset model recorded in the preset algorithm library is a complete historical model of historical training, the candidate training task is a model training task for training the historical model; and then determining the task similarity between the model training task and the candidate training task includes: determining the similarity between the current model training task and the model training task for training the historical model.

[0071] Optionally, when the preset model recorded in the preset algorithm library is a historical sub-model split from a historical model, the candidate training task is a model training sub-task split from the model training task for training the historical model, and further determining the task similarity between the model training task and the candidate training task includes: determining the similarity between the model training sub-task in the model training task that indicates querying the historical sub-model and the model training sub-task for training the historical sub-model.

[0072] As an optional embodiment, a training data set is used to train a candidate model whose model performance meets the preset adaptability standard, and obtaining a target model includes: determining the candidate model whose model performance meets the preset adaptability standard as the model to be trained; using a representative data set in the training data set to perform transfer learning on the model to be trained to obtain the target model.

[0073] In the above-mentioned embodiment of the present application, after screening out candidate models that meet the preset adaptability standards, the target model can be trained based on the candidate models by using transfer learning, wherein transfer learning improves new learned tasks by transferring knowledge from related tasks that have been learned. Therefore, the model parameters in the candidate model can be migrated to the target model, and then a small amount of training data in the training data set, i.e., the representative data set, can be used to adjust the model parameters of the migrated target model, thereby achieving the purpose of training the target model using a small amount of training data, thereby improving the training efficiency of the target model.

[0074] As an optional embodiment, a training data set is used to train a candidate model whose model performance meets the preset adaptability standard, and a target model is obtained, including: when the candidate model is determined based on multiple model training subtasks, that is, the candidate model is a combination of multiple preset models, a distributed training method can be used to train each preset model that constitutes the candidate model separately, and then the training results of each preset model are combined to obtain the target model.

[0075] The above-mentioned embodiments of the present application can improve the training efficiency of the target model through distributed training.

[0076] As an optional embodiment, before querying at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library, the method also includes: obtaining a historical model trained historically, wherein the historical model is a combination of multiple historical sub-models with different functions; decomposing the historical model to obtain multiple historical sub-models and the dependency relationship corresponding to each historical sub-model; containerizing each historical sub-model and the dependency relationship corresponding to the historical sub-model to obtain a model container corresponding to each historical sub-model; and storing each model container in the preset algorithm library, wherein the historical sub-model encapsulated in each model container is a preset model in the preset algorithm library.

[0077] In the above-mentioned embodiments of the present application, the preset models recorded in the preset algorithm library can be pre-containerized using the containerization packaging technology. Therefore, when the preset model is a historical model trained historically, the corresponding model container can be obtained after the historical model is containerized, and then the function of the corresponding historical model can be realized by calling the model container through containerization; when the preset model is a historical sub-model split from a historical model trained historically, each historical sub-model can be obtained after containerization to be called in a model container, and then the function of the corresponding historical sub-model can be realized by calling the model container through containerization, thereby containerizing and packaging the multiple preset models recorded in the preset algorithm library, so that each preset model can be directly called, thereby reducing the difficulty of model calling, and then the training of the target model can be realized faster through the preset model that is easier to schedule.

[0078] The present invention also provides a preferred embodiment, which provides an automated full-link method for computer vision (CV) scene problems, aiming to lower the application threshold of CV technology, improve resource utilization efficiency, and achieve horizontal expansion of demand.

[0079] Figure 3 is a schematic diagram of an automated full-link method for computer vision (CV) scene problems according to an embodiment of the present invention, such as Figure 3 As shown, it includes: the entire process of CV model automated training, deployment and reasoning, full-process monitoring and early warning, and horizontal expansion on demand.

[0080] Alternatively, if Figure 3 As shown in the figure, the entire process of CV model automated training, deployment, and reasoning includes: automated training, automated deployment, and reasoning.

[0081] Alternatively, if Figure 3 As shown, automated training includes the following steps:

[0082] Step S301, data input.

[0083] Step S302: automated data preprocessing.

[0084] Step S303: model architecture search.

[0085] Step S304: hyperparameter optimization.

[0086] Step S305: distributed training.

[0087] Step S306, adaptive learning rate adjustment, training is completed.

[0088] In the above embodiment of the present application, through the above steps S301 to S306, the automatic training of the CV task model can be realized, and the user only needs to select the training data on the interface.

[0089] In the above step S302, data preprocessing is automated for data verification and analysis using TensorFlow Data Validation (TFDV), and the AutoAugment algorithm is used to automatically find the optimal data augmentation strategy.

[0090] Optionally, the AutoAugment algorithm is a search algorithm, and its core principle is as follows:

[0091] Step S3021, define the search space: including various data enhancement operations (such as rotation, cropping, color transformation, etc.) and their intensities.

[0092] Step S3022, using a reinforcement learning method: using a recursive neural network (RNN) as a controller to generate a data enhancement strategy.

[0093] Step S3023, training process:

[0094] The controller generates a data augmentation policy.

[0095] ii Use this strategy to enhance the training data.

[0096] iii. Train a small proxy network on the augmented data.

[0097] iv. Use the performance of the proxy network on the validation set as a reward signal.

[0098] v Update controller parameters to produce a better policy.

[0099] Step S3024, iterative optimization: repeat the above process to continuously improve the enhancement strategy.

[0100] Step S3025, final selection: select the strategy with the best performance on the validation set as the final result.

[0101] Optionally, the core formula of the AutoAugment algorithm is: S*=argmax S R(M(D a ug(S))), where S is the enhancement strategy, R is the reward function, M is the agent model, and D a ug(S) is the enhanced dataset.

[0102] In the above step S303, the model architecture search adopts the Neural Architecture Search (NAS) technology architecture, and the specific implementation algorithm of the architecture is ENAS (Efficient Neural Architecture Search), and the search efficiency is optimized by the ENAS algorithm. Neural Architecture Search (NAS) is a technology for automatically designing neural network architectures, and Efficient Neural Architecture Search (ENAS) is an efficient implementation of NAS.

[0103] As an optional example, the implementation principle of model architecture search includes:

[0104] Step S3031, defining the search space, specifically includes: creating a set of all possible network operations (such as convolution, pooling, etc.); defining a flexible unit structure that can combine these operations through different connection methods.

[0105] Step S3032, controller design, specifically includes: using a recurrent neural network (RNN) as a controller, wherein the controller is responsible for generating a description of the network architecture and determining the operation and connection method of each node.

[0106] Step S3033, parameter sharing mechanism, specifically includes: during the search process, all sampled models share parameters, which greatly reduces the number of parameters that need to be trained and improves the search efficiency.

[0107] Step S3034, training process, specifically includes:

[0108] a. Controller sampling architecture: The RNN controller generates a description of the network architecture.

[0109] b. Construct sub-network: According to the output of the controller, construct the corresponding sub-network from the shared parameters.

[0110] c. Evaluate performance: Train the sub-network on a small batch of the training set. Evaluate the performance of the sub-network on the validation set.

[0111] d. Update the controller: Use the policy gradient method to update the controller based on the performance of the sub-network.

[0112] e. Update shared parameters: Use gradient descent method to update shared parameters.

[0113] Step S3035, iterative optimization, specifically includes: repeating step S3034, continuously improving the network architecture and model parameters.

[0114] Step S3036, final selection, specifically includes: selecting the architecture with the best performance in the search process; and training the complete model from scratch using the selected architecture.

[0115] Figure 4 is a schematic diagram of an Efficient Neural Architecture Search (ENAS) according to an embodiment of the present invention, such as Figure 4 As shown, the steps include:

[0116] Step S401, controller RNN.

[0117] Step S402, obtaining a hypernetwork with shared parameters by sampling the architecture.

[0118] Step S403: Perform training and evaluation to obtain a reward signal, wherein the reward signal is used to update the controller.

[0119] In the above step S304, the hyperparameter optimization specifically includes: combining the Bayesian Optimization and Hyperband algorithms to perform hyperparameter optimization.

[0120] As an optional embodiment, the principle of Bayesian Optimization (BO) includes:

[0121] Step S3041, define the objective function: f(x), where x is the hyperparameter configuration.

[0122] Step S3042, select a proxy model: usually a Gaussian process (GP) is used.

[0123] Step S3043, select an acquisition function: such as Expected Improvement (EI).

[0124] Step S3044, iterative optimization, specifically includes:

[0125] a. Use GP to fit the observed points.

[0126] b. Use the acquisition function to select the next evaluation point.

[0127] c. Evaluate new points and update the GP model.

[0128] Optionally, the core formula of the Bayesian Optimization (BO) algorithm is: t = argmax a EI(a|D 1:t-1 ), EI(a)=E[max(f(a)-f(a + ),0)], where a tis the hyperparameter selected at time t, D 1:t-1 is the previous observation, f(a + ) is the current best observed value.

[0129] As an optional embodiment, the principle of the Hyperband algorithm includes:

[0130] Step S3045, resource allocation: evenly allocate the total computing budget R to different branch brackets.

[0131] Step S3046, elimination round by round: In each bracket, train and evaluate the configurations round by round, and eliminate the configurations with poor performance.

[0132] Step S3047, resource reallocation: allocating the saved resources to configurations with better performance for more adequate training.

[0133] Figure 5 is a schematic diagram of a Hyperband algorithm process according to an embodiment of the present invention, such as Figure 5 As shown, the following steps are included:

[0134] Step S501, initialization configuration.

[0135] Step S502: Allocate resources.

[0136] Step S503: Evaluate performance.

[0137] Step S504: Eliminate configurations with poor performance.

[0138] Step S505, determine whether the budget limit is reached, if yes, execute step S506, if no, return to step S502.

[0139] Step S506, returning to the optimal configuration.

[0140] In the above step S305, the distributed training is implemented using the Horovod framework, and the Ring-AllReduce algorithm is used to optimize communication efficiency.

[0141] As an optional embodiment, the principle of the Ring-AllReduce algorithm includes:

[0142] Step S3051, data parallelism: each worker node has a complete copy of the model but processes a different subset of the data. Implemented in TensorFlow using tf.distribute.Strategy.

[0143] Step S3052, model parallelism: split the model onto different devices. Use tf.device() to specify computing devices for different layers.

[0144] Step S3053, parameter server architecture: centralized parameter storage and update. Use tf.distribute.experimental.ParameterServerStrategy.

[0145] Step S3054, decentralized training: such as Ring-AllReduce algorithm. Use Horovod framework and integrate it into TensorFlow.

[0146] Figure 6 is a schematic diagram of a Ring-AllReduce algorithm according to an embodiment of the present invention, such as Figure 6 As shown, the following steps are included:

[0147] Step S601: Node 1 transfers a gradient to node 2.

[0148] Step S602: Node 2 transfers the gradient to node 3.

[0149] Step S603: Node 3 transfers the gradient to node 4.

[0150] Step S604: Node 4 transfers the gradient to node 1.

[0151] In the above step S306, the learning rate is adaptively adjusted, and a Dynamic Learning Rate Adaptation (DLRA) method is proposed. The DLRA method adapts to changes in the training process by dynamically adjusting the learning rate.

[0152] As an optional embodiment, the DLRA method allows the learning rate to be automatically adjusted according to the change of the loss function, and the learning speed is adaptively changed during the training process. The implementation steps are as follows:

[0153] Step S3061, initialization: set the initial learning rate η 0 and the adaptive coefficient β.

[0154] Step S3062, each training step:

[0155] a. Calculate the current loss L t .

[0156] b. Calculate the loss change rate: δ L =(L t -L t-1 ) / L t-1 .

[0157] c. Update learning rate: η t =η 0 ×exp(-β×δL ).

[0158] Step S3063, apply the new learning rate to the optimizer.

[0159] Alternatively, if Figure 3 As shown in the figure, automated deployment and reasoning include the following steps:

[0160] Step S307, compressing and quantizing the model obtained through automatic training.

[0161] Step S308: model conversion.

[0162] Step S309: containerized deployment.

[0163] Step S310: cloud deployment and edge device deployment.

[0164] In the above step S307, the model is compressed and quantized, and TensorFlow Model Optimization Toolkit is used to compress the model, and Quantization-Aware Training (QAT) technology is adopted.

[0165] As an optional embodiment, the principle of model compression includes:

[0166] Step S3071, weight pruning: remove unimportant connections (weights) in the model and convert the sparse matrix into a dense matrix.

[0167] Step S3072, Quantization: The weight value in the original model is a floating point 32-bit number, which is converted to 8 bits.

[0168] Step S3073, knowledge distillation: transfer the knowledge of the original large model (teacher, with complex model structure and many weight parameters) to the target small model (student, with relatively simple model structure and few weight parameters). This process is called knowledge distillation.

[0169] Optionally, the core formula of knowledge distillation is: L = αH(y,σ(z s / T))+(1-α)H(σ(z t / T),σ(z s / T)), where y is the true label, z s / and z t / are the logits of the student and teacher models respectively, T is the temperature parameter, and α is the weight coefficient.

[0170] Step S3074, low-rank factorization: decompose the large weight matrix into the product of multiple small matrices.

[0171] In the above step S308, the model is converted, ONNX (Open Neural Network Exchange) is used as the intermediate representation, and TensorRT is used for model optimization and acceleration.

[0172] Optionally, model conversion is the process of converting a deep learning model from one framework to another or to a common intermediate representation. ONNX (Open Neural Network Exchange) is a widely used open format for representing machine learning models. Below I will explain in detail how to achieve model conversion, especially using ONNX as an intermediate representation.

[0173] As an optional example, we take ONNX model conversion as an example. ONNX defines a set of standard operators and data types that can represent the computational graphs of most deep learning models. The conversion process mainly includes: a) mapping the computational graph of the source framework to ONNX operators; b) converting the model parameters to a format supported by ONNX; c) generating an ONNX model file.

[0174] Through the above conversion process, the model conversion, verification and optimization are achieved. ONNX is used as an intermediate representation, so that the model can be migrated between different frameworks and hardware platforms.

[0175] Figure 7 is a schematic diagram of an ONNX graph structure according to an embodiment of the present invention, such as Figure 7 As shown, the following steps are included:

[0176] Step S701, input.

[0177] Step S702, convolutional layer.

[0178] Step S703, activating the layer.

[0179] Step S704, pooling layer.

[0180] Step S705, fully connected layer.

[0181] Step S706, output.

[0182] In the above step S309, containerized deployment is performed, using Docker container technology to encapsulate the model and dependencies, and Kubernetes is used for container orchestration and management.

[0183] As an optional embodiment, the principles of Docker container technology include:

[0184] Step S3091, containerization.

[0185] Optionally, Docker uses the Linux kernel's namespaces and control groups (cgroups) technology to create an independent container environment. Namespaces provide isolation of resources such as processes, networks, and file systems, while control groups are responsible for limiting the system resources that each container can use.

[0186] Step S3092, mirror layering.

[0187] Optionally, Docker images use a layered storage architecture. Each image consists of multiple read-only layers that are stacked together using a Union File System. This structure supports fast image building and efficient storage.

[0188] Step S3093, lightweight virtualization.

[0189] Optionally, unlike traditional virtual machines, Docker containers share the host's operating system kernel without running a full operating system. This makes containers start quickly and use less resources.

[0190] Step S3094, standardized packaging.

[0191] Alternatively, Docker provides a standardized way to package applications and their dependencies. Dockerfile defines the steps to build an image, ensuring consistency of the application in different environments.

[0192] Figure 8 is a schematic diagram of a Kubernetes deployment process according to an embodiment of the present invention, such as Figure 8 As shown in the figure, Kubernetes is an open source container orchestration platform for deploying, managing and scaling containerized applications, including the following steps:

[0193] Step S801, Docker image.

[0194] Step S802: Kubernetes Deployment, where Kubernetes Deployment is a core concept in Kubernetes and is used to manage Pods and Replica Sets to ensure high availability and stability of applications.

[0195] Step S803, Kubernetes Service, where Kubernetes Service is one of the core resource objects in Kubernetes, defines an access entry address for a service, and provides a unified access method for Pod.

[0196] Step S804, Ingress Controller, where Ingress Controller is a key component in Kubernetes, used to manage and configure routing rules for inbound network traffic.

[0197] Step S805: load balancing.

[0198] In the above embodiment of the present application, the role of Kubernetes is to manage and monitor the online running Docker objects. The operator only needs to configure, start, and stop Docker through the UI, and an alarm notification will be issued when the operation is abnormal.

[0199] In the above step S310, edge computing is deployed, using Tensor Flow Lite to implement mobile and embedded device deployment, and using Edge TPU to accelerate reasoning on edge devices.

[0200] As an optional embodiment, the implementation principle of TensorFlow Lite includes:

[0201] Step S3101, model conversion, specifically includes:

[0202] Convert the trained TensorFlow model to TensorFlow Lite format;

[0203] Use TOCO (TensorFlow Lite Optimizing Converter) tool for conversion;

[0204] Graph optimization is performed during the conversion process, such as operator fusion and constant folding.

[0205] Step S3102, quantification, specifically includes:

[0206] Convert model weights from 32-bit floating point numbers to 8-bit integers;

[0207] Supports multiple quantization methods: a) Post-training quantization: directly quantize the pre-trained model b) Quantization-aware training: simulate quantization effects during training.

[0208] Step S3103, operator optimization, specifically includes:

[0209] Operator implementation optimized for mobile devices;

[0210] Use SIMD instruction sets such as NEON and SSE to accelerate computing;

[0211] Optimizations for specific hardware (such as ARM processors).

[0212] Step S3104, memory optimization, specifically includes:

[0213] Use memory pools to manage temporary tensors to reduce memory allocation and deallocation overhead;

[0214] Supports memory mapping to load models and reduce memory usage.

[0215] Step S3105, inference engine, specifically includes:

[0216] A lightweight interpreter for performing model inference;

[0217] Supports selective registration of required operators to reduce binary file size.

[0218] Alternatively, if Figure 3 As shown in the figure, the whole process monitoring and early warning includes the following steps:

[0219] In step S311, the Open Telemetry collector collects information about cloud deployment and edge device deployment, and then executes steps S312, S313, and S314 respectively.

[0220] Step S312: distributed tracing.

[0221] Step S313: log aggregation.

[0222] Step S314, indicator collection.

[0223] Step S315, performing anomaly detection based on the output contents of step S312, step S313 and step S314.

[0224] Step S316, generating an early warning based on anomaly detection.

[0225] As an optional example, the whole process monitoring includes step S311 to step S314, using distributed tracing and log aggregation technology to achieve real-time monitoring of the entire CV model life cycle.

[0226] As an optional embodiment, the technical implementation of the whole process monitoring includes:

[0227] Distributed tracing: Use the Open Telemetry framework (also known as open telemetry).

[0228] Log aggregation: Use ELK (Elasticsearch, Logstash, Kibana) stack.

[0229] Metrics collection: Use Prometheus to collect time series data.

[0230] Fig. 9 is a schematic diagram of a monitoring system architecture according to an embodiment of the present invention, such as Fig. 9 As shown, the following steps are included:

[0231] Step S901, application service.

[0232] Step S902, Open Telemetry Collector, then execute step S903, step S904 and step S905.

[0233] Step S903: Use Jaeger to track data.

[0234] Step S904: Use Prometheus to collect monitoring indicators.

[0235] Step S905: use Elsticsearch to aggregate logs.

[0236] Step S906: Use Grafana to analyze and display the output results of step S903 and step S904.

[0237] Step S907: Use Kibana to analyze and display the output result of step S905.

[0238] In the above embodiment of the present application, tracking points will be performed on the application service (data will be recorded through specific code), and then the OpenTelemetryCollector will be used to centrally collect the data at the tracking point location, and then the subscription function will be used to send it to different types of storage and front-end combinations to display the data. The data stored in Jageer and Prometheus are mostly structured data and the real-time requirements are not very high. For example, a delay of several minutes is acceptable, and it can be displayed with Grafana. Elasticsearch is a high-speed search storage, which can be used with Kibana to display simple indicators in real time without delay.

[0239] As an optional example, the full flow control warning includes steps S315 and S3164, which use machine learning algorithms to perform anomaly detection and predictive maintenance based on the collected monitoring data.

[0240] As an optional embodiment, the technical implementation of full flow control warning includes:

[0241] Anomaly detection: Use Isolation Forest algorithm and AutoEncode algorithm.

[0242] Time series forecasting: using the Prophet model.

[0243] Alarm distribution: Alarms are distributed end-to-end through tools.

[0244] Fig.10 is a schematic diagram of an early warning process according to an embodiment of the present invention, such as Fig.10 As shown, the following steps are included:

[0245] Step S1001, monitoring data input.

[0246] Step S1002, abnormality detection, if the detection result is normal, execute step S1003, if the detection result is abnormal, execute step S1004.

[0247] Step S1003, continue monitoring.

[0248] Step S1004: alarm generation.

[0249] Step S1005: alarm aggregation.

[0250] Step S1006, notification is sent.

[0251] As an optional embodiment, the core principles of the Isolation Forest algorithm include:

[0252] 1) Basic idea: Outliers are more easily "isolated". Normal data points are usually densely distributed together, while outliers are often far away from these dense areas.

[0253] 2) Working mechanism: randomly select a feature; randomly select a split point between the maximum and minimum values ​​of the feature; divide the data into two parts according to this split point; repeat this process until each point is isolated.

[0254] 3) Anomaly score calculation: For each data point, record the average number of steps required to isolate it; anomalies usually require fewer steps to be isolated; convert the average number of steps into anomaly scores.

[0255] Fig.11 is a schematic diagram of an Isolation Forest algorithm according to an embodiment of the present invention, such as Fig.11 As shown, the following steps are included:

[0256] Step S1101, obtaining the original data set.

[0257] Step S1102, randomly select features.

[0258] Step S1103: randomly select a segmentation point.

[0259] Step S1104, split the data.

[0260] Step S1105, determine whether all points are isolated, if so, execute step S1106, if not, return to step S1102.

[0261] Step S1106, calculating the isolated path length of each point.

[0262] Step S1107: Convert to anomaly score.

[0263] Step S1108, identifying abnormal points.

[0264] Fig.12 is a schematic diagram of an AutoEncoder architecture according to an embodiment of the present invention, such as Fig.12 As shown, including:

[0265] a. Input Layer: Receives raw data.

[0266] b. Encoder Hidden Layers: Gradually compress data.

[0267] c. Bottleneck Layer: The most compressed representation of data.

[0268] d. Decoder Hidden Layers: gradually recover data.

[0269] e. Output Layer: reconstructed data.

[0270] It should be noted that if Fig.12 As shown in Figure 1, the bottleneck layer (highlighted in green) is the key part of the AutoEncoder, which contains the compressed representation of the input data. The input layer and the output layer (highlighted in pink) have the same dimension because the goal of the AutoEncoder is to reconstruct the input data.

[0271] In the above embodiments of the present application, the AutoEncoder architecture can learn the key features of the data and play an important role in anomaly detection. Normal data can usually be well reconstructed, while the reconstruction effect of abnormal data is poor, so it can be detected.

[0272] Alternatively, if Figure 3 As shown in the figure, the horizontal expansion of demand includes the following steps:

[0273] Step S317: After the warning is generated, new demand is input.

[0274] Step S318: Algorithm adaptability evaluation.

[0275] Step S319, determine whether the requirements are met, if yes, execute step S320, if not, execute step S321.

[0276] Step S320, directly apply.

[0277] Step S321, the algorithm library is expanded, and then steps S322 and S323 are executed.

[0278] Step S322, transfer learning.

[0279] Step S323: few-sample learning.

[0280] Step S324, based on the learning results of step S322 and step S323, the algorithm library is updated, and the updated algorithm library can be used as data input in the automated training process.

[0281] As an optional embodiment, the horizontal expansion requirement is realized through the algorithm library, and the design and update of the algorithm library are realized through the algorithm library design and management, as well as the version control and update mechanism.

[0282] As an optional example, the algorithm library structure is a basic CV task algorithm library. Each algorithm in the algorithm library supports the smallest granularity task, and each task is classified, managed and stored. These smallest task unit algorithms can be used in a "building blocks" manner to solve external complex requirements.

[0283] Fig.13 is a schematic diagram of an algorithm library structure according to an embodiment of the present invention. Fig.13 As shown, the algorithm library adopts a hierarchical structure design, in which the algorithm library includes at least: a basic model layer, a task-specific layer and a domain-specific layer, in which the basic model layer includes at least: ResNet, VGG and Inception; the task-specific layer includes at least: a classification model, a detection model and a segmentation model; the domain-specific layer includes at least: a medical impact model and an autonomous driving model.

[0284] As an optional example, version control is mainly divided into three levels: the first level corresponds to the algorithm task, the second level corresponds to the specific scenario, and the third level corresponds to the most fine-grained algorithm version. In the algorithm warehouse, the finest granularity of the algorithm is the version number (abstractly understood as the unique name of the algorithm), which is composed of non-repeating English characters plus numbers. For example: face_id.1.0.11.

[0285] Fig.14 is a schematic diagram of a version control structure according to an embodiment of the present invention, such as Fig.14 As shown, scenario analysis can be performed for the algorithm task to identify scenario 1, scenario 2, and scenario 3, where scenario 1 corresponds to algorithm version 3, scenario 2 corresponds to algorithm version 4, and scenario 3 corresponds to algorithm version 1 and algorithm version 2.

[0286] As an optional example, the update mechanism is used to update the algorithm model in the existing task scenario: you only need to upload the offline trained model to save the current model, and the system will automatically generate an algorithm id. Select the existing first-level and second-level tags for the algorithm. If there is no first-level or second-level tag, you can create a new first-level and second-level tag.

[0287] In the above step S318, the algorithm adaptability evaluation includes the following steps:

[0288] Step S3181, task similarity calculation.

[0289] It should be noted that the purpose of task similarity calculation is to evaluate the similarity between a new task and existing tasks in the algorithm library in order to select the most suitable basic algorithm or model for adaptation.

[0290] As an optional embodiment, the calculation principle of task similarity calculation is: a) feature space comparison: compare the feature distribution of new tasks and existing tasks; b) data statistical properties: compare the statistical characteristics of data sets (such as mean, variance, skewness, etc.); c) task goal similarity: evaluate the similarity of task goals (such as classification, regression, clustering, etc.).

[0291] As an optional example, the calculation formula for task similarity is: S(T1,T2)=w1*Sf(T1,T2)+w2*Sd(T1,T2)+w3*So(T1,T2), where S(T1,T2) is the overall similarity between tasks T1 and T2; Sf is the feature space similarity; Sd is the data statistical similarity; So is the task target similarity; w1, w2, w3 are weight coefficients.

[0292] Optionally, the feature space similarity Sf can be calculated using cosine similarity: Sf(T1, T2) = (F1·F2) / (||F1||*||F2||), where F1 and F2 are feature vectors of tasks T1 and T2.

[0293] Step S3182, rapid adaptability evaluation, aims to quickly determine whether the selected algorithm can adapt to the new task within a reasonable time and resource range.

[0294] Optionally, the steps of rapid adaptability evaluation include: a) sampling: extracting representative samples from the new task dataset; b) model adaptation: quickly adapting the model using transfer learning or fine-tuning techniques; c) performance evaluation: evaluating the performance of the adapted model on the validation set; d) resource consumption analysis: evaluating the time and computing resource consumption of the adaptation process.

[0295] Optionally, the evaluation indexes of the rapid adaptability evaluation include at least: a performance index, an efficiency index and an adaptability index.

[0296] Optionally, the performance indicators include at least: accuracy, F1 score, mean square error (MSE) and area under the curve (AUC).

[0297] Optionally, the efficiency indicators include at least: training time, inference time, memory usage, and GPU / CPU utilization.

[0298] Optionally, the adaptability index includes at least: learning curve slope, convergence speed and generalization ability (consistency of performance on different subsets).

[0299] Fig.15 is a core architecture diagram of a rapid adaptability evaluation according to an embodiment of the present invention. Fig.15 As shown, the steps include:

[0300] Step S1501, obtaining a new task / data set.

[0301] Step S1502: task similarity calculation.

[0302] Step S1503: Select the most similar existing algorithm.

[0303] Step S1504, data sampling.

[0304] Step S1505: fast model adaptation.

[0305] Step S1506: performance evaluation.

[0306] Step S1507: resource consumption analysis.

[0307] Step S1508, determine whether the adaptability standard is met, if so, execute step S1509, if not, execute step S1511.

[0308] Step S1509, perform complete adaptation.

[0309] Step S1510, deployment and monitoring.

[0310] Step S1511, try the next similarity algorithm, and then return to step S1503.

[0311] In the above step S322, transfer learning uses the knowledge learned in the source domain to improve the learning effect of the target domain. It is based on the assumption that there is transferable knowledge between different domains.

[0312] As an optional embodiment, the loss function formula of transfer learning can be a typical transfer learning loss function, expressed as: L = Ltask(θ) + λLtransfer(θ), where Ltask(θ) is the loss of the target task, Ltransfer(θ) is the transfer-related loss (such as domain adaptation loss), and λ is a hyperparameter for balancing the two losses.

[0313] Fig.16 is a schematic diagram of transfer learning according to an embodiment of the present invention, such as Fig.16 As shown, the steps include:

[0314] Step S1601, obtaining a pre-trained model.

[0315] Step S1602, the special case extractor is frozen.

[0316] Step S1603, adding a new task-related layer.

[0317] Step S1604, fine-tuning on the target data.

[0318] Step S1605, evaluating model performance.

[0319] Step S1606, determine whether the performance meets the requirements, if yes, execute step S1607, otherwise execute step S1608.

[0320] Step S1607, deploy the model.

[0321] Step S1608, adjust the migration strategy, and then return to step S1602.

[0322] In the above step S323, few-sample learning aims to learn from a very small number of labeled samples, usually using the idea of ​​meta-learning to learn "how to learn".

[0323] As an optional embodiment, the few-shot learning takes the prototypical network as an example, and its loss function can be expressed as: L = -log(p(y|x)), where p(y|x) = exp(-d(f(x),c y )) / ∑ k exp(-d(f(x),c k ), f(x) is the embedding of sample x, cy is the prototype of category y, and d is the distance function (such as Euclidean distance).

[0324] Fig.17 is a schematic diagram of a few-sample learning according to an embodiment of the present invention, such as Fig.17 As shown, the steps include:

[0325] Step S1701, obtaining a large-scale auxiliary data set.

[0326] Step S1702: meta-learning stage.

[0327] Step S1703, learning generalization ability.

[0328] Step S1704, a small amount of target task data.

[0329] Step S1705, rapid adaptation.

[0330] Step S1706, evaluating model performance.

[0331] Step S1707, determine whether the performance meets the requirements, if so, execute step S1708, otherwise execute step S1709.

[0332] Step S1708, deploy the model.

[0333] Step S1709, adjust the meta-learning strategy, and then return to step S1702.

[0334] It should be noted that both transfer learning and few-shot learning are effective methods for dealing with data scarcity, and are often used in combination in practice to achieve the best results. For example, you can first use transfer learning to build a good feature extractor, and then apply few-shot learning techniques on this basis to adapt to new categories or tasks.

[0335] Optional example 1: Smart campus security management.

[0336] Application background: In smart park scenarios, it is necessary to identify safety hazards in real time, such as crowd gathering, helmet wearing, flame detection, etc., and respond to emergencies in a timely manner. The automated identification system can significantly improve the response speed and accuracy.

[0337] Implementation process:

[0338] Data preprocessing and annotation: We used TensorFlow Data Validation to verify the data and collected 100,000 images of scenes containing people and flames. We automatically enhanced the data based on AutoAugment, which increased the data enhancement accuracy to 93%.

[0339] Automatic model training: Combined with the Neural Architecture Search (NAS) method, the network architecture is optimized and searched. Through the ENAS algorithm, the number of sub-network training times generated by the controller is reduced by 20%, and the accuracy rate is finally achieved. Bayesian Optimization is used to optimize the hyperparameters, adjust the learning rate and batch size, and the training time is shortened by 30%.

[0340] Distributed deployment: On edge devices (such as NVIDIA edge boxes), the Ring-AllReduce algorithm is used to implement distributed training, improve communication efficiency, and control inference latency within 50ms.

[0341] Model compression: Using Quantization-Aware Training (QAT) technology, the model size is compressed to 60% of the original size and the inference speed is increased by 2 times.

[0342] Optional example 2: Forklift identification in warehouse management.

[0343] Application background: In large-scale warehouse management scenarios, the storage and access of goods (such as steel coils) need to be monitored in real time to ensure the accuracy and safety of forklift operations.

[0344] Implementation process:

[0345] Data collection and labeling: Automatically collect and label 50,000 forklift operation video frames, and use semi-supervised learning to reduce 80% of manual labeling time; data enhancement strategies include rotation, color transformation, etc., and the accuracy of enhanced data is increased to 95.3%.

[0346] Automated model training: Based on the adaptive learning rate adjustment (DLRA) method, the model converges quickly with a dynamic learning rate. The learning rate decays by 30% after each training, and the model optimization is completed within 10 hours, with a recognition accuracy of 95.5%.

[0347] Hyperparameter optimization: Hyperband algorithm is used to allocate resources, automatically eliminate poorly performing model parameter configurations, and use the optimal configuration for final model training. This process reduces hyperparameter search time by about 40% and improves resource utilization by 60%.

[0348] Edge inference and compressed deployment: Using quantized models on SE5 edge devices, the inference speed is shortened to 30ms / frame. Through ONNX conversion and TensorRT optimization, the model can be deployed across devices and adapted to more than 99% of hardware.

[0349] Optional example 3: Safety warning in petrochemical park.

[0350] Application background: Petrochemical parks have major safety hazards such as fire. To achieve fire warning through video algorithms and automated AI platforms, real-time detection of helmet wearing and flame recognition are required.

[0351] Implementation process:

[0352] Data collection and enhancement: 120,000 images of fire and human behavior were collected and enhanced using AutoAugment. The validation set performance of the enhancement strategy increased by 6%. The data volume increased to 1.5 times the original amount.

[0353] Transfer learning and distillation training: Using pre-trained models for transfer learning, the amount of data training was reduced to 1 / 3 of the original, and the model accuracy reached 96%. Using knowledge distillation, the feature extraction of the large teacher model was transferred to the lightweight student model, reducing the model size by 40%.

[0354] Automated deployment: Model deployment is done through Docker containers and Kubernetes, and deployment and updates can be completed within 5 seconds. The average inference time is within 45ms, which is adapted to the computing power of edge box devices.

[0355] Real-time monitoring and abnormal warning: Based on the Prophet model, time series prediction and real-time warning are realized, and indicator changes are detected 3 seconds before abnormal events occur, with a warning accuracy rate of 93%.

[0356] According to an embodiment of the present invention, a model training device embodiment is also provided. It should be noted that the model training device can be used to execute the model training method in the embodiment of the present invention, and the model training method in the embodiment of the present invention can be executed in the model training device.

[0357] Fig.18 is a schematic diagram of a model training device according to an embodiment of the present invention. Fig.18As shown, the device may include: an acquisition module 1802, used to acquire a model training task, wherein the model training task includes at least: a training data set and a task requirement; a query module 1804, used to query at least one candidate model that meets the task requirement from a plurality of preset models recorded in a preset algorithm library; an evaluation module 1806, used to evaluate the model performance of each candidate model using a representative data set in the training data set, wherein the training data set includes a plurality of preset sample data, the representative data set includes a plurality of representative sample data, the representative sample data is preset sample data that can represent the data features of the training data set, the number of representative sample data in the representative data set is less than the number of preset sample data in the training data set, the plurality of representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain model performance; a training module 1808, used to use the training data set to train a candidate model whose model performance meets a preset adaptability standard to obtain a target model.

[0358] It should be noted that the acquisition module 1802 in this embodiment can be used to execute step S102 in the embodiment of the present application, the query module 1804 in this embodiment can be used to execute step S104 in the embodiment of the present application, the evaluation module 1806 in this embodiment can be used to execute step S106 in the embodiment of the present application, and the training module 1808 in this embodiment can be used to execute step S108 in the embodiment of the present application. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments.

[0359] In an embodiment of the present invention, when it is necessary to perform model training based on a training data set and task requirements in a model training task, at least one candidate model can be screened from a plurality of preset models recorded in a preset algorithm library according to the task requirements in the model training task, and then a small number of samples that can represent the data features in the training data set are selected from the training data set as a representative data set, and then the representative data set is split into a pre-training set used to train the candidate models, and a pre-verification set for performance evaluation of the trained candidate models, and then based on the model performance of each candidate model, the candidate models can be screened according to the model performance, and then the selected candidate models are simply trained using the training data set to obtain the target model requested for training by the model training task, thereby using a representative data set with a small amount of data to screen the candidate models, and the candidate models that meet the requirements of the model training task can be quickly determined, and then training based on the candidate models can ensure training efficiency and accuracy, thereby achieving the technical effect of reducing the complexity of model training, and thus solving the technical problem of low model training efficiency.

[0360] As an optional embodiment, the query module includes: a first identification unit, used to identify the model requirement label indicated by the model training task, wherein the model requirement label is used to indicate the algorithm model required for the model training task; a first query unit, used to query at least one candidate model that matches the model requirement label among multiple preset models recorded in a preset algorithm library, wherein the preset algorithm library is also used to record the model preset label corresponding to each preset model.

[0361] As an optional embodiment, the query module includes: an analysis unit, which is used to analyze the task requirements and decompose the model training task into multiple model training subtasks, wherein the task requirements are at least used to describe multiple target sub-models that constitute the target model, and each model training subtask is used to indicate the selection of a target sub-model in a preset algorithm library; a second identification unit, which is used to identify the sub-model requirement label indicated by each model training subtask, wherein the sub-model requirement label is at least used to indicate the target sub-model type required for the model training subtask; a second query unit, which is used to query the target sub-model type corresponding to each sub-model requirement label in the multiple preset sub-model types recorded in the preset algorithm library, wherein the multiple preset models of the preset algorithm library are pre-divided into multiple preset sub-model types, and each preset sub-model type has a corresponding sub-model preset label; a combination unit, which is used to combine the preset models in the multiple target sub-model types to obtain at least one candidate model.

[0362] As an optional embodiment, the combination unit includes: a determination subunit, used to determine the combination order of multiple target sub-model types according to multiple model training sub-tasks decomposed from the model training task, wherein the combination order represents the structural order of the multiple model training sub-tasks in the model training task; a combination subunit, used to randomly select preset models from at least one preset model corresponding to each target sub-model type according to the combination order, and combine them to obtain at least one candidate model.

[0363] As an optional embodiment, the device also includes: an identification submodule, which is used to identify the candidate training tasks performed by training each candidate model after querying at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library, wherein the candidate training tasks are at least used to describe the feature space, data statistical attributes and task target similarity used to train the candidate models; a determination submodule, which is used to determine the task similarity between the model training task and the candidate training task, wherein the model training task is also used to describe the feature space, data statistical attributes and task target similarity used by the training target model, and the task similarity is based on the feature space similarity, data statistical similarity and task target similarity of the model training task and the candidate training task, the feature space similarity is determined based on the feature space of the model training task and the candidate training task, the data statistical similarity is determined based on the data statistical attributes of the model training task and the candidate training task, and the task target similarity is determined based on the task target similarity of the model training task and the candidate training task; a deletion submodule, which is used to delete the candidate models whose task similarity does not meet the preset task similarity threshold.

[0364] As an optional embodiment, the training module includes: a determination unit, used to determine a candidate model whose model performance meets a preset adaptability standard as a model to be trained; a migration unit, used to use a representative data set in the training data set to perform transfer learning on the model to be trained to obtain a target model.

[0365] As an optional embodiment, the device also includes: an acquisition submodule, which is used to acquire a historical model trained historically before querying at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library, wherein the historical model is a combination of multiple historical sub-models with different functions; a decomposition submodule, which is used to decompose the historical model to obtain multiple historical sub-models and the dependency relationship corresponding to each historical sub-model; an encapsulation submodule, which is used to containerize each historical sub-model and the dependency relationship corresponding to the historical sub-model to obtain a model container corresponding to each historical sub-model; and a storage submodule, which is used to store each model container in the preset algorithm library, wherein the historical sub-model encapsulated in each model container is a preset model in the preset algorithm library.

[0366] An embodiment of the present invention may provide an electronic device, which may be a computer terminal, and the computer terminal may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.

[0367] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.

[0368] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the model training method: obtaining a model training task, wherein the model training task includes at least: a training data set and task requirements; querying at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library; using a representative data set in the training data set to evaluate the model performance of each candidate model, wherein the training data set includes multiple preset sample data, the representative data set includes multiple representative sample data, the representative sample data is preset sample data that can represent the data characteristics of the training data set, the number of representative sample data in the representative data set is less than the number of preset sample data in the training data set, and the multiple representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain model performance; using the training data set, training the candidate model whose model performance meets the preset adaptability standard to obtain the target model.

[0369] Fig.19 is a structural block diagram of a computer terminal according to an embodiment of the present invention. Fig.19 As shown, the computer terminal 1900 may include: one or more (only one is shown in the figure) processors 1902 , and a memory 1904 .

[0370] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the model training method and device in the embodiment of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned model training method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories can be connected to the terminal 1900 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0371] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain a model training task, wherein the model training task includes at least: a training data set and task requirements; query at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library; use a representative data set in the training data set to evaluate the model performance of each candidate model, wherein the training data set includes multiple preset sample data, the representative data set includes multiple representative sample data, the representative sample data is preset sample data that can represent the data characteristics of the training data set, the number of representative sample data in the representative data set is less than the number of preset sample data in the training data set, and the multiple representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain model performance; use the training data set to train candidate models whose model performance meets the preset adaptability standard to obtain a target model.

[0372] Optionally, the processor may also execute the program code of the following steps: identifying the model requirement label indicated by the model training task, wherein the model requirement label is used to indicate the algorithm model required for the model training task; and querying at least one candidate model that matches the model requirement label among multiple preset models recorded in the preset algorithm library, wherein the preset algorithm library is also used to record the model preset label corresponding to each preset model.

[0373] Optionally, the processor may also execute the program code of the following steps: analyzing task requirements, decomposing the model training task into multiple model training subtasks, wherein the task requirements are at least used to describe multiple target sub-models constituting the target model, and each model training subtask is used to indicate the selection of a target sub-model in a preset algorithm library; identifying a sub-model requirement label indicated by each model training sub-task, wherein the sub-model requirement label is at least used to indicate the target sub-model type required for the model training sub-task; querying the target sub-model type corresponding to each sub-model requirement label among the multiple preset sub-model types recorded in the preset algorithm library, wherein the multiple preset models of the preset algorithm library are pre-divided into multiple preset sub-model types, and each preset sub-model type has a corresponding sub-model preset label; combining the preset models in the multiple target sub-model types to obtain at least one candidate model.

[0374] Optionally, the processor may also execute the program code of the following steps: determining a combination order of multiple target sub-model types based on multiple model training sub-tasks decomposed from the model training task, wherein the combination order represents the structural order of the multiple model training sub-tasks in the model training task; and randomly selecting a preset model from at least one preset model corresponding to each target sub-model type according to the combination order to combine and obtain at least one candidate model.

[0375] Optionally, the processor may also execute program code for the following steps: identifying candidate training tasks performed for training each candidate model, wherein the candidate training tasks are at least used to describe the feature space, data statistical attributes, and task objective similarity used to train the candidate model; determining task similarity between the model training task and the candidate training task, wherein the model training task is also used to describe the feature space, data statistical attributes, and task objective similarity used to train the target model, the task similarity is based on the feature space similarity, data statistical similarity, and task objective similarity between the model training task and the candidate training task, the feature space similarity is determined based on the feature space of the model training task and the candidate training task, the data statistical similarity is determined based on the data statistical attributes of the model training task and the candidate training task, and the task objective similarity is determined based on the task objective similarity between the model training task and the candidate training task; and deleting candidate models whose task similarity does not meet a preset task similarity threshold.

[0376] Optionally, the processor may also execute the program code of the following steps: determining a candidate model whose model performance meets a preset adaptability standard as a model to be trained; performing transfer learning on the model to be trained using a representative data set in the training data set to obtain a target model.

[0377] Optionally, the processor may also execute the program code of the following steps: obtaining a historical model trained historically, wherein the historical model is a combination of multiple historical sub-models with different functions; decomposing the historical model to obtain multiple historical sub-models and dependency relationships corresponding to each historical sub-model; containerizing each historical sub-model and the dependency relationships corresponding to the historical sub-models to obtain a model container corresponding to each historical sub-model; storing each model container in a preset algorithm library, wherein the historical sub-model encapsulated in each model container is a preset model in the preset algorithm library.

[0378] It can be understood by those skilled in the art that Fig.19 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig.19The structure of the electronic device is not limited. For example, the computer terminal 1900 may also include Fig.19 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig.19 Different configurations are shown.

[0379] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a computer program. The computer program can be stored in a non-volatile medium. The non-volatile storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0380] The embodiment of the present invention further provides a non-volatile storage medium. Optionally, in this embodiment, the non-volatile storage medium can be used to store the program code executed by the model training method provided in the above embodiment.

[0381] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0382] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a model training task, wherein the model training task includes at least: a training data set and task requirements; querying at least one candidate model that meets the task requirements among multiple preset models recorded in a preset algorithm library; using a representative data set in the training data set to evaluate the model performance of each candidate model, wherein the training data set includes multiple preset sample data, the representative data set includes multiple representative sample data, the representative sample data is preset sample data that can represent the data characteristics of the training data set, the number of representative sample data in the representative data set is less than the number of preset sample data in the training data set, and the multiple representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain model performance; using the training data set, training candidate models whose model performance meets the preset adaptability standard to obtain a target model.

[0383] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: identifying a model requirement tag indicated by a model training task, wherein the model requirement tag is used to indicate an algorithm model required for the model training task; and querying at least one candidate model that matches the model requirement tag among multiple preset models recorded in a preset algorithm library, wherein the preset algorithm library is also used to record a model preset tag corresponding to each preset model.

[0384] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: analyzing task requirements, decomposing the model training task into multiple model training subtasks, wherein the task requirements are at least used to describe multiple target sub-models that constitute the target model, and each model training subtask is used to indicate the selection of a target sub-model in a preset algorithm library; identifying a sub-model requirement label indicated by each model training subtask, wherein the sub-model requirement label is at least used to indicate the target sub-model type required for the model training subtask; querying the target sub-model type corresponding to each sub-model requirement label among the multiple preset sub-model types recorded in the preset algorithm library, wherein the multiple preset models of the preset algorithm library are pre-divided into multiple preset sub-model types, and each preset sub-model type has a corresponding sub-model preset label; combining the preset models in the multiple target sub-model types to obtain at least one candidate model.

[0385] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: determining a combination order of multiple target sub-model types based on multiple model training sub-tasks decomposed from the model training task, wherein the combination order represents the structural order of the multiple model training sub-tasks in the model training task; and randomly selecting preset models from at least one preset model corresponding to each target sub-model type according to the combination order to combine and obtain at least one candidate model.

[0386] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: identifying candidate training tasks performed for training each candidate model, wherein the candidate training tasks are at least used to describe the feature space, data statistical attributes and task objective similarity used to train the candidate model; determining the task similarity between the model training task and the candidate training task, wherein the model training task is also used to describe the feature space, data statistical attributes and task objective similarity used to train the target model, the task similarity is based on the feature space similarity, data statistical similarity and task objective similarity of the model training task and the candidate training task, the feature space similarity is determined based on the feature space of the model training task and the candidate training task, the data statistical similarity is determined based on the data statistical attributes of the model training task and the candidate training task, and the task objective similarity is determined based on the task objective similarity of the model training task and the candidate training task; deleting candidate models whose task similarity does not meet a preset task similarity threshold.

[0387] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: determining a candidate model whose model performance meets a preset adaptability standard as a model to be trained; using a representative data set in the training data set to perform transfer learning on the model to be trained to obtain a target model.

[0388] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: obtaining a historical model trained historically, wherein the historical model is a combination of multiple historical sub-models with different functions; decomposing the historical model to obtain multiple historical sub-models and dependency relationships corresponding to each historical sub-model; containerizing each historical sub-model and the dependency relationships corresponding to the historical sub-models to obtain a model container corresponding to each historical sub-model; storing each model container in a preset algorithm library, wherein the historical sub-model encapsulated in each model container is a preset model in the preset algorithm library.

[0389] The embodiment of the present invention further provides a computer program product, including a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the steps of the model training method provided in the above embodiment are implemented.

[0390] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0391] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0392] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0393] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0394] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0395] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a non-volatile storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned non-volatile storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0396] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A model training method, characterized in that: include: Obtaining a model training task, wherein the model training task at least includes: a training data set and a task requirement; Query at least one candidate model that meets the task requirements among multiple preset models recorded in the preset algorithm library; Use a representative data set in the training data set to evaluate the model performance of each candidate model, wherein the training data set includes a plurality of preset sample data, the representative data set includes a plurality of representative sample data, the representative sample data is the preset sample data that can represent the data features of the training data set, the number of the representative sample data in the representative data set is less than the number of the preset sample data in the training data set, the plurality of representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain the model performance; The candidate model whose model performance meets the preset adaptability standard is trained using the training data set to obtain a target model.

2. The method according to claim 1, characterized in that Among the multiple preset models recorded in the preset algorithm library, searching for at least one candidate model that meets the task requirement includes: Identify a model requirement tag indicated by the model training task, wherein the model requirement tag is used to indicate an algorithm model required by the model training task; Among the multiple preset models recorded in the preset algorithm library, at least one candidate model matching the model requirement label is queried, wherein the preset algorithm library is also used to record the model preset label corresponding to each preset model.

3. The method according to claim 1, characterized in that: Among the multiple preset models recorded in the preset algorithm library, searching for at least one candidate model that meets the task requirement includes: Analyze the task requirements and decompose the model training task into a plurality of model training subtasks, wherein the task requirements are at least used to describe a plurality of target submodels constituting the target model, and each of the model training subtasks is used to indicate the selection of the target submodel in the preset algorithm library; Identify a sub-model requirement tag indicated by each of the model training sub-tasks, wherein the sub-model requirement tag is at least used to indicate a target sub-model type required by the model training sub-task; Among the multiple preset sub-model types recorded in the preset algorithm library, query the target sub-model type corresponding to each sub-model requirement tag, wherein the multiple preset models in the preset algorithm library are pre-divided into multiple preset sub-model types, and each preset sub-model type has a corresponding sub-model preset tag; The preset models in a plurality of the target sub-model types are combined to obtain at least one candidate model.

4. The method according to claim 3, characterized in that: Combining the preset models in the plurality of target sub-model types to obtain at least one candidate model comprises: Determine a combination order of a plurality of the target sub-model types according to a plurality of the model training sub-tasks decomposed from the model training task, wherein the combination order represents a structural order of the plurality of the model training sub-tasks in the model training task; According to the combination order, the preset models are randomly selected from at least one preset model corresponding to each target sub-model type to be combined to obtain at least one candidate model.

5. The method according to claim 1, characterized in that After searching for at least one candidate model that meets the task requirement among a plurality of preset models recorded in the preset algorithm library, the method further includes: Identifying candidate training tasks performed to train each of the candidate models, wherein the candidate training tasks are used to at least describe the feature space, data statistical properties, and task goal similarity used to train the candidate models; Determine the task similarity between the model training task and the candidate training task, wherein the model training task is also used to describe the feature space, data statistical attributes and task target similarity used to train the target model, and the task similarity is based on the feature space similarity, data statistical similarity and task target similarity between the model training task and the candidate training task, the feature space similarity is determined based on the feature space of the model training task and the candidate training task, the data statistical similarity is determined based on the data statistical attributes of the model training task and the candidate training task, and the task target similarity is determined based on the task target similarity between the model training task and the candidate training task; The candidate models whose task similarity does not meet the preset task similarity threshold are deleted.

6. The method according to claim 1, characterized in that Using the training data set, training the candidate model whose model performance meets the preset adaptability standard, and obtaining the target model includes: Determine the candidate model whose model performance meets the preset adaptability standard as the model to be trained; Using the representative data set in the training data set, transfer learning is performed on the model to be trained to obtain the target model.

7. The method according to claim 1, characterized in that Before searching for at least one candidate model that meets the task requirement among a plurality of preset models recorded in the preset algorithm library, the method further includes: Acquire a historical model trained historically, wherein the historical model is a combination of multiple historical sub-models with different functions; Decomposing the historical model to obtain a plurality of historical sub-models and dependency relationships corresponding to each of the historical sub-models; Containerize and encapsulate each of the historical sub-models and the dependency relationship corresponding to the historical sub-models to obtain a model container corresponding to each of the historical sub-models; Each of the model containers is stored in the preset algorithm library, wherein the historical sub-model encapsulated in each of the model containers is the preset model in the preset algorithm library.

8. A model training device, characterized in that: include: An acquisition module, used to acquire a model training task, wherein the model training task at least includes: a training data set and a task requirement; A query module, used to query at least one candidate model that meets the task requirements among multiple preset models recorded in a preset algorithm library; An evaluation module, used to evaluate the model performance of each candidate model using a representative data set in the training data set, wherein the training data set includes a plurality of preset sample data, the representative data set includes a plurality of representative sample data, the representative sample data is the preset sample data that can represent the data features of the training data set, the number of the representative sample data in the representative data set is less than the number of the preset sample data in the training data set, the plurality of representative sample data in the representative data set are divided into a pre-training set and a pre-verification set, the representative sample data in the pre-training set is used to train each candidate model, and the pre-verification set is used to perform performance evaluation on each trained candidate model to obtain the model performance; The training module is used to use the training data set to train the candidate model whose model performance meets the preset adaptability standard to obtain the target model.

9. An electronic device, comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the model training method described in any one of claims 1 to 7 through the computer program.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the model training method described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Model training method and device, computer equipment, storage medium and program product

    CN121459092A