Ai model reuse method and apparatus
By introducing category mapping costs and model multiplexing weights in AI model multiplexing, the problem of poor model multiplexing in the prior art is solved, and better model multiplexing effect and faster training process are achieved.
Patent Information
- Application Number
- PCT/CN2024/102579
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-06-28
- Publication Date
- 2025-05-30
AI Technical Summary
Among the existing AI model reuse methods, the model reuse effect is poor, resulting in poor processing effect.
By introducing category mapping costs, model multiplexing weights are determined, and model parameters are updated based on the weights and target training data sets, an AI model suitable for target tasks is obtained.
The effect of model reuse is improved, making the processing effect of the target model better, and can accelerate the model training process, saving computing power and time overhead.
Smart Images

Figure CN2024102579_30052025_PF_FP_ABST
Abstract
Description
AI model reuse method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 24, 2023, with application number 202311594582.6 and application name “A method and device for reusing an AI model”, all of the contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence technology, and in particular to a method and device for reusing an AI model. Background Art
[0003] In recent years, the application of artificial intelligence (AI) technology has become more and more extensive, and how to obtain a better AI model is crucial.
[0004] Currently, many companies and research centers have built large-scale model libraries that include a large number of pre-trained models. Pre-trained models are AI models trained based on previous tasks (both general and specific). For new tasks, pre-trained models in the model library can be reused to obtain an AI model for the new task. In other words, a pre-trained model in the model library is used as the initialization model for the new task, and then this initialization model is optimized to obtain an AI model suitable for the new task.
[0005] In some existing AI model reuse methods, the processing effect of the AI model obtained based on the pre-trained model is low, that is, the existing AI model reuse method has the problem of poor model reuse effect.
[0006] Summary of the Invention
[0007] This application provides a method and device for reusing an AI model, which can improve the effect of model reuse.
[0008] This application adopts the following technical solutions:
[0009] In the first aspect, the present application provides a method for reusing an AI model, including: using a first AI model as an initialization model for a target task, and using the first AI model to process each sample data in a target training data set, the first AI model is an AI model trained according to a source training data set of a source task; and based on the processing result of each sample data in the target training data set, determining the category mapping cost in which the category of the target task is mapped to the category of the source task; and determining, based on the category mapping cost, a model reuse weight from the first AI model to a second AI model used to process the target task; and then obtaining the second AI model based on the model reuse weight and the target training data set.
[0010] The AI model reuse method provided in this application, since the category mapping cost between the category of the source task and the category of the target task can reflect the reusability between the source model used to process the source task (i.e., the above-mentioned first AI model) and the target model used to process the target task (i.e., the above-mentioned second AI model), therefore, using the category mapping cost as a consideration factor for model reuse can guide the model training process to better reuse the features of the pre-trained model and obtain a target model with better processing effect, that is, this method can improve the effect of model reuse.
[0011] Furthermore, model reuse based on category mapping cost can also guide the model training process to converge faster, which can save computing power and time overhead to a certain extent.
[0012] In one possible implementation, obtaining the second AI model based on the model reuse weight and the target training data set specifically includes: determining the model loss based on the model reuse weight and the target training data set; and updating the parameters of the first AI model based on the model loss to obtain the second AI model.
[0013] In one possible implementation, determining the category mapping cost for mapping the target task category to the source task category based on the processing results of each sample data in the target training dataset specifically includes: determining the detection accuracy of detecting the target task category as the source task category based on the category confidence of the target object in each sample data in the processing results; and determining the category mapping cost for mapping the target task category to the source task category based on the detection accuracy. Detection accuracy is positively correlated with category confidence, i.e., the higher the category confidence, the higher the detection accuracy.
[0014] In one possible implementation, the cost of mapping the target task's category to the source task's category is negatively correlated with the accuracy of detecting the target task's category as the source task's category. That is, the higher the accuracy of detecting the target task's category as the source task's category, the lower the cost of mapping the target task's category to the source task's category.
[0015] In one possible implementation, the category of the target task is mapped to the category of the source task with a category mapping cost that satisfies:
[0016] Among them, C represents the category mapping cost, which is a matrix of size d1×d2, d1 represents the number of categories of the source task (the number of detection types contained in the source task), and d2 represents the number of categories of the target task (that is, the number of detection types contained in the target task). ijIndicates the mapping cost of mapping the target task category j to the source task category i, i = 1, 2, .... d.1., j = 1, 2, ..., d2.
[0017] In one possible implementation, the model reuse weight is:
[0018] Where d1 represents the number of categories of the source task, that is, the number of detection types contained in the source task, d2 represents the number of categories of the target task, that is, the number of detection types contained in the target task, i = 1, 2, ..., d1, j = 1, 2, ..., d2, w ij Indicates the model reuse weight from the first AI model to the second AI model when the category j of the target task is detected as the category i of the source task.
[0019] The above model reuse weight W satisfies:
[0020] stW×1=p,1×W=q
[0021] Among them, c ij It represents the category mapping cost of mapping the target task category j to the source task category i, p represents the probability distribution of the importance of the sample data of the target objects of each category in the source training data set, and q represents the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model.
[0022] In this application, the process of reusing the first AI model of the source task to the target task is actually to find a transmission scheme with the minimum overall transmission cost (i.e., the minimum mapping cost) from one distribution to another. The transmission scheme is the above-mentioned model reuse weight (W). The meaning of W×1=p in the above-mentioned constraint condition is: to enable the first AI model's detection capability for various categories of the source task to continue to be used in the second AI model, that is, the knowledge learned in the first AI model for detecting various categories is transferred to the second AI model to the greatest extent of the source distribution. The meaning of the above-mentioned constraint condition 1×W=q is: after the first AI model is reused in the second AI model, the second AI model's detection effect on various categories of the target task is as similar as possible, that is, it is expected that the second AI model can achieve the best possible level in each category of the target task.
[0023] In one possible implementation, the above-mentioned determination of model loss based on model reuse weight and target training data set specifically includes: mapping the first category confidence of the target object to the second category confidence of the target object based on the model reuse weight, where the first category confidence of the target object is obtained by processing the sample data in the target training data set using the first AI model; and determining the model loss based on the second category confidence.
[0024] The above-mentioned model loss can be determined based on the second category confidence of the target object and the true category confidence of the target object, and the value of the category loss function (ie, the category loss value) is determined. That is, the model loss is the category loss value.
[0025] Optionally, in a target detection scenario, the processing results of the sample data include the category confidence of the target object in the sample data and the location information of the target object. The location information of the target object can be the coordinate information of the detection box used to mark the target object, and the model loss can be the sum of the category loss and the location loss. The location loss is the value of the location loss function (i.e., the location loss value) determined based on the predicted location information of the target object and the actual location information of the target object in the sample data.
[0026] In one possible implementation, the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model is uniformly distributed, so that the second AI model can achieve the best possible level in each category of the target task.
[0027] In one possible implementation, the AI model reuse method provided in the present application also includes: determining the reuse accuracy of the first AI model, the reuse accuracy of the first AI model is used to determine whether to use the first AI model as the initialization model for the target task, and the reuse accuracy of the first AI model is the detection accuracy of the target task category being detected as the source task category when the first AI model is used to process the target training data set.
[0028] When the detection accuracy of the first AI model is greater than the accuracy threshold, the first AI model is used as the initialization model for the second AI model. When the detection accuracy of the first AI model is less than or equal to the accuracy threshold, the first AI model is not used as the initialization model for the second AI model, and another model (such as a randomly initialized model) can be selected as the initialization model for the second AI model. Determining the initialization model for the second AI model based on the detection accuracy helps to obtain a second AI model with better performance.
[0029] In a second aspect, the present application provides a computing device comprising various modules for implementing the method described in the first aspect and one of its possible implementations, such as a processing module, a first determination module, a second determination module, etc.
[0030] The computing device has the functionality to implement the behaviors in the method examples of any one of the first aspect and its possible implementations. The functionality can be implemented by hardware or by hardware executing corresponding software implementations. The hardware or software includes one or more modules corresponding to the functionality.
[0031] In a third aspect, an embodiment of the present application provides a computing device comprising a memory and at least one processor connected to the memory, wherein the memory is used to store computer program code, and the computer program code comprises computer instructions, which, when executed by at least one processor, cause the processor to execute the method described in any one of the methods of the first aspect and its possible implementations.
[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions. When the computer instructions are run on a computer, the method of the first aspect and any one of its possible implementations is executed.
[0033] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions are run on a computer, the method of the first aspect and any one of its possible implementations is executed.
[0034] In the sixth aspect, an embodiment of the present application provides a chip or chip system, comprising: a processor, a memory and an interface; the processor is used to read computer instructions stored in the memory through the interface, and run the computer instructions to execute the method of the first aspect and any one of its possible implementation methods.
[0035] It should be understood that the beneficial effects achieved by the technical solutions of the second to sixth aspects of this application and the corresponding possible implementation methods can be referred to the technical effects of the first aspect and its corresponding possible implementation methods mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG1 is a schematic diagram of a reuse principle of a model provided in an embodiment of the present application;
[0037] FIG2 is a hardware schematic diagram of a computing device provided in an embodiment of the present application;
[0038] FIG3 is a flow chart of a model reuse method according to an embodiment of the present application;
[0039] FIG4 is a second flow chart of a model reuse method provided in an embodiment of the present application;
[0040] FIG5 is a third flow chart of a model reuse method provided in an embodiment of the present application;
[0041] FIG6 is a fourth flow chart of a model reuse method provided in an embodiment of the present application;
[0042] FIG7 is a fifth flow chart of a model reuse method provided in an embodiment of the present application;
[0043] FIG8 is a sixth flow chart of a model reuse method provided in an embodiment of the present application;
[0044] FIG9 is a schematic diagram of a structure of a computing device according to an embodiment of the present application;
[0045] FIG10 is a second schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0047] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, "first AI model" and "second AI model" are used to distinguish different AI models, rather than to describe a specific order of AI models.
[0048] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0049] In the description of the embodiments of the present application, unless otherwise specified, “plurality” means two or more.
[0050] The embodiments of the present application relate to AI model reuse. In the field of artificial intelligence technology, an AI model is a mathematical model constructed to achieve a certain function (such as target detection, image processing, natural language processing, etc.). The structure of the AI model is complex and the number of parameters is large. Usually, a large number of training samples are required to train an AI model, and training an AI model requires a large resource overhead and takes a long time. At present, by building a model library (that is, building some pre-trained models into a model library), and using the AI model in the model library as the initial model for subsequent other tasks, the target model is obtained based on the initial model, that is, model reuse. Specifically, model reuse refers to using the AI model in the model library as a pre-trained model. Subsequently, for new tasks, based on the training data set and pre-trained model of the new task, the knowledge learned from the pre-trained model is transferred to the target model that can be used for the new task. Model reuse technology can significantly improve the efficiency of model development and deployment.
[0051] Model reuse technology can be applied in areas such as autonomous driving systems, medical diagnosis, financial risk management, natural language processing, and image recognition. For example, it can be used in computing service units and production application centers to quickly generate high-performance models by leveraging pre-trained models. In these applications, by reusing the knowledge and capabilities of existing models and migrating them to new tasks or domains, the workload of retraining models can be greatly reduced, enabling faster adaptation to new environments and tasks, improving system performance and accuracy, and ultimately providing high-quality services and decision-making.
[0052] In computing service units, model reuse technology can leverage the rich knowledge and representation capabilities of pre-trained models to reduce the time and resource costs of training models from scratch, improve computing resource utilization, and provide customized computing services to customers. Production application centers can use model reuse technology to quickly add new features and provide more accurate and personalized services.
[0053] The AI model reuse method provided in the embodiments of the present application can be used in scenarios related to category detection, such as object detection scenarios, category recognition scenarios, or classification scenarios, etc.
[0054] Object detection is a crucial and burgeoning research branch in computer vision. Object detection technology can efficiently identify and precisely locate multiple objects in an image, enabling computer vision models to gain a deeper understanding of the image. Object detection has been successfully applied in various production and life scenarios across both industrial and civilian sectors. For example, in logistics and warehousing, it can automatically detect product defects; in applications such as facial recognition and autonomous driving, it can track the location of target objects in real time.
[0055] In recent years, many companies and research centers have built large-scale target detection model libraries, which include a large number of pre-trained models for detecting various types of targets. The pre-trained models are models trained based on previous detection tasks (including general tasks and specific tasks).
[0056] Referring to Figure 1, for a new detection task (referred to as the target task), a suitable pre-trained model (called the source model) can be selected from the target detection model library. The pre-trained model is a detection model trained for the source task using the source training dataset; then, the source model is fine-tuned using the target training dataset to obtain a target model that can handle the target task.
[0057] The source task mentioned above can also be called an upstream task, and the target task can also be called a downstream task. The source and target tasks can be tasks in the same field or different fields. For example, if the source task is item recognition on a shelf and the target task is food recognition on a shelf, the source and target tasks belong to the same field. Another example is if the source task is remote sensing image recognition and the target task is medical image recognition, the source and target tasks belong to different fields.
[0058] The above-mentioned pre-trained models usually contain the characteristics of multi-domain knowledge and have extensive reusability. Using the pre-trained model as the initialization model for the target task, the model parameters of the pre-trained model are a better optimization starting point for the target task, so that the downstream tasks can converge within a limited training time, which can improve the training efficiency of the target model, achieve high versatility and adaptability, and promote the efficient and accurate implementation of target detection technology in various application scenarios.
[0059] Currently, one method of model reuse is transfer learning. Transfer learning is the process of transferring knowledge acquired based on the source task to the target task. For a pre-trained model from the source task (the source training dataset cannot be accessed), the pre-trained model is trained using the training dataset of the target task, and all parameters of the pre-trained model are fine-tuned to obtain the target model. This method is a transfer learning method with full parameter fine-tuning.
[0060] The migration method of full parameter fine-tuning requires sufficient labeled samples to fine-tune all parameters. As a result, the model training overhead is high, including the cost of computing resources and time. This limits the flexibility and scalability of model reuse in practical applications. Furthermore, the migration method of full parameter fine-tuning may cause the model to overfit due to insufficient labeled samples, resulting in poor detection performance.
[0061] Another approach to model reuse is knowledge distillation, a technique that extracts knowledge from a deep model (the teacher model, or source model) and transfers it to a shallower model (the student model, or target model). Knowledge distillation requires manual tuning of a large number of hyperparameters, which is time-consuming and expensive.
[0062] Another model reuse method is the parameter-efficient transfer learning method, which achieves knowledge transfer by minimizing the number of parameters in the target task, that is, only adjusting a small number of parameters in the source model. Since only a small number of parameters in the source model are adjusted, in some cases it may not be possible to change the inherent knowledge of the source model and to discover potential deep semantic information, resulting in poor processing performance of the target model.
[0063] In summary, the above three methods of model reuse all have the problem of poor model reuse effect. To address the above problem, the embodiment of the present application provides a method for reusing an AI model. In the process of reusing a pre-trained model (source model), a category mapping cost is introduced in which the target task is mapped to the category of the source task, and then the model reuse weight from the source model to the target model is determined based on the category mapping cost. The model loss is then determined based on the model reuse weight and the target training data set, and the parameters of the source model are updated based on the model loss to obtain the target model. In this method, the category mapping cost of the source task category and the target task category can reflect the reusability between the source model used to process the source task and the target model used to process the target task. Therefore, taking the category mapping cost as a consideration for model reuse can guide the model training process to better reuse the features of the pre-trained model and obtain a target model with better processing effect. That is, this method can improve the effect of model reuse.
[0064] Furthermore, model reuse based on category mapping cost can also guide the model training process to converge faster, which can save computing power and time overhead to a certain extent.
[0065] The execution subject of the reuse method of the AI model provided in the embodiment of the present application can be a computing device, such as a server, a desktop computer, etc. For example, Figure 2 is a schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application. The various components shown in Figure 2 can be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application-specific integrated circuits. As shown in Figure 2, the computing device may include: one or more processors (for example, the general-purpose processor 201 and the heterogeneous processor 205 shown in Figure 2), a memory 202, and a communication interface 203. Among them, the general-purpose processor 201, the memory 202, the communication interface 203, and the heterogeneous processor 205 can be connected through a bus 204, or connected to each other in other ways. Optionally, the various components contained in the computing device can be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application-specific integrated circuits.
[0066] The computing device may include one or more general-purpose processors 201. The processor 201 is the control center of the computing device. The processor 201 may be a CPU or other general-purpose processors. The general-purpose processor may be a microprocessor or any conventional processor.
[0067] The controller in general-purpose processor 201 is the nerve center and command center of the computing device. Based on instruction opcodes and timing signals, the controller generates operational control signals to control instruction fetching and execution. Optionally, general-purpose processor 201 may also include a memory for storing instructions and data. Exemplarily, processor 201 may include one or more general-purpose central processing units (CPUs), such as CPU 0 and CPU 1 shown in FIG2 .
[0068] Heterogeneous processor 205 is a processor that is heterogeneous from general-purpose processor 201. Heterogeneous processor 205 may include, for example, a graphics processing unit (GPU), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Heterogeneous processor 205 may include one or more processing cores, and typically includes multiple processing cores.
[0069] The memory 202 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. In the embodiment of the present application, the memory 202 can store information such as computer instructions.
[0070] In one possible implementation, the memory 202 may exist independently of the processor (e.g., the general-purpose processor 201 or the heterogeneous processor 205). The memory 202 may be connected to the processor via a bus 204 and used to store data, instructions, or program codes. When the processor calls and executes the instructions or program codes stored in the memory 202, the relevant steps of the method provided in the embodiments of the present application can be implemented.
[0071] In another possible implementation, the memory 202 may also be integrated with the processor.
[0072] The communication interface 203 may be a transceiver module for communicating with other devices or communication networks, such as Ethernet, RAN, or wireless local area networks (WLAN). The communication interface 203 may receive instructions, messages, or data. The transceiver module may be a device such as a transceiver or a transceiver. Alternatively, the communication interface 203 may be a transceiver circuit located within the processor 201, for implementing signal input and output of the processor. The communication interface 203 may be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface, or the communication interface 203 may be a wireless interface.
[0073] The general-purpose processor 201 (i.e., central processing unit (CPU)) is a general-purpose computing module that is primarily responsible for load logic calculations and logic control functions. It can efficiently handle single, complex computational tasks, but its performance in large-scale computations is relatively low. Therefore, the general-purpose processor 201 can distribute large-scale computational tasks (such as AI model training tasks) to the heterogeneous processor 205. After the heterogeneous processor 205 completes the computation, it returns the computation results to the general-purpose processor 201.
[0074] Bus 204 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be classified as an address bus, a data bus, a control bus, etc. Buses can also be classified as serial buses and parallel buses. For ease of illustration, FIG2 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0075] It should be noted that the computing device shown in FIG. 2 is merely an example of a computing device, and the computing device may have more or fewer components than those shown in FIG. 2 , may combine two or more components, or may have a different component configuration.
[0076] In conjunction with the above, the AI model reuse method provided in the embodiments of the present application reuses a first AI model used to process a source task to obtain a second AI model used to process a target task. In the following embodiments, the training dataset used to train the first AI model is referred to as the source training dataset, and the training dataset used to train the second AI model is referred to as the target training dataset.
[0077] In the embodiment of the present application, the training data set (including the source training data set and the target training data set) includes a plurality of sample data. Optionally, the sample data may be an image (the sample data may also be referred to as a sample image).
[0078] It should be noted that, since the model reuse method of the embodiment of the present application is based on the category mapping cost between the category of the target task and the category of the source task, the first AI model used to process the source task is reused to obtain the second AI model used to process the target task. It can be seen that the source task and the target task involve categories. Therefore, the AI model reuse method provided by the embodiment of the present application is mainly used for scenarios related to category detection. Taking the target detection scenario as an example, the category of the source task in the following embodiment refers to the detection type included in the source task, and the category of the target task refers to the detection type included in the target task. For example, the detection types included in the source task are 2 kinds of reptiles (reptiles of category 1 and reptiles of category 2), then the category of the source task includes two categories, namely category 1 and category 2.
[0079] The following is a detailed introduction to the AI model reuse method provided in an embodiment of the present application. As shown in Figure 3, the method includes S301-S304.
[0080] S301: Use the first AI model as the initialization model for the target task and use the first AI model to process each sample data in the target training data set.
[0081] The first AI model is an AI model trained based on the source training data set of the source task, and the target training data set is used to train the second AI model for processing the target task.
[0082] Optionally, the first AI model is from a model library. The user can select an appropriate pre-trained model from the model library based on the requirements of the target task and use it as the first AI model. For example, if the target task is to detect reptiles, a model for animal recognition can be selected from the model library as the pre-trained model for the target task. For another example, if the target task is to detect sandwiches and ice cream, a model for identifying hamburgers and bread can be selected from the model library as the pre-trained model for the target task.
[0083] It should be noted that both the first AI model and the second AI model can be models for single-target processing or models for multi-target processing. Among them, single target or multi-target can be understood as the category of things that the AI model can process is one category or multiple categories. For example, in the target detection scenario, the first AI model can detect two categories. For example, the first AI model can detect that the image to be detected includes objects of category A1 and / or objects of category B1; the second AI model can detect three categories. For example, the second AI model can detect that an image to be detected includes at least one of objects of category A2, objects of category B2, or objects of category C2.
[0084] S302 : Based on the processing result of each sample data in the target training data set, determine the category mapping cost of mapping the category of the target task to the category of the source task.
[0085] In an embodiment of the present application, a first AI model is used to process all sample data in a target training data set to obtain a processing result for each sample data. The processing result of processing the sample data by the AI model may include the category confidence of the target object in the sample data.
[0086] It should be understood that the category confidence of an object is the confidence that the object belongs to each category. A higher category confidence indicates a higher degree of confidence that the object belongs to that category. The category with the highest category confidence is determined as the category of the object. For example, the categories detectable by the first AI model include category A1 and category B1. For a sample data, the category confidence in its processing result includes the confidence that the object's category is category A1 and the confidence that the object's category is category B1. If the confidence for category A1 is higher than the confidence for category B1, the category of the object in the sample data is determined to be category A1.
[0087] Optionally, in a target detection scenario, the processing result of the sample data includes the category confidence of the target object in the sample data and the position information of the target object.
[0088] The position information of the target object may be the coordinate information of a detection frame used to mark the target object. For example, if the detection frame is a rectangular frame, the position information of the target object is the coordinates of the four vertices of the detection frame.
[0089] In an embodiment of the present application, the target training data set is a training data set corresponding to the target task. Assuming that the categories of the target task include things of category A2, things of category B2, and things of category C2, the sample data in the target training data set includes at least one of the above-mentioned things of category A2, things of category B2, or things of category C2.
[0090] For the processing results of the sample data in the target training data set using the first AI model, the category of the above-mentioned target task is mapped to the category of the source task, which can be understood as: a certain category in the sample data of the target task is identified by the first AI model as a certain category of the source task. For example, the target object in the sample data is actually a sandwich. After the sample data is input into the first AI model, it is detected that the target object is a hamburger (with the highest confidence), then the sandwich category of the target task is mapped to the hamburger category of the source task.
[0091] It should be understood that the category mapping cost is used to measure the reusability between the first AI model and the second AI model. It can also be understood that the category mapping cost is used to measure the similarity between the category features of the source task and the category features of the target task. The higher the category similarity, the higher the reusability.
[0092] In combination with FIG3 , as shown in FIG4 , in one implementation, the above S302 can be implemented through the following S3021 - S3022 .
[0093] S3021. Determine the detection accuracy of the target task category being detected as the source task category based on the category confidence of the target object in each sample data in the processing result.
[0094] In the embodiment of the present application, there is a positive correlation between detection accuracy and category confidence, that is, the higher the category confidence, the higher the detection accuracy.
[0095] In one implementation, the detection accuracy and the category confidence may satisfy a functional relationship y=f(x), where y represents the detection accuracy and x represents the category confidence, for example, y=ax, a>0.
[0096] S3022: Determine, based on the detection accuracy, a category mapping cost for mapping the category of the target task to the category of the source task.
[0097] In the embodiment of the present application, the category mapping cost of mapping the target task category to the source task category is negatively correlated with the detection accuracy when the target task category is detected as the source task category. In other words, the higher the detection accuracy when the target task category is detected as the source task category, the lower the category mapping cost of mapping the target task category to the source task category.
[0098] According to the above description, the category mapping cost is determined based on the detection accuracy of each category of the target task being detected as a category of the source task when the target training dataset is processed using the first AI model. Therefore, the category mapping cost is a matrix, which can also be called a cost matrix or overhead matrix. In other words, the category mapping cost is the reuse cost or reuse overhead of the AI model used to process the source task to the AI model used to process the target task.
[0099] In some embodiments, the category mapping cost of mapping the target task category to the source task category satisfies:
[0100] Among them, C represents the category mapping cost, which is a matrix of size d1×d2, d1 represents the number of categories of the source task (the number of detection types contained in the source task), and d2 represents the number of categories of the target task (that is, the number of detection types contained in the target task). ij Indicates the mapping cost of mapping the target task category j to the source task category i, i = 1, 2, ..., d1, j = 1, 2, ..., d2.
[0101] The calculation process of the above category mapping cost C is as follows:
[0102] Step 1: For each sample data in the target training dataset, calculate the detection accuracy when the first AI model is used to process the sample data, and the category j of the target task is detected as the category i of the source task.
[0103] In one implementation, the detection accuracy of the target task category j being detected as the source task category i may be AP (mean average precision).
[0104] The detection accuracy of the target task category is detected as the source task category, which is the following precision matrix AP:
[0105] Among them, a ij It represents the detection accuracy when the category j of the target task is mapped to the category i of the source task.
[0106] Step 2: The inverse of the detection accuracy is used as the category mapping cost of mapping the target task category j to the source task category i.
[0107] Right now That is to say, The reuse cost value from the i-th category of the source task to the j-th category of the target task is filled into the cost matrix (ie, the category mapping cost matrix C).
[0108] In the above step 1, the detection accuracy of the target task category j being detected as the source task category i is calculated based on the processing results of the sample data.
[0109] In the target detection scenario, the first AI model is used to detect each sample data in the target training data set. The detection result (i.e., the processing result) of each sample data includes the category confidence of the detected target object belonging to each category (the category of the source task) and the location information of the target object (which can be called a detection box or a prediction box).
[0110] Based on the description of the category mapping cost, it can be seen that the detection accuracy a when the category j of the target task is mapped to the category i of the source task ij The higher the value, the higher the similarity between the category features of the source task category i and the target task category j. The target task category j is mapped to the mapping cost c of the source task category i. ij The smaller it is, the smaller the reuse overhead between the first AI model and the second AI model.
[0111] The above process of determining the detection accuracy of the target task category j being detected as the source task category i based on the detection results of the sample data is as follows:
[0112] S1. According to the detection results of each target object in each sample data, determine the intersection over union (IoU) of the predicted box and the true box of each target object in each sample data.
[0113] IoU is used to indicate the degree of overlap between the predicted box and the true box, which can measure the accuracy of the detection results.
[0114] Taking a target object in the sample data as an example, based on the detection results of the target object being detected as various categories of the source task, the detection box corresponding to the category of the source task with the maximum confidence in the detection results is compared with the true detection box of the target object to obtain the intersection-union ratio.
[0115] Exemplarily, taking the case where the source task categories include 2 categories (respectively denoted as category O1 and category O2), and the target task categories include 2 categories (respectively denoted as category T1 and category T2), for a sample data (i.e., a sample image), assuming that the sample data includes 7 targets, the 7 targets include one or more targets of category T1, and one or more targets of category T2. For one of the targets, the confidence of the detection result includes the confidence that the target is detected as category O1 and the confidence that the target is detected as category O2. The category corresponding to the larger confidence value of the two confidences is determined as the predicted category of the target. Referring to Table 1 below, the true category, true box, detection result (confidence and predicted box), predicted category, and the intersection-union ratio of the predicted box and the true box of the target in the sample data are illustrated.
[0116] Table 1
[0117] Combined with Table 1, for example, for the target object 1 in the sample data, the true category of the target object is category T1, the true box is K1, and the confidence level when it is detected as category O1 in the detection result is Z 11 , the confidence level when it is detected as category O2 is Z 12 , where the maximum confidence is Z 11 , then the predicted category of target object 1 is category O1, the corresponding prediction box is M1, and the intersection over union ratio between M1 and K1 is calculated as IoU 11 According to Table 1 above, for this sample data, the calculated IoUs include IoU1, IoU2, IoU3, IoU4, IoU5, IoU6, and IoU7.
[0118] S2. According to the IoU corresponding to all target objects of all samples in the target training dataset, the detection accuracy of each category of the target task is counted as each category of the source task.
[0119] The calculation process of the above detection accuracy is described by continuing to use the example shown in Table 1 as an example.
[0120] First, according to the contents of Table 1, we can count the IoUs when each category of the target task is detected as each category of the source task. For example, according to the detection results of the 7 targets in the above sample data, among the 7 targets, the target task category T1 is detected as the source task category O1 with 3 IoUs, namely IoU1, IoU3, and IoU7; the target task category T1 is detected as the source task category O2 with 1 IoU, namely IoU6; the target task category T2 is detected as the source task category O1 with 1 IoU, namely IoU4; the target task category T2 is detected as the source task category O2 with 2 IoUs, namely IoU2 and IoU5.
[0121] Similarly, for each sample data in the target training data set, statistical results similar to those in Table 1 can be obtained. The difference is that the number of target objects in other sample data may be different from the number of target objects in the sample data shown in Table 1, and / or the other sample data include one or two categories of target objects.
[0122] Secondly, the number of IoUs of various types is counted for the target training dataset.
[0123] Continuing with the above example, for the target training dataset, the IoU statistics include four types of IoU, namely: the IoU of the target task category T1 being detected as the source task category O1, the IoU of the target task category T1 being detected as the source task category O2, the IoU of the target task category T2 being detected as the source task category O1, and the IoU of the target task category T2 being detected as the source task category O2. Assuming that the target training dataset includes N sample data, optionally, the number of four types of IoUs counted for the target training dataset can be recorded as the following matrix U.
[0124] Among them, for sample data 1, the number of the four types of IoU calculated is IoU 11 、IoU 12 、IoU 13 、IoU 14 ; For sample data 2, the number of the four IoUs calculated is IoU 21 、IoU 22 、IoU 23 、IoU 24 ; For sample data N, the number of the four IoUs calculated is IoU N1 、IoU N2 、IoU N3 、IoU N4 .
[0125] Combined with the above U matrix, the total number of each of the four types of IOU in the target training data set is determined, and recorded as IoU1, IoU2, IoU3, and IoU4 respectively.
[0126] IoU1=IoU 11 +IoU 21 +……+Ou N1 ;
[0127] IoU2=IoU 11 +IoU 21 +……+Ou N1 ;
[0128] IoU3=IoU 11 +IoU 21 +……+Ou N1 ;
[0129] IoU4=IoU 11 +IoU 21 +……+Ou N1 .
[0130] Finally, the detection accuracy when the target task category is detected as the source task category is determined based on the number of IoUs that exceed the preset threshold among the numbers of IoUs of each category.
[0131] After counting the total number of each of the four categories of IOU in the above target training data set, the number of IoUs exceeding the preset threshold in each category of IoU is determined according to the preset threshold (the threshold specified by the AP mechanism). It should be understood that IoU exceeding the threshold indicates that the detection accuracy is higher. Assume that among the above IoU1 IoUs, the number of IoUs exceeding the preset threshold is recorded as n1; among the IoU2 IoUs, the number of IoUs exceeding the preset threshold is recorded as n2; among the IoU3 IoUs, the number of IoUs exceeding the preset threshold is recorded as n3; among the IoU4 IoUs, the number of IoUs exceeding the preset threshold is recorded as n4. The number of target objects belonging to category T1 in all sample data in the target training data set is X, and the number of target objects belonging to category T2 is Y. Since the number of categories of the source task is 2 and the number of categories of the target task is also 2, the category mapping cost is a 2×2 matrix. If it is recorded as matrix AP:
[0132] Among them, a 11 It represents the detection accuracy when the target task category T1 is detected as the source task category O1, a 12 It represents the detection accuracy when the target task category T1 is detected as the source task category O2, a 21 It represents the detection accuracy when the target task category T2 is detected as the source task category O1, a 22 It represents the detection accuracy when the target task category T2 is detected as the source task category O2,
[0133] S303: Determine a model reuse weight from the first AI model to the second AI model based on the category mapping cost.
[0134] Among them, the second AI model is used to process the target task, and the model reuse weight is the weight of mapping the processing result of the first AI model to the processing result of the second AI model. The model reuse weight is a matrix, so the model reuse weight can also be called the model transfer matrix from the first AI model to the second AI model. The model transfer matrix can reflect the relationship between the output result of the first AI model and the output result of the second AI model for the same input. When the model reuse weight is known, the expected processing result when the second AI model is used to process the sample data in the target training data set can be determined based on the processing result of the first AI model on the sample data.
[0135] For example, for a sample data in the target training data set, after being processed by the first AI model, the processing result of the sample data is output, that is, the category confidence (the highest category confidence) is Z1, and the model reuse weight is W. Then the category confidence Z2 in the expected processing result of the sample data using the second AI model is: Z2=Z1×W.
[0136] The above model reuse weight can be expressed as:
[0137] Among them, w ij Indicates the model reuse weight from the first AI model to the second AI model when the category j of the target task is detected as the category i of the source task.
[0138] In this embodiment of the present application, the process of reusing the first AI model of the source task to the target task is actually to find a transmission scheme with the minimum overall transmission cost (i.e., the minimum mapping cost) from one distribution to another. The transmission scheme is the model reuse weight (W). Based on the known category mapping cost (the above-mentioned category mapping cost C) from each category of the source task to each category of the target task, the overall mapping cost of reusing the first AI model to the second AI model can be expressed as:
[0139] The model reuse weight W that minimizes the overall category mapping cost satisfies the following formula (1):
[0140] stW×1=p,1×W=q Formula (1)
[0141] in, is the target optimization function, W×1=p and 1×W=q are the constraints.
[0142] In the constraints, p represents the probability distribution of the importance of sample data of each category in the source training data set, and q represents the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model.
[0143] Optionally, the probability distribution p of the importance of each category of sample data in the source training dataset is the ratio of the number of sample data in each category to the total amount of sample data in the source training dataset, and the probability value corresponding to each category in the probability distribution p reflects the detection capability of the first AI model for each category of the source task. The probability distribution q of the importance of the expected detection accuracy of each category of the target task using the second AI model is a uniform distribution.
[0144] For example, assuming that the number of categories of the source task is d1=3, and the number of categories of the target task is d2=2, q=[q1 q2],
[0145] Then the first constraint condition W×1=p is:
[0146] That is, p1 = w 11 +w 12 , p2=w 21 +w 22 , p3=w 31 +w 32 .
[0147] The constraint W × 1 = p in the above conditions ensures that the first AI model's detection capabilities for each category of the source task are continuously applied to the second AI model. In other words, the knowledge learned by the first AI model for detecting each category is transferred to the second AI model to the greatest extent possible, using the source distribution. For example, if the first AI model has a strong detection capability for category A1 of the source task, then if the above constraints are met, the first AI model can be reused to further leverage its learned knowledge of category A1 for the target task.
[0148] Then the second constraint 1×W=q is:
[0149] That is, q1 = w 11 +w 21 +w 31 ,q2=w 12 +w 22 +w 32 .
[0150] As can be seen from the above embodiments, q is uniformly distributed, i.e., q1 = q2 in the above example. The constraint 1 × W = q implies that, after the first AI model is reused in the second AI model, the second AI model's detection performance for each category of the target task is as similar as possible. In other words, the second AI model is expected to achieve the best possible performance across all categories of the target task. For example, if the target task includes two categories, the purpose of the constraint is to ensure that the second AI model's detection performance for the first category is similar to that for the second category.
[0151] Optionally, for the above formula (1), the sinkhorn algorithm can be used to solve the model reuse weight W. The detailed process of solving the sinkhorn algorithm can refer to relevant information of the prior art and will not be described in detail here.
[0152] S304: Obtain a second AI model based on the model reuse weight and the target training data set.
[0153] Optionally, in combination with FIG3 , as shown in FIG5 , the above S304 may be implemented through S3041 - S3042 .
[0154] S3041. Determine the model loss based on the model reuse weight and the target training data set.
[0155] 5 , as shown in FIG6 , S3041 specifically includes S3041a - S3041b .
[0156] S3041a. Based on the model reuse weight, map the first category confidence of the target object to the second category confidence of the target object.
[0157] Among them, the first category confidence of the target object is obtained by processing the sample data in the target training data set using the first AI model.
[0158] It can be understood that the category confidence output by the AI model is a vector, which can also be called a classification weight vector.
[0159] The specific process of mapping the first category confidence of the target object to the second category confidence of the target object is: multiplying the first category confidence of the target object by the model reuse weight to obtain the second category confidence of the target object. The second category confidence includes the confidence of each category of the sample data detected as the target task. The second category confidence is the expected processing result of the second AI model obtained after the current first AI model is reused.
[0160] In the embodiment of the present application, the above-mentioned mapping of the first category confidence of the target object to the second category confidence of the target object is a process of changing the processing result of the first AI model. The process of changing the category confidence involves category change (i.e., category mapping), that is, modifying the category of the target object in the sample data in the processing result (the category in the source task) to the actual category of the target object. For example, the target object in the sample data is identified by the first AI model as category A1 in the source task (such as a hamburger). In fact, the target object in the sample data is category A2 (sandwich) in the target task. Then, the confidence Z1 that the target object belongs to category A1 in the processing result is modified to the confidence Z2 that the target object belongs to category A2, Z2=Z1×w, w represents the model reuse weight from the first AI model to the second AI model when the category of the target task is detected as the category of the source task.
[0161] S3041b. Determine the model loss based on the second category confidence.
[0162] It is understandable that the second category confidence of the target object obtained in S3041b is an estimated result of processing the sample data based on the second AI model. Optionally, the value of the category loss function (i.e., the category loss value) is determined based on the second category confidence of the target object and the true category confidence of the target object (e.g., 100%). This category loss value is the model loss.
[0163] In some embodiments, for target detection scenarios, the processing results obtained by using the first AI model to process the sample data in the target training data set may also include the location information of the target object, which is an estimated result of processing the sample data based on the second AI model. In this case, the above-mentioned model loss can be the sum of the category loss and the position loss, the category loss is determined based on the second category confidence of the target object, and the position loss is determined based on the location information of the target object. Optionally, the value of the position loss function (i.e., the position loss value) is determined based on the predicted location information of the target object and the actual location information of the target object in the sample data.
[0164] S3042. Update the parameters of the first AI model based on the model loss to obtain a second AI model.
[0165] Optionally, an error back propagation (BP) algorithm can be used to perform gradient back propagation on the model loss, and then the gradient information can be used to update the parameters of the current first AI model.
[0166] It should be understood that using the first AI model as the initialization model for the target task and continuously updating the first AI model with the target training dataset to obtain the second AI model is a multi-iteration process (a cyclic process). Referring to Figure 7, using the first AI model as the initialization model, after obtaining the above-mentioned model reuse weights, the process of obtaining the second AI model based on the model reuse weights and the target training dataset includes the following steps a to e.
[0167] Step a: Input the i-th sample data in the target training data set into the current first AI model to obtain the processing result of the i-th sample data.
[0168] The value of i is one of 1, 2, ..., M, where M is the number of sample data contained in the target training dataset and M is a positive integer greater than or equal to 1. When i=1, the current first AI model is the initialization model selected from the model library.
[0169] The processing result of the i-th sample data includes the confidence level of each target object in the sample data obtained by processing the sample data using the current first AI model. In the target detection scenario, the processing result also includes the location information of each target object.
[0170] Step b: Map the processing results of the i-th sample data according to the model reuse weight to obtain the prediction results of the i-th sample data.
[0171] The prediction result is an estimated result based on the processing of the sample data by the second AI model.
[0172] Step c: Determine the model loss.
[0173] Step d: Update the parameters of the current first AI model based on the model loss.
[0174] For the relevant contents of steps a to d, please refer to the description of the above embodiment and will not be repeated here.
[0175] Step e: Determine whether the training end condition is met.
[0176] Optionally, the number of model training times can be set as the training end condition. Of course, other conditions can also be used as training end conditions.
[0177] In this embodiment of the present application, if the training end condition is met, the first AI model after the parameter update in step e is used as the second AI model. If the training end condition is not met, i is updated (i=i+1), that is, the next sample data is used and the next round of training is performed again from step a until the training end condition is met, and the last updated first AI model is used as the second AI model.
[0178] In summary, based on the AI model reuse method provided in the embodiment of the present application, in the process of reusing the pre-trained model (source model), the category of the source task is mapped to the category of the target task. The category mapping cost is introduced, and then the model reuse weight of the source model to the target model is determined based on the category mapping cost, and the model loss is determined according to the model reuse weight and the target training data set, and the parameters of the source model are updated according to the model loss to obtain the target model. In this method, since the category mapping cost of the source task category and the target task category can reflect the reusability between the source model used to process the source task and the target model used to process the target task, therefore, taking the category mapping cost as a consideration for model reuse can guide the model training process to better reuse the features of the pre-trained model and obtain a target model with better processing effect, that is, this method can improve the effect of model reuse.
[0179] Furthermore, model reuse based on category mapping cost can also guide the model training process to converge faster, which can save computing power and time overhead to a certain extent.
[0180] In conjunction with the above, with reference to FIG8 , the complete process of the AI model reuse method provided in the embodiment of the present application is described below. The method includes:
[0181] S701 : Count the number of sample data of each category in the source training data set, and generate a probability distribution p of the importance of the sample data in the source training data set.
[0182] S702: Generate a probability distribution q of the importance of prediction detection accuracy of sample data in the target training dataset.
[0183] For detailed description of S701 and S702 , reference may be made to the description of S303 in the above embodiment, which will not be repeated here.
[0184] S703: Input each sample data in the target training data set into the first AI model to obtain a predicted output.
[0185] S704: Calculate the reuse accuracy of the first AI model based on the predicted output of the first AI model.
[0186] The reuse accuracy of the first AI model is used to determine whether to use the first AI model as the initialization model for the target task. The reuse accuracy of the first AI model is the detection accuracy of the target task category being detected as the source task category when the first AI model is used to process sample data in the target training dataset.
[0187] In one implementation, the reuse accuracy of the first AI model may be the sum of all elements in the accuracy matrix AP determined in S302. Assuming that the reuse accuracy of the first AI model is denoted as Y, then
[0188] It should be noted that when the detection accuracy of the first AI model is greater than the accuracy threshold, the first AI model is used as the initialization model for the second AI model, and steps S705-S708 and subsequent to S710 are executed. When the detection accuracy of the first AI model is less than or equal to the accuracy threshold, the first AI model is not used as the initialization model for the second AI model. In this case, a randomly initialized model can be used as the initialization model for the second AI model and trained based on the target training dataset, that is, steps S709 and subsequent to S710 are executed.
[0189] Determining the initialization model of the second AI model based on the detection accuracy described above helps to obtain a second AI model with better performance.
[0190] S705. Determine the category mapping cost C based on the prediction output of the first AI model.
[0191] S706 : Determine a model reuse weight W according to the category mapping cost C, the probability distribution p, and the probability distribution q.
[0192] S707: Input the sample data i in the target training data set into the first AI model to obtain the predicted output of the sample data i.
[0193] S708. Map the predicted output of the sample data i to the predicted output of the second AI model according to the model reuse weight W.
[0194] S709: Input the sample data i in the target training data set into the randomly initialized model to obtain the predicted output.
[0195] The randomly initialized model may be a model obtained by randomly initializing the parameters of the first AI model, or may be other models, which is not limited in the embodiments of the present application.
[0196] S710: Determine model loss.
[0197] S711. Update the parameters of the first AI model according to the model loss.
[0198] S712: Determine whether the end condition is met.
[0199] If the end condition of model training is met, the first AI model updated this time will be used as the second AI model; if the end condition of model training is not met, i (i=i+1) will be updated, and then return to S707 to continue execution until the end condition is met, and the first AI model updated for the last time will be used as the second AI model.
[0200] Through the above S701-S712, a second model with better processing effect can be obtained, and model training based on category mapping cost can also guide the model training process to converge faster, which can save computing power and time overhead to a certain extent.
[0201] It is understandable that, in order to implement the above functions, the above computing device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the method steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0202] The embodiment of the present application can divide the functional modules of the above-mentioned computing device according to the above-mentioned method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0203] 9 shows a possible structural diagram of the computing device involved in the above embodiment, in the case of dividing each functional module according to each function. The computing device includes a processing module 801, a first determining module 802 and a second determining module 803.
[0204] Among them, the processing module 801 is used to execute S301, S703, S707, S708, and S709 in the above method embodiment; the first determination module 802 is used to execute S302 (including S3021-S3022), S303, S704, S705, and S706 in the above method embodiment; the second determination module 803 is used to execute S304 (including S3041-S3042), S710, and S711 in the above method embodiment.
[0205] In the case of using an integrated unit, Figure 10 shows another possible structural diagram of the computing device involved in the above embodiments. The computing device may include: a processing module 901 and a communication module 902. The processing module 901 can be used to control and manage the operations of the computing device. For example, the processing module 901 can be used to support the computing device in executing the steps performed by the processing module 801, the first determination module 802, and the second determination module 803 in the above method embodiment, and / or other processes used in the technology described herein. The communication module 902 can be used to support the computing device in communicating with other network entities. Optionally, as shown in Figure 10, the computing device may also include a storage module 903 for storing the program code and data of the computing device.
[0206] The processing module 901 may be a processor, for example, the general-purpose processor 201 or the heterogeneous processor 205 in FIG. The communication module 902 may be a transceiver, a transceiver circuit, or a communication interface, for example, the communication interface 203 in FIG. The storage module 903 may be a memory, for example, the memory 202 in FIG. 2 .
[0207] The various modules of the above-mentioned computing device can also be used to perform other actions in the above-mentioned method embodiment. All relevant contents of each step involved in the above-mentioned method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0208] For more details on how the modules included in the computing device implement the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here. The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments.
[0209] Optionally, an embodiment of the present application further provides a computing device (for example, the computing device may be a chip or a chip system), which includes at least a processor (a general-purpose processor and / or a heterogeneous processor) and at least one interface. The interface can be used to receive signals from other devices (for example, memories of other electronic devices). For another example, the interface can be used to send signals to other devices (for example, processors). The interface can read instructions stored in the memory and send the instructions to the processor, which is used to read the instructions to execute the method in any of the above method embodiments.
[0210] In one possible design, the computing device further includes a memory. The memory is used to store necessary program instructions and data. The processor can read the computer instructions stored in the memory through an interface to cause the computing device to execute the method in any of the above-mentioned method embodiments. Of course, the memory may not be in the computing device. When the computing device is a chip system, it may be composed of a chip or may include a chip and other discrete devices. This embodiment of the application is not specifically limited to this.
[0211] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions in accordance with the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a magnetic disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state drive (SSD)).
[0212] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0213] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0214] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0215] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0216] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0217] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for reusing an artificial intelligence (AI) model, characterized in that: include: Using the first AI model as the initialization model of the target task, and using the first AI model to process each sample data in the target training data set; wherein the first AI model is an AI model trained according to the source training data set of the source task; Determining, based on a processing result of each sample data in the target training data set, a category mapping cost at which the category of the target task is mapped to the category of the source task; Determining a model reuse weight from the first AI model to the second AI model according to the category mapping cost; the second AI model is used to process the target task; The second AI model is obtained according to the reuse weight and the target training data set.
2. The method according to claim 1, characterized in that The obtaining the second AI model according to the reuse weight and the target training data set includes: Determining a model loss based on the model reuse weight and the target training data set; According to the model loss, the parameters of the first AI model are updated to obtain the second AI model.
3. The method according to claim 1 or 2, characterized in that: The determining, based on the processing result of each sample data in the target training data set, a category mapping cost of mapping the category of the target task to the category of the source task comprises: Determining the detection accuracy of the target task category being detected as the source task category according to the category confidence of the target object in each sample data in the processing result; A category mapping cost at which the category of the target task is mapped to the category of the source task is determined according to the detection accuracy.
4. The method according to claim 3, characterized in that There is a negative correlation between a category mapping cost at which the category of the target task is mapped to the category of the source task and a detection accuracy when the category of the target task is detected as the category of the source task.
5. The method according to any one of claims 1 to 4, characterized in that: The model reuse weight is: Where d1 represents the number of categories of the source task, d2 represents the number of categories of the target task, i = 1, 2, ..., d1, j = 1, 2, ..., d2, w ij represents a model reuse weight from the first AI model to the second AI model when the category j of the target task is detected as the category i of the source task; The model reuse weight W satisfies: stW×1=p,1×W=q Among them, c ij It represents the category mapping cost of mapping the category j of the target task to the category i of the source task, p represents the probability distribution of the importance of the sample data of the target objects of each category in the source training data set, and q represents the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model.
6. The method according to any one of claims 1 to 5, characterized in that: The determining of the model loss based on the model reuse weight and the target training data set includes: Based on the model reuse weight, mapping the first category confidence of the target object to the second category confidence of the target object; the first category confidence of the target object is obtained by processing the sample data in the target training data set using the first AI model; A model loss is determined based on the second category confidence.
7. The method according to any one of claims 1 to 6, characterized in that: The probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model is uniformly distributed.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Determine the reuse accuracy of the first AI model; the reuse accuracy of the first AI model is used to determine whether to use the first AI model as the initialization model of the target task, and the reuse accuracy of the first AI model is the detection accuracy of the category of the target task being detected as the category of the source task when the target training data set is processed using the first AI model.
9. A computing device, characterized in that include: a processing module, a first determining module, and a second determining module; The processing module is used to use the first AI model as the initialization model of the target task, and use the first AI model to process each sample data in the target training data set; wherein the first AI model is an AI model trained according to the source training data set of the source task; The first determination module is used to determine the category mapping cost of mapping the category of the target task to the category of the source task based on the processing result of each sample data in the target training data set; The first determination module is further used to determine a model reuse weight from the first AI model to the second AI model according to the category mapping cost; the second AI model is used to process the target task; The second determination module is used to obtain the second AI model according to the reuse weight and the target training data set.
10. The computing device according to claim 9, characterized in that The second determination module is specifically used to determine the model loss based on the model reuse weight and the target training data set; and according to the model loss, update the parameters of the first AI model to obtain the second AI model.
11. The computing device according to claim 9 or 10, characterized in that: The first determination module is specifically configured to determine the detection accuracy of the category of the target task being detected as the category of the source task according to the category confidence of the target object in each sample data in the processing result; And according to the detection accuracy, a category mapping cost of mapping the category of the target task to the category of the source task is determined.
12. The computing device according to claim 11, characterized in that There is a negative correlation between a category mapping cost at which the category of the target task is mapped to the category of the source task and a detection accuracy when the category of the target task is detected as the category of the source task.
13. The computing device according to any one of claims 9 to 12, characterized in that: The model reuse weight is: Where d1 represents the number of categories of the source task, d2 represents the number of categories of the target task, i = 1, 2, ..., d1, j = 1, 2, ..., d2, w ij represents a model reuse weight from the first AI model to the second AI model when the category j of the target task is detected as the category i of the source task; The model reuse weight W satisfies: stW×1=p,1×W=q Among them, c ij It represents the category mapping cost of mapping the category j of the target task to the category i of the source task, p represents the probability distribution of the importance of the sample data of the target objects of each category in the source training data set, and q represents the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model.
14. The computing device according to any one of claims 9 to 13, characterized in that: The second determination module is specifically used to map the first category confidence of the target object to the second category confidence of the target object based on the model reuse weight; the first category confidence of the target object is obtained by processing the sample data in the target training data set using the first AI model; and determining the model loss based on the second category confidence.
15. The computing device according to any one of claims 9 to 14, characterized in that: The probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model is uniformly distributed.
16. The computing device according to any one of claims 9 to 15, characterized in that: The first determination module is also used to determine the reuse accuracy of the first AI model; the reuse accuracy of the first AI model is used to determine whether to use the first AI model as the initialization model of the target task, and the reuse accuracy of the first AI model is the detection accuracy of the category of the target task being detected as the category of the source task when the target training data set is processed using the first AI model.
17. A computing device, characterized in that The method comprises a memory and at least one processor connected to the memory, wherein the memory is used to store computer program code, and the computer program code comprises computer instructions. When the computer instructions are executed by the at least one processor, the processor executes the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that: Computer instructions are stored, and when the computer instructions are executed on a computer, the method according to any one of claims 1 to 8 is executed.
19. A chip, characterized in that: The method comprises a processor, a memory and an interface; the processor is used to read the computer instructions stored in the memory through the interface, and run the computer instructions to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for multiplexing AI model
CN120047709A
Method and system for reusing deep neural network model training model
CN110428051A
Model data processing method, related device, equipment and storage medium
CN116629338A
Machine learning system and machine learning method
US20230229965A1
Evaluating target domain machine learning model for deployment
WO2023222185A1