Method and device for multiplexing AI model

By introducing category mapping cost and model multiplexing weights into the AI ​​model multiplexing method, the problem of poor model multiplexing in the prior art is solved, and more efficient model multiplexing and faster training convergence are achieved.

CN120047709APending Publication Date: 2025-05-27HUAWEI TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311594582.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Among the existing AI model multiplexing methods, the model multiplexing effect is poor, and it is difficult to effectively utilize the features of the pre-trained model.

Method used

By introducing category mapping costs, model multiplexing weights are determined, and model parameters are updated based on the weights and target training datasets, an AI model suitable for new tasks is obtained.

Benefits of technology

The effect of model reuse is improved, the features of pre-trained models can be better reused, and the model training process can be accelerated to a certain extent, saving computing power and time overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047709A_ABST
    Figure CN120047709A_ABST
Patent Text Reader

Abstract

The invention provides an AI model multiplexing method and device, relates to the technical field of artificial intelligence, and can improve the model multiplexing effect. The method comprises the steps that a first AI model serves as an initialization model of a target task, the first AI model is adopted to process each piece of sample data in a target training data set, and the first AI model is an AI model obtained through training according to a source training data set of a source task; based on a processing result of each piece of sample data in the target training data set, determining a category mapping cost that the category of the target task is mapped into the category of the source task; according to the category mapping cost, determining a model multiplexing weight from the first AI model to a second AI model used for processing the target task; and obtaining a second AI model according to the multiplexing weight and the target training data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and device for reusing an AI model. Background Art

[0002] In recent years, the application of artificial intelligence (AI) technology has become more and more extensive, and how to obtain a better AI model is crucial.

[0003] At present, many enterprises and research centers have built large-scale model libraries, which include a large number of pre-trained models. Pre-trained models are AI models trained based on previous tasks (including general tasks and specific tasks). For new tasks, the pre-trained models in the model library can be reused to obtain AI models for new tasks. In other words, the pre-trained models in the model library are used as the initialization models for new tasks, and then the initialization models are optimized to obtain AI models suitable for new tasks.

[0004] In some existing AI model reuse methods, the processing effect of the AI ​​model obtained based on the pre-trained model is low, that is, the existing AI model reuse method has the problem of poor model reuse effect. Summary of the invention

[0005] The present application provides a method and device for reusing an AI model, which can improve the effect of model reuse.

[0006] This application adopts the following technical solutions:

[0007] In a first aspect, the present application provides a method for reusing an AI model, comprising: using a first AI model as an initialization model for a target task, and using the first AI model to process each sample data in a target training data set, the first AI model is an AI model trained according to a source training data set of a source task; and based on the processing result of each sample data in the target training data set, determining a category mapping cost in which a category of the target task is mapped to a category of the source task; and determining a model reuse weight from the first AI model to a second AI model for processing the target task based on the category mapping cost; and then obtaining the second AI model based on the model reuse weight and the target training data set.

[0008] The AI ​​model reuse method provided in the present application, since the category mapping cost between the category of the source task and the category of the target task can reflect the reusability between the source model used to process the source task (i.e., the first AI model mentioned above) and the target model used to process the target task (i.e., the second AI model mentioned above), therefore, taking the category mapping cost as a consideration factor for model reuse can guide the model training process to better reuse the features of the pre-trained model and obtain a target model with better processing effect, that is, this method can improve the effect of model reuse.

[0009] Furthermore, model reuse based on category mapping cost can also guide the model training process to converge faster, which can save computing power and time overhead to a certain extent.

[0010] In one possible implementation, obtaining the second AI model according to the model reuse weight and the target training data set specifically includes: determining the model loss based on the model reuse weight and the target training data set; and updating the parameters of the first AI model according to the model loss to obtain the second AI model.

[0011] In a possible implementation, the above-mentioned determination of the category mapping cost of the target task being mapped to the category of the source task based on the processing result of each sample data in the target training data set specifically includes: determining the detection accuracy of the target task category being detected as the category of the source task according to the category confidence of the target object in each sample data in the processing result; and determining the category mapping cost of the target task category being mapped to the category of the source task according to the detection accuracy. There is a positive correlation between the detection accuracy and the category confidence, that is, the higher the category confidence, the higher the detection accuracy.

[0012] In a possible implementation, there is a negative correlation between the category mapping cost of the target task category being mapped to the source task category and the detection accuracy when the target task category is detected as the source task category. That is, the higher the detection accuracy when the target task category is detected as the source task category, the lower the category mapping cost of the target task category being mapped to the source task category.

[0013] In one possible implementation, the category of the target task is mapped to the category of the source task with a category mapping cost satisfying:

[0014]

[0015] Among them, C represents the category mapping cost, which is a size d 1 ×d 2 The matrix, d 1 Indicates the number of categories of the source task (the number of detection types contained in the source task), d 2Indicates the number of categories of the target task (that is, the number of detection types contained in the target task). ij Indicates the mapping cost of the target task category j being mapped to the source task category i, i = 1, 2, ..., d 1 , j=1,2,......,d 2 .

[0016] In one possible implementation, the model reuse weight is:

[0017]

[0018] Among them, d 1 represents the number of categories of the source task, that is, the number of detection types contained in the source task, d 2 Indicates the number of categories of the target task, that is, the number of detection types contained in the target task, i = 1, 2, ..., d 1 , j=1,2,......,d 2 , w ij Indicates the model reuse weight from the first AI model to the second AI model when the category j of the target task is detected as the category i of the source task.

[0019] The above model reuse weight W satisfies:

[0020]

[0021] stW×1=p,1×W=q

[0022] Among them, c ij It represents the category mapping cost of mapping the category j of the target task to the category i of the source task, p represents the probability distribution of the importance of the sample data of the target objects of each category in the source training data set, and q represents the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model.

[0023] In the present application, the process of reusing the first AI model of the source task to the target task is actually to find a transmission scheme with the overall minimum transmission cost (i.e., the minimum mapping cost) from one distribution to another, and the transmission scheme is the above-mentioned model reuse weight (W). The meaning of W×1=p in the above-mentioned constraint condition is: to allow the first AI model's detection capability for various categories of the source task to continue to be used in the second AI model, that is, the knowledge learned in the first AI model for detecting various categories is transferred to the second AI model to the greatest extent with the source distribution. The meaning of the above-mentioned constraint condition 1×W=q is: after the first AI model is reused in the second AI model, the detection effect of the second AI model on each category of the target task is as similar as possible, that is, it is expected that the second AI model can achieve the best possible level in each category of the target task.

[0024] In one possible implementation, the above-mentioned determination of the model loss based on the model reuse weight and the target training data set specifically includes: based on the model reuse weight, mapping the first category confidence of the target object to the second category confidence of the target object, the first category confidence of the target object is obtained by processing the sample data in the target training data set using the first AI model; and determining the model loss based on the second category confidence.

[0025] The above-mentioned model loss can be determined according to the second category confidence of the target object and the true category confidence of the target object, and the value of the category loss function (ie, the category loss value) is determined, that is, the model loss is the category loss value.

[0026] Optionally, in the target detection scenario, the processing result of the sample data includes the category confidence of the target in the sample data and the location information of the target. The location information of the target may be the coordinate information of the detection frame used to mark the target, and the model loss may be the sum of the category loss and the location loss, and the location loss is the value of the location loss function (i.e., the location loss value) determined based on the predicted location information of the target and the actual location information of the target in the sample data.

[0027] In one possible implementation, the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model is uniformly distributed, so that the second AI model can achieve the best possible level in each category of the target task.

[0028] In one possible implementation, the AI ​​model reuse method provided in the present application also includes: determining the reuse accuracy of the first AI model, the reuse accuracy of the first AI model is used to determine whether to use the first AI model as the initialization model of the target task, the reuse accuracy of the first AI model is the detection accuracy of the target task category being detected as the source task category when the target training data set is processed using the first AI model.

[0029] When the detection accuracy of the first AI model is greater than the accuracy threshold, the first AI model is used as the initialization model of the second AI model. When the detection accuracy of the first AI model is less than or equal to the accuracy threshold, the first AI model is not used as the initialization model of the second AI model, and other models (such as randomly initialized models) can be selected as the initialization model of the second AI model. Determining the initialization model of the second AI model based on the detection accuracy helps to obtain a second AI model with better performance.

[0030] In a second aspect, the present application provides a computing device, which includes various modules for implementing the method described in the first aspect and one of its possible implementation methods, such as a processing module, a first determination module, a second determination module, etc.

[0031] The computing device has the function of implementing the behavior in the method example of any one of the first aspect and possible implementations thereof. The function can be implemented by hardware, or by hardware executing corresponding software implementation. The hardware or software includes one or more modules corresponding to the above functions.

[0032] In a third aspect, an embodiment of the present application provides a computing device, including a memory and at least one processor connected to the memory, the memory is used to store computer program code, the computer program code includes computer instructions, when the computer instructions are executed by at least one processor, the processor executes the method described in any one of the methods of the first aspect and its possible implementations

[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions, which, when executed on a computer, execute the method of the first aspect and any one of its possible implementations.

[0034] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions are run on a computer, the method of the first aspect and any one of its possible implementations is executed.

[0035] In the sixth aspect, an embodiment of the present application provides a chip or a chip system, comprising: a processor, a memory and an interface; the processor is used to read computer instructions stored in the memory through the interface, and run the computer instructions to execute the method of the first aspect and any one of its possible implementation methods.

[0036] It should be understood that the beneficial effects achieved by the technical solutions of the second to sixth aspects of the present application and the corresponding possible implementation methods can be referred to the technical effects of the first aspect and its corresponding possible implementation methods mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of the reuse principle of a model provided in an embodiment of the present application;

[0038] Figure 2 A hardware schematic diagram of a computing device provided in an embodiment of the present application;

[0039] Figure 3 One of the flow diagrams of a model reuse method provided in an embodiment of the present application;

[0040] Figure 4 A second flow chart of a model reuse method provided in an embodiment of the present application;

[0041] Figure 5 A third flow chart of a model reuse method provided in an embodiment of the present application;

[0042] Figure 6 A fourth flow chart of a model reuse method provided in an embodiment of the present application;

[0043] Figure 7 A fifth flow chart of a model reuse method provided in an embodiment of the present application;

[0044] Figure 8 A sixth flow chart of a model reuse method provided in an embodiment of the present application;

[0045] Fig. 9 One of the structural schematic diagrams of a computing device provided in an embodiment of the present application;

[0046] Fig.10 The second structural diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0047] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0048] The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects rather than to describe a specific order of objects. For example, the first AI model and the second AI model are used to distinguish different AI models rather than to describe a specific order of AI models.

[0049] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0050] In the description of the embodiments of the present application, unless otherwise specified, “plurality” means two or more.

[0051] The embodiments of the present application relate to AI model reuse. In the field of artificial intelligence technology, an AI model is a mathematical model constructed to achieve a certain function (such as target detection, image processing, natural language processing, etc.). The structure of the AI ​​model is complex and the number of parameters is large. Usually, a large number of training samples are required to train an AI model, and training an AI model requires a large resource overhead and takes a long time. At present, by building a model library (that is, building some pre-trained models into a model library), and using the AI ​​model in the model library as the initial model for subsequent other tasks, the target model is obtained based on the initial model, that is, model reuse. Specifically, model reuse refers to using the AI ​​model in the model library as a pre-trained model. Subsequently, for new tasks, based on the training data set and pre-trained model of the new task, the knowledge learned from the pre-trained model is transferred to the target model that can be used for the new task. The model reuse technology can significantly improve the efficiency of model development and deployment.

[0052] Model reuse technology can be applied in the fields of autonomous driving systems, medical diagnosis, financial risk management, natural language processing, and image recognition. For example, it can be applied in computing service units and production application centers. By using pre-trained models, high-performance models can be quickly generated. In these applications, by reusing the knowledge and capabilities of existing models and migrating them to new tasks or fields, the workload of re-model training can be greatly reduced, and new environments and tasks can be adapted more quickly, improving the performance and accuracy of the system, thereby providing high-quality services and decisions.

[0053] In computing service units, model reuse technology can utilize the rich knowledge and representation capabilities of pre-trained models to reduce the time and resource costs of training models from scratch, improve the utilization of computing resources, and provide customized computing services to customers. Production application centers can quickly add new functions through model reuse technology and provide more accurate and personalized services.

[0054] The AI ​​model reuse method provided in the embodiments of the present application can be used in scenarios related to category detection, such as object detection scenarios, category recognition scenarios, or classification scenarios, etc.

[0055] Among them, object detection is a crucial and booming research branch in the field of computer vision. Object detection technology can efficiently identify multiple objects in an image and accurately locate the positions of multiple objects in an image, so that computer vision models can understand the image more deeply. Object detection has been successfully applied to various production and life scenarios in industrial and civil fields. For example, in logistics warehousing, it can automatically detect product defects; in applications such as face recognition and autonomous driving, it can track the position of target objects in real time.

[0056] In recent years, many companies and research centers have built large-scale target detection model libraries, which include a large number of pre-trained models for detecting various types of targets. The pre-trained models are models trained based on previous detection tasks (including general tasks and specific tasks).

[0057] refer to Figure 1 For a new detection task (referred to as the target task), a suitable pre-trained model (referred to as the source model) can be selected from the target detection model library. The pre-trained model is a detection model trained for the source task using the source training dataset. Then, the source model is fine-tuned using the target training dataset to obtain a target model that can handle the target task.

[0058] The above source tasks can also be called upstream tasks, and the target tasks can also be called downstream tasks. The source tasks and target tasks can be tasks in the same field or tasks in different fields. For example, the source task is item recognition on the shelf, and the target task is food recognition on the shelf. The source task and the target task belong to the same field. For another example, the source task is remote sensing image recognition, and the target task is medical image recognition. The source task and the target task belong to different fields.

[0059] The above-mentioned pre-trained models usually contain the characteristics of multi-domain knowledge and have extensive reusability. The pre-trained model is used as the initialization model of the target task. For the target task, the model parameters of the pre-trained model are a better optimization starting point, so that the downstream tasks can converge within a limited training time, which can improve the training efficiency of the target model, achieve high versatility and adaptability, and promote the efficient and accurate implementation of target detection technology in various application scenarios.

[0060] At present, one method of model reuse is transfer learning. Transfer learning is to transfer the knowledge obtained based on the source task to the target task. For a pre-trained model from the source task (the source training dataset cannot be accessed), the training dataset of the target task is used to train the pre-trained model, and all parameters of the pre-trained model are fine-tuned to obtain the target model. This method is a transfer learning method with full parameter fine-tuning.

[0061] The migration method of fine-tuning all parameters requires sufficient labeled samples to fine-tune all parameters. Therefore, the cost of model training is relatively large, including the cost of computing resources and time. This limits the flexibility and scalability of model reuse in practical applications. In addition, the migration method of fine-tuning all parameters may cause the model to overfit due to insufficient labeled samples, that is, the detection effect of the model is poor.

[0062] Another method of model reuse is knowledge distillation, which is a technique for extracting knowledge from a deep model (teacher model, i.e. source model) and transferring it to a shallow model (student model, i.e. target model). Knowledge distillation requires manual adjustment of a large number of hyperparameters, which takes a lot of time.

[0063] Another method of model reuse is parameter-efficient transfer learning, which achieves knowledge transfer by minimizing the number of parameters in the target task, that is, only adjusting a small number of parameters in the source model. Since only a small number of parameters in the source model are adjusted, in some cases the inherent knowledge of the source model may not be changed, and the potential deep semantic information cannot be discovered, resulting in poor processing effect of the target model.

[0064] In summary, the above three methods of model reuse all have the problem of poor model reuse effect. To address the above problem, the embodiment of the present application provides a method for reusing an AI model. In the process of reusing a pre-trained model (source model), a category mapping cost is introduced in which the target task is mapped to the category of the source task, and then the model reuse weight from the source model to the target model is determined based on the category mapping cost. The model loss is then determined based on the model reuse weight and the target training data set, and the parameters of the source model are updated based on the model loss to obtain the target model. In this method, the category mapping cost of the source task category and the target task category can reflect the reusability between the source model used to process the source task and the target model used to process the target task. Therefore, taking the category mapping cost as a consideration for model reuse can guide the model training process to better reuse the features of the pre-trained model and obtain a target model with better processing effect, that is, this method can improve the effect of model reuse.

[0065] Furthermore, model reuse based on category mapping cost can also guide the model training process to converge faster, which can save computing power and time overhead to a certain extent.

[0066] The execution subject of the AI ​​model reuse method provided in the embodiment of the present application may be a computing device, such as a server, a desktop computer, etc. For example, Figure 2A schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application is provided. Figure 2 The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits. Figure 2 As shown, the computing device may include: one or more processors (e.g. Figure 2 201 and heterogeneous processor 205), memory 202, and communication interface 203. The general processor 201, memory 202, communication interface 203, and heterogeneous processor 205 may be connected to each other via bus 204, or in other ways. Optionally, the various components included in the computing device may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0067] The computing device may include one or more general-purpose processors 201. The processor 201 is the control center of the computing device. The processor 201 may be a CPU or other general-purpose processors. The general-purpose processor may be a microprocessor or any conventional processor.

[0068] The controller in the general processor 201 is the nerve center and command center of the computing device. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions. Optionally, a memory can also be set in the general processor 201 to store instructions and data. Exemplarily, the processor 201 may include one or more general central processing units (CPUs), such as Figure 2 CPU 0 and CPU1 are shown in .

[0069] The heterogeneous processor 205 is a processor that is heterogeneous with the general processor 201. For example, the heterogeneous processor 205 may include a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc. The heterogeneous processor 205 may include one or more processing cores, and usually the heterogeneous processor 205 includes multiple processing cores.

[0070] The memory 202 includes, but is not limited to, a random access memory (RAM), a read only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or an optical memory, a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. In the embodiment of the present application, the memory 202 can store information such as computer instructions.

[0071] In a possible implementation, the memory 202 may exist independently of the processor (such as the general-purpose processor 201 or the heterogeneous processor 205). The memory 202 may be connected to the processor via the bus 204 for storing data, instructions, or program codes. When the processor calls and executes the instructions or program codes stored in the memory 202, the relevant steps in the method provided in the embodiment of the present application can be implemented.

[0072] In another possible implementation, the memory 202 may also be integrated with the processor.

[0073] The communication interface 203 may be a transceiver module, which is used to communicate with other devices or communication networks, such as Ethernet, RAN, wireless local area networks (WLAN), etc. The communication interface 203 may receive instructions, messages or data, etc. The transceiver module may be a device such as a transceiver or a transceiver. Optionally, the communication interface 203 may also be a transceiver circuit located in the processor 201, which is used to implement signal input and signal output of the processor. The communication interface 203 may be a wired interface (port), such as a fiber distributed data interface (FDDI), a gigabit Ethernet (GE) interface, or the communication interface 203 may also be a wireless interface.

[0074] The above-mentioned general processor 201 (i.e., central processing unit CPU) is a general computing module, which is mainly responsible for the logical calculation and logical control functions of the load, and can efficiently process single complex calculation sequence tasks, but its performance in large-scale calculations is relatively low. Therefore, the general processor 201 can distribute large-scale calculation tasks (such as AI model training tasks) to the heterogeneous processor 205, and after the heterogeneous processor 205 completes the calculation, it returns the calculation result to the general processor 201.

[0075] The bus 204 may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus may also be divided into a serial bus and a parallel bus. For ease of representation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0076] It should be noted that Figure 2 The computing device shown is only one example of a computing device that can have more than Figure 2 More or fewer components may be shown, two or more components may be combined, or there may be a different configuration of components.

[0077] In combination with the above, the AI ​​model reuse method provided in the embodiment of the present application obtains a second AI model for processing a target task by reusing the first AI model for processing a source task. In the following embodiments, the training data set used to train the first AI model is referred to as the source training data set, and the training data set used to train the second AI model is referred to as the target training data set.

[0078] In the embodiment of the present application, the training data set (including the source training data set and the target training data set) includes a plurality of sample data. Optionally, the sample data may be an image (the sample data may also be referred to as a sample image).

[0079] It should be noted that, since the model reuse method of the embodiment of the present application is based on the category mapping cost between the category of the target task and the category of the source task, the first AI model used to process the source task is reused to obtain the second AI model used to process the target task. It can be seen that the source task and the target task involve categories. Therefore, the AI ​​model reuse method provided in the embodiment of the present application is mainly used for scenarios related to category detection. Taking the target detection scenario as an example, the category of the source task in the following embodiment refers to the detection type included in the source task, and the category of the target task refers to the detection type included in the target task. For example, the detection types included in the source task are 2 kinds of reptiles (reptiles of category 1 and reptiles of category 2), then the category of the source task includes two categories, namely category 1 and category 2.

[0080] The following is a detailed description of the reuse method of the AI ​​model provided in the embodiment of the present application. Figure 3As shown, the method includes S301-S304.

[0081] S301: Use the first AI model as the initialization model of the target task, and use the first AI model to process each sample data in the target training data set.

[0082] The first AI model is an AI model trained based on a source training data set of a source task, and the target training data set is used to train a second AI model for processing a target task.

[0083] Optionally, the first AI model comes from a model library, and the user can select a suitable pre-trained model from the model library based on the requirements of the target task and use it as the first AI model. For example, if the target task is to detect reptiles, a model for animal recognition can be selected from the model library as a pre-trained model for the target task; for another example, if the target task is to detect sandwiches and ice cream, a model for identifying hamburgers and bread can be selected from the model library as a pre-trained model for the target task.

[0084] It should be noted that both the first AI model and the second AI model can be models for single-target processing or models for multi-target processing. Among them, single target or multi-target can be understood as the category of objects that the AI ​​model can process is one category or multiple categories. For example, in the target detection scenario, the first AI model can detect two categories, such as using the first AI model to detect objects of category A1 and / or objects of category B1 in the image to be detected; the second AI model can detect three categories, such as using the second AI model to detect an image to be detected that includes at least one of objects of category A2, objects of category B2, or objects of category C2.

[0085] S302: Based on the processing result of each sample data in the target training data set, determine the category mapping cost of mapping the category of the target task to the category of the source task.

[0086] In an embodiment of the present application, a first AI model is used to process all sample data in a target training data set to obtain a processing result for each sample data. The processing result of the sample data processed by the AI ​​model may include the category confidence of the target object in the sample data.

[0087] It should be understood that the category confidence of the target object is the confidence that the target object belongs to each category. The higher the category confidence, the more reliable the target object belongs to the category. The category with the highest category confidence is determined as the category of the target object. For example, the categories detectable by the first AI model include category A1 and category B1. For a sample data, the category confidence in its processing result includes the confidence that the category of the target object is category A1 and the confidence that the category of the target object is category B1. If the confidence of category A1 is higher than the confidence of category B1, the category of the target object in the sample data is determined to be category A1.

[0088] Optionally, in a target detection scenario, the processing result of the sample data includes the category confidence of the target object in the sample data and the position information of the target object.

[0089] The target object's location information may be coordinate information of a detection frame used to mark the target object. For example, if the detection frame is a rectangular frame, the target object's location information may be coordinates of four vertices of the detection frame.

[0090] In an embodiment of the present application, the target training data set is a training data set corresponding to the target task. Assuming that the categories of the target task include things of category A2, things of category B2, and things of category C2, the sample data in the target training data set includes at least one of the above-mentioned things of category A2, things of category B2, or things of category C2.

[0091] For the processing results of the sample data in the target training data set using the first AI model, the category of the above-mentioned target task is mapped to the category of the source task, which can be understood as: a category in the sample data of the target task is recognized by the first AI model as a category of the source task. For example, the target object in the sample data is actually a sandwich. After the sample data is input into the first AI model, it is detected that the target object is a hamburger (with the highest confidence), then the sandwich category of the target task is mapped to the hamburger category of the source task.

[0092] It should be understood that the category mapping cost is used to measure the reusability between the first AI model and the second AI model. It can also be understood that the category mapping cost is used to measure the similarity between the category characteristics of the source task and the category characteristics of the target task. The higher the category similarity, the higher the reusability.

[0093] Combination Figure 3 ,like Figure 4 As shown, in one implementation, the above S302 can be implemented through the following S3021-S3022.

[0094] S3021. Determine the detection accuracy of the target task category being detected as the source task category according to the category confidence of the target object in each sample data in the processing result.

[0095] In the embodiment of the present application, there is a positive correlation between the detection accuracy and the category confidence, that is, the higher the category confidence, the higher the detection accuracy.

[0096] In one implementation, the detection accuracy and the category confidence may satisfy a functional relationship y=f(x), where y represents the detection accuracy and x represents the category confidence, for example, y=ax, a>0.

[0097] S3022: Determine, according to the detection accuracy, a category mapping cost at which the category of the target task is mapped to the category of the source task.

[0098] In the embodiment of the present application, the category mapping cost of the target task category being mapped to the category of the source task is negatively correlated with the detection accuracy when the target task category is detected as the category of the source task. In other words, the higher the detection accuracy when the target task category is detected as the category of the source task, the lower the category mapping cost of the target task category being mapped to the category of the source task.

[0099] According to the above description, the category mapping cost is determined based on the detection accuracy of each category of the target task being detected as a category of the source task when the target training data set is processed using the first AI model. Therefore, the category mapping cost is a matrix, which can also be called a cost matrix or an overhead matrix. In other words, the category mapping cost is the reuse cost or reuse overhead of the AI ​​model used to process the source task being reused to the AI ​​model used to process the target task.

[0100] In some embodiments, the category of the target task is mapped to the category of the source task by a category mapping cost satisfying:

[0101]

[0102] Among them, C represents the category mapping cost, which is a size d 1 ×d 2 The matrix, d 1 Indicates the number of categories of the source task (the number of detection types contained in the source task), d 2 Indicates the number of categories of the target task (that is, the number of detection types contained in the target task). ij Indicates the mapping cost of the target task category j being mapped to the source task category i, i = 1, 2, ..., d 1 , j=1,2,......,d 2 .

[0103] The calculation process of the above category mapping cost C is as follows:

[0104] Step 1: For each sample data in the target training data set, calculate the detection accuracy when the category j of the target task is detected as the category i of the source task when the sample data is processed by the first AI model.

[0105] In one implementation, the detection accuracy of the target task category j being detected as the source task category i may be AP (mean average precision).

[0106] The detection accuracy of the target task category is detected as the source task category, which is the following accuracy matrix AP:

[0107]

[0108] Among them, a ij It represents the detection accuracy when the category j of the target task is mapped to the category i of the source task.

[0109] Step 2: Take the inverse of the detection accuracy as the category mapping cost of mapping the target task category j to the source task category i.

[0110] Right now That is to say, The reuse cost value from the i-th category of the source task to the j-th category of the target task is filled into the cost matrix (ie, the category mapping cost matrix C).

[0111] In the above step 1, the detection accuracy of the target task category j being detected as the source task category i is calculated based on the processing results of the sample data.

[0112] In the target detection scenario, the first AI model is used to detect each sample data in the target training data set. The detection result (i.e., processing result) of each sample data includes the category confidence of the detected target object belonging to each category (the category of the source task) and the location information of the target object (which can be called a detection box or a prediction box).

[0113] Based on the description of the category mapping cost, it can be seen that the detection accuracy a when the category j of the target task is mapped to the category i of the source task ij The higher the value, the higher the similarity between the category features of the source task category i and the target task category j. The target task category j is mapped to the mapping cost c of the source task category i. ij The smaller it is, the smaller the reuse overhead between the first AI model and the second AI model is.

[0114] The above process of determining the detection accuracy of the target task category j being detected as the source task category i based on the detection results of the sample data is as follows:

[0115] S1. According to the detection results of each target object in each sample data, determine the intersection over union (IoU) of the predicted box and the real box of each target object in each sample data.

[0116] IoU is used to indicate the degree of overlap between the predicted box and the true box, which can measure the accuracy of the detection results.

[0117] Taking a target object in the sample data as an example, based on the detection results of the target object being detected as various categories of the source task, the detection box corresponding to the category of the source task with the maximum confidence in the detection results is compared with the real detection box of the target object to obtain the intersection-union ratio.

[0118] For example, the source task category includes two categories (respectively denoted as category O and category C). 1 and category O 2 ), the target task categories include 2 categories (respectively denoted as category T 1 and category T 2 ) as an example, for a sample data (i.e., a sample image), assuming that the sample data includes 7 targets, and the 7 targets include one or more categories of T 1 , and one or more categories of T 2 For one of the targets, the confidence of the detection result includes that the target is detected as category O 1 The confidence and the target object is detected as category O 2 The confidence of the two confidences is determined as the category corresponding to the larger confidence value of the two confidences as the predicted category of the target object. Referring to the following Table 1, the true category, true box, detection result (confidence and predicted box), predicted category and intersection-over-union ratio of the predicted box and the true box of the target object in the sample data are shown.

[0119] Table 1

[0120]

[0121] Combined with Table 1, for example, for the target object 1 in the sample data, the true category of the target object is category T 1 , the ground-truth box is K 1 , detected as category O in the detection results 1 The confidence level is Z 11 , detected as category O 2 The confidence level is Z 12 , where the maximum confidence is Z11 , then the predicted category of target object 1 is category O 1 , the corresponding prediction box is M 1 , calculate M 1 and K 1 The intersection over union ratio is IoU 11 According to Table 1 above, for this sample data, the calculated IoU includes IoU 1 、IoU 2 、IoU 3 、IoU 4 、IoU 5 、IoU 6 、IoU 7 .

[0122] S2. According to the IoU corresponding to all target objects of all samples in the target training dataset, the detection accuracy of each category of the target task is detected as each category of the source task.

[0123] The calculation process of the above detection accuracy is described by continuing to use the example shown in Table 1 as an example.

[0124] First, according to the contents of Table 1, we can calculate the IoU of each category of the target task when it is detected as each category of the source task. For example, according to the detection results of the 7 targets in the above sample data, among the 7 targets, the category T of the target task 1 The category O detected as the source task 1 There are 3 IoUs, namely IoU 1 、IoU 3 、IoU 7 ; Category of target task T 1 The category O detected as the source task 2 There is 1 IoU, which is IoU 6 ; Category of target task T 2 The category O detected as the source task 1 There is 1 IoU, which is IoU 4 ; Category of target task T 2 The category O detected as the source task 2 There are 2 IoUs, which are 2 and IoU 5 .

[0125] Similarly, for each sample data in the target training data set, statistical results similar to those in Table 1 can be obtained. The difference is that the number of target objects in other sample data may be different from the number of target objects in the sample data shown in Table 1, and / or the other sample data include one or two categories of target objects.

[0126] Secondly, the number of IoUs of various categories is counted for the target training dataset.

[0127] Continuing with the above example, for the target training data set, the IoU statistics include four types of IoU, namely: the category T of the target task 1 The category O detected as the source task 1 IoU, the category T of the target task 1 The category O detected as the source task 2 IoU, the category T of the target task 2 The category O detected as the source task 1 IoU, and the category T of the target task 2 The category O detected as the source task 2 Assuming that the target training data set includes N sample data, optionally, the number of four types of IoUs counted for the target training data set can be recorded as the following matrix U.

[0128]

[0129] Among them, for sample data 1, the number of the four types of IoU calculated is IoU 11 、IoU 12 、IoU 13 、IoU 14 ; For sample data 2, the number of the four IoUs calculated is IoU 21 、IoU 22 、IoU 23 、IoU 24 ; For sample data N, the number of the four IoUs calculated is IoU N1 、IoU N2 、IoU N3 、IoU N4 .

[0130] Combined with the above U matrix, the total number of each of the four types of IOU in the target training data set is determined, which are recorded as IoU1, IoU2, IoU3, and IoU4 respectively. Among them:

[0131] IoU1=IoU 11 +IoU 21 +……+Ou N1 ;

[0132] IoU2=IoU 11 +IoU 21 +……+Ou N1 ;

[0133] IoU3=IoU 11 +IoU21 +……+Ou N1 ;

[0134] IoU4=IoU 11 +IoU 21 +……+Ou N1 .

[0135] Finally, the detection accuracy when the category of the target task is detected as the category of the source task is determined according to the number of IoUs of each category that exceeds the preset threshold.

[0136] After counting the total number of each of the four types of IOUs in the target training data set, the number of IoUs exceeding the preset threshold in each type of IoU is determined according to the preset threshold (threshold specified by the AP mechanism). It should be understood that IoU exceeding the threshold indicates that the detection accuracy is high. Assume that the number of IoUs exceeding the preset threshold in the above IoU1 IoU is recorded as n 1 ; Among the 2 IoUs, the number of IoUs exceeding the preset threshold is recorded as n 2 ; Among the three IoUs, the number of IoUs exceeding the preset threshold is recorded as n 3 ; Among the 4 IoUs, the number of IoUs exceeding the preset threshold is recorded as n 4 , all sample data in the target training data set belong to category T 1 The number of objects is X, belonging to category T 2 The number of target objects is Y. Since the number of categories of the source task is 2 and the number of categories of the target task is also 2, the category mapping cost is a 2×2 matrix. For example, it is denoted as matrix AP:

[0137]

[0138] Among them, a 11 Represents the category of the target task T 1 The category O detected as the source task 1 The detection accuracy is a 12 Represents the category of the target task T 1 The category O detected as the source task 2 The detection accuracy is a 21 Represents the category of the target task T 2 The category O detected as the source task 1 The detection accuracy is a 22 Represents the category of the target task T 2 The category O detected as the source task 2 The detection accuracy is

[0139] S303: Determine a model reuse weight from the first AI model to the second AI model according to the category mapping cost.

[0140] Among them, the second AI model is used to process the target task, and the model reuse weight is the weight of mapping the processing result of the first AI model to the processing result of the second AI model. The model reuse weight is a matrix, so the model reuse weight can also be called the model transfer matrix from the first AI model to the second AI model, and the model transfer matrix can reflect the relationship between the output result of the first AI model and the output result of the second AI model for the same input. When the model reuse weight is known, the expected processing result when the second AI model is used to process the sample data in the target training data set can be determined based on the processing result of the first AI model on the sample data.

[0141] For example, for a sample data in the target training data set, after being processed by the first AI model, the processing result of the sample data is output, that is, the category confidence (the highest category confidence) is Z 1 , the model reuse weight is W, then the second AI model is used to process the sample data in the expected category confidence Z 2 is: Z 2 =Z 1 ×W.

[0142] The above model reuse weight can be expressed as:

[0143]

[0144] Among them, w ij Indicates the model reuse weight from the first AI model to the second AI model when the category j of the target task is detected as the category i of the source task.

[0145] In the embodiment of the present application, the process of reusing the first AI model of the source task to the target task is actually to find a transmission scheme with the minimum overall transmission cost (i.e., the minimum mapping cost) from one distribution to another, and the transmission scheme is the above-mentioned model reuse weight (W). Based on the known category mapping cost (the above-mentioned category mapping cost C) from each category of the source task to each category of the target task, the overall mapping cost of the first AI model reused to the second AI model can be expressed as:

[0146]

[0147] The model reuse weight W that minimizes the overall category mapping cost satisfies the following formula (1):

[0148]

[0149] stW×1=p,1×W=q Formula (1)

[0150] in, is the target optimization function, W×1=p and 1×W=q are the constraints.

[0151] In the constraints, p represents the probability distribution of the importance of sample data of each category in the source training data set, and q represents the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model.

[0152] Optionally, the probability distribution p of the importance of each category of sample data in the above-mentioned source training data set is the ratio of the number of sample data of each category in the source training data set to the total amount of sample data in the source training data set, and the probability value corresponding to each category in the probability distribution p reflects the detection ability of the first AI model for each category of the source task. The probability distribution q of the importance of the expected detection accuracy of each category of the target task using the second AI model is a uniform distribution.

[0153] For example, assume that the number of categories of the source task is d 1 =3, the number of categories of the target task is d 2 =2, q=[q 1 q 2 ],

[0154] Then the first constraint condition W×1=p mentioned above is:

[0155]

[0156] That is, p 1 =w 11 +w 12 , p 2 =w 21 +w 22 , p 3 =w 31 +w 32 .

[0157] The meaning of W×1=p in the above constraint condition is: the detection capability of the first AI model for each category of the source task is continued to be used in the second AI model, that is, the knowledge learned in the first AI model for detecting each category is transferred to the second AI model to the greatest extent possible with the source distribution. For example, the first AI model has a strong detection capability for category A1 of the source task. If the above constraint condition is met, when the first AI model is reused, the category knowledge about category A1 learned by the first AI model can continue to play a strong role in the target task.

[0158] Then the second constraint 1×W=q is:

[0159]

[0160] That is q 1 =w 11 +w 21 +w 31 ,q 2 =w 12 +w 22 +w 32 .

[0161] According to the content of the above embodiment, it can be known that q is uniformly distributed, that is, q in the above example 1 =q 2 , the above constraint 1×W=q means that after the first AI model is reused in the second AI model, the detection effect of the second AI model on each category of the target task is as similar as possible, that is, it is expected that the second AI model can achieve the best possible level in each category of the target task. For example, the target task includes two categories, and the purpose of the above constraint is to make the detection effect of the second AI model on the first category close to the detection effect of the second AI model on the second category.

[0162] Optionally, for the above formula (1), a sinkhorn algorithm may be used to solve the model reuse weight W. The detailed process of solving the sinkhorn algorithm may refer to relevant information of the prior art and will not be described in detail here.

[0163] S304. Obtain a second AI model according to the model reuse weight and the target training data set.

[0164] Optionally, combined Figure 3 ,like Figure 5 As shown, the above S304 can be implemented through S3041-S3042.

[0165] S3041. Determine the model loss based on the model reuse weight and the target training data set.

[0166] Combination Figure 5 ,like Figure 6 As shown, S3041 specifically includes S3041a-S3041b.

[0167] S3041a. Based on the model reuse weight, map the first category confidence of the target object to the second category confidence of the target object.

[0168] Among them, the first category confidence of the target object is obtained by processing the sample data in the target training data set using the first AI model.

[0169] It can be understood that the category confidence output by the AI ​​model is a vector, which can also be called a classification weight vector.

[0170] The specific process of mapping the first category confidence of the target object to the second category confidence of the target object is: multiplying the first category confidence of the target object by the model reuse weight to obtain the second category confidence of the target object. The second category confidence includes the confidence of each category of the sample data detected as the target task. The second category confidence is the expected processing result of the second AI model obtained after the current first AI model is reused.

[0171] In the embodiment of the present application, the above-mentioned mapping of the first category confidence of the target object to the second category confidence of the target object is a process of changing the processing result of the first AI model. The process of changing the category confidence involves category change (i.e., category mapping), that is, modifying the category of the target object in the sample data in the processing result (the category in the source task) to the actual category of the target object. For example, the target object in the sample data is identified by the first AI model as category A1 in the source task (such as hamburgers), but in fact the target object in the sample data is category A2 (sandwiches) in the target task. Then the confidence Z of the target object belonging to category A1 in the processing result is changed to 1 Modified to confidence Z that the target belongs to category A2 2 , Z 2 =Z 1 ×w, w represents the model reuse weight from the first AI model to the second AI model when the category of the target task is detected as the category of the source task.

[0172] S3041b. Determine the model loss based on the second category confidence.

[0173] It can be understood that the second category confidence of the target object obtained in S3041b is an estimated result of processing the sample data based on the second AI model. Optionally, the value of the category loss function (i.e., the category loss value) is determined based on the second category confidence of the target object and the true category confidence of the target object (e.g., 100%), and the category loss value is the model loss.

[0174] In some embodiments, for the target detection scenario, the processing result obtained by using the first AI model to process the sample data in the target training data set may also include the location information of the target object, and the location information is an estimated result of processing the sample data based on the second AI model. In this case, the above-mentioned model loss may be the sum of the category loss and the position loss, the category loss is determined based on the second category confidence of the target object, and the position loss is determined based on the location information of the target object. Optionally, the value of the position loss function (i.e., the position loss value) is determined based on the predicted location information of the target object and the actual location information of the target object in the sample data.

[0175] S3042. Update the parameters of the first AI model according to the model loss to obtain a second AI model.

[0176] Optionally, an error back propagation (BP) algorithm can be used to perform gradient back propagation on the model loss, and then the gradient information is used to update the parameters of the current first AI model.

[0177] It should be understood that the process of using the first AI model as the initialization model of the target task and continuously updating the first AI model with the target training data set to obtain the second AI model is a multi-iteration process (a cyclic process). Figure 7 , taking the first AI model as the initialization model, after obtaining the above-mentioned model reuse weight, the process of obtaining the second AI model based on the model reuse weight and the target training data set includes the following steps a to e.

[0178] Step a: input the i-th sample data in the target training data set into the current first AI model to obtain the processing result of the i-th sample data.

[0179] The value of i is one of 1, 2, ..., M, M is the number of sample data included in the target training data set, and M is a positive integer greater than or equal to 1. When i=1, the current first AI model is an initialization model selected from the model library.

[0180] The processing result of the i-th sample data includes the confidence of each target object in the sample data obtained by processing the sample data using the current first AI model. In the target detection scenario, the processing result also includes the location information of each target object.

[0181] Step b: Map the processing result of the i-th sample data according to the model reuse weight to obtain the prediction result of the i-th sample data.

[0182] The prediction result is an estimated result based on the processing of the sample data by the second AI model.

[0183] Step c: Determine the model loss.

[0184] Step d: Update the parameters of the current first AI model according to the model loss.

[0185] For the relevant contents of step a to step d, reference can be made to the description of the above embodiment and will not be repeated here.

[0186] Step e: determine whether the training end condition is met.

[0187] Optionally, the number of model training times can be set as the training end condition, and of course other conditions can also be used as the training end condition.

[0188] In the embodiment of the present application, if the training end condition is met, the first AI model after the parameters are updated in step e is used as the second AI model. If the training end condition is not met, i (i=i+1) is updated, that is, the next sample data is used, and the next round of training process is performed again from step a until the training end condition is met, and the first AI model updated for the last time is used as the second AI model.

[0189] In summary, based on the AI ​​model reuse method provided in the embodiment of the present application, in the process of reusing the pre-trained model (source model), the category of the source task is mapped to the category of the target task. The category mapping cost is introduced, and then the model reuse weight from the source model to the target model is determined based on the category mapping cost, and the model loss is determined according to the model reuse weight and the target training data set, and the parameters of the source model are updated according to the model loss to obtain the target model. In this method, since the category mapping cost of the source task category and the target task category can reflect the reusability between the source model used to process the source task and the target model used to process the target task, therefore, taking the category mapping cost as a consideration for model reuse can guide the model training process to better reuse the features of the pre-trained model and obtain a target model with better processing effect, that is, this method can improve the effect of model reuse.

[0190] Furthermore, model reuse based on category mapping cost can also guide the model training process to converge faster, which can save computing power and time overhead to a certain extent.

[0191] Combined with the above content, reference Figure 8 The complete process of the AI ​​model reuse method provided in the embodiment of the present application is described below. The method includes:

[0192] S701 . Count the number of sample data of each category in the source training data set, and generate a probability distribution p of the importance of the sample data of the source training data set.

[0193] S702: Generate a probability distribution q of the importance of prediction detection accuracy of sample data in the target training data set.

[0194] For the detailed description of S701 and S702, reference may be made to the description of S303 in the above embodiment, which will not be repeated here.

[0195] S703: Input each sample data in the target training data set into the first AI model to obtain a predicted output.

[0196] S704: Calculate the reuse accuracy of the first AI model according to the predicted output of the first AI model.

[0197] The reuse accuracy of the first AI model is used to determine whether to use the first AI model as the initialization model for the target task. The reuse accuracy of the first AI model is the detection accuracy of the category of the target task being detected as the category of the source task when the sample data in the target training data set is processed using the first AI model.

[0198] In one implementation, the reuse accuracy of the first AI model may be the sum of all elements in the accuracy matrix AP determined in S302. Assuming that the reuse accuracy of the first AI model is denoted as Y, then

[0199] It should be noted that when the detection accuracy of the first AI model is greater than the accuracy threshold, the first AI model is used as the initialization model of the second AI model, and steps S705-S708 and S710 are continued. When the detection accuracy of the first AI model is less than or equal to the accuracy threshold, the first AI model is not used as the initialization model of the second AI model. At this time, a randomly initialized model can be used as the initialization model of the second AI model, and the randomly initialized model is trained based on the target training data set, that is, S709 and steps after S710 are executed.

[0200] Determining the initialization model of the second AI model according to the detection accuracy described above helps to obtain a second AI model with better performance.

[0201] S705. Determine the category mapping cost C based on the predicted output of the first AI model.

[0202] S706. Determine a model reuse weight W according to the category mapping cost C, the probability distribution p, and the probability distribution q.

[0203] S707. Input the sample data i in the target training data set into the first AI model to obtain the predicted output of the sample data i.

[0204] S708. Map the predicted output of sample data i to the predicted output of the second AI model according to the model reuse weight W.

[0205] S709: Input the sample data i in the target training data set into the randomly initialized model to obtain the predicted output.

[0206] The randomly initialized model may be a model obtained by randomly initializing the parameters of the first AI model, or may be other models, which is not limited in the embodiments of the present application.

[0207] S710: Determine model loss.

[0208] S711. Update the parameters of the first AI model according to the model loss.

[0209] S712: Determine whether the end condition is met.

[0210] If the end condition of model training is met, the first AI model updated this time will be used as the second AI model; if the end condition of model training is not met, i (i=i+1) is updated, and then return to S707 to continue execution until the end condition is met, and the first AI model updated for the last time will be used as the second AI model.

[0211] Through the above S701-S712, a second model with better processing effect can be obtained, and model training based on category mapping cost can also guide the model training process to converge faster, which can save computing power and time overhead to a certain extent.

[0212] It is understandable that, in order to realize the above functions, the above computing device includes hardware structures and / or software modules corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the method steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0213] The embodiment of the present application can divide the functional modules of the above-mentioned computing device according to the above-mentioned method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0214] In the case of dividing each functional module into corresponding functional modules, Fig. 9 A possible structural diagram of the computing device involved in the above embodiment is shown. The computing device includes a processing module 801 , a first determining module 802 and a second determining module 803 .

[0215] Among them, the processing module 801 is used to execute S301, S703, S707, S708, and S709 in the above method embodiment; the first determination module 802 is used to execute S302 (including S3021-S3022), S303, S704, S705, and S706 in the above method embodiment; the second determination module 803 is used to execute S304 (including S3041-S3042), S710, and S711 in the above method embodiment.

[0216] In the case of an integrated unit, Fig.10 Another possible structural diagram of the computing device involved in the above embodiment is shown. The computing device may include: a processing module 901 and a communication module 902. The processing module 901 may be used to control and manage the actions of the computing device. For example, the processing module 901 may be used to support the computing device to execute the steps executed by the processing module 801, the first determination module 802, and the second determination module 803 in the above method embodiment, and / or other processes for the technology described herein. The communication module 902 may be used to support the communication between the computing device and other network entities. Optionally, as Fig.10 As shown, the computing device may further include a storage module 903 for storing program codes and data of the computing device.

[0217] The processing module 901 may be a processor, for example, the processor may be Figure 2 The general purpose processor 201 and the heterogeneous processor 205 in the communication module 902 can be a transceiver, a transceiver circuit or a communication interface, for example Figure 2 The communication interface 203 in the storage module 903 may be a memory, for example Figure 2 Memory 202 in.

[0218] The various modules of the above-mentioned computing device can also be used to perform other actions in the above-mentioned method embodiment. All relevant contents of each step involved in the above-mentioned method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.

[0219] For more details on how the modules included in the computing device implement the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here. Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.

[0220] Optionally, an embodiment of the present application further provides a computing device (for example, the computing device may be a chip or a chip system), which includes at least a processor (a general-purpose processor and / or a heterogeneous processor) and at least one interface. The interface may be used to receive signals from other devices (for example, memories of other electronic devices). For another example, the interface may be used to send signals to other devices (for example, processors). The interface may read instructions stored in the memory and send the instructions to the processor, which is used to read instructions to execute the method in any of the above method embodiments.

[0221] In one possible design, the computing device further includes a memory. The memory is used to store necessary program instructions and data, and the processor can read the computer instructions stored in the memory through an interface to enable the computing device to execute the method in any of the above method embodiments. Of course, the memory may not be in the computing device. When the computing device is a chip system, it may be composed of a chip, or may include a chip and other discrete devices, which is not specifically limited in the embodiments of the present application.

[0222] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instruction is loaded and executed on a computer, the process or function in accordance with the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instruction can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instruction can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) mode or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that includes one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a magnetic disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state drive (SSD)).

[0223] Through the description of the above implementation methods, technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0224] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0225] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0226] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0227] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, or all or part of the technical solution. The computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: flash memory, mobile hard disk, read-only memory, random access memory, disk or optical disk, and other media that can store program codes.

[0228] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for reusing artificial intelligence AI models. It is characterized in that include: Using the first AI model as the initialization model of the target task, and using the first AI model to process each sample data in the target training data set; wherein the first AI model is an AI model trained according to the source training data set of the source task; Determining, based on a processing result of each sample data in the target training data set, a category mapping cost at which the category of the target task is mapped to the category of the source task; Determining a model reuse weight from the first AI model to the second AI model according to the category mapping cost; the second AI model is used to process the target task; The second AI model is obtained according to the reuse weight and the target training data set.

2. The method according to claim 1, It is characterized in that The obtaining the second AI model according to the reuse weight and the target training data set includes: Determining a model loss based on the model reuse weight and the target training data set; According to the model loss, the parameters of the first AI model are updated to obtain the second AI model.

3. The method according to claim 1 or 2, It is characterized in that The determining, based on the processing result of each sample data in the target training data set, a category mapping cost of mapping the category of the target task to the category of the source task comprises: Determining the detection accuracy of the target task category being detected as the source task category according to the category confidence of the target object in each sample data in the processing result; A category mapping cost at which the category of the target task is mapped to the category of the source task is determined according to the detection accuracy.

4. The method according to claim 3, It is characterized in that There is a negative correlation between a category mapping cost at which the category of the target task is mapped to the category of the source task and a detection accuracy when the category of the target task is detected as the category of the source task.

5. The method according to any one of claims 1 to 4, It is characterized in that The model reuse weight is: Among them, d 1 represents the number of categories of the source task, d 2 Indicates the number of categories of the target task, i=1,2,......,d 1 , j=1,2,......,d 2 , w ij represents a model reuse weight from the first AI model to the second AI model when the category j of the target task is detected as the category i of the source task; The model reuse weight W satisfies: stW×1=p,1×W=q Among them, c ij It represents the category mapping cost of mapping the category j of the target task to the category i of the source task, p represents the probability distribution of the importance of the sample data of the target objects of each category in the source training data set, and q represents the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model.

6. The method according to any one of claims 1 to 5, It is characterized in that The determining of the model loss based on the model reuse weight and the target training data set includes: Based on the model reuse weight, mapping the first category confidence of the target object to the second category confidence of the target object; the first category confidence of the target object is obtained by processing the sample data in the target training data set using the first AI model; A model loss is determined based on the second category confidence.

7. The method according to any one of claims 1 to 6, It is characterized in that The probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model is uniformly distributed.

8. The method according to any one of claims 1 to 7, It is characterized in that The method further comprises: Determine the reuse accuracy of the first AI model; the reuse accuracy of the first AI model is used to determine whether to use the first AI model as the initialization model of the target task, and the reuse accuracy of the first AI model is the detection accuracy of the category of the target task being detected as the category of the source task when the target training data set is processed using the first AI model.

9. A computing device, It is characterized in that include: a processing module, a first determining module, and a second determining module; The processing module is used to use the first AI model as the initialization model of the target task, and use the first AI model to process each sample data in the target training data set; wherein the first AI model is an AI model trained according to the source training data set of the source task; The first determination module is used to determine the category mapping cost of mapping the category of the target task to the category of the source task based on the processing result of each sample data in the target training data set; The first determination module is further used to determine a model reuse weight from the first AI model to the second AI model according to the category mapping cost; the second AI model is used to process the target task; The second determination module is used to obtain the second AI model according to the reuse weight and the target training data set.

10. The computing device according to claim 9, It is characterized in that The second determination module is specifically used to determine the model loss based on the model reuse weight and the target training data set; and according to the model loss, update the parameters of the first AI model to obtain the second AI model.

11. A computing device according to claim 9 or 10, It is characterized in that The first determination module is specifically configured to determine the detection accuracy of the category of the target task being detected as the category of the source task according to the category confidence of the target object in each sample data in the processing result; And according to the detection accuracy, a category mapping cost of mapping the category of the target task to the category of the source task is determined.

12. The computing device according to claim 11, It is characterized in that There is a negative correlation between a category mapping cost at which the category of the target task is mapped to the category of the source task and a detection accuracy when the category of the target task is detected as the category of the source task.

13. A computing device according to any one of claims 9 to 12, It is characterized in that The model reuse weight is: Among them, d 1 represents the number of categories of the source task, d 2 Indicates the number of categories of the target task, i=1,2,......,d 1 , j=1,2,......,d 2 , w ij represents a model reuse weight from the first AI model to the second AI model when the category j of the target task is detected as the category i of the source task; The model reuse weight W satisfies: stW×1=p,1×W=q Among them, c ij It represents the category mapping cost of mapping the category j of the target task to the category i of the source task, p represents the probability distribution of the importance of the sample data of the target objects of each category in the source training data set, and q represents the probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model.

14. A computing device according to any one of claims 9 to 13, It is characterized in that The second determination module is specifically used to map the first category confidence of the target object to the second category confidence of the target object based on the model reuse weight; the first category confidence of the target object is obtained by processing the sample data in the target training data set using the first AI model; and determining the model loss based on the second category confidence.

15. A computing device according to any one of claims 9 to 14, It is characterized in that The probability distribution of the importance of the expected detection accuracy of each category of the target task using the second AI model is uniformly distributed.

16. A computing device according to any one of claims 9 to 15, It is characterized in that The first determination module is also used to determine the reuse accuracy of the first AI model; the reuse accuracy of the first AI model is used to determine whether to use the first AI model as the initialization model of the target task, and the reuse accuracy of the first AI model is the detection accuracy of the category of the target task being detected as the category of the source task when the target training data set is processed using the first AI model.

17. A computing device, It is characterized in that The method comprises a memory and at least one processor connected to the memory, wherein the memory is used to store computer program code, and the computer program code comprises computer instructions. When the computer instructions are executed by the at least one processor, the processor executes the method according to any one of claims 1 to 8.

18. A computer-readable storage medium, It is characterized in that Computer instructions are stored, and when the computer instructions are executed on a computer, the method according to any one of claims 1 to 8 is executed.

19. A chip, It is characterized in that The method comprises a processor, a memory and an interface; the processor is used to read the computer instructions stored in the memory through the interface, and run the computer instructions to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Remote sensing feature extraction deep reinforcement learning method based on multi-modal knowledge expression

    CN121682183A

  • Ai model reuse method and apparatus

    EP4807689A1

  • Ai model reuse method and apparatus

    WO2025107640A1