Multi-task model fusion method and device, storage medium and program product
Through the combination of particle swarm optimization algorithm and sparse technology, the multi-functional intelligent assistant has solved the problem of computing performance bottlenecks and storage space occupation on multi-task data sets, and achieved efficient model fusion and rapid response, which is suitable for scenarios such as intelligent customer service and intelligent code generation.
Patent Information
- Application Number
- CN202510381795.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
AI Technical Summary
When building multifunctional intelligent assistants, the existing technology faces computing performance bottlenecks and storage space occupation problems, especially when joint fine-tuning on multi-task data sets, it requires a large amount of video memory and high computing power resources, and model switching operations lead to storage bandwidth contention, affecting inference delay.
The particle swarm optimization algorithm (PSO) is used in combination with sparse technology, and by using the expert model and its sparse version as the initial particles, using the multi-task scoring mean as the optimization target, dynamically adjusting the particle velocity and parameters, generating a fusion model, and optimizing the calculation diagram and storage mapping table to reduce redundant parameters.
It significantly improves model fusion efficiency, reduces computer storage space usage and consumption of high computing resources, improves inference response speed and model applicability, and is suitable for multi-task scenarios.
Smart Images

Figure CN120354340A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer science and technology, and in particular, to a multi-task model fusion method, apparatus, storage medium, and program product. Background Art
[0002] Currently, large language models are widely used in various scenarios, such as intelligent customer service, machine translation, automated document generation, etc. However, these application scenarios often require fine-tuning the pre-trained model using specific datasets before being deployed to specific application scenarios. Currently, the construction method of multi-functional intelligent assistants based on large language models is mainly based on fine-tuning, and joint fine-tuning on the datasets of multiple tasks is performed to obtain a multi-task model. However, the traditional joint fine-tuning method for constructing multi-functional intelligent assistants still faces significant computer performance bottlenecks. The traditional joint fine-tuning method needs to perform intensive calculations on multiple task datasets simultaneously, which not only requires a video memory capacity of dozens of GB to support large-scale parameter updates, but also consumes high computing power resources of thousands of GPU hours. More critically, this process generates multiple independently stored copies of the expert model, and a single parameter model requires a large amount of storage space. When deploying a multi-task system, the I / O speed of the storage medium becomes the main source of response latency, and frequent model switching operations easily cause storage bandwidth contention, resulting in a significant increase in the actual inference latency. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art, the present invention proposes a multi-task model fusion method, apparatus, storage medium, and program product, which improves the efficiency of model fusion and the processing ability of multi-tasks, and reduces the occupation of computer storage space.
[0004] On the one hand, the present invention provides a multi-task model fusion method, including the following steps:
[0005] Obtain a number of expert models for different task scenarios;
[0006] Take the parameters of each expert model and the corresponding sparsified version of each expert model as the initial particle swarm; wherein, the sparsified version is generated by randomly masking the parameters of each expert model according to a preset discard rate;
[0007] Calculate the performance score of the fusion model for each task according to the corresponding preset evaluation index, and take the mean value of the performance scores of all tasks as the objective function of particle swarm optimization;
[0008] Use the particle swarm optimization algorithm to iteratively update the initial particle swarm, calculate the historical optimal solution and the global optimal solution of each particle based on the objective function, and dynamically adjust the particle velocity and parameters;
[0009] Take the global optimal particle parameters after iterative completion as the fused model parameters to generate a fused model.
[0010] In an embodiment of the present invention, using the sparsification technique, the parameters of each expert model are combined with the parameters of a given pre-trained model. By randomly masking the parameters of each expert model according to a preset dropout rate to generate sparsified expert model parameters, the corresponding sparsified version is obtained, which is expressed as:
[0011]
[0012] where m t is a mask tensor composed of Bernoulli random variables generated according to the dropout rate p, θ t are the parameters of the expert model, θ0 are the parameters of the pre-trained model, are the sparsified expert model parameters.
[0013] In an embodiment of the present invention, the particle velocity update formula is:
[0014]
[0015] where w is a parameter for controlling momentum, c1 and c2 are weight parameters for controlling the attention degree of the particle swarm optimization algorithm to global information and individual information, r1 and r2 are uniformly distributed random variables, and the update formula of the particle is θ gbest is the global optimal solution, θ t,pbest is the historical optimal solution of the particle.
[0016] In an embodiment of the present invention, deploying the fused model to a target terminal includes:
[0017] Receiving the global optimal particle parameter set generated by the fused model and performing mixed-precision quantization processing on the parameter set;
[0018] According to the computing architecture characteristics of the target terminal, reconstruct the computational graph of the fused model, merge the weight matrices with matching dimensions into a single tensor operation unit, and generate a lightweight inference engine including the SIMD instruction set;
[0019] Establish a task-parameter index mapping table, where the mapping table records the storage address ranges of parameter blocks corresponding to each task and the binary activation masks. When the target terminal receives a task request, according to the mapping table, slice and load the target parameter block to the video memory from the storage medium by address range, and dynamically activate the corresponding computing path by performing a bitwise AND operation with the binary activation mask.
[0020] In an embodiment of the present invention, multitasks are divided into text tasks and image tasks, and expert models corresponding to the text tasks and the image tasks are obtained respectively. The expert models at least include a visual expert model and a text expert model. The visual expert model is used to process image data input by a user; the text expert model is used to process text data input by the user.
[0021] In an embodiment of the present invention, the text task and the image task are respectively decomposed into a number of sub-text tasks and sub-image tasks, and the corresponding expert models are determined according to each sub-text task and sub-image task.
[0022] In an embodiment of the present invention, the performance scores corresponding to each sub-text task are calculated respectively and all the sub-text tasks are summarized to obtain the average score of the text task;
[0023] The performance scores corresponding to each sub-image task are calculated respectively and all the sub-image tasks are summarized to obtain the average score of the image task;
[0024] According to the average score of the text task and the average score of the image task, the objective function is determined.
[0025] In an embodiment of the present invention, the objective function is expressed as:
[0026]
[0027] where f(θ) is the objective function, is the performance score of the model parameter θ on the i-th sub-text task, n is the number of sub-text tasks, is the performance score of the model parameter θ on the j-th sub-image task, m is the number of sub-image tasks, and α and β are weight coefficients.
[0028] In an embodiment of the present invention, the method is applicable to intelligent code generation. The image data includes code screenshots, charts, and / or mathematical formula images; the text data includes: text instructions and / or code text;
[0029] The sub-image tasks include: extracting code text from code screenshots, analyzing the logical relationship of charts, and / or extracting formulas from mathematical formula images;
[0030] The visual expert model includes: an OCR model, a chart understanding model, and a mathematical formula recognition model; wherein, the OCR model is used to extract code text from code screenshots; the chart understanding model is used to analyze the logical relationship of charts to generate corresponding executable code; the mathematical formula recognition model is used to extract formulas from mathematical formula images and convert them into executable code;
[0031] The sub-text task includes parsing text instructions and / or parsing code text;
[0032] The text expert model includes: a code generation model, a mathematical reasoning model, and an instruction understanding model; wherein, the code generation model is used to parse the semantics of text instructions and generate executable code, or reconstruct the input code snippet into executable code; the mathematical reasoning model is used to parse the mathematical logic involved in text instructions and generate code for mathematical calculations or formulas; the instruction understanding model is used to parse text instructions, extract key requirements and logical structures, and provide context information for code generation.
[0033] In an embodiment of the present invention, the method is applicable to intelligent customer service generation; wherein, the text data includes dialogue text, and the image data includes screenshots of operation interfaces;
[0034] The sub-text task includes: identifying the text intention of dialogue text and / or generating multi-turn dialogue responses; the sub-image task includes: parsing screenshots of operation interfaces;
[0035] The text model includes: a fusion intention recognition model, which is used to identify text intentions and output user intention classification results; a dialogue generation model, which is used to generate multi-turn dialogue response texts in combination with user intention classification results and control type labels;
[0036] The visual expert model includes an interface parsing model, which is used to parse screenshots of operation interfaces, identify operable controls in the screenshots, and generate control type labels.
[0037] Another aspect of the present invention further provides a multi-task model fusion device, including:
[0038] An expert model acquisition module, which is used to acquire several expert models for different task scenarios;
[0039] An initial particle swarm generation module, which is used to use the parameters of each expert model and the sparse version corresponding to each expert model as the initial particle swarm; wherein, the sparse version is generated by randomly masking the parameters of each expert model according to a preset discard rate;
[0040] A target function construction module, which is used to calculate the performance score of the fusion model for each task according to the corresponding preset evaluation index, and use the mean value of the performance scores of all tasks as the target function for particle swarm optimization;
[0041] An optimization module, which is used to iteratively update the initial particle swarm using the particle swarm optimization algorithm, calculate the historical optimal solution and the global optimal solution of each particle based on the target function, and dynamically adjust the particle velocity and parameters;
[0042] A fusion module is used to take the globally optimal particle parameters after iteration as the fused model parameters, and generate a fusion model for collaborative processing of text input and image input.
[0043] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above multi-task model fusion method are implemented.
[0044] On the other hand, a computer program product of the present invention includes a computer program. When the computer program is executed by a processor, the steps of the above multi-task model fusion method are implemented.
[0045] As can be seen from the above solutions, the advantages of the present invention are as follows:
[0046] The multi-task model fusion method provided by the present invention, on the one hand, introduces the particle swarm optimization (PSO) algorithm into model fusion. By utilizing the swarm intelligence and gradient-free optimization characteristics of PSO, it avoids the computational bottleneck of traditional gradient methods in large-scale model fusion, significantly improves the fusion efficiency, and is applicable to multi-task scenarios. On the other hand, the expert model and its sparse version are used as the initial particles, which expands the scale and diversity of the particle swarm. This combination significantly improves the performance and stability of the fusion model, and at the same time provides more efficient support for processing large-scale expert models. This method not only improves the efficiency of model fusion, but also significantly improves the performance and applicability of the fusion model, reduces the dependence on large-scale data and computing resources, reduces the occupation of computer storage space, reduces the consumption of high computing power resources, and improves the inference response speed. Description of the Drawings
[0047] Figure 1 It is a schematic diagram of the overall process of the multi-task model fusion method provided by an embodiment of the present invention;
[0048] Figure 2 It shows the comparison of the technical effects between the present invention and other methods on simple tasks of small models;
[0049] Figure 3 It shows the comparison of the technical effects between the present invention and other methods on complex tasks of large models;
[0050] Figure 4 It is a schematic diagram of the structure of the multi-task model fusion device provided by an embodiment of the present invention.
[0051] Reference Signs:
[0052] 400: Multi-task model fusion device;
[0053] 410: Expert model acquisition module;
[0054] 420: Initial particle swarm generation module;
[0055] 430: Objective function construction module;
[0056] 440: Optimization module;
[0057] 450: Fusion module. Detailed implementation manner
[0058] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0059] Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.
[0060] Model fusion is a method that can integrate the knowledge of existing different expert models to construct a general artificial intelligence assistant that can perform well in the fields where each expert is good at. The purpose of model fusion is to build a fusion model with multi-task capabilities by integrating multiple fine-tuned expert language models, so as to reduce the dependence on large-scale data and computing resources. This method can efficiently utilize the existing expert models and avoid the high cost of fine-tuning pre-trained models from scratch. However, there are still deficiencies in the existing technologies in terms of fusion efficiency and computing resource requirements, especially when facing large-scale models and complex tasks, it is necessary to consider achieving the optimal fusion effect under limited resources. In this regard, the present invention considers applying the Particle Swarm Optimization (PSO) algorithm to model fusion. PSO does not require calculating gradients, but optimizes the objective through an evaluation function, which enables it to efficiently utilize limited data resources for optimization. In addition, PSO uses swarm information to guide the movement of particles in each iteration and can quickly converge to high-quality solutions, which is particularly important in multi-task model fusion. Therefore, the present invention applies the Particle Swarm Optimization (PSO) algorithm to model fusion. Specifically, it can adapt to the requirements of multi-task model fusion by initializing particles as the parameters of expert models and taking the average score of multiple tasks as the optimization objective. This method not only avoids the complexity of gradient calculation but also can make full use of task data to optimize the fusion process. At the same time, combined with the sparsification technology, by removing some parameters to reduce the redundancy between models, the performance of the fusion model can be improved and the fusion effect can be enhanced. The sparsification technology can not only be used to reduce parameter conflicts but also to construct more particles, thus expanding the scale of the particle swarm. By introducing the sparsification technology, more diverse initial particles can be generated, further enhancing the efficiency and effect of PSO in model fusion.
[0061] Based on the above considerations, the invention innovatively combines the Particle Swarm Optimization algorithm with the sparsification technology and proposes a multi-task model fusion method. By using the expert model and its sparsified version as the initial particles and taking the average of multi-task scores as the optimization objective, this method can quickly converge to a high-quality fusion model in a limited number of iteration steps. It not only improves the efficiency of model fusion but also significantly improves the performance and applicability of the fusion model, reduces the dependence on large-scale data and computing resources, reduces the occupancy of computer storage space, reduces the consumption of high computing power resources, and improves the inference response speed.
[0062] Refer to Figure 1 As shown in Figure 1 FIG. shows a schematic flowchart of the multi-task model fusion method provided by an embodiment of the present invention.
[0063] A multi-task model fusion method includes the following steps:
[0064] Step S1: Obtain a number of expert models for different task scenarios.
[0065] According to the functional requirements of the target multi-task system (such as code generation, machine translation, text classification, question answering tasks, etc.), select pre-trained models fine-tuned for specific tasks from the open-source community, enterprise private model libraries, or third-party API services to obtain a number of expert models for different task scenarios.
[0066] Step S2: Use the parameters of each expert model and the corresponding sparsified version of each expert model as the initial particle swarm; where the sparsified version is generated by randomly masking the parameters of each expert model according to a preset dropout rate.
[0067] Analyze the parameters of each expert model. The parameters of an expert model usually refer to the adjustable variables learned by the model during training, such as the weight matrix: including attention mechanism parameters, feed-forward network parameters, word embedding matrix, etc., the bias term: including the bias vectors of each layer, and the dynamically calculated parameters: including gating mechanism parameters, sparse activation parameters, etc.
[0068] Then, combined with the sparsification technology, calculate the sparsified version corresponding to each expert model, while reducing parameter conflicts, construct more particles, thereby expanding the scale of the particle swarm. Specifically, combine the parameters of each expert model with the parameters of a given pre-trained model, and generate sparsified expert model parameters by randomly masking the parameters of each expert model according to a preset dropout rate to obtain the corresponding sparsified version, expressed as:
[0069]
[0070] where, m t is a mask tensor composed of Bernoulli random variables generated according to the dropout rate p, θ t are the parameters of the expert model, θ0 are the parameters of the pre-trained model, are the sparsified expert model parameters.
[0071] The finally obtained initial particle swarm is the set of all expert models, sparsified expert models, and pre-trained models, that is
[0072] Step S3: Calculate the performance score of the fusion model for each task according to the corresponding preset evaluation index, and use the mean value of the performance scores of all tasks as the objective function for particle swarm optimization.
[0073] In this embodiment, calculate the performance score of the fusion model for each task according to the corresponding preset evaluation index, obtain the mean value of the performance scores of all tasks as the optimization objective, and calculate the average performance of each particle (model parameter) on all tasks through the evaluation function. The optimization objective function is defined as: where n is the number of tasks, and score i (θ) is the score of the model parameter θ on the i-th task. Based on the optimization objective, the historical best of each particle and the global best of all particles can be defined.
[0074] where the historical best solution and the global best solution are respectively
[0075] It should be noted here that the preset evaluation metrics for different tasks are not the same. For example, the preset evaluation metrics for the question-and-answer task are usually: EM value (Exact Match) and F1 score. The EM value is set to 1 when the generated answer is exactly the same as the reference answer, otherwise it is 0; the F1 score is calculated based on the overlap degree at the word level of the answer. The preset evaluation metric for the code generation task is usually Execution Accuracy, which is calculated by the proportion of the generated code passing the test cases. The preset evaluation metric for the machine translation task usually adopts BLEU (Bilingual Evaluation Understudy), which is calculated by comparing the n-gram overlap degree between the generated translation and the reference translation.
[0076] Step S4: Use the particle swarm optimization algorithm to iteratively update the initial particle swarm, calculate the historical best solution and the global best solution of each particle based on the objective function, and dynamically adjust the particle velocity and parameters.
[0077] The formula for updating the particle velocity is:
[0078]
[0079] where w is the parameter controlling the momentum, c1 and c2 are the weight parameters used to control the attention degree of the particle swarm optimization algorithm to the global information and individual information; r1 and r2 are uniformly distributed random variables used to introduce randomness to avoid falling into the local optimal solution; the update formula for the particle is θ gbest is the global best solution, and θ t,pbest is the historical best solution of the particle.
[0080] Step S5: Use the global best particle parameters after iteration as the fused model parameters, generate a fused model and deploy it to the target terminal.
[0081] Starting from the initial particle swarm Θ initial iterate S times, and the final global best particle is used as the fused model parameters, denoted as Then a fused model that can handle multiple tasks can be generated.
[0082] In one embodiment, the fusion model is further deployed to a target terminal, specifically as follows: receiving the globally optimal particle parameter set generated by the fusion model, and performing mixed-precision quantization processing on the parameter set. According to the computing architecture characteristics of the target terminal, reconstruct the computational graph of the fusion model, merge the weight matrices with matching dimensions into a single tensor operation unit, and generate a lightweight inference engine including the SIMD instruction set. Establish a task-parameter index mapping table, where the mapping table records the storage address range of the parameter block corresponding to each task and the binary activation mask. When the target terminal receives a task request, load the target parameter block into the video memory in slices according to the address range from the storage medium according to the mapping table, and dynamically activate the corresponding computing path by performing a bitwise AND operation with the binary activation mask. In this way, the fusion model can be deployed to the target terminal.
[0083] The multi-task model fusion method provided in this embodiment, on the one hand, introduces the Particle Swarm Optimization (PSO) algorithm into model fusion. By utilizing the swarm intelligence and gradient-free optimization characteristics of PSO, it avoids the computational bottleneck of traditional gradient methods in large-scale model fusion. At the same time, PSO accelerates convergence through group information sharing, significantly improving the fusion efficiency and being able to quickly achieve a high-quality fusion effect in a limited number of iteration steps, especially suitable for multi-task scenarios. It has a stable improvement compared with existing methods on different tasks. On the other hand, the expert model and its sparsified version are used as the initial particles, expanding the scale and diversity of the particle swarm. The sparsification technology not only effectively reduces parameter conflicts but also enhances the global search ability of PSO by generating more diverse initial solutions. This combination method significantly improves the performance and stability of the fusion model, while providing more efficient support for processing large-scale expert models. It has a stable improvement compared with existing methods on different tasks.
[0084] In one embodiment, when obtaining multi-tasks according to the functional requirements of the target multi-task system, the multi-tasks can be further decomposed into text tasks and image tasks, and the expert models corresponding to the text tasks and the image tasks are obtained respectively. The expert model includes at least a visual expert model and a text expert model. The visual expert model is used to process the image data input by the user; the text expert model is used to process the text data input by the user. Then, the text task and the image task are respectively decomposed into several sub-text tasks and sub-image tasks, and the corresponding expert models are determined according to each sub-text task and sub-image task. Furthermore, by calculating the performance scores corresponding to each sub-text task and summarizing all sub-text tasks, the average score of the text task is obtained; by calculating the performance scores corresponding to each sub-image task and summarizing all sub-image tasks, the average score of the image task is obtained; and then, according to the average score of the text task and the average score of the image task, the objective function can be determined. Among them, the objective function is expressed as:
[0085]
[0086] Among them, f(θ) is the objective function, is the performance score of the model parameter θ on the i-th sub-text task, n is the number of sub-text tasks, is the performance score of the model parameter θ on the j-th sub-image task, m is the number of sub-image tasks, and α, β are weight coefficients.
[0087] In this embodiment, multiple tasks are separated into text and image tasks, allowing vision expert models (such as ViT, ResNet) to focus on image feature extraction (such as edge detection, object localization), and text expert models (such as BERT, GPT) to focus on semantic understanding. Furthermore, the text tasks and image tasks are further divided into multiple sub-tasks, and the optimal model (such as mBART for translation, YOLOv7 for detection) is selected for each sub-task, avoiding the performance compromise of a single model on multiple sub-tasks. The scores of each sub-task are calculated independently (such as the translation BLEU and the detection mAP do not affect each other). During optimization, the task priorities can be balanced through weighted averaging to be applicable to different application scenarios. For example, in a medical scenario, the segmentation IoU weight is set to 0.6, and the classification weight is 0.4. In addition, in this embodiment, dynamic resource allocation can also be achieved, and expert model instances are dynamically scheduled according to the task load (such as expanding the NLP container during high-concurrency text requests). The vision model is preferentially deployed on the GPU (using CUDA to accelerate convolution), and the text model can be partially offloaded to the NPU (to accelerate attention calculation). Storage optimization can also be achieved, sharing the basic layer parameters (such as the vision and text models sharing the Embedding layer), reducing redundant parameters by more than 30%. The expert models are loaded on demand, and through LRU cache management (such as only retaining the models of high-frequency tasks in the video memory), the training cost is reduced, and the occupancy of the inference video memory is reduced. When adding new tasks, only the corresponding expert models need to be extended (such as inserting the Whisper model when adding speech tasks), without reconstructing the overall architecture, enhancing flexibility.
[0088] In one embodiment, taking the intelligent code generation system as an example, the system requirements are parsed to determine that the image data includes code screenshots, charts, and / or mathematical formula images. The text data includes text instructions and / or code text. The sub-image tasks include: extracting the code text in the code screenshots, parsing the logical relationships of the charts, and / or extracting the formulas in the mathematical formula images. The corresponding visual expert models include: an OCR model, a chart understanding model, and a mathematical formula recognition model; wherein, the OCR model is used to extract the code text in the code screenshots; the chart understanding model is used to parse the logical relationships of the charts to generate corresponding executable code; the mathematical formula recognition model is used to extract the formulas in the mathematical formula images and convert them into executable code. The sub-text tasks include parsing the text instructions and / or parsing the code text. The corresponding text expert models include: a code generation model, a mathematical reasoning model, and an instruction understanding model; wherein, the code generation model is used to parse the semantics of the text instructions and generate executable code, or reconstruct the input code snippets into executable code; the mathematical reasoning model is used to parse the mathematical logic involved in the text instructions and generate code for mathematical calculations or formulas; the instruction understanding model is used to parse the text instructions, extract the key requirements and logical structures, and provide context information for code generation.
[0089] In another embodiment, taking the intelligent customer service generation system as an example, the system requirements are parsed to determine that the text data includes dialogue text and the image data includes screenshots of the operation interface. The sub-text tasks include: identifying the text intent of the dialogue text and / or generating multi-turn dialogue responses; the sub-image task includes parsing the screenshots of the operation interface. The text models include: a fusion intent recognition model, which is used to identify the text intent and output the user intent classification result; a dialogue generation model, which is used to generate multi-turn dialogue response text by combining the user intent classification result and the control type label. The visual expert model includes an interface parsing model, which is used to parse the screenshots of the operation interface, identify the operable controls in the screenshots, and generate control type labels.
[0090] Next, the method of the present invention is compared with the prior art, referring to Figure 2 、 Figure 3 as shown in Figure 2 shows the comparison of the technical effects of the present invention and other methods on the simple tasks (GLUE) of the small model (Flan-T5-Base), Figure 3 shows the comparison of the technical effects of the present invention and other methods on the complex tasks (AlpacaEval, MBPP, GSM8K) of the large models (Llama-2-13B, Llama-3-8B, Mistral-7B-v0.3). As shown by Figure 2It can be seen that, compared with the baseline model, the average score of this method on the simple tasks (GLUE) of the small model (Flan-T5-Base) has increased by 2% - 5% compared with the baseline model. From Figure 3 It can be seen that on the complex tasks (AlpacaEval instruction following, MBPP code, GSM8K mathematics) of large models (Llama-2-13B, Llama-3-8B, Mistral-7B-v0.3), the maximum increase can reach 65%. Compared with the prior art, the present invention can significantly reduce the video memory occupancy. In summary, the present invention not only improves the efficiency of model fusion, but also significantly improves the performance and applicability of the fusion model, reduces the dependence on large-scale data and computing resources, reduces the occupancy of computer storage space, reduces the consumption of high computing power resources, and improves the inference response speed.
[0091] In addition, it should also be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0092] The following is a device embodiment corresponding to the above method embodiment. As Figure 4 shown, Figure 4 shows a schematic structural diagram of a multi-task model fusion device. The implementation manner of this device can be implemented in cooperation with the above method implementation manner. The relevant technical details mentioned in the above method implementation manner are still valid in the implementation manner of this device. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in the implementation manner of this device can also be applied to the above method implementation manner.
[0093] A multi-task model fusion device 400 includes at least:
[0094] An expert model acquisition module 410, configured to acquire a number of expert models for different task scenarios.
[0095] An initial particle swarm generation module 420, configured to use the parameters of each expert model and the sparse version corresponding to each expert model as the initial particle swarm; wherein, the sparse version is generated by randomly masking the parameters of each expert model according to a preset discard rate.
[0096] A target function construction module 430, configured to calculate the performance score of the fusion model for each task according to the corresponding preset evaluation index, and use the mean value of the performance scores of all tasks as the target function for particle swarm optimization.
[0097] An optimization module 440 is used to iteratively update the initial particle swarm by using a particle swarm optimization algorithm, calculate the historical optimal solution and the global optimal solution of each particle based on the objective function, and dynamically adjust the particle velocity and parameters.
[0098] A fusion module 450 is used to use the global optimal particle parameters after the iteration is completed as the fused model parameters to generate a fused model.
[0099] The device embodiments described above are merely illustrative. For example, the division of functional modules is only a logical functional division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present invention, the functional modules may be integrated in a processing unit, or each module may exist physically alone, or two or more modules may be integrated in a processing unit.
[0100] In addition, the above method embodiments can be implemented in whole or in part by software. When implemented using software, the above method embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-task model fusion method provided above. When the computer instructions or computer programs are loaded or executed on the computer, the processes or functions described in the method embodiments of the present invention are generated in whole or in part.
[0101] In addition, if the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium for storing a computer program for executing the multi-task model fusion method, so that a computer device (which may be a personal computer, a server, or a network device, etc.) can execute all or part of the steps of the methods described in the embodiments of the present invention.
[0102] It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0103] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.
Claims
1. A multi-task model fusion method, characterized in that, Including the following steps: Obtain a number of expert models for different task scenarios; Use the parameters of each expert model and the corresponding sparsified version of each expert model as the initial particle swarm; wherein, the sparsified version is generated by randomly masking the parameters of each expert model according to a preset dropout rate; For each task, calculate the performance score of the fusion model according to the corresponding preset evaluation index, and use the average value of the performance scores of all tasks as the objective function for particle swarm optimization; Use the particle swarm optimization algorithm to iteratively update the initial particle swarm, calculate the historical optimal solution and the global optimal solution of each particle based on the objective function, and dynamically adjust the particle velocity and parameters; Use the parameters of the global optimal particle after iteration as the parameters of the fused model, generate the fused model and deploy it to the target terminal.
2. The method according to claim 1, wherein Using the sparsification technology, combine the parameters of each expert model with the parameters of a given pre-trained model, and generate the sparsified expert model parameters by randomly masking the parameters of each expert model according to a preset dropout rate, to obtain the corresponding sparsified version, expressed as: where m t is a mask tensor composed of Bernoulli random variables generated according to the dropout rate p, and θ t are the parameters of the expert model, and θ0 are the parameters of the pre-trained model, are the sparsified expert model parameters.
3. The method according to claim 1, wherein The formula for updating the particle velocity is: Among them, w is a parameter for controlling momentum, c1 and c2 are weight parameters for controlling the degree of attention of the particle swarm optimization algorithm to global information and individual information, r1 and r2 are uniformly distributed random variables, and the update formula of the particle is θ gbest is the global optimal solution, and θ t,pbest is the historical optimal solution of the particle.
4. The method according to claim 1, wherein Deploying the fusion model to the target terminal includes: Receiving the set of global optimal particle parameters generated by the fusion model, and performing mixed-precision quantization processing on the parameter set; According to the computing architecture characteristics of the target terminal, reconstruct the computational graph of the fusion model, merge the weight matrices with matching dimensions into a single tensor operation unit, and generate a lightweight inference engine including the SIMD instruction set; Establish a task-parameter index mapping table, which records the storage address range of the parameter blocks corresponding to each task and the binary activation mask. When the target terminal receives a task request, load the target parameter block into the video memory in slices according to the address range from the storage medium, and dynamically activate the corresponding computing path by performing a bitwise AND operation with the binary activation mask.
5. The method according to any one of claims 1-4, characterized in that Divide the multi-task into text tasks and image tasks, and respectively obtain the expert models corresponding to the text tasks and the image tasks. The expert models at least include a visual expert model and a text expert model. The visual expert model is used to process the image data input by the user; the text expert model is used to process the text data input by the user.
6. The method according to claim 5, wherein Decompose the text task and the image task into a number of sub-text tasks and sub-image tasks respectively, and determine the corresponding expert models according to each sub-text task and sub-image task.
7. The method according to claim 6, wherein Calculate the performance scores corresponding to each sub-text task respectively and summarize all sub-text tasks to obtain the average text task score; Calculate the performance scores corresponding to each sub-image task respectively and summarize all sub-image tasks to obtain the average image task score; Determine the objective function according to the average text task score and the average image task score, wherein the objective function is expressed as: where \(f(\theta)\) is the objective function, is the performance score of the model parameter \(\theta\) on the \(i\)-th sub-text task, \(n\) is the number of sub-text tasks, is the performance score of the model parameter \(\theta\) on the \(j\)-th sub-image task, \(m\) is the number of sub-image tasks, and \(\alpha\), \(\beta\) are weight coefficients.
8. The method according to claim 7, wherein The method is applicable to intelligent code generation, and the image data includes code screenshots, charts, and / or mathematical formula images; the text data includes: text instructions and / or code text; The sub-image tasks include: extracting code text from code screenshots, parsing the logical relationships of charts, and / or extracting formulas from mathematical formula images; The visual expert models include: an OCR model, a chart understanding model, and a mathematical formula recognition model; among them, the OCR model is used to extract code text from code screenshots; the chart understanding model is used to parse the logical relationships of charts to generate corresponding executable code; the mathematical formula recognition model is used to extract formulas from mathematical formula images and convert them into executable code; The sub-text tasks include parsing text instructions and / or parsing code text; The text expert models include: a code generation model, a mathematical reasoning model, and an instruction understanding model; among them, the code generation model is used to parse the semantics of text instructions and generate executable code, or reconstruct the input code snippets into executable code; the mathematical reasoning model is used to parse the mathematical logic involved in text instructions and generate code for mathematical calculations or formulas; the instruction understanding model is used to parse text instructions, extract key requirements and logical structures, and provide context information for code generation.
9. The method according to claim 7, wherein The method is applicable to intelligent customer service generation; among them, the text data includes dialogue text, and the image data includes screenshots of operation interfaces; The sub-text tasks include: identifying the text intent of dialogue text and / or generating multi-round dialogue responses; the sub-image tasks include: parsing screenshots of operation interfaces; The text models include: a fusion intent recognition model, used to identify text intent and output the classification result of user intent; a dialogue generation model, used to generate multi-round dialogue response texts in combination with the classification result of user intent and control type labels; The visual expert model includes an interface parsing model, used to parse screenshots of operation interfaces, identify operable controls in the screenshots, and generate control type labels.
10. A multi-task model fusion device, characterized in that, Including: An expert model acquisition module, used to acquire a number of expert models for different task scenarios; An initial particle swarm generation module, used to use the parameters of each expert model and the corresponding sparse version of each expert model as the initial particle swarm; among them, the sparse version is generated by randomly masking the parameters of each expert model according to a preset discard rate; A target function construction module, used to calculate the performance score of the fusion model for each task according to the corresponding preset evaluation index, and use the mean value of the performance scores of all tasks as the target function for particle swarm optimization; An optimization module, used to iteratively update the initial particle swarm using the particle swarm optimization algorithm, calculate the historical optimal solution and global optimal solution of each particle based on the target function, and dynamically adjust the particle velocity and parameters; A fusion module, used to use the global optimal particle parameters after iteration as the model parameters after fusion to generate a fusion model.
11. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of claims 1-9 are implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-9.
Citation Information
Cited By
Intelligent customer service method and system based on sparse expert model
CN121597717A
Multi-modal large model merging method and device for operation and maintenance of industrial equipment
CN121682454A