Method and apparatus for fusing model parameters

By evaluating and splicing the local and global compatibility scores of deep neural network models, the model parameter incompatibility problem is solved, and the performance and stability of the fusion model are improved.

CN119720101BActive Publication Date: 2025-06-13SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510221599.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In application scenarios such as multitask learning, transfer learning, and model integration, model parameters often have incompatibility, affecting the performance of the fusion model and resulting in instability in output.

Method used

By obtaining the model parameters of each parameter position of the model to be fused, local compatibility and global compatibility are evaluated, and the model parameters are spliced ​​based on these scores to obtain the fusion parameters of each parameter position of the fusion model.

Benefits of technology

An accurate and intelligent parameter splicing strategy is realized to retain information to the greatest extent and improve the performance and robustness of the fusion model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119720101B_ABST
    Figure CN119720101B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and provides a method and device for fusing model parameters. The method includes: obtaining the model parameters at each parameter position in the model to be fused, where the model to be fused is constructed based on a deep neural network; evaluating each model parameter in the model to be fused respectively based on the model parameters of the model to be fused to obtain the local-level compatibility scores of each model parameter; constructing a histogram distribution corresponding to the model parameters of the model to be fused, and quantifying the global-level compatibility score of the model to be fused based on the histogram distribution; splicing the model parameters at each parameter position based on the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter to obtain the fused parameters at each parameter position in the fused model. The method provided by the present invention realizes an accurate and intelligent parameter splicing strategy, retains information to the greatest extent, and greatly improves the model performance and robustness of the fused model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method and device for fusing model parameters. Background Art

[0002] Deep neural networks have achieved remarkable results in many fields such as image recognition, natural language processing, and recommendation systems. In the practical application of complex tasks, in order to save training costs and improve training efficiency, generally, an initial model that performs well in each field based on the existing deep neural network is used for multi-task learning, transfer learning, and model integration to obtain a final fusion model that can be competent for complex tasks.

[0003] However, in application scenarios such as multi-task learning, transfer learning, and model integration, the problem of incompatible model parameters often occurs. Due to differences in data distribution, network structure, and optimization strategies during the training process of different models, the learned parameters often have incompatibilities, which not only affect the performance of the fusion model but also may lead to unstable output of the fusion model. Summary of the Invention

[0004] The present invention provides a method and device for fusing model parameters to solve the defect in the prior art that the model parameters to be fused often have incompatibilities, which not only affect the performance of the fusion model but also may lead to unstable output of the fusion model.

[0005] The present invention provides a method for fusing model parameters, including:

[0006] Obtaining the model parameters at each parameter position in the model to be fused, where the model to be fused is constructed based on a deep neural network;

[0007] Based on the model parameters of the model to be fused, respectively evaluating each model parameter in the model to be fused to obtain the local-level compatibility scores of the model parameters;

[0008] Constructing a histogram distribution corresponding to the model parameters of the model to be fused, and quantifying the global-level compatibility score of the model to be fused based on the histogram distribution;

[0009] Based on the local-level compatibility scores of the model parameters and the global-level compatibility scores corresponding to the model parameters, splicing the model parameters at each parameter position to obtain the fusion parameters corresponding to each parameter position in the fusion model.

[0010] A method for fusing model parameters provided by the present invention, which splices the model parameters at each parameter position based on the local-level compatibility scores of the respective model parameters and the global-level compatibility scores corresponding to the respective model parameters to obtain the fused parameters corresponding to each parameter position in the fused model, includes:

[0011] Performing score fusion based on the local-level compatibility scores of the respective model parameters and the global-level compatibility scores corresponding to the respective model parameters to obtain the dual compatibility scores of the respective model parameters;

[0012] Based on the dual compatibility scores of the respective model parameters, splicing the model parameters at each parameter position to obtain the respective fused parameters of the fused model.

[0013] A method for fusing model parameters provided by the present invention, which splices the model parameters at each parameter position based on the dual compatibility scores of the respective model parameters to obtain the respective fused parameters of the fused model, includes:

[0014] Determining the task complexity scores of each network layer;

[0015] When the task complexity score is greater than a preset score threshold, based on the dual compatibility scores of the model parameters under each network layer, performing weighted splicing on the model parameters at each parameter position under each network layer to obtain the respective fused parameters.

[0016] A method for fusing model parameters provided by the present invention, after determining the task complexity scores of each network layer, further includes:

[0017] When the task complexity score is not greater than the preset score threshold, comparing the dual compatibility scores corresponding to the model parameters under any network layer in the model to be fused, and using the model parameters corresponding to the dual compatibility score with the larger comparison result as the respective fused parameters.

[0018] A method for fusing model parameters provided by the present invention, determining the task complexity scores of each network layer, includes:

[0019] Based on at least one of the network layer number and network layer type to which the parameter position corresponding to the respective model parameter belongs, determining the task complexity score;

[0020] The network layer type includes at least one of a convolutional layer, a fully connected layer, and a pooling layer.

[0021] A method for fusing model parameters provided by the present invention, based on the model parameters of the model to be fused, respectively evaluating each model parameter in the model to be fused to obtain the local-level compatibility scores of each model parameter, including:

[0022] Extract the parameter matrix of the model parameters of the model to be fused;

[0023] Based on the uncertainty evaluation network and the parameter matrix of the model to be fused, respectively evaluate each model parameter to obtain the local-level compatibility scores of each model parameter.

[0024] A method for fusing model parameters provided by the present invention, based on the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter, splicing the model parameters at each parameter position to obtain the fused parameters corresponding to each parameter position in the fused model, and then including:

[0025] Obtain the fine-tuning sample data of the fused model;

[0026] Compare the fine-tuning sample data with the initial sample data of the model to be fused to obtain a sample distribution comparison result;

[0027] In the case where the sample distribution comparison result shows a difference, based on the double compatibility scores of each model parameter and the sample distribution comparison result, calculate the optimization weights of the fused parameters belonging to the same parameter position as each model parameter;

[0028] Based on the optimization weights of each fused parameter and the previous learning rates of each fused parameter in the previous iteration round, determine the learning rates of each fused parameter in the current iteration round;

[0029] Based on the optimization weights and learning rates of each fused parameter, update each fused parameter in the current iteration round to obtain updated fused parameters.

[0030] A method for fusing model parameters provided by the present invention, based on the optimization weights of each fused parameter and the previous learning rates of each fused parameter in the previous iteration round, determining the learning rates of each fused parameter in the current iteration round, expressed as:

[0031] ;

[0032] Wherein, represents the learning rates of each fused parameter in the current iteration round; represents the previous learning rates of each fused parameter in the previous iteration round; represents the dynamic adjustment coefficient; indicating the optimization weights of the respective fusion parameters; indicating the number of rows of the parameter matrix corresponding to the fusion parameter; indicating the number of columns of the parameter matrix corresponding to the fusion parameter.

[0033] The present invention also provides a fusion device for model parameters, including:

[0034] An acquisition unit that acquires the model parameters at each parameter position in the model to be fused, where the model to be fused is constructed based on a deep neural network;

[0035] A local-level evaluation unit that evaluates each model parameter in the model to be fused based on the model parameters of the model to be fused, and obtains the local-level compatibility scores of the respective model parameters;

[0036] A global-level evaluation unit that constructs a histogram distribution corresponding to the model parameters of the model to be fused, and quantifies the global-level compatibility score of the model to be fused based on the histogram distribution;

[0037] A fusion unit that splices the model parameters at each parameter position based on the local-level compatibility scores of the respective model parameters and the global-level compatibility scores corresponding to the respective model parameters, to obtain the fusion parameters corresponding to each parameter position in the fusion model.

[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the model parameter fusion method as described in any one of the above.

[0039] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the model parameter fusion method as described in any one of the above.

[0040] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the model parameter fusion method as described in any one of the above.

[0041] The model parameter fusion method and device provided by the present invention quantify the local-level compatibility scores of each model parameter of the model to be fused, as well as the global-level compatibility score of the model to be fused, and splice the model parameters at each parameter position based on the local-level compatibility scores of the respective model parameters and the global-level compatibility scores corresponding to the respective model parameters, to obtain the fusion parameters corresponding to each parameter position in the fusion model, implementing an accurate and intelligent parameter splicing strategy, retaining information to the greatest extent, and greatly improving the model performance and robustness of the fusion model. Description of the Drawings

[0042] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0043] Figure 1 is one of the flow schematic diagrams of the model parameter fusion method provided by the present invention;

[0044] Figure 2 is the second of the flow schematic diagrams of the model parameter fusion method provided by the present invention;

[0045] Figure 3 is the structural schematic diagram of the model parameter fusion device provided by the present invention;

[0046] Figure 4 is the structural schematic diagram of the electronic device provided by the present invention. Detailed Embodiments

[0047] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0048] To address the above problems, the present invention provides a model parameter fusion method to achieve efficient model parameter fusion, thereby improving the robustness and generalization ability of the model. Figure 1 is one of the flow schematic diagrams of the model parameter fusion method provided by the present invention, as Figure 1 shown, the method includes:

[0049] Step 110, obtaining the model parameters at each parameter position in the model to be fused, where the model to be fused is constructed based on a deep neural network;

[0050] Here, the model to be fused can refer to a deep neural network model pre-trained on a large-scale dataset. The features learned by the model to be fused on a general task can be transferred to other specific tasks, thereby accelerating the training process of the models corresponding to other specific tasks and improving the performance of the models.

[0051] Specifically, two pre-trained models related to the target task can be obtained and used as the models to be fused. Here, the models to be fused can be different instances of the same type, such as image recognition models, audio recognition models, etc. Additionally, the model parameters at each parameter position in the models to be fused can be accessed through the APIs provided by the deep learning framework, and the model parameters at each parameter position in multiple pre-trained models are obtained, thus obtaining the model parameters at each parameter position in the models to be fused. Here, the deep learning framework can be TensorFlow, PyTorch. It should be noted that there can be multiple models to be fused, at least two.

[0052] Step 120: Based on the model parameters of the models to be fused, evaluate each model parameter in the models to be fused respectively to obtain the local-level compatibility scores of each model parameter.

[0053] Here, the local-level compatibility scores can be used to measure the applicability of a single model parameter among all the model parameters of the models to be fused, and can reflect the performance of a single model parameter on a specific layer, task, or dataset, that is, to reflect the compatibility of a single model parameter with all the model parameters of the models to be fused.

[0054] Specifically, for the local-level compatibility scores of each model parameter of a single model to be fused, it can be obtained by locally evaluating the parameter values of the model parameters at each parameter position. For example, through a parameter uncertainty evaluation network, the uncertainty matrix of the parameter value of each model parameter can be calculated to obtain the local-level compatibility score of this model parameter, and the compatibility score of a single parameter dimension is obtained. Similarly, the local-level compatibility scores of each model parameter of another fused model can be calculated by the same method as described above.

[0055] Step 130: Construct the histogram distribution corresponding to the model parameters of the models to be fused, and quantify the global-level compatibility score of the models to be fused based on the histogram distribution.

[0056] Here, the global-level compatibility score can be used to measure the consistency and applicability of the model parameters of a single model to be fused among all the model parameters of the models to be fused, and can be used to reflect the overall distribution characteristics of the model parameters of a single model to be fused and the matching degree with the ideal distribution, that is, to reflect the compatibility of the model parameters of a single model to be fused as a whole with all the model parameters of the models to be fused.

[0057] Specifically, for the global-level compatibility score of a single model to be fused, a histogram distribution corresponding to the model parameters of the model to be fused is constructed, and based on the histogram distribution, the global-level compatibility score of the model to be fused is quantified. Specifically, first, the parameter values of the model parameters at each parameter position can be statistically counted, and the parameter value range is equally divided into u intervals, and the number of parameters in each interval is counted , and a histogram distribution corresponding to the model parameters of a single model to be fused is constructed. Then, the probability that the parameter value of the model parameter falls into the i th interval can be calculated , and it can be calculated through the following calculation formula:

[0058] = ;

[0059] In the formula, represents the probability that the parameter value of the parameter falls into the i th interval; represents the number of parameters in the i th interval; represents the number of rows of the parameter matrix corresponding to the model parameters of the model to be fused; represents the number of columns of the parameter matrix corresponding to the model parameters of the model to be fused.

[0060] Next, by calculating the information entropy of the model to be fused, in order to consider the model complexity and overfitting risk, a regularization term is introduced. Here, the information entropy of the model to be fused can be calculated through the following formula:

[0061] ;

[0062] In the formula, represents the information entropy of the fused model; represents the th model parameter, k represents the total number of model parameters of the fused model; represents the regularization coefficient; represents the L2 norm of the model parameters of the fused model, which is used to control the model complexity.

[0063] Then, through the method of calculating the information entropy of a single model to be fused described above, the information entropy of another model to be fused can be calculated. Then, through the information entropy difference between the two models to be fused, the global-level compatibility score of a single model to be fused can be calculated. Here, the global-level compatibility score of a single model to be fused can be calculated through the following formula:

[0064] ;

[0065] In the formula, and respectively represent the global - level compatibility scores of the models A and B to be fused, where G represents global; represents the global - uncertainty evaluation network; represents the information entropy of the model A to be fused; represents the information entropy of the model B to be fused. It can be seen from the formula that the global - level compatibility scores of the models to be fused are the same.

[0066] Step 140: Based on the local - level compatibility scores of the model parameters and the global - level compatibility scores corresponding to the model parameters, splice the model parameters at each parameter position to obtain the fused parameters corresponding to each parameter position in the fused model.

[0067] Specifically, for the model parameters corresponding to any same parameter position of the models to be fused, to calculate the fused parameter at this parameter position, first, the model parameters at this parameter position can be evaluated through the local - level compatibility score of this model parameter and the global - level compatibility score of the model to be fused to which this model parameter belongs, obtaining a double - compatibility score. Then, according to the double - compatibility score, the model parameters at this parameter position are spliced to obtain the fused parameter corresponding to this parameter position in the fused model. For example, the local - level compatibility score of this model parameter and the global - level compatibility score of the model to be fused to which this model parameter belongs can be weighted to calculate the double - compatibility score of this model parameter.

[0068] Then, based on the magnitude of the double - compatibility score, a parameter splicing method corresponding to the double - compatibility score can be selected, and according to the parameter splicing method, the model parameters of the models to be fused corresponding to this parameter position are spliced to obtain the fused parameter corresponding to this parameter position in the fused model, so as to implement an accurate and intelligent parameter splicing strategy, retain information to the greatest extent, and without increasing the inference cost. Similarly, for the model parameters corresponding to any same parameter position of the models to be fused, by calculating the fused parameter at this parameter position, the fused parameters corresponding to other parameter positions can be calculated.

[0069] It should be explained that the parameter splicing method here can include directly selecting the model parameter of any model to be fused at this parameter position as the fused parameter. Or, through the double - compatibility scores of the model parameters of each model to be fused at this parameter position, the model parameters of each model to be fused at this parameter position are weighted and calculated, and the result of the weighted calculation is used as the fused parameter.

[0070] In addition, the obtained fusion model can be regarded as a model with the learning ability of the models to be fused, which can complete the tasks that the models to be fused can complete respectively. More importantly, it can complete complex tasks composed of the tasks that the models to be fused can complete respectively.

[0071] It can be understood that deep neural networks have achieved remarkable achievements in many fields such as image recognition, natural language processing, and recommendation systems. Among them, the models to be fused in the field of image recognition can be image classification models, object detection models, image segmentation models, face recognition models, etc.; the models to be fused in the field of natural language processing can be translation models, text classification models, dialogue generation models, etc.; the models to be fused in the field of recommendation systems can be content-based recommendation models, sequential recommendation models, etc.

[0072] In one embodiment, taking the object detection model as an example of the model to be fused, first, a vehicle detection model for vehicle detection and a pedestrian detection model for pedestrian detection can be used as the models to be fused, and the model parameters at each parameter position in the vehicle detection model and the pedestrian detection model can be obtained. Then, the model parameters in the vehicle detection model and the pedestrian detection model can be evaluated respectively through the model parameters of the vehicle detection model and the pedestrian detection model to obtain the local-level compatibility scores of each model parameter, that is, the local-level compatibility scores of each model parameter in the vehicle detection model and the pedestrian detection model are obtained.

[0073] Furthermore, construct the histogram distributions corresponding to the vehicle detection model and the pedestrian detection model. Then, through the two constructed histogram distributions, quantify the global-level compatibility scores of the vehicle detection model and the pedestrian detection model, that is, quantify the global-level compatibility scores of all model parameters in the vehicle detection model and the global-level compatibility scores of all model parameters in the pedestrian detection model. Next, the model parameters corresponding to the vehicle detection model and the pedestrian detection model at each parameter position can be spliced through the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter to obtain a fusion model that can finally be used for autonomous driving.

[0074] The method provided by the embodiment of the present invention quantifies the local-level compatibility scores of each model parameter of the model to be fused and the global-level compatibility scores of the model to be fused. Through the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter, the model parameters at each parameter position are spliced to obtain the fusion parameters corresponding to each parameter position in the fusion model, realizing an accurate and intelligent parameter splicing strategy, retaining information to the greatest extent, and greatly improving the model performance and robustness of the fusion model.

[0075] Based on any of the above embodiments, step 140 includes:

[0076] Performing score fusion based on the local-level compatibility scores of the respective model parameters and the global-level compatibility scores corresponding to the respective model parameters to obtain the dual compatibility scores of the respective model parameters;

[0077] Based on the dual compatibility scores of the respective model parameters, splicing the model parameters at each parameter position to obtain the respective fusion parameters of the fusion model.

[0078] Specifically, score fusion can be performed by combining the local-level compatibility scores of the respective model parameters and the global-level compatibility scores corresponding to the models to be fused to which the respective model parameters belong. For example, weighted calculation can be performed to calculate the dual compatibility scores of the respective model parameters.

[0079] The dual compatibility scores here can be calculated through the following formula:

[0080] ;

[0081] ;

[0082] In the formula, , respectively represent the dual compatibility scores of any model parameters at the same parameter position in the models to be fused A and B; , respectively represent the global-level compatibility scores of the models to be fused A and B; , represent the local-level compatibility scores of any model parameters at the same parameter position in the models to be fused A and B, where L represents the local level; represents the weight coefficient, which is used to adjust the relative importance of the local-level compatibility score and the global-level compatibility score.

[0083] Then, through the dual compatibility scores of the respective model parameters, the model parameters at each parameter position can be spliced to obtain the respective fusion parameters of the fusion model.

[0084] It can be understood that the larger the value of the dual compatibility score, the stronger the compatibility and the better the performance of the model parameters of the model to be fused at this parameter position. Therefore, the model parameters with larger dual compatibility scores can be selected as the fusion parameters.

[0085] It can also be understood that the dual compatibility scores corresponding to the model parameters of different models to be fused at the same parameter position can be used as the weights for splicing and fusing the model parameters at this parameter position, and weighted splicing of the model parameters at this parameter position can be performed to obtain more comprehensive, accurate, and stronger-performance and more robust fusion parameters.

[0086] The method provided by the embodiments of the present invention can more flexibly adapt to tasks of different complexities by comprehensively considering local-level and global-level compatibility scores, thereby improving the performance of the fusion model. Moreover, the introduction of the global-level compatibility score takes into account the consistency of parameters in the entire model, which helps to enhance the generalization ability of the fusion model on unseen data.

[0087] Based on any of the above embodiments, the method of splicing the model parameters at each parameter position to obtain the respective fusion parameters of the fusion model according to the dual compatibility scores of the model parameters includes:

[0088] Determine the task complexity scores of each network layer;

[0089] When the task complexity score is greater than a preset score threshold, based on the dual compatibility scores of the model parameters under each network layer, perform weighted splicing on the model parameters at each parameter position under each network layer to obtain the respective fusion parameters.

[0090] Here, the task complexity score can be used to measure the complexity of the tasks undertaken by the model parameters and guide the selection of the fusion strategy. In addition, the preset score threshold here can be a preset score value, which can be used to measure the complexity of the tasks undertaken by the model parameters at this parameter position.

[0091] Specifically, by evaluating the complexity of the tasks undertaken by each network layer, factors such as the model structure, task requirements, and dataset characteristics can be analyzed. For example, it can be based on factors such as the depth, width, number of parameters, task type (such as classification, regression, generation, etc.), dataset size, and quality to determine the task complexity scores of each network layer of the model to be fused.

[0092] Then, the respective task complexity scores corresponding to the same network layer of different models to be fused can be averaged, or the larger task complexity score can be selected to obtain the task complexity scores of each network layer. Then, compare the task complexity score of this network layer with the preset score threshold.

[0093] When the task complexity score of the network layer is greater than the preset score threshold, it indicates that the task complexity of the network layer is relatively large and the task difficulty is relatively high. When dealing with high-complexity tasks, soft splicing is preferentially adopted to ensure the retention of information. Therefore, the "soft splicing" method is preferentially selected for splicing. Specifically, for the splicing of the model parameters at each parameter position under any network layer, the dual compatibility scores of the model parameters at each position under the network layer can be used as the fusion weights, and the model parameters corresponding to each parameter position under the network layer are weighted and spliced to obtain the fusion parameters at each parameter position under the network layer. Similarly, the method for obtaining the fusion parameters at each parameter position under any network layer can be used to calculate the fusion parameters at each parameter position when the task complexity scores of other network layers are greater than the preset score threshold.

[0094] Here, the fusion parameter at this parameter position can be calculated by the following formula:

[0095] + ;

[0096] In the formula, represents the fusion parameter at the position of the Nth layer of the fusion model; represents the model parameters at each parameter position under the Nth layer of the model A to be fused; represents the dual compatibility score of the model parameters at the Nth layer of the model A to be fused; represents the model parameters at each parameter position under the Nth layer of the model B to be fused; represents the dual compatibility score of the model parameters at the Nth layer of the model B to be fused.

[0097] The method provided by the embodiments of the present invention, by introducing the task complexity score of the model parameters and comparing the task complexity score of the model parameters with the preset score threshold, realizes an adaptive parameter splicing method with the model position, improves the accuracy and performance of the fusion parameters. And when the task complexity score is greater than the preset score threshold, based on the dual compatibility scores of the model parameters under each network layer, the model parameters under each network layer are weighted and spliced to obtain each fusion parameter, which can more flexibly adapt to tasks with different complexities, thereby improving the performance of the fusion model. By weighting and splicing the parameters of multiple models, this method can integrate the advantages of different models, reduce the possible biases or overfitting risks of a single model, and thus enhance the robustness of the fusion model.

[0098] Based on any of the above embodiments, after determining the task complexity scores of each network layer, it further includes:

[0099] When the task complexity score is not greater than a preset score threshold, compare the dual compatibility scores corresponding to the model parameters under any network layer in the to-be-fused model, and use the model parameters corresponding to the dual compatibility score with the larger comparison result as the respective fusion parameters.

[0100] Specifically, when the task complexity score is not greater than a preset score threshold, it indicates that the task difficulty and complexity at the parameter position are relatively low. For example, when dealing with tasks of low complexity, a "hard splicing" method may be selected to improve speed and reduce computational overhead. In detail, for the "hard splicing" method, by comparing the dual compatibility scores corresponding to the model parameters under any network layer in the to-be-fused model, the model parameters corresponding to the dual compatibility score with the larger comparison result are used as the respective fusion parameters. For example, "hard splicing" can be implemented through the following formula:

[0101] ;

[0102] In the formula, represents the fusion parameter under the Nth layer of the fusion model; represents the dual compatibility score of the model parameters of the to-be-fused model A under the Nth layer , and when it is greater than the dual compatibility score of the model parameters of the to-be-fused model B under the Nth layer, the fusion parameter under the Nth layer of the fusion model is the model parameter of the to-be-fused model A under the Nth layer ; otherwise, the fusion parameter under the Nth layer of the fusion model is the model parameter of the to-be-fused model B under the Nth layer .

[0103] It can be understood that the larger the value of the dual compatibility score, the better the parameter performance of the model parameters corresponding to the dual compatibility score. Then, the model parameters can be shared with the fusion model as the fusion parameters of the fusion model to improve the model performance of the fusion model.

[0104] The method provided by the embodiments of the present invention, when the task complexity score is not greater than a preset score threshold, compares the dual compatibility scores corresponding to the model parameters under any network layer in the to-be-fused model, and uses the model parameters corresponding to the dual compatibility score with the larger comparison result as the respective fusion parameters, achieving efficient, accurate, and performance-optimal model parameter selection, and reducing the computational amount for obtaining the fusion parameters.

[0105] Generally speaking, the method provided by the embodiments of the present invention, by introducing the task complexity score of model parameters, when the task complexity score is greater than the preset score threshold, based on the dual compatibility scores of each model parameter, weights and splices the model parameters at each parameter position to obtain each fusion parameter, so as to achieve "soft splicing". When the task complexity score is not greater than the preset score threshold, compare the dual compatibility scores corresponding to the model parameters at any parameter position in the model to be fused, and use the model parameters corresponding to the dual compatibility score with the larger comparison result as each fusion parameter, so as to achieve "hard splicing". That is, when it is detected that the complexity of the task increases, such as the data distribution becomes more complex, it is automatically adjusted to the "soft splicing" method to ensure higher model accuracy. On the contrary, in the case of simple tasks and high requirements for real-time performance, the hard splicing method will be preferentially selected to reduce the computational overhead.

[0106] It should be noted that an additional parameter sharing mechanism is introduced, so that the parameters of different models to be fused can "borrow" information from each other during the inference process. For example, when a certain model to be fused performs well on a specific task, the fusion model can improve the overall performance by sharing its parameter updates, realizing cross-model compatible optimization and optimizing through knowledge sharing and information complementarity between cross-models.

[0107] Based on any of the above embodiments, determining the task complexity scores of each network layer includes:

[0108] Determine the task complexity score based on at least one of the network layer number and network layer type to which the parameter position corresponding to each model parameter belongs;

[0109] The network layer type includes at least one of a convolutional layer, a fully connected layer, and a pooling layer.

[0110] Specifically, through a preset task complexity scoring rule, for example, the higher the network layer number, the higher the corresponding task complexity score; conversely, the lower the corresponding task complexity score. The task complexity scores corresponding to the convolutional layer, the fully connected layer, and the pooling layer gradually increase. For example, for lower layers, such as the input layer and the convolutional layer, hard splicing is preferably used.

[0111] Then, based on at least one of the network layer number and network layer type to which the parameter position corresponding to the model parameter belongs, calculate the task complexity score corresponding to the model parameter. Similarly, for other model parameters, calculate them in the same way.

[0112] In one embodiment, assume there are two pre-trained models (models to be fused) and , and the two models have multiple levels of parameters and , where N represents the serial number of the network layer, for example, the first layer, the second layer, etc. Thus, the task complexity score can be determined by at least one of the network layer number and the network layer type to which the parameter position corresponding to the model parameter belongs, and by selecting an appropriate splicing strategy for each layer, more efficient fusion can be achieved.

[0113] In addition to performing compatibility evaluation based on the local-level compatibility score and the global-level compatibility score, the method provided by the embodiments of the present invention also introduces the concept of multi-level splicing, that is, different splicing strategies are applied at different network layer numbers or network layer types of the fusion model. Thus, according to the complexity, parameter size, and task requirements reflected by the network layer number and / or network layer type, the splicing method is flexibly adjusted, further optimizing the model performance and fusion efficiency of the fusion model.

[0114] Based on any of the above embodiments, step 120 includes:

[0115] Extract the parameter matrix of the model parameters of the to-be-fused model;

[0116] Based on the uncertainty evaluation network and the parameter matrix of the to-be-fused model, evaluate each model parameter respectively to obtain the local-level compatibility score of each model parameter.

[0117] Specifically, the parameter matrix corresponding to the model parameters of each to-be-fused model can be extracted respectively. For example, the parameter matrix corresponding to the to-be-fused model A can be expressed as . Then, by inputting the parameter matrices of each to-be-fused model into the uncertainty evaluation network, and evaluating each model parameter through the uncertainty evaluation network respectively, the local-level compatibility score of each model parameter can be obtained. For example, the local-level compatibility scores of each model parameter of the to-be-fused models A and B are obtained respectively. Here, the local-level compatibility score of any model parameter can be calculated by the following formula:

[0118] ;

[0119] In the formula, represents the local-level compatibility score of the model parameter at the parameter position of the to-be-fused model A; represents the local-level uncertainty evaluation network, and L represents the local level; represents the parameter matrix of the to-be-fused model A; represents the parameter matrix of the to-be-fused model B.

[0120] The method provided by the embodiment of the present invention performs parameter-level uncertainty evaluation on the parameter matrix corresponding to the model parameters of the model to be fused through an uncertainty evaluation network, and obtains the accurate parameter performance of each model parameter among all model parameters, laying a foundation for the optimal splicing of subsequent model parameters.

[0121] It should be noted that traditional methods such as pruning, model integration, and transfer learning are mostly static and do not take into account the dynamics of the input data distribution changes. Therefore, when facing the continuously changing data in actual applications, the model performance of the fused model obtained after fusion often performs poorly. In view of this problem, based on any of the above embodiments, after step 140, it includes:

[0122] Obtain the fine-tuning sample data of the fused model;

[0123] Compare the fine-tuning sample data with the initial sample data of the model to be fused to obtain a sample distribution comparison result;

[0124] In the case where the sample distribution comparison result shows a difference, based on the dual compatibility scores of the model parameters and the sample distribution comparison result, calculate the optimization weights of the fusion parameters belonging to the same parameter position as each model parameter;

[0125] Based on the optimization weights of the fusion parameters and the previous learning rates of the fusion parameters in the previous iteration round, determine the learning rates of the fusion parameters in the current iteration round;

[0126] Based on the optimization weights and learning rates of the fusion parameters, update the fusion parameters in the current iteration round to obtain updated fusion parameters.

[0127] Here, the fine-tuning sample data refers to the training sample data required for fine-tuning and optimizing the fused model. The initial sample data of the model to be fused here refers to the training data set used to pre-train the model to be fused. It should be noted that the fine-tuning sample data is usually the sample data collected in the specific scenario where the fused model is applied, while the initial sample data of the model to be fused is usually the sample data related to the task performed by the fused model. The initial sample data is simpler and less complex than the fine-tuning sample data.

[0128] Specifically, the sample data under the specific task performed by the fused model can be obtained as the fine-tuning sample data. For example, it can be image data for the autonomous driving task. In addition, the initial sample data corresponding to the model to be fused can also be obtained. For example, it can be image data used for pedestrian detection and image data used for vehicle detection.

[0129] Next, by comparing the fine-tuned sample data with the initial sample data of the model to be fused, a sample distribution comparison result can be obtained. For example, it can be calculating the statistical features of the fine-tuned sample data and the initial sample data, such as mean, variance, distribution form, etc. Then, by comparing the statistical features of the fine-tuned sample data and the initial sample data, a sample distribution comparison result is obtained. The sample distribution comparison result here can be calculated by the following formula:

[0130] ;

[0131] In the formula, represents the sample distribution comparison result; represents the distribution of the fine-tuned sample data; represents the distribution of the initial sample data. Thus, by calculating the difference between these two distributions, it can be identified whether there is a large deviation between the fine-tuned sample data and the initial training data, which can be regarded as when it is greater than the preset difference threshold, the sample distribution comparison result is that there is a difference.

[0132] Furthermore, in the case where the sample distribution comparison result shows a difference, based on the double compatibility scores of each model parameter and the sample distribution comparison result, the optimization weight of the fusion parameter belonging to the same parameter position as each model parameter is calculated. For example, for the optimization weight of the fusion parameter at any parameter position in the fusion model, it can be calculated through the double compatibility score and the sample distribution comparison result of any model to be fused at the same parameter position, and the optimization weight at this parameter position is calculated. Similarly, the optimization weights of other parameter positions in the fusion model can be calculated by the method of calculating the optimization weight at this parameter position. The optimization weight here can be calculated by the following formula:

[0133] ;

[0134] In the formula, represents the optimization weight of the fusion parameter at the parameter position in the fusion model; represents the double compatibility score of the model to be fused A at the parameter position ; represents the adjustment coefficient, which is used to control the intensity of dynamic adjustment.

[0135] Thus, according to the data distribution difference, the optimization weight of the fusion parameter can be flexibly adjusted to cope with the change of the fine-tuned sample data.

[0136] Further, in the fine-tuning stage, the learning rate of each fusion parameter in the current iteration can be determined by the optimization weight of each fusion parameter and the previous learning rate of each fusion parameter in the previous iteration, so as to realize the dynamic adjustment of the learning rate based on parameter compatibility and improve the fine-tuning effect.

[0137] Finally, update each fusion parameter in the current iteration through the optimization weight and learning rate of each fusion parameter to obtain updated fusion parameters, and cycle in this way until the iteration reaches the preset number of times, end the fine-tuning training of the fusion model, and obtain a fusion model with accurate, reliable prediction results, better model performance, and strong robustness.

[0138] The method provided by the embodiment of the present invention calculates the optimization weight of the fusion parameter belonging to the same parameter position as each model parameter based on the double compatibility score of each model parameter and the sample distribution comparison result in the case that the sample distribution comparison result shows a difference; determines the learning rate of each fusion parameter in the current iteration based on the optimization weight of each fusion parameter and the previous learning rate of each fusion parameter in the previous iteration; updates each fusion parameter in the current iteration based on the optimization weight and learning rate of each fusion parameter to obtain updated fusion parameters, which alleviates the problem of poor optimization effect of model parameters in the case that there is a distribution difference between the fine-tuning sample data and the initial sample data of the pre-trained model, and greatly improves the training effect of the fine-tuning training of the fusion model.

[0139] Based on any of the above embodiments, determining the learning rate of each fusion parameter in the current iteration based on the optimization weight of each fusion parameter and the previous learning rate of each fusion parameter in the previous iteration is expressed as:

[0140] ;

[0141] wherein, represents the learning rate of each fusion parameter in the current iteration; represents the previous learning rate of each fusion parameter in the previous iteration; represents the dynamic adjustment coefficient; represents the optimization weight of each fusion parameter; represents the number of rows of the parameter matrix corresponding to the fusion parameter; represents the number of columns of the parameter matrix corresponding to the fusion parameter.

[0142] Based on any of the above embodiments, Figure 2 is the second flow diagram of the fusion method of the model parameters provided by the present invention, as Figure 2 shown, the method includes:

[0143] First, obtain the model parameters at each parameter position in the model to be fused. Then, perform local-level parameter uncertainty assessment based on the model parameters, that is, based on the model parameters of the model to be fused, evaluate each model parameter in the model to be fused respectively to obtain the local-level compatibility scores of each model parameter; here, the local compatibility matrix of each network layer can also be obtained through the local-level compatibility scores of each model parameter, which can also be regarded as the local-level compatibility scores. In addition, perform global model information content assessment based on the model parameters, that is, construct the histogram distribution corresponding to the model parameters of the model to be fused, and quantify the global-level compatibility score of the model to be fused based on the histogram distribution.

[0144] Next, dual-weight parameter compatibility assessment can be performed through the local-level compatibility scores and the global-level compatibility scores to obtain the dual compatibility scores. Further, perform adaptive splicing of multi-level parameters on the model parameters of the model to be fused, that is, for the model parameters of different network layers, select the parameter splicing method adapted to it to splice the model parameters, obtain more accurate fusion parameters, and then obtain a fusion model with better performance and stronger robustness. It should be noted that the model parameter fusion method provided by the embodiments of the present invention can be applied to the fusion of multiple pre-trained models. For example, for multiple pre-trained models M 1 ,M 2 ,…,M n .

[0145] After obtaining the fusion model, fine-tuning of the fusion model can be performed. Specifically, obtain the fine-tuning sample data of the fusion model; compare the fine-tuning sample data with the initial sample data of the model to be fused to obtain the sample distribution comparison result; in the case where the sample distribution comparison result shows a difference, calculate the optimization weights of the fusion parameters belonging to the same parameter position as each model parameter based on the dual compatibility scores of each model parameter and the sample distribution comparison result; determine the learning rates of each fusion parameter in the current iteration round based on the optimization weights of each fusion parameter and the previous learning rates of each fusion parameter in the previous iteration round; update each fusion parameter in the current iteration round based on the optimization weights and learning rates of each fusion parameter to obtain the updated fusion parameters.

[0146] The method provided by the embodiments of the present invention proposes two intelligent parameter splicing strategies of precise splicing and weighted fusion, which retain information to the greatest extent and do not require an increase in the inference cost. Through multi-model extension, the generalization ability and robustness of the model are further enhanced on the basis of multiple pre-trained models, which is particularly suitable for complex tasks and multi-scenario applications, and enhances the flexibility and scalability of the model. This method is applicable to multiple fields such as computer vision and natural language processing, significantly improving the overall performance of the model and having broad application prospects.

[0147] Based on any of the above embodiments,Figure 3 is a schematic structural diagram of a fusion device for model parameters provided by the present invention. As Figure 3 shown, the device includes:

[0148] An acquisition unit 310 that acquires model parameters at each parameter position in the model to be fused, where the model to be fused is constructed based on a deep neural network;

[0149] A local-level evaluation unit 320 that evaluates each model parameter in the model to be fused based on the model parameters of the model to be fused, and obtains local-level compatibility scores for each model parameter;

[0150] A global-level evaluation unit 330 that constructs a histogram distribution corresponding to the model parameters of the model to be fused, and quantifies the global-level compatibility score of the model to be fused based on the histogram distribution;

[0151] A fusion unit 340 that splices the model parameters at each parameter position based on the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter, and obtains fusion parameters corresponding to each parameter position in the fusion model.

[0152] The device provided by the embodiments of the present invention quantifies the local-level compatibility scores of each model parameter of the model to be fused and the global-level compatibility score of the model to be fused. By using the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter, the model parameters at each parameter position are spliced to obtain fusion parameters corresponding to each parameter position in the fusion model, realizing an accurate and intelligent parameter splicing strategy, retaining information to the greatest extent, and greatly improving the model performance and robustness of the fusion model.

[0153] Based on any of the above embodiments, the fusion unit is specifically applied to:

[0154] Perform score fusion based on the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter to obtain double compatibility scores for each model parameter;

[0155] Splice the model parameters at each parameter position based on the double compatibility scores of each model parameter to obtain each fusion parameter of the fusion model.

[0156] Based on any of the above embodiments, the fusion unit is further specifically applied to:

[0157] Determine the task complexity scores of each network layer;

[0158] When the task complexity score is greater than a preset score threshold, based on the dual compatibility scores of the model parameters at each network layer, the model parameters at each parameter position under each network layer are weighted and concatenated to obtain the respective fusion parameters.

[0159] Based on any of the above embodiments, the fusion unit is further specifically applied to:

[0160] When the task complexity score is not greater than the preset score threshold, compare the dual compatibility scores corresponding to the model parameters under any network layer in the model to be fused, and use the model parameters corresponding to the dual compatibility score with the larger comparison result as the respective fusion parameters.

[0161] Based on any of the above embodiments, the fusion unit is further specifically applied to:

[0162] Determine the task complexity score based on at least one of the network layer number and network layer type to which the parameter positions corresponding to the respective model parameters belong;

[0163] The network layer type includes at least one of a convolutional layer, a fully connected layer, and a pooling layer.

[0164] Based on any of the above embodiments, the local-level evaluation unit is specifically applied to:

[0165] Extract the parameter matrix of the model parameters of the model to be fused;

[0166] Based on the uncertainty evaluation network and the parameter matrix of the model to be fused, evaluate each of the model parameters to obtain the local-level compatibility scores of the respective model parameters.

[0167] Based on any of the above embodiments, an optimization unit is included after the fusion unit, and the optimization unit is specifically used for:

[0168] Obtain the fine-tuning sample data of the fusion model;

[0169] Compare the fine-tuning sample data with the initial sample data of the model to be fused to obtain a sample distribution comparison result;

[0170] When the sample distribution comparison result shows a difference, calculate the optimization weights of the fusion parameters belonging to the same parameter position as each model parameter based on the dual compatibility scores of the respective model parameters and the sample distribution comparison result;

[0171] Based on the optimization weights of the respective fusion parameters and the previous learning rates of the respective fusion parameters in the previous iteration round, determine the learning rates of the respective fusion parameters in the current iteration round;

[0172] Based on the optimized weights and learning rates of the respective fusion parameters, update the respective fusion parameters in the current iteration round to obtain updated fusion parameters.

[0173] Based on any of the above embodiments, the optimization unit is further specifically configured to:

[0174] ;

[0175] Wherein, represents the learning rate of the respective fusion parameters in the current iteration round; represents the previous learning rate of the respective fusion parameters in the previous iteration round; represents the dynamic adjustment coefficient; represents the optimized weight of the respective fusion parameters; represents the number of rows of the parameter matrix corresponding to the fusion parameter; represents the number of columns of the parameter matrix corresponding to the fusion parameter.

[0176] Figure 4 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the method for fusing model parameters. The method includes: obtaining the model parameters at each parameter position in the model to be fused, where the model to be fused is constructed based on a deep neural network; respectively evaluating the model parameters at each model parameter position in the model to be fused based on the model parameters of the model to be fused to obtain the local-level compatibility scores of the respective model parameters; constructing a histogram distribution corresponding to the model parameters of the model to be fused, and quantifying the global-level compatibility score of the model to be fused based on the histogram distribution; and splicing the model parameters at each parameter position based on the local-level compatibility scores of the respective model parameters and the global-level compatibility scores corresponding to the respective model parameters to obtain the fusion parameters corresponding to each parameter position in the fusion model.

[0177] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0178] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the model parameter fusion method provided by the above-mentioned various methods. The method includes: obtaining the model parameters at each parameter position in the model to be fused, where the model to be fused is constructed based on a deep neural network; based on the model parameters of the model to be fused, respectively evaluating each model parameter in the model to be fused to obtain the local-level compatibility scores of each model parameter; constructing a histogram distribution corresponding to the model parameters of the model to be fused, and based on the histogram distribution, quantifying the global-level compatibility score of the model to be fused; based on the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter, splicing the model parameters at each parameter position to obtain the fusion parameters corresponding to each parameter position in the fusion model.

[0179] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the model parameter fusion method provided by the above-mentioned various methods. The method includes: obtaining the model parameters at each parameter position in the model to be fused, where the model to be fused is constructed based on a deep neural network; based on the model parameters of the model to be fused, respectively evaluating each model parameter in the model to be fused to obtain the local-level compatibility scores of each model parameter; constructing a histogram distribution corresponding to the model parameters of the model to be fused, and based on the histogram distribution, quantifying the global-level compatibility score of the model to be fused; based on the local-level compatibility scores of each model parameter and the global-level compatibility scores corresponding to each model parameter, splicing the model parameters at each parameter position to obtain the fusion parameters corresponding to each parameter position in the fusion model.

[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0181] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for fusing model parameters, characterized in that: include: Obtaining model parameters of each parameter position in the model to be fused, wherein the model to be fused is constructed based on a deep neural network, and the model to be fused includes a vehicle detection model for vehicle detection and a pedestrian detection model for pedestrian detection; Based on the model parameters of the model to be fused, each model parameter in the model to be fused is evaluated respectively to obtain a local level compatibility score of each model parameter, specifically comprising: extracting a parameter matrix of the model parameters of the model to be fused; based on an uncertainty evaluation network and the parameter matrix of the model to be fused, each model parameter is evaluated respectively to obtain a local level compatibility score of each model parameter; Constructing a histogram distribution corresponding to the model parameters of the model to be fused, and quantifying the global compatibility score of the model to be fused based on the histogram distribution; Based on the local compatibility score of each model parameter and the global compatibility score corresponding to each model parameter, score fusion is performed to obtain a dual compatibility score of each model parameter; Based on the dual compatibility scores of the model parameters, the model parameters at the parameter positions are concatenated to obtain fusion parameters corresponding to the parameter positions in the fusion model; Based on the local compatibility score of each model parameter and the global compatibility score corresponding to each model parameter, the model parameters at each parameter position are spliced ​​to obtain the fusion parameters corresponding to each parameter position in the fusion model, and then include: Acquire fine-tuning sample data of the fusion model, where the fine-tuning sample data is image data used for an autonomous driving task; Comparing the fine-tuning sample data with the initial sample data of the model to be fused to obtain a sample distribution comparison result, wherein the initial sample data is image data for pedestrian detection and image data for vehicle detection; When the sample distribution comparison result shows that there is a difference, based on the dual compatibility score of each model parameter and the sample distribution comparison result, the optimization weight of the fusion parameter belonging to the same parameter position as each model parameter is calculated; Determining the learning rate of each fusion parameter in the current iteration round based on the optimization weight of each fusion parameter and the previous learning rate of each fusion parameter in the previous iteration round; Based on the optimization weights and learning rates of the fusion parameters, the fusion parameters are updated in the current iteration round to obtain updated fusion parameters.

2. The method for fusing model parameters according to claim 1, characterized in that: The step of splicing the model parameters at the positions of the parameters based on the dual compatibility scores of the model parameters to obtain the fusion parameters of the fusion model includes: Determine the task complexity score for each network layer; When the task complexity score is greater than a preset score threshold, based on the dual compatibility scores of the model parameters under each network layer, the model parameters at each parameter position under each network layer are weightedly spliced ​​to obtain the fusion parameters.

3. The method for fusing model parameters according to claim 2, characterized in that: The step of determining the task complexity score of each network layer further includes: When the task complexity score is not greater than a preset score threshold, the dual compatibility scores corresponding to the model parameters under any network layer in the model to be fused are compared, and the model parameters corresponding to the dual compatibility scores with larger comparison results are used as the fusion parameters.

4. The method for fusing model parameters according to claim 2, characterized in that: The step of determining the task complexity score of each network layer includes: Determine the task complexity score based on at least one of the number of network layers and the type of network layers to which the parameter positions corresponding to the model parameters belong; The network layer type includes at least one of a convolutional layer, a fully connected layer, and a pooling layer.

5. The method for fusing model parameters according to claim 1, characterized in that: The learning rate of each fusion parameter in the current iteration round is determined based on the optimization weight of each fusion parameter and the previous learning rate of each fusion parameter in the previous iteration round, which is expressed as: ; in, Represents the learning rate of each fusion parameter of the current iteration round; represents the previous learning rate of each fusion parameter in the previous iteration round; Indicates the dynamic adjustment coefficient; represents the optimization weight of each fusion parameter; Indicates the number of rows of the parameter matrix corresponding to the fusion parameters; Indicates the number of columns of the parameter matrix corresponding to the fusion parameters.

6. A model parameter fusion device, characterized in that: include: An acquisition unit, which acquires model parameters of each parameter position in a model to be fused, wherein the model to be fused is constructed based on a deep neural network, and the model to be fused includes a vehicle detection model for vehicle detection and a pedestrian detection model for pedestrian detection; The local level evaluation unit evaluates each model parameter in the model to be fused based on the model parameters of the model to be fused, and obtains the local level compatibility score of each model parameter, specifically comprising: extracting the parameter matrix of the model parameters of the model to be fused; based on the uncertainty evaluation network and the parameter matrix of the model to be fused, evaluating each model parameter, and obtaining the local level compatibility score of each model parameter; A global level evaluation unit, constructing a histogram distribution corresponding to the model parameters of the model to be fused, and quantitatively obtaining a global level compatibility score of the model to be fused based on the histogram distribution; A fusion unit, which performs score fusion based on the local compatibility score of each model parameter and the global compatibility score corresponding to each model parameter to obtain a dual compatibility score of each model parameter; based on the dual compatibility score of each model parameter, splices the model parameters at each parameter position to obtain a fusion parameter corresponding to each parameter position in the fusion model; The fusion unit then includes an optimization unit, and the optimization unit is specifically used for: Acquire fine-tuning sample data of the fusion model, where the fine-tuning sample data is image data used for an autonomous driving task; Comparing the fine-tuning sample data with the initial sample data of the model to be fused to obtain a sample distribution comparison result, wherein the initial sample data is image data for pedestrian detection and image data for vehicle detection; When the sample distribution comparison result shows that there is a difference, based on the dual compatibility score of each model parameter and the sample distribution comparison result, the optimization weight of the fusion parameter belonging to the same parameter position as each model parameter is calculated; Determining the learning rate of each fusion parameter in the current iteration round based on the optimization weight of each fusion parameter and the previous learning rate of each fusion parameter in the previous iteration round; Based on the optimization weights and learning rates of the fusion parameters, the fusion parameters are updated in the current iteration round to obtain updated fusion parameters.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for fusing model parameters as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Multi-modal image fusion method based on high-order degradation model

    CN117197627A

  • Point cloud image fusion method and system for vehicle navigation

    CN117911829A