Large Model Multi-Preference Alignment Method and Device Based on Hierarchical Mixture of Experts Model

Through the large-model multi-preference alignment method based on the hierarchical hybrid expert model, the complexity and efficiency problems of the large-model alignment method in the prior art during multi-objective and user preference processing are solved, and efficient parameter reduction, calculation optimization and user preference adaptation are achieved.

CN119862423BActive Publication Date: 2025-06-17HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510340570.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-17
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

When facing multiple targets and user preferences, existing large-model alignment methods are difficult to balance conflict needs, resulting in high training costs, long model inference time and high system complexity, and the inability to effectively handle conflicts and dynamic changes between different targets.

Method used

The large model multi-preference alignment method based on the hierarchical hybrid expert model is adopted. By obtaining the pre-trained single-objective fine-tuning large language model, the target vector of the single-objective strategy is extracted for singular value decomposition, a low-rank adapter is generated as a LoRA expert, and a multi-objective LoRA expert model is constructed through the PCB-merging and Free-merging fusion method. At the same time, a linear routing layer is generated and a weight router is designed, the reward loss function is optimized to obtain multi-objective routing experts, and a hierarchical hybrid expert model is built to process user input prompt words and preference vectors.

Benefits of technology

It reduces the parameter scale and computational overhead of the model, reduces training and storage costs, improves inference efficiency, and can dynamically adapt to different user preferences to achieve the optimal Pareto frontier for multi-objective alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862423B_ABST
    Figure CN119862423B_ABST
Patent Text Reader

Abstract

The present invention provides a large model multi-preference alignment method and device based on a hierarchical mixture of experts model, which relates to the technical field of natural language processing. The method includes: obtaining a pre-trained single-objective fine-tuning model; extracting the target vectors of each single-objective strategy in the model, decomposing the target vectors by the task vector singular value decomposition method to generate low-rank adapters as the LoRA experts for each single-objective; using PCB-merging and Free-merging fusion models for processing to obtain multi-objective LoRA experts; generating a linear routing layer and constructing a reward loss function; optimizing the loss function using mirror gradient descent and smooth Chebyshev scalarization to obtain multi-objective routing experts; designing a weight router; constructing a hierarchical mixture of experts model according to the multi-objective LoRA experts, the multi-objective routing experts and the weight router; inputting the obtained user input prompt words and preference vectors into the hierarchical mixture of experts model, and outputting results that conform to the user's preferences. Using the present invention can improve the inference efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a large-model multi-preference alignment method and device based on a hierarchical hybrid expert model. Background Art

[0002] With the rapid development of artificial intelligence, generative large language models (LLMs) have made significant progress in the field of natural language processing with their powerful language understanding and generation capabilities. LLMs models are pre-trained on massive text data to learn speech patterns and rules, and can generate natural, fluent and creative texts. However, in practical applications, LLMs models often need to be adjusted according to the personalized needs of users and different usage scenarios to meet diverse task requirements. When user preferences involve multiple target dimensions, single target alignment methods often cannot balance conflicting needs, so the model is difficult to cope with diverse user needs and preferences. Therefore, in large model alignment, how to balance multiple goals and meet the preferences of different users at the same time has become a major challenge at present.

[0003] Most traditional multi-target alignment methods are based on reinforcement learning with human feedback or direct preference optimization techniques, which linearly weight different targets into a single target according to given preferences for RL optimization one by one. At present, a multi-target decoding method and a multi-target parameter fusion method during reasoning have been proposed, which are more flexible and efficient than reinforcement learning for target alignment methods. At present, a multi-target correction model or a reward model that guides decoding is used to solve the multi-target alignment problem. However, the above methods still face various shortcomings. The above methods have long reasoning time, high storage cost, and difficulty in fully handling conflicts and dynamic changes between different targets. Therefore, how to balance the trade-offs between multiple targets while maintaining efficiency is still a difficult problem that needs to be solved urgently in current technology. Summary of the invention

[0004] In order to solve the technical problems of high training cost, large model inference time and system complexity in the prior art, and the inability to capture the connection between objectives when facing multiple objective functions that may conflict with different objectives, the embodiment of the present invention provides a large model multi-preference alignment method and device based on a hierarchical hybrid expert model. The technical solution is as follows:

[0005] On the one hand, a large model multi-preference alignment method based on a hierarchical hybrid expert model is provided, the method is implemented by a large model multi-preference alignment device based on a hierarchical hybrid expert model, and the method includes:

[0006] S1. Obtain the pre-trained large language model fine-tuned for single-objective; extract the target vectors of each single-objective policy in the model, decompose the target vectors of each single-objective through the task vector singular value decomposition method, and generate low-rank adapters as the LoRA experts for each single-objective; according to the LoRA experts for each objective, use the PCB-merging and Free-merging fusion methods to construct a multi-objective LoRA expert model;

[0007] S2. Generate a linear routing layer, and construct a reward loss function corresponding to the linear routing layer; use the mirror gradient descent method and the smoothed Chebyshev scalarization method to optimize the loss function to obtain multi-objective routing experts;

[0008] S3. Design a weight router; construct a hierarchical mixture-of-experts model based on the multi-objective LoRA expert model, the multi-objective routing experts, and the weight router;

[0009] S4. Obtain the prompt words and preference vectors input by the user, input the prompt words and preference vectors input by the user into the mixture-of-experts model for alignment processing, and output to meet the user's preferences.

[0010] Optionally, for extracting the target vectors of each single-objective policy in the model in S1, the target vectors are sparsely processed through the task vector singular value decomposition method, and low-rank adapters are generated as the LoRA experts for each single-objective; according to the LoRA experts for each objective, use the PCB-merging fusion model and the Free-merging fusion model for processing to obtain a multi-objective LoRA expert model, including:

[0011] S11. Define the parameters of the pre-trained large language model fine-tuned for single-objective; according to the parameters of the pre-trained large language model fine-tuned for single-objective, extract the target vectors of each single-objective policy in the model;

[0012] S12. Perform sparse processing using the task vector singular value decomposition method, obtain low-rank adapters by decomposing the target vectors of each single-objective; use the low-rank adapters as the LoRA experts for each single-objective;

[0013] S13. Set the preference vectors, and use the PCB-merging fusion model and the Free-merging fusion model for processing according to the LoRA experts for each single-objective to obtain a multi-objective LoRA expert model.

[0014] Optionally, the reward loss function in S2 is represented by the following formula (1):

[0015] (1)

[0016] Among them, represents the linear weighted combination of the reward signals of N targets under the preference ; represents the hidden state; represents the parameters of the pre-trained model; represents the routing expert; represents the hierarchical mixture-of-experts model; represents all LoRA experts.

[0017] Optionally, the optimization process of the loss function by the S2 using the mirror gradient descent method and the smoothed Chebyshev scalarization method is represented by the following formula (2):

[0018] (2)

[0019] Among them, represents the reference point of each target, that is, the expected performance level; represents the preference weight of the i-th target; represents the trainable parameters of the model.

[0020] Optionally, the weight router is used to map the continuous preference vector of the user to the nearest expert and assign voting weights to each selected expert; among them, the nearest experts include: LoRA experts and routing experts;

[0021] Among them, the calculation process of the weight router is represented by the following formula (3):

[0022] (3)

[0023] Among them, represents the preference vector given by the user; represents the numbers of the N experts closest to the preference vector given by the user.

[0024] Optionally, the S4 inputs the prompt word and preference vector input by the user into the mixture-of-experts model for alignment processing and outputs in line with the user's preference, including:

[0025] S41. Input the prompt word and preference vector input by the user into the hierarchical mixture-of-experts model. The weight router selects multiple routing experts for activation according to the preference vector and assigns corresponding weights to each routing expert;

[0026] S42. Input the corresponding weights assigned to each routing expert into the multi-target routing expert. According to the received hidden layer state input, perform dynamic voting on each single-target LoRA expert, and perform weighted summation on the voting values of all routing experts according to the corresponding weights assigned to each routing expert to obtain the final voting summary result;

[0027] S43. Input the final vote summary result into the multi-objective LoRA expert model. Through the final vote summary result, dynamically activate the LoRA expert parameters, and combine them with the parameters of the pre-trained large language model fine-tuned for a single objective to generate results that conform to the user's preferences.

[0028] Optionally, the process of the output in S4 conforming to the user's preferences is represented by the following formula (4):

[0029] (4)

[0030] Where: represents the final output when given the input x under the preference ; represents the original output of the pre-trained parameters; represents the voting weight given to each LoRA expert; represents the B matrix of each LoRA expert; represents the A matrix of each LoRA expert; represents the input hidden state.

[0031] On the other hand, a large model multi-preference alignment device based on a hierarchical mixture of experts model is provided. This device is applied to the large model multi-preference alignment method based on the hierarchical mixture of experts model. The device includes:

[0032] A first construction unit for obtaining a pre-trained large language model fine-tuned for a single objective; extracting the target vectors of each single objective strategy in the model, sparsifying the target vectors through the task vector singular value decomposition method, and generating low-rank adapters as LoRA experts for each single objective; processing according to the LoRA experts for each objective using the PCB-merging fusion model and the Free-merging fusion model to obtain a multi-objective LoRA expert model;

[0033] An acquisition unit for generating a linear routing layer and constructing a reward loss function corresponding to the linear routing layer; optimizing the loss function using the mirror gradient descent method and the smoothed Chebyshev scalarization method to obtain multi-objective routing experts;

[0034] A second construction unit for designing a weight router; constructing a hierarchical mixture of experts model based on the multi-objective LoRA expert model, the multi-objective routing experts, and the weight router;

[0035] An output unit for obtaining the prompt words and preference vectors input by the user, inputting the prompt words and preference vectors input by the user into the mixture of experts model for alignment processing, and outputting results that conform to the user's preferences.

[0036] Optionally, the first construction unit is used for:

[0037] Define the parameters of the pre-trained large language model after single-object fine-tuning; according to the parameters of the pre-trained large language model after single-object fine-tuning, extract the target vectors of each single-object policy in the model;

[0038] Perform sparse processing using the task vector singular value decomposition method, and obtain low-rank adapters by decomposing the target vectors of each single object, and use the low-rank adapters as the LoRA experts for each single object;

[0039] Set the preference vector, and process it using the PCB-merging fusion model and the Free-merging fusion model according to the LoRA experts of each single object to obtain a multi-object LoRA expert model.

[0040] Optionally, the reward loss function is represented by the following formula (1):

[0041] (1)

[0042] Wherein, represents the linear weighted combination of the reward signals of N targets under the preference ; represents the hidden state; represents the parameters of the pre-trained model; represents the routing expert; represents the hierarchical hybrid expert model; represents all LoRA experts.

[0043] Optionally, the process of optimizing the loss function using the mirror gradient descent method and the smoothed Chebyshev scalarization method is represented by the following formula (2):

[0044] (2)

[0045] Wherein, represents the reference point of each target, that is, the expected performance level; represents the preference weight of the i-th target; represents the trainable parameters of the model.

[0046] Optionally, the weight router is used to map the continuous preference vector of the user to the nearest expert and assign voting weights to each selected expert; wherein, the nearest experts include: LoRA experts and routing experts;

[0047] Wherein, the calculation process of the weight router is represented by the following formula (3):

[0048] (3)

[0049] Among them, represents the preference vector given by the user; represents the numbers of the N experts closest to the preference vector given by the user.

[0050] Optionally, the output unit is used for:

[0051] Input the prompt words and preference vector input by the user into the hierarchical mixture-of-experts model. The weight router selects multiple routing experts for activation according to the preference vector and assigns corresponding weights to each routing expert;

[0052] Input the corresponding weights assigned to each routing expert into the multi-objective routing expert. According to the received hidden layer state input, perform dynamic voting on each single-object LoRA expert, and perform weighted summation on the voting values of all routing experts according to the corresponding weights assigned to each routing expert to obtain the final voting summary result;

[0053] Input the final voting summary result into the multi-objective LoRA expert model. Through the final voting summary result, dynamically activate the LoRA expert parameters, and combine the parameters of the pre-trained single-object fine-tuned large language model to generate preferences that meet the user's requirements.

[0054] Optionally, the process of outputting preferences that meet the user's requirements is represented by the following formula (4):

[0055] (4)

[0056] Among them, represents the final output when the input x is given under the preference ; represents the original output of the pre-trained parameters; represents the voting weight given to each LoRA expert; represents the B matrix of each LoRA expert; represents the A matrix of each LoRA expert; represents the input hidden state.

[0057] On the other hand, a large model multi-preference alignment device based on a hierarchical mixture-of-experts model is provided. The large model multi-preference alignment device based on a hierarchical mixture-of-experts model includes: a processor; a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned large model multi-preference alignment method based on a hierarchical mixture-of-experts model is implemented.

[0058] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned large model multi-preference alignment methods based on the hierarchical mixture of experts model.

[0059] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0060] In the embodiments of the present invention, first, a pre-trained large language model after single-target fine-tuning is obtained; the target vectors of each single-target policy in the model are extracted, and the target vector of each single-target is decomposed by the task vector singular value decomposition method to generate a low-rank adapter as the LoRA expert of each single-target; according to the LoRA experts of each target, PCB-merging and Free-merging fusion methods are used to construct a multi-target LoRA expert model; a linear routing layer is generated, and a reward loss function corresponding to the linear routing layer is constructed; the mirror gradient descent method and the smoothed Chebyshev scalarization method are used to optimize the loss function to obtain a multi-target routing expert; a weight router is designed; according to the multi-target LoRA expert model, the multi-target routing expert and the weight router, a hierarchical mixture of experts model is constructed; the prompt words and preference vectors input by the user are obtained, and the prompt words and preference vectors input by the user are input into the hierarchical mixture of experts model for alignment processing, and the output conforms to the user's preferences.

[0061] In the embodiments of the present invention, first, by generating low-rank adapters and constructing multi-target LoRA experts, and by designing routing experts, the parameter scale of the model is reduced, and the storage and calculation costs are reduced; secondly, in the embodiments of the present invention, it remains frozen during inference, and only a small amount of training is required for the routing experts, reducing the training cost; thirdly, in the embodiments of the present invention, only a small number of expert models are activated during inference, reducing the calculation overhead and improving the inference efficiency; finally, through the hierarchical expert model design, it can dynamically adapt to different user preferences and achieve the optimal Pareto frontier of multi-target alignment. Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0063] Figure 1 It is a flowchart of a large model multi-preference alignment method based on a hierarchical mixture of experts model provided by the embodiments of the present invention;

[0064] Figure 2It is a structural schematic diagram of a hierarchical hybrid expert framework provided by an embodiment of the present invention;

[0065] Figure 3 It is an overall flow chart of a hierarchical hybrid expert model provided by an embodiment of the present invention;

[0066] Figure 4 It is a block diagram of a large model multi-preference alignment device based on a hierarchical hybrid expert model provided by an embodiment of the present invention;

[0067] Figure 5 It is a structural schematic diagram of a large model multi-preference alignment device based on a hierarchical hybrid expert model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0068] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0069] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0070] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.

[0071] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0072] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0073] The embodiment of the present invention provides a large model multi-preference alignment method based on a hierarchical hybrid expert model, which can be implemented by a large model multi-preference alignment device based on a hierarchical hybrid expert model, and the large model multi-preference alignment device based on a hierarchical hybrid expert model can be a terminal or a server. Figure 1 The flowchart of the large model multi-preference alignment method based on the hierarchical hybrid expert model is shown. The processing flow of the method may include the following steps:

[0074] S1. Obtain a pre-trained large language model fine-tuned for single-objectives; extract the target vectors of each single-objective strategy in the model, perform sparse processing on the target vectors by the task vector singular value decomposition method, and generate low-rank adapters as the LoRA experts for each single-objective; according to the LoRA experts for each objective, process them using the PCB-merging fusion model and the Free-merging fusion model to obtain a multi-objective LoRA expert model.

[0075] Among them, the multi-objective LoRA expert model is the first-level basic component of the hierarchical mixture-of-experts model, and each objective LoRA expert is responsible for handling the alignment problem of a single preference.

[0076] Optionally, the specific implementation process of S1 includes S11 - S13:

[0077] S11. Define the parameters of the pre-trained large language model fine-tuned for single-objectives; according to the parameters of the pre-trained large language model fine-tuned for single-objectives, extract the target vectors of each single-objective strategy in the model;

[0078] Among them, each single-objective strategy is obtained by fine-tuning on the target corresponding to each strategy.

[0079] In a feasible implementation, each single-objective fine-tuning model is initialized with a same set of pre-trained parameters and fine-tuned on their respective corresponding objectives to obtain the parameters of the pre-trained large language model fine-tuned for single-objectives.

[0080] In a feasible implementation, the process of extracting the target vectors of each single-objective strategy is represented by the following formula (1):

[0081] (1)

[0082] Among them, represents the target vector of each single-objective strategy; represents the fine-tuned parameters; represents the parameters of the pre-trained model.

[0083] Among them, by extracting the target vectors, the ability of the model on a single objective can be captured.

[0084] S12. Perform sparse processing using the task vector singular value decomposition method. By decomposing the target vectors of each single-objective, obtain low-rank adapters; take the low-rank adapters as the LoRA experts for each single-objective;

[0085] Among them, each LoRA expert specializes in handling a specific objective. To enhance the preference coverage of LoRA experts, this application samples preferences from within an N-dimensional simplex to construct LoRA experts dedicated to corresponding preferences.

[0086] S13. Set the preference vector, and process it using the PCB-merging fusion model and the Free-merging fusion model according to each single-objective LoRA expert to obtain a multi-objective LoRA expert model.

[0087] Among them, the PCB-merging fusion model uses internal balance to measure the importance of parameters in a single task, and uses external balance to evaluate the parameter similarity between different tasks, and selects significant parameters that are beneficial to all tasks.

[0088] Among them, the Free-merging fusion model effectively identifies and filters specific parameters that are harmful to multi-tasks using the frequency domain information of parameters, so as to minimize the impact of task conflicts on backbone parameters at the lowest cost.

[0089] Among them, using the PCB-merging and Free-merging fusion methods can identify and delete redundant parameters within each task vector, and at the same time reduce the parameter competition between multiple task vectors; therefore, a set of model parameters can achieve performance improvement on multiple tasks without sacrificing performance on specific tasks; it can obtain performance comparable to multi-task training without introducing additional training.

[0090] In a feasible implementation, obtain the target vectors of all objectives and a specific preference vector, and process them using the PCB-merging fusion model and the Free-merging fusion model to obtain a multi-objective LoRA expert for a specific preference.

[0091] In a feasible implementation, this application extracts N single-objective LoRA experts and constructs L multi-objective LoRA experts; by linearly combining the outputs of multiple LoRA experts and attaching the original output of the pre-trained parameters as the final output of each module with a mixture of experts in the hierarchical mixture of experts model, an output that conforms to the user's preference is generated.

[0092] S2. Generate a linear routing layer, construct a reward loss function corresponding to the linear routing layer; use the mirror gradient descent method and the smoothed Chebyshev scalarization method to optimize the loss function to obtain a multi-objective routing expert.

[0093] Among them, the multi-objective routing expert is the second-level intermediate layer component of the hierarchical hybrid expert model, which is used to dynamically select and combine among multiple LoRA experts according to the user's preference vector to generate an output that meets the user's preferences. This application uses a router focused on specific preferences as an expert, which can achieve a balance between the number of parameters and performance.

[0094] Among them, each routing expert is represented by a linear layer, with the input being the hidden state and the output being the voting weights for all LoRA experts, which is represented by the following formula (2):

[0095] (2)

[0096] Among them, represents the first parameter of the routing expert; represents the second parameter of the routing expert; represents the voting weights for all LoRA experts; represents the hidden layer state.

[0097] Among them, the routing expert is located in each Transformer trainable module, which can make fine-grained voting selections at the Transformer module level and the Transformer hierarchy level, and the routing expert can fully analyze the input hidden state and dynamically select and combine LoRA experts according to different inputs. Among them, each routing expert is only responsible for a specific preference, and its optimization goal is to minimize the reward loss corresponding to its preference, so it is called an expert.

[0098] Optionally, the reward loss function of S2 is represented by the following formula (3):

[0099] (3)

[0100] Among them, represents the linear weighted combination of the reward signals of N targets under the preference ; represents the hidden state; represents the parameters of the pre-trained model; represents the routing expert; represents the hierarchical hybrid expert model; represents all LoRA experts.

[0101] Among them, in order to capture the non-convex region of the Pareto front and eliminate the inherent disadvantages of linear scalarization, this application uses the Chebyshev scalarization method to calibrate the current strategy through a reference point and further optimize it.

[0102] Optionally, the optimization process of S2 using the mirror gradient descent method and the smoothed Chebyshev scalarization method for the loss function is represented by the following formula (4):

[0103] (4)

[0104] where, represents the reference point of each objective, i.e., the desired performance level; represents the preference weight of the i-th objective; represents the trainable parameters of the model.

[0105] S3. Design a weight router; construct a hierarchical hybrid expert model based on the multi-objective LoRA expert model, the multi-objective routing expert, and the weight router.

[0106] Among them, the weight router is the third-level intermediate component of the hierarchical hybrid expert model. To ensure the computational efficiency during inference, for N-object alignment, this application only selects and activates N nearest experts, including LoRA experts and routing experts.

[0107] Among them, the weight router is used to decompose the multi-objective Pareto optimization problem into multiple multi-objective sub-problems and map the continuous preference vector of the user to a discrete single preference expert.

[0108] Optionally, the weight router is used to map the continuous preference vector of the user to the nearest expert and assign voting weights to each selected expert; among them, the nearest experts include: LoRA experts and routing experts;

[0109] Among them, the calculation process of the weight router is represented by the following formula (5):

[0110] (5)

[0111] where, represents the preference vector given by the user; represents the numbers of the N nearest experts to the preference vector given by the user.

[0112] Among them, to achieve the assignment of voting weights to each selected expert, this application uses the representative vector of the LoRA expert to divide the N-dimensional simplex into several regions, and uses the representative vector of the routing expert to further divide the regions into smaller regions. Since the preference vector given by the user is regarded as a vector sampled from the N-dimensional simplex with a sum of 1, each given preference will fall into one of the fine-grained N-dimensional simplices divided by the experts, which can be regarded as a more fine-grained N-object alignment. Among them, the N sub-objectives are supported by N experts as basis vectors.

[0113] S4. Obtain the prompt words and preference vectors input by the user, input the prompt words and preference vectors input by the user into the mixture-of-experts model for alignment processing, and output to match the user's preferences.

[0114] Among them, Figure 2 is a schematic structural diagram of a hierarchical mixture-of-experts framework provided by an embodiment of the present invention; in a feasible implementation manner, the hierarchical mixture-of-experts framework includes: zero-level pre-trained original parameters, first-level LoRA experts, second-level routing experts, and third-level weight routers. Obtain the preference weights and hidden layer state inputs given by the user; input the preference weights given by the user into the third-level weight router, activate the nearest N experts according to the preference weights, where the N activated experts form an N-dimensional simplex, determine the votes of each expert, obtain the voting scores for all experts, and input the voting scores for all experts into the second-level routing expert; input the hidden layer input state into the second-level routing expert, dynamically select and combine the voting weights of each LoRA expert according to the input hidden layer state granularity, and obtain the voting scores for the LoRA experts; among them, each routing expert specializes in a specific preference; input the voting scores of the LoRA experts and the hidden layer input state into the first-level LoRA experts, and generate to match the specific user preferences according to the input hidden layer state and the original parameters of the pre-trained model.

[0115] Optionally, the specific implementation process of S4 includes S41 - S43:

[0116] S41. Input the prompt words and preference vectors input by the user into the hierarchical mixture-of-experts model. The weight router selects multiple routing experts for activation according to the preference vectors and assigns corresponding weights to each routing expert.

[0117] S42. Input the corresponding weights assigned to each routing expert into the multi-objective routing expert. According to the received hidden layer state input, perform dynamic voting on each single-objective LoRA expert, and perform weighted summation on the voting values of all routing experts according to the corresponding weights assigned to each routing expert to obtain the final voting summary result.

[0118] S43. Input the final voting summary result into the multi-objective LoRA expert model. Through the final voting summary result, dynamically activate the LoRA expert parameters, and combine the parameters of the pre-trained single-objective fine-tuned large language model to generate to match the user's preferences.

[0119] Optionally, the process of the output of S4 matching the user's preferences is represented by the following formula (6):

[0120] (6)

[0121] Among them, Denotes the final output given the input x under the preference ; Denotes the original output of the pre-trained parameters; Denotes the voting weights given to each LoRA expert; Denotes the B matrix of each LoRA expert; Denotes the A matrix of each LoRA expert; Denotes the hidden state of the input.

[0122] Among them, Figure 3 is an overall flowchart of a hierarchical mixture-of-experts model provided by an embodiment of the present invention; in a feasible implementation manner, taking three objectives as an example, N trained single-object models are obtained, and sparse processing is performed by decomposing using the task vector singular value decomposition method to obtain N single-object LoRA experts; the PCB-merging fusion method and the Free-merging fusion method are used for processing to obtain L multi-object LoRA experts; multi-object reinforcement learning is performed based on the L multi-object LoRA experts to obtain R multi-object routing experts; a hierarchical mixture-of-experts model is constructed according to the N single-object LoRA experts, the L multi-object LoRA experts, and the R multi-object routing experts; the prompt words input by the user and the preference weights given by the user are obtained; the prompt words input by the user and the preference weights given by the user are input into the hierarchical mixture-of-experts model for processing to obtain an aligned output, that is, in line with the user's preferences.

[0123] Among them, to verify the effectiveness of the present application, the present application has conducted experiments on multiple benchmark datasets, covering 14 different objectives and 200 types of user preferences. The hierarchical mixture-of-experts framework proposed in the present application is superior to 15 existing baseline methods in the multi-object alignment task. Specifically, the hierarchical mixture-of-experts reaches the optimal Pareto front in multiple tasks. Especially in the two-object, three-object, and multi-object alignment tasks, the hierarchical mixture-of-experts model performs outstandingly. For example, in the Helpful Assistant task, the Pareto front of the hierarchical mixture-of-experts model is superior to the existing baseline methods, and the best performance is achieved under multiple weight settings. In addition, the hierarchical mixture-of-experts model also performs well in the five-object alignment task, and the average score is significantly higher than other methods. On the HelpSteer dataset, the average score of the hierarchical mixture-of-experts model reaches 61.74, higher than the MOD and RiC baseline methods.

[0124] Among them, the hierarchical mixture-of-experts model proposed in this application only needs to store 7.64B of parameters, reducing the storage and computing costs; the training parameters of the hierarchical mixture-of-experts model proposed in this application are only 8M, far lower than other methods, such as 0.64B of RiC and 0.16B of PAD, reducing the training costs. During the inference stage of this application, the hierarchical mixture-of-experts framework only activates a small number of expert models, reducing the computational overhead. The inference cost of the hierarchical mixture-of-experts model proposed in this application is only 1.23 times that of normal inference, far lower than other methods, such as 3.10 of MOD and 2.98 of PAD, improving the inference efficiency.

[0125] In the embodiment of the present invention, first, a large language model fine-tuned for a single target after pre-training is obtained; the target vector of each single target policy in the model is extracted, and the target vector of each single target is decomposed by the task vector singular value decomposition method to generate a low-rank adapter as the LoRA expert for each single target; according to the LoRA experts of each target, PCB-merging and Free-merging fusion methods are used to construct a multi-target LoRA expert model; a linear routing layer is generated, and a reward loss function corresponding to the linear routing layer is constructed; the mirror gradient descent method and the smoothed Chebyshev scalarization method are used to optimize the loss function to obtain a multi-target routing expert; a weight router is designed; according to the multi-target LoRA expert model, the multi-target routing expert, and the weight router, a hierarchical mixture-of-experts model is constructed; the prompt word and preference vector input by the user are obtained, and the prompt word and preference vector input by the user are input into the hierarchical mixture-of-experts model for alignment processing, and the output conforms to the user's preference.

[0126] By decomposing the multi-objective alignment problem into a series of single-preference sub-problems and employing specialized experts to handle each sub-problem, this application can effectively cover the entire Pareto front while maintaining a low training overhead. This application can not only balance the trade-offs between different objectives but also achieve superior Pareto-optimal results on multiple benchmarks. This application uses lightweight and parameter-efficient LoRA adapters as specific-object experts, reducing the number of trainable parameters required and the storage cost. In addition, the routing expert mechanism further optimizes the model selection process, improves the inference efficiency, and achieves a cost-effective balance among the number of parameters, inference cost, and performance. This application uses model sparsification and model fusion to extract LoRA experts without any training. At the same time, the hierarchical mixture-of-experts model proposed in this application can perform dynamic balancing at multiple granularity levels, including the model, target vector, Transformer layer, and Transformer module. This application can flexibly adapt to diverse user preferences and provide more personalized output results. The three basic components in the hierarchical mixture-of-experts model of this application only focus on specific objectives and preferences, enabling the model to dynamically switch between different preferences. This application can insert new experts at any time to adapt to new objectives, thus better meeting the growing diverse needs of users.

[0127] In the embodiments of the present invention, first, by generating low-rank adapters to construct multi-object LoRA experts and designing routing experts, the parameter scale of the model is reduced, and the storage and calculation costs are lowered. Second, during inference, the embodiments of the present invention remain frozen and only require a small amount of training for the routing experts, reducing the training cost. Third, during inference, only a small number of expert models are activated, reducing the computational overhead and improving the inference efficiency. Finally, through the hierarchical expert model design, it can dynamically adapt to different user preferences and achieve the optimal Pareto front for multi-objective alignment.

[0128] Figure 4 is a block diagram of a large model multi-preference alignment device based on a hierarchical mixture-of-experts model shown according to an exemplary embodiment. This device is used for the large model multi-preference alignment method based on the hierarchical mixture-of-experts model. Referring to Figure 4 , this device includes a first construction unit 410, an acquisition unit 420, a second construction unit 430, and an output unit 440. Among them:

[0129] The first construction unit 410 is used to obtain a pre-trained large language model after single-object fine-tuning; extract the target vectors of each single-object policy in the model, perform sparse processing on the target vectors by the task vector singular value decomposition method, and generate low-rank adapters as the LoRA experts for each single object; according to the LoRA experts for each object, use the PCB-merging fusion model and the Free-merging fusion model for processing to obtain a multi-object LoRA expert model;

[0130] The acquisition unit 420 is used to generate a linear routing layer, construct a reward loss function corresponding to the linear routing layer; optimize the loss function by using the mirror gradient descent method and the smooth Chebyshev scalarization method to obtain multi-object routing experts;

[0131] The second construction unit 430 is used to design a weight router; construct a hierarchical mixture-of-experts model according to the multi-object LoRA expert model, the multi-object routing experts, and the weight router;

[0132] The output unit 440 is used to obtain the prompt words and preference vectors input by the user, input the prompt words and preference vectors input by the user into the mixture-of-experts model for alignment processing, and output in line with the user's preferences.

[0133] Optionally, the first construction unit 410 is used to:

[0134] Define the parameters of the pre-trained large language model after single-object fine-tuning; extract the target vectors of each single-object policy in the model according to the parameters of the pre-trained large language model after single-object fine-tuning;

[0135] Perform sparse processing by using the task vector singular value decomposition method, obtain low-rank adapters by decomposing the target vectors of each single object; use the low-rank adapters as the LoRA experts for each single object;

[0136] Set the preference vectors, and use the PCB-merging fusion model and the Free-merging fusion model for processing according to the LoRA experts for each single object to obtain a multi-object LoRA expert model.

[0137] Optionally, the reward loss function is represented by the following formula (1):

[0138] (1)

[0139] Where, represents the linear weighted combination of the reward signals of N targets under the preference ; represents the hidden state; represents the parameters of the pre-trained model; Denote the routing expert; Denote the hierarchical mixture-of-experts model; Denote all LoRA experts.

[0140] Optionally, the optimization process of the loss function using the mirror gradient descent method and the smoothed Chebyshev scalarization method is represented by the following formula (2):

[0141] (2)

[0142] where Denote the reference point of each objective, i.e., the desired performance level; Denote the preference weight of the i-th objective; Denote the trainable parameters of the model.

[0143] Optionally, the weight router is used to map the user's continuous preference vector to the nearest expert and assign voting weights to each selected expert; where the nearest experts include: LoRA experts and routing experts;

[0144] where the calculation process of the weight router is represented by the following formula (3):

[0145] (3)

[0146] where Denote the preference vector given by the user; Denote the numbers of the N nearest experts to the preference vector given by the user.

[0147] Optionally, the output unit 440 is used to:

[0148] Input the prompt word and preference vector input by the user into the hierarchical mixture-of-experts model. The weight router selects multiple routing experts for activation according to the preference vector and assigns corresponding weights to each routing expert;

[0149] Input the corresponding weights assigned to each routing expert into the multi-objective routing expert. According to the received hidden layer state input, dynamically vote on each single-objective LoRA expert, and according to the corresponding weights assigned to each routing expert, weight-sum the voting values of all routing experts to obtain the final voting summary result;

[0150] Input the final voting summary result into the multi-objective LoRA expert model. Through the final voting summary result, dynamically activate the LoRA expert parameters and combine the parameters of the pre-trained single-objective fine-tuned large language model to generate the preferences that meet the user.

[0151] Optionally, the process of making the output conform to the user's preferences is represented by the following formula (4):

[0152] (4)

[0153] where represents the final output given the input x under the preference ; represents the original output of the pre-trained parameters; represents the voting weight given to each LoRA expert; represents the B matrix of each LoRA expert; represents the A matrix of each LoRA expert; represents the hidden state of the input.

[0154] In the embodiment of the present invention, first, a large language model after single-target fine-tuning of pre-training is obtained; the target vectors of each single-target strategy in the model are extracted, and the target vectors of each single-target are decomposed by the task vector singular value decomposition method to generate low-rank adapters as LoRA experts for each single-target; according to the LoRA experts of each target, PCB-merging and Free-merging fusion methods are adopted to construct a multi-target LoRA expert model; a linear routing layer is generated, and a reward loss function corresponding to the linear routing layer is constructed; the mirror gradient descent method and the smoothed Chebyshev scalarization method are used to optimize the loss function to obtain multi-target routing experts; a weight router is designed; according to the multi-target LoRA expert model, the multi-target routing experts and the weight router, a hierarchical hybrid expert model is constructed; the prompt words and preference vectors input by the user are input into the hierarchical hybrid expert model for alignment processing, and the output conforms to the user's preferences.

[0155] In the embodiment of the present invention, first, by generating low-rank adapters and constructing multi-target LoRA experts, and by designing routing experts, the parameter scale of the model is reduced, and the storage and calculation costs are reduced; second, in the embodiment of the present invention, it remains frozen during inference, and only a small amount of training is required for the routing experts, reducing the training cost; third, in the embodiment of the present invention, only a small number of expert models are activated during inference, reducing the calculation overhead and improving the inference efficiency; finally, through the hierarchical expert model design, it can dynamically adapt to different user preferences and achieve the optimal Pareto frontier of multi-target alignment.

[0156] Figure 5 is a schematic structural diagram of a large model multi-preference alignment device based on a hierarchical hybrid expert model provided by an embodiment of the present invention. As Figure 5 shown, the large model multi-preference alignment device based on the hierarchical hybrid expert model may include the above Figure 4The large model multi-preference alignment device based on the hierarchical mixture of experts model shown. Optionally, the large model multi-preference alignment device 510 based on the hierarchical mixture of experts model may include a first processor 2001.

[0157] Optionally, the large model multi-preference alignment device 510 based on the hierarchical mixture of experts model may also include a memory 2002 and a transceiver 2003.

[0158] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, for example, through a communication bus.

[0159] Next, in combination with Figure 5 Specific introductions will be made to the various components of the large model multi-preference alignment device 510 based on the hierarchical mixture of experts model:

[0160] Among them, the first processor 2001 is the control center of the large model multi-preference alignment device 510 based on the hierarchical mixture of experts model, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, for example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0161] Optionally, the first processor 2001 can execute various functions of the large model multi-preference alignment device 510 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0162] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 shown in

[0163] In a specific implementation, as an embodiment, the large model multi-preference alignment device 510 based on the hierarchical mixture of experts model may also include multiple processors, such as Figure 5The first processor 2001 and the second processor 2004 shown in []. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0164] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.

[0165] Optionally, the memory 2002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through an interface circuit of the large model multi-preference alignment device 510 based on the hierarchical hybrid expert model ( Figure 5 not shown in []). The embodiments of the present invention do not make specific limitations on this.

[0166] The transceiver 2003 is used to communicate with a network device or with a terminal device.

[0167] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 5 not shown separately in []). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0168] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through an interface circuit of the large model multi-preference alignment device 510 based on the hierarchical hybrid expert model ( Figure 5 not shown in []). The embodiments of the present invention do not make specific limitations on this.

[0169] It should be noted that Figure 5 the structure of the large model multi-preference alignment device 510 based on the hierarchical mixture of experts model shown in does not limit the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0170] In addition, the technical effects of the large model multi-preference alignment device 510 based on the hierarchical mixture of experts model can refer to the technical effects of the large model multi-preference alignment method based on the hierarchical mixture of experts model described in the above method embodiments, which will not be elaborated here.

[0171] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0172] It should also be understood that the memory in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0173] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0174] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.

[0175] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0176] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0177] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0178] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0179] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0180] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0181] In addition, the functional units in various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0182] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0183] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A large model multi-preference alignment method based on a hierarchical hybrid expert model, characterized in that: The method comprises: S1. Obtain a pre-trained single-target fine-tuned large language model; extract the target vector of each single-target strategy in the model, perform sparse processing on the target vector by task vector singular value decomposition method, and generate a low-rank adapter as the LoRA expert of each single target; according to the LoRA expert of each target, use PCB-merging fusion model and Free-merging fusion model for processing to obtain a multi-target LoRA expert model; Among them, the S1 extracts the target vector of each single-target strategy in the model, performs sparse processing on the target vector by the task vector singular value decomposition method, and generates a low-rank adapter as the LoRA expert of each single target; according to the LoRA expert of each target, the PCB-merging fusion model and the Free-merging fusion model are used for processing to obtain a multi-target LoRA expert model, including: S11, defining parameters of the pre-trained single-objective fine-tuned large language model; according to the parameters of the pre-trained single-objective fine-tuned large language model, extracting the target vector of each single-objective strategy in the model; S12. Use the task vector singular value decomposition method to perform sparse processing, and obtain a low-rank adapter by decomposing the target vector of each single target; use the low-rank adapter as the LoRA expert of each single target; S13, setting the preference vector, using the PCB-merging fusion model and the Free-merging fusion model to process each single-target LoRA expert, and obtaining a multi-target LoRA expert model; S2. Generate a linear routing layer and construct a reward loss function corresponding to the linear routing layer; optimize the loss function using a mirror gradient descent method and a smoothed Chebyshev scalarization method to obtain a multi-objective routing expert; S3. Design weighted routers; construct a hierarchical hybrid expert model based on the multi-objective LoRA expert model, multi-objective routing experts and weighted routers; S4, obtaining the prompt word and preference vector input by the user; inputting the prompt word and preference vector input by the user into the hybrid expert model for alignment processing, and outputting the information in accordance with the user's preference.

2. The large model multi-preference alignment method based on hierarchical hybrid expert model according to claim 1 is characterized in that: The reward loss function of S2 is expressed by the following formula (1): (1) in, Reward signals representing N targets are preferred A linear weighted combination of Indicates hidden state; Represents the parameters of the pre-trained model; represents routing expert; Represents a hierarchical mixture of experts model; Represents all LoRA experts.

3. The large model multi-preference alignment method based on hierarchical hybrid expert model according to claim 1 is characterized in that: The process of optimizing the loss function using the mirror gradient descent method and the smoothed Chebyshev scalar quantization method in S2 is expressed by the following formula (2): (2) in, Indicate the reference point for each objective, i.e. the desired level of performance; represents the preference weight of the i-th target; Represents the trainable parameters of the model.

4. The large model multi-preference alignment method based on hierarchical hybrid expert model according to claim 1 is characterized in that: The weight router is used to map the user's continuous preference vector to the nearest expert and assign a voting weight to each selected expert; wherein the nearest experts include: LoRA experts and routing experts; The calculation process of the weighted router is expressed by the following formula (3): (3) in, Represents the preference vector given by the user; Indicates the numbers of the N experts closest to the preference vector given by the user.

5. The large model multi-preference alignment method based on hierarchical hybrid expert model according to claim 1 is characterized in that: The step S4 inputs the prompt word and the preference vector input by the user into the hybrid expert model for alignment processing, and outputs an alignment that meets the user's preference, including: S41, inputting the prompt word and preference vector input by the user into the hierarchical hybrid expert model, and the weight router selects multiple routing experts to activate according to the preference vector, and assigns a corresponding weight to each routing expert; S42, assigning a corresponding weight to each routing expert and inputting it into the multi-target routing expert, dynamically voting for each single-target LoRA expert according to the received hidden layer state input, assigning a corresponding weight to each routing expert, and weighted summing the voting values ​​of all routing experts to obtain the final voting summary result; S43. Input the final voting summary result into the multi-objective LoRA expert model. Through the final voting summary result, dynamically activate the LoRA expert parameters, and combine the parameters of the pre-trained single-objective fine-tuned large language model to generate a preference that meets the user.

6. The large model multi-preference alignment method based on hierarchical hybrid expert model according to claim 1 is characterized in that: The process of the output of S4 conforming to the user's preference is expressed by the following formula (4): (4) in, Indicate in preference The final output for a given input x; represents the raw output of the pre-trained parameters; Indicates the voting weight given to each LoRA expert; The B matrix represents each LoRA expert; A matrix representing each LoRA expert; represents the hidden state of the input; N represents the number of activated LoRA experts.

7. A large model multi-preference alignment device based on a hierarchical hybrid expert model, the large model multi-preference alignment device based on a hierarchical hybrid expert model is used to implement the large model multi-preference alignment method based on a hierarchical hybrid expert model as claimed in any one of claims 1 to 6, characterized in that: The device comprises: The first construction unit is used to obtain a pre-trained single-target fine-tuned large language model; extract the target vector of each single-target strategy in the model, perform sparse processing on the target vector by the task vector singular value decomposition method, and generate a low-rank adapter as the LoRA expert of each single target; according to the LoRA expert of each target, the PCB-merging fusion model and the Free-merging fusion model are used for processing to obtain a multi-target LoRA expert model; An acquisition unit is used to generate a linear routing layer and construct a reward loss function corresponding to the linear routing layer; the loss function is optimized by using a mirror gradient descent method and a smoothed Chebyshev scalarization method to obtain a multi-objective routing expert; The second construction unit is used to design a weighted router; a hierarchical hybrid expert model is constructed based on the multi-objective LoRA expert model, the multi-objective routing experts and the weighted router; The output unit is used to obtain the prompt word and preference vector input by the user, input the prompt word and preference vector input by the user into the hybrid expert model for alignment processing, and output in accordance with the user's preference.

8. A large model multi-preference alignment device based on a hierarchical hybrid expert model, characterized in that: The large model multi-preference alignment device based on the hierarchical hybrid expert model includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-task processing method and system for large language model fused with LoRA

    CN119416143A