Target large language model determination method and device, equipment, storage medium and program product
By determining the target network module in the large language model and cropping or migrating, the problem of high resource consumption of large language model is solved, and efficient training and low-cost applications are achieved.
Patent Information
- Application Number
- CN202510527969.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
Large language models require a large amount of resources during training and application, resulting in high research and application costs and low learning efficiency, making it difficult to effectively apply in resource-constrained environments.
In the target knowledge learning scenario, the target network module is determined based on the data set of the target knowledge inference function, and is cropped or transferred, the amount of parameters and calculation is reduced, and the cropped or migrated models are used for training, so as to achieve efficient training of the target large language model.
It reduces hardware training and application costs, improves the training efficiency and learning efficiency of large language models, and adapts to resource-constrained environments.
Smart Images

Figure CN120449982A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, storage medium, and program product for determining a target large language model. Background Art
[0002] In recent years, the field of artificial intelligence has made significant progress. Large language models have demonstrated excellent performance in tasks such as natural language processing and computer vision. By leveraging massive amounts of data and powerful computing power, large language models have demonstrated excellent performance in many fields. However, large language models often require a large amount of training data and massive computing resources, which increases research and application costs and is particularly difficult to apply in resource-constrained environments. Although there are currently a variety of learning methods for target large language models, these methods still have many problems, such as low learning efficiency of the target large language model (such as low training efficiency) and high hardware learning costs (such as high hardware training costs and high development costs). Summary of the Invention
[0003] The present application provides a method, apparatus, device, storage medium, and program product for determining a target large language model, which can reduce the hardware learning cost of the target large language model and improve the learning efficiency of the target large language model.
[0004] In a first aspect, an embodiment of the present application provides a method for determining a target large language model, comprising: in a target knowledge learning scenario, determining a target network module corresponding to the target knowledge reasoning function in a first large language model based on a data set corresponding to the target knowledge reasoning function; if the target knowledge learning scenario is a knowledge tailoring scenario, tailoring the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function to obtain the tailored first large language model; training the tailored first large language model based on a first training set to obtain a first target large language model; if the target knowledge learning scenario is a knowledge transfer scenario, migrating the target network module to a second large language model based on the data set corresponding to the target knowledge reasoning function to obtain the migrated second large language model; training the migrated second large language model based on a second training set to obtain a second target large language model.
[0005] In the second aspect, an embodiment of the present application also provides a device for determining a target large language model, including: a target network module determination module, which is used to determine the target network module corresponding to the target knowledge reasoning function in the first large language model based on the data set corresponding to the target knowledge reasoning function in a target knowledge learning scenario; a cropping module, which is used to crop the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function if the target knowledge learning scenario is a knowledge cropping scenario, to obtain the cropped first large language model; a first training module, which is used to train the cropped first large language model based on the first training set to obtain the first target large language model; a migration module, which is used to migrate the target network module to the second large language model based on the data set corresponding to the target knowledge reasoning function if the target knowledge learning scenario is a knowledge migration scenario, to obtain the migrated second large language model; a second training module, which is used to train the migrated second large language model based on the second training set to obtain the second target large language model.
[0006] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining the target large language model as described in the embodiment of the present application.
[0007] In a fourth aspect, an embodiment of the present application further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the method for determining a target large language model as described in an embodiment of the present application.
[0008] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the method for determining the target large language model as described in the embodiment of the present application.
[0009] The technical solution of the embodiment of the present application is as follows: in a target knowledge learning scenario, based on the data set corresponding to the target knowledge reasoning function, the target network module corresponding to the target knowledge reasoning function in the first large language model is determined; if the target knowledge learning scenario is a knowledge tailoring scenario, the target network module in the first large language model is tailored based on the data set corresponding to the target knowledge reasoning function to obtain the tailored first large language model; based on the first training set, the tailored first large language model is trained to obtain the first target large language model; if the target knowledge learning scenario is a knowledge transfer scenario, based on the data set corresponding to the target knowledge reasoning function, the target network module is transferred to the second large language model to obtain the transferred second large language model; based on the second training set, the transferred second large language model is trained to obtain the second target large language model. In an embodiment of the present application, in a knowledge tailoring scenario, by tailoring the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function, the number of parameters and the amount of calculation can be reduced. Thus, the tailored first large language model is trained based on the first training set to obtain the first target large language model. This method can improve the training efficiency of the first large language model and reduce the hardware training cost of the first large language model, and also reduce the hardware application cost of the first target large language model. In the case of knowledge migration, the target network module is migrated to the second large language model based on the data set corresponding to the target knowledge reasoning function, so that the reuse of the target network module can be achieved. Thus, the migrated second large language model is trained based on the second training set to obtain the second target large language model. This method can avoid repeated training of the corresponding parameters of the target network module (or target knowledge reasoning function), reduce the amount of calculation, and thus improve the training efficiency of the second large language model and reduce the hardware training cost of the second large language model. At the same time, it can also reduce the research cost and development cost of the second target large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 A flowchart of a method for determining a target large language model provided in an embodiment of the present application;
[0012] Figure 2 A flowchart of another method for determining a target large language model provided in an embodiment of the present application;
[0013] Figure 3 A flowchart of another method for determining a target large language model provided in an embodiment of the present application;
[0014] Figure 4 A schematic diagram of the structure of a device for determining a target large language model provided in an embodiment of the present application;
[0015] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0016] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0017] It should be understood that the various steps described in the method implementation of the present disclosure can be performed in different orders and / or in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. The term "including" and its variations used herein are open inclusions, that is, "including but not limited to". It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative and not restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more". It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and relevant provisions.
[0018] Figure 1 This is a flow chart of a method for determining a target large language model provided in an embodiment of the present application. The method can be executed by a device for determining a target large language model, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, PC or server. Figure 1 As shown, the method includes:
[0019] S110 . In a target knowledge learning scenario, determining a target network module corresponding to the target knowledge reasoning function in the first language model based on a data set corresponding to the target knowledge reasoning function.
[0020] Among them, the target knowledge learning scenario can include knowledge tailoring scenario and knowledge transfer scenario.
[0021] In this embodiment, the knowledge tailoring scenario can be to prune the target network module corresponding to the target knowledge reasoning function in the first language model to make it lighter and more efficient, so as to be suitable for scenarios in resource-constrained environments.
[0022] In this embodiment, the knowledge transfer scenario can be to migrate the target network module corresponding to the target knowledge reasoning function in the first language model (equivalent to the source model) to the second language model (equivalent to the target model), so as to improve the performance of the second language model in related tasks or adapt to new fields by reusing the knowledge corresponding to the target knowledge reasoning function.
[0023] In this embodiment, there is no limitation on the knowledge reasoning function. For example, it can be logical reasoning ability (such as the ability to derive mathematical formulas), procedural reasoning ability (such as the ability to write program code), spatiotemporal reasoning (such as, "It is 9:00 AM in country A, what time is it in country B?"), etc. The target knowledge reasoning function can be understood as one of the multiple knowledge reasoning functions, such as logical reasoning ability.
[0024] The data set can be understood as the training set of the first language model. In this embodiment, there is no restriction on the source of the data set, for example, it can be a public data set, or a data set constructed by a model or artificially.
[0025] In this embodiment, the first language model is not limited and can be, for example, the Tongyi Qianwen 2.5-72B model (Qwen2.5-72B), a Generative Pre-trained Transformer model (GPT), or a self-developed model. The first language model has a target knowledge reasoning function.
[0026] The network module can be understood as a neural network structural unit, such as an embedding layer, encoder, decoder, self-attention layer, feedforward neural network (FFN) layer, residual connection, and other components (or combinations of components). Multiple network modules can together constitute the underlying computing architecture of the first language model. The target network module can be the network module corresponding to the target knowledge reasoning function.
[0027] In this embodiment, in the target knowledge learning scenario, a first knowledge reasoning result set is determined based on a data set corresponding to the target knowledge reasoning function through a first large language model without any intervention; a plurality of network modules in the first large language model are intervened one by one to obtain a plurality of intervened network modules; a new knowledge reasoning result set is determined based on a data set corresponding to the target knowledge reasoning function through each of the plurality of intervened network modules; and a target network module that has a significant impact on the target knowledge reasoning function is screened out based on the new knowledge reasoning result set and the first knowledge reasoning result set.
[0028] S120: If the target knowledge learning scenario is a knowledge tailoring scenario, tailor the target network module in the first language model based on the data set corresponding to the target knowledge reasoning function to obtain the tailored first language model.
[0029] In this embodiment, for any one of the multiple target network modules, the target network module is pruned, and the parameter dimensions of the network modules connected to it are adjusted accordingly to obtain a candidate pruned first-largest language model. The performance of the candidate pruned first-largest language model is evaluated. If the candidate pruned first-largest language model no longer has the target knowledge reasoning function and maintains the integrity and validity of its remaining knowledge reasoning function (i.e., non-target knowledge reasoning function), the candidate pruned first-largest language model is used as the pruned first-largest language model. If the candidate pruned first-largest language model still has the target knowledge reasoning function or cannot maintain the integrity and validity of its remaining knowledge reasoning function (i.e., non-target knowledge reasoning function), the pruning of the target network module is canceled, and the next target network module is pruned, and so on, until the pruned first-largest language model is obtained.
[0030] S130: Based on the first training set, train the pruned first large language model to obtain a first target large language model.
[0031] In this embodiment, there is no restriction on the specific content of the first training set, but it is necessary to remove content related to the target knowledge reasoning function. For example, the first training set can cover industries such as technology (accounting for 30%), finance (accounting for 20%), medical care (accounting for 15%), education (accounting for 15%), retail (accounting for 10%), manufacturing (accounting for 5%), and entertainment (accounting for 5%). The weight ratio can be adjusted according to actual conditions. For example, focusing on the business field can increase the proportion of finance and retail, while focusing on innovation can increase the weight of technology and medical care. In this embodiment, the quality and format of the first training set can also be verified. In this embodiment, there is no restriction on the verification method. For example, it can be data integrity verification (such as checking whether there are missing values or empty fields, such as missing text, incomplete labels, etc.), format consistency verification (such as checking whether the format is unified), content compliance verification (such as filtering sensitive information), manual sampling (such as random sampling review to ensure that data labeling is accurate and the content is reasonable), etc.
[0032] In this embodiment, there is no restriction on the training method of the pruned first language model, for example, fine-tuning, incremental training, etc.
[0033] In this embodiment, since the first training set covers data from multiple industries, the model parameters can be adjusted to retain key knowledge and achieve the effect of knowledge integration; through training, the structural damage caused by cropping can be repaired, the logic and generalization capabilities can be improved, and the reasoning ability of the first target large language model can be strengthened.
[0034] S140: If the target knowledge learning scenario is a knowledge transfer scenario, the target network module is migrated to the second largest language model based on the data set corresponding to the target knowledge reasoning function to obtain the migrated second largest language model.
[0035] In this embodiment, there is no restriction on the second largest language model. For example, it can be the Tongyi Qianwen 2.5-72B model (Qwen2.5-72B), the Generative Pre-trained Transformer model (GPT), etc., or it can be a self-developed model. The first largest language model is different from the second largest language model, and the composition structure and knowledge reasoning function can be different. The second largest language model does not have the target knowledge reasoning function. The migrated second largest language model has the target knowledge reasoning function of the first largest language model migration.
[0036] In this embodiment, there is no restriction on the migration method. For example, the entire target network module can be embedded in the architecture of the second largest language model; the target network module can also be integrated into the target operator in the second largest language model to form a new composite operator; or the parameter set of the target network module can be integrated and fused with the parameter set of the target operator in the second largest language model to form a unified parameter space.
[0037] S150: Based on the second training set, train the migrated second largest language model to obtain a second target large language model.
[0038] In this embodiment, the specific content of the second training set is not limited; the content can cover a wide range of industries, and the weighting of samples from each industry can be adjusted based on actual circumstances. The second training set can be richer in content than the first training set. For example, training samples can be manually written or automatically generated using natural language generation technology or template filling methods. Data augmentation strategies can be used to increase the diversity and richness of training samples, resulting in new training samples. These training samples and new training samples can then be combined to form the second training set.
[0039] In this embodiment, the quality and format of the second training set can also be verified. In this embodiment, there is no limitation on the verification method, and examples include data integrity verification (such as checking for missing values or empty fields, such as missing text, incomplete labels, etc.), format consistency verification (such as checking whether the format is unified), content compliance verification (such as filtering sensitive information), manual sampling (such as random sampling review to ensure that data labeling is accurate and content is reasonable), etc.
[0040] In this embodiment, there is no restriction on the training method of the second largest language model after migration, for example, fine-tuning, incremental training, etc.
[0041] In this embodiment, the second training set is richer in content, resulting in a second language model with stronger generalization and domain adaptability after migration. Through training, model parameters can be adjusted to adapt to the new data distribution, improving the reasoning ability of the second language model after migration.
[0042] The technical solution of the embodiment of the present application is as follows: in a target knowledge learning scenario, based on the data set corresponding to the target knowledge reasoning function, the target network module corresponding to the target knowledge reasoning function in the first large language model is determined; if the target knowledge learning scenario is a knowledge tailoring scenario, the target network module in the first large language model is tailored based on the data set corresponding to the target knowledge reasoning function to obtain the tailored first large language model; based on the first training set, the tailored first large language model is trained to obtain the first target large language model; if the target knowledge learning scenario is a knowledge transfer scenario, based on the data set corresponding to the target knowledge reasoning function, the target network module is migrated to the second large language model to obtain the migrated second large language model; based on the second training set, the migrated second large language model is trained to obtain the second target large language model. In an embodiment of the present application, in a knowledge tailoring scenario, by tailoring the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function, the number of parameters and the amount of computation can be reduced. Thus, the tailored first large language model is trained based on the first training set to obtain the first target large language model. This method can improve the training efficiency of the first large language model and reduce the hardware training cost of the first large language model, and also reduce the hardware application cost of the first target large language model. In the case of knowledge migration, by migrating the target network module to the second large language model based on the data set corresponding to the target knowledge reasoning function, the target network module can be reused across models. Thus, the migrated second large language model is trained based on the second training set to obtain the second target large language model. This method can avoid repeated training of the corresponding parameters of the target network module (or target knowledge reasoning function), reduce the amount of computation, and thus improve the training efficiency of the second large language model and reduce the hardware training cost of the second large language model. At the same time, it can also reduce the research cost and development cost of the second target large language model.
[0043] Figure 2 This is a flow chart of another method for determining a target large language model provided by the embodiment of the present application. This embodiment is applied to the knowledge tailoring scenario. This embodiment of the present application is a specific implementation based on the above invention embodiment. Figure 2 The method provided in the embodiment of the present application specifically includes the following steps:
[0044] S201. Perform knowledge reasoning on the dataset based on questions and answers using the first language model to obtain a first knowledge reasoning result set.
[0045] Among them, the first language model includes multiple network modules.
[0046] In this embodiment, there is no restriction on the source of the question-answer dataset. For example, it can be a public dataset, or a dataset constructed through a model or artificial means.
[0047] For example, the question-answer pair dataset is recorded as baselineSet_orgin, and any sample form in the question-answer pair dataset can be simply recorded as S={Q, A}, where Q is the prompt word and A is the standard answer.
[0048] Exemplarily, without any intervention, knowledge reasoning operations can be performed on all samples in the question-answer pair dataset based on the first language model to obtain a first knowledge reasoning result set.
[0049] S202: The question-answer pair dataset and the first knowledge reasoning result set are correspondingly combined into a question-answer pair result dataset.
[0050] Among them, the question-answer result dataset includes a dataset corresponding to the target knowledge reasoning function and a dataset corresponding to the non-target knowledge reasoning function.
[0051] For example, for any sample S={Q,A} in the question-answer pair dataset, the first knowledge inference result is generated as A obj (The first knowledge reasoning result set consists of multiple first knowledge reasoning results). obj Merge sample S, then the sample form in the question-answer pair result dataset is S1={Q,A,A obj}, the question-answer result dataset can be recorded as baselineSet.
[0052] S203: Execute the intervention strategy on each of the multiple network modules to obtain multiple intervened network modules.
[0053] In this embodiment, there is no limitation on the intervention strategy, and for example, various strategies may be included, including zeroing, noise injection, and weight adjustment, etc. The intervention strategy of each network module may be different.
[0054] In this embodiment, there is no restriction on the order in which each network module executes the intervention strategy. For example, the intervention strategy can be executed one by one according to the sequence number of each network module to obtain multiple network modules after intervention. The network module after intervention is recorded as LLM. obj seq , seq is the network module sequence number.
[0055] S204. Determine the target network module corresponding to the target knowledge reasoning function in the first language model based on the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function through the multiple intervened network modules.
[0056] In this embodiment, the reasoning performance of each of the multiple intervened network modules on the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function is determined (respectively recorded as the second knowledge reasoning result set and the third knowledge reasoning result set), and based on the second knowledge reasoning result set, the third knowledge reasoning result set and the first knowledge reasoning result set, the target network module that has a significant impact on the target knowledge reasoning function is screened out.
[0057] Optionally, through multiple intervened network modules, the target network module corresponding to the target knowledge reasoning function in the first large language model is determined based on the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function, including: determining a second knowledge reasoning result set based on the data set corresponding to the target knowledge reasoning function through each intervened network module; wherein the data set corresponding to the target knowledge reasoning function includes a question-answer pair data set corresponding to the target knowledge reasoning function and a first knowledge reasoning result set corresponding to the target knowledge reasoning function; determining a third knowledge reasoning result set based on the data set corresponding to the non-target knowledge reasoning function through each intervened network module; determining the similarity between the second knowledge reasoning result set and the elements with the same position index in the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtaining multiple first similarity results; determining the similarity between the third knowledge reasoning result set and the elements with the same position index in the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function, and obtaining multiple second similarity results; determining the target network module corresponding to the target knowledge reasoning function in the first large language model based on the multiple first similarity results and the multiple second similarity results.
[0058] Among them, the data set corresponding to the non-target knowledge reasoning function includes the question-answer pair data set corresponding to the non-target knowledge reasoning function and the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function.
[0059] For example, a dataset corresponding to the target knowledge reasoning function is extracted from the question-answer pair result dataset baselineSet, which is denoted as S set =[S1={Q1,A 1, A obj1},S2={Q2,A2,A obj2},...,S n ={Q n ,A n ,A o bjn}], and randomly extract the data set corresponding to the non-target knowledge reasoning function, denoted as S' set =[S'1={Q'1,A'1,A' obj1},S'2={Q'2,A'2,A' obj2},...,S'n ={Q' n ,A' n ,A' objn Based on each intervention network module LLM obj seq , respectively for the dataset S corresponding to the target knowledge reasoning function set , the dataset S' corresponding to the non-target knowledge reasoning function set Execute the knowledge reasoning operation and generate the results as the second knowledge reasoning result set A set seq =[A seq1 ,A seq2 ,....,A seqn ], the third knowledge reasoning result set A' set seq =[A' seq1 ,A' seq2, ....,A' seqn ] Calculate A respectively seqi With A obji 、A' seqi With A' obji where i represents the position index.
[0060] In this embodiment, there is no restriction on the similarity algorithm, and any similarity calculation method in the field of artificial intelligence can be used, such as the cosine similarity algorithm, which is denoted as R. set seq ={...,R(A obji ,A seqi ),...}, multiple second similarity results T' set seq ={...,R(A' obji ,A' seqi ),...}.
[0061] In this embodiment, the LLM can be determined by analyzing multiple first similarity results and multiple second similarity results. obj Whether the network module after intervention with sequence number seq has a causal relationship with the target knowledge reasoning function, the network module after intervention that has a causal relationship with the target knowledge reasoning function is used as the target network module.
[0062] In this embodiment, a second knowledge reasoning result set and a third knowledge reasoning result set are generated based on the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function for each intervened network module, respectively. The target network module corresponding to the target knowledge reasoning function is located based on multiple first similarity results between the second knowledge reasoning result set and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and based on multiple second similarity results between the third knowledge reasoning result set and the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function. This can improve the accuracy of target network module identification.
[0063] Optionally, determining a target network module corresponding to a target knowledge reasoning function in a first large language model based on multiple first similarity results and multiple second similarity results includes: performing a mean operation on multiple first similarity results to obtain a first mean result; performing a mean operation on multiple second similarity results to obtain a second mean result; determining a first mean difference result based on the first mean result and the second mean result; if the first mean difference result is greater than a first set threshold, taking the corresponding intervened network module as the target network module to obtain at least one target network module.
[0064] In this embodiment, LLM obj Whether the network module after intervention with the sequence number seq has a causal relationship with the target knowledge reasoning function can be analyzed through the causal effect between the two. Based on the double difference analysis principle, the average causal effect is ACE(X seq T )=|AVG(T set seq )-AVG(T' set seq )|, AVG is the mean operation, where AVG(T set seq ) is the first mean result, AVG(T' set seq ) is the second mean value result. ACE(X seq T ) can represent the first mean difference result. If ACE(X seq T ) is greater than a first set threshold (e.g., 0.5), then it is determined that the network module after intervention with the sequence number seq has a causal relationship with the target knowledge reasoning function. According to the sequence number of the network module after intervention, the causal analysis of all network modules after intervention and the target knowledge reasoning function is completed in sequence, and all target network modules with causal relationships are recorded as LLM obj (X)={...,[seq,ACE(X seq T )],...}.
[0065] In this embodiment, by determining the mean difference between multiple first similarity results and multiple second similarity results and comparing them with the first set threshold, the target network module that has a significant impact on the target knowledge reasoning function is screened out, which can further improve the accuracy of target network module identification.
[0066] S205 . Sort the at least one target network module according to the first mean difference result corresponding to each target network module in the at least one target network module to obtain a first sorting result.
[0067] For example, the first mean difference result ACE(X seq T ), sort all target network modules from large to small, and obtain the first sorting result.
[0068] S206: Prune the first language model based on the current target network module corresponding to the first sorting result to obtain a candidate pruned model.
[0069] In this embodiment, the target network module in the first large language model may be pruned according to the order of the first sorting result. For the current target network module, the current target network module in the first large language model may be pruned to obtain a candidate pruned model.
[0070] S207 : Determine whether the candidate trimmed model is the largest language model after trimming based on the data set corresponding to the target knowledge reasoning function.
[0071] In this embodiment, through the candidate trimming model, a trimmed knowledge reasoning set (which can be recorded as the trimmed second knowledge reasoning result set and the trimmed third knowledge reasoning result set, respectively) can be generated according to the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function, and based on the trimmed second knowledge reasoning result set and the trimmed third knowledge reasoning result set and the first knowledge reasoning result set, it is judged whether the candidate trimming model is the trimmed first language model. If the candidate trimming model is the trimmed first language model, the trimming process is terminated. If the candidate trimming model is not the trimmed first language model, the trimming of the current target network module is canceled, and the trimming operation of the target network module in the first language model continues in the order of the first sorting result (that is, the next target network module is trimmed) until the trimmed first language model is obtained.
[0072] Among them, the first language model after trimming can be understood as a model that does not have the target knowledge reasoning function and maintains the integrity and effectiveness of the non-target reasoning function (that is, it has correct and complete non-target knowledge reasoning function).
[0073] Optionally, judging whether the candidate cropped model is the first largest language model after cropping based on the data set corresponding to the target knowledge reasoning function includes: determining the cropped second knowledge reasoning result set based on the data set corresponding to the target knowledge reasoning function through the candidate cropped model; determining the cropped third knowledge reasoning result set based on the data set corresponding to the non-target knowledge reasoning function through the candidate cropped model; determining the similarity between the cropped second knowledge reasoning result set and the elements with the same position index in the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtaining multiple cropped first similarity results; determining the similarity between the cropped third knowledge reasoning result set and the elements with the same position index in the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function, and obtaining multiple cropped second similarity results; judging whether the candidate cropped model is the first largest language model after cropping based on the multiple cropped first similarity results and the multiple cropped second similarity results.
[0074] For example, a dataset corresponding to the target knowledge reasoning function is extracted from the question-answer pair result dataset baselineSet, which is denoted as S set =[S1={Q1,A 1, A obj1},S2={Q2,A2,A obj2},...,S n ={Q n ,A n ,A o bjn}], and randomly extract the data set corresponding to the non-target knowledge reasoning function, denoted as S' set =[S'1={Q'1,A'1,A' obj1},S'2={Q'2,A'2,A' obj2},...,S' n ={Q' n ,A' n ,A' objn Based on the candidate cropping model, the dataset S corresponding to the target knowledge reasoning function is set , the dataset S' corresponding to the non-target knowledge reasoning function set Execute the knowledge reasoning operation and generate the results as the trimmed second knowledge reasoning result set B set seq =[B seq1 ,B seq2 ,....,B seqn ], the trimmed third knowledge reasoning result set B' set seq =[B' seq1 ,B' seq2, ....,B' seqn] Calculate B respectively seqi With A obji , B' seqi With A' obji where i represents the position index.
[0075] In this embodiment, there is no restriction on the similarity algorithm, and any similarity calculation method in the field of artificial intelligence can be used, such as the cosine similarity algorithm, which is denoted as R. s et seq ={...,R(A obji ,B seqi ),...}, multiple cropped second similarity results K' set seq ={...,R(A' obji ,B' seq i ),...}.
[0076] In this embodiment, whether the candidate cropped model is the cropped first large language model may be determined by comparing a plurality of cropped first similarity results and a plurality of cropped second similarity results.
[0077] In this embodiment, a candidate trimming model is used to generate a trimmed second knowledge reasoning result set and a trimmed third knowledge reasoning result set based on the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function, respectively. The candidate trimming model is judged to be the trimmed first largest language model based on multiple trimmed first similarity results between the trimmed second knowledge reasoning result set and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and based on multiple trimmed second similarity results between the trimmed third knowledge reasoning result set and the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function. This method can improve the accuracy of recognition of the trimmed first largest language model.
[0078] Optionally, determining whether a candidate trimmed model is the trimmed first largest language model based on multiple trimmed first similarity results and multiple trimmed second similarity results includes: performing a mean operation on the multiple trimmed first similarity results to obtain a trimmed first mean result; performing a mean operation on the multiple trimmed second similarity results to obtain a trimmed second mean result; if the trimmed first mean result is less than a second set threshold and the trimmed second mean result is greater than a third set threshold, determining the candidate trimmed model as the trimmed first largest language model; if the trimmed first mean result is greater than or equal to the second set threshold, or the trimmed second mean result is less than or equal to the third set threshold, canceling the trimming operation on the first largest language model by the current target network module, and trimming the first largest language model based on the next target network module corresponding to the first sorting result to obtain a new candidate trimmed model; using the new candidate trimmed model as the candidate trimmed model, and returning to the step of determining whether the candidate trimmed model is the trimmed first largest language model based on the data set corresponding to the target knowledge reasoning function, and re-executing.
[0079] For example, the first similarity results K after multiple clipping set seq ={...,R(A obji ,B seqi ),...}, multiple cropped second similarity results K' set seq ={...,R(A' obji ,B' seqi ),...}. AVG is the average operation. If the first average result after clipping is AVG(K set seq ) is less than the second set threshold (the second set threshold is not limited, for example, it can be 0.3), and the clipped second mean result AVG(K' set seq ) is greater than the third set threshold (there is no limit on the third set threshold, for example, it can be 0.7), the candidate trimmed model is determined as the first largest language model after trimming, which means that the candidate trimmed model at this time has removed the target knowledge reasoning function and maintains the integrity and effectiveness of its remaining knowledge reasoning function (non-target knowledge reasoning function). If AVG(K set seq ) is greater than or equal to the second set threshold, or AVG(K' set seq) is less than or equal to the third set threshold, the pruning of the current target network module is canceled, and the first large language model is pruned based on the next target network module corresponding to the first sorting result to obtain a new candidate pruned model; the new candidate pruned model is used as the candidate pruned model, and the step of determining whether the candidate pruned model is the pruned first large language model based on the data set corresponding to the target knowledge reasoning function is returned and re-executed.
[0080] In this embodiment, the accuracy of recognition of the cropped first largest language model is further improved by determining the average result of multiple cropped first similarity results to obtain the cropped first average result, determining the average result of multiple cropped second similarity results to obtain the cropped second average result, comparing the cropped first average result with the second set threshold, and comparing the cropped second average result with the third set threshold to determine whether the candidate cropped model is the cropped first largest language model.
[0081] S208: Based on the first training set, train the pruned first large language model to obtain a first target large language model.
[0082] In this embodiment, the target network module corresponding to the target knowledge reasoning function in the first large language model is determined based on the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function through multiple intervened network modules, so that the target network module can be accurately located. At least one target network module is sorted according to the first mean difference result corresponding to each target network module in at least one target network module to obtain a first sorting result. The method of trimming the target network module in the first large language model based on the first sorting result can improve the efficiency and accuracy of trimming, and can reduce repeated training and development time, improve the learning efficiency of the trimmed first large language model, and reduce the overall cost of artificial intelligence research and application. After obtaining the trimmed first large language model, the method of training the trimmed first large language model based on the first training set to obtain the first target large language model can further optimize the performance, so that a lightweight and high-precision first target large language model can be obtained.
[0083] Figure 3 This is a flow chart of another method for determining a target large language model provided by the embodiment of the present application. This embodiment is applied to the knowledge transfer scenario. This embodiment of the present application is a specific implementation based on the above invention embodiment. Figure 3 The method provided in the embodiment of the present application specifically includes the following steps:
[0084] S301. Perform knowledge reasoning on the dataset based on questions and answers using the first language model to obtain a first knowledge reasoning result set.
[0085] Among them, the first language model includes multiple network modules.
[0086] S302: The question-answer pair dataset and the first knowledge reasoning result set are correspondingly combined into a question-answer pair result dataset.
[0087] Among them, the question-answer result dataset includes a dataset corresponding to the target knowledge reasoning function and a dataset corresponding to the non-target knowledge reasoning function.
[0088] S303: Execute the intervention strategy on each of the multiple network modules to obtain multiple intervened network modules.
[0089] S304. Determine the target network module corresponding to the target knowledge reasoning function in the first language model based on the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function through the multiple intervened network modules.
[0090] It should be noted that in the knowledge transfer scenario, the method of determining the target network module corresponding to the target knowledge reasoning function in the first language model is similar to the method of determining the target network module corresponding to the target knowledge reasoning function in the first language model in the knowledge tailoring scenario, and will not be repeated in this embodiment.
[0091] S305 . Sort the at least one target network module according to the first mean difference result corresponding to each target network module in the at least one target network module to obtain a second sorting result.
[0092] Exemplarily, all target network modules may be sorted from large to small according to the size of the first mean difference result to obtain a second sorting result.
[0093] S306: Migrate the current target network module corresponding to the second sorting result to the second largest language model to obtain a candidate migration model.
[0094] In this embodiment, the target network module in the first language model can be migrated to the second language model according to the order of the second sorting result. For the current target network module, the current target network module in the first language model can be migrated to the second language model to obtain a candidate migration model.
[0095] S307 : Determine whether the candidate transfer model is the second largest language model after transfer based on the data set corresponding to the target knowledge reasoning function.
[0096] In this embodiment, through the candidate migration model, a knowledge reasoning set after migration (which can be recorded as the second knowledge reasoning result set after migration) can be generated according to the data set corresponding to the target knowledge reasoning function, and based on the second knowledge reasoning result set after migration and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, it is judged whether the candidate migration model is the second largest language model after migration. If the candidate migration model is the second largest language model after migration, the migration process is terminated. If the candidate migration model is not the second largest language model after migration, the migration of the current target network module is canceled, and the next target network module is migrated to the second largest language model in the order of the second sorting result, until the second largest language model after migration is obtained.
[0097] Among them, the second largest language model after migration can be understood as a model that obtains the target knowledge reasoning function through migration.
[0098] Optionally, judging whether the candidate migration model is the second largest language model after migration based on the data set corresponding to the target knowledge reasoning function includes: determining the second knowledge reasoning result set after migration based on the data set corresponding to the target knowledge reasoning function through the candidate migration model; determining the similarity between the elements with the same position index in the second knowledge reasoning result set after migration and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtaining multiple first similarity results after migration; performing a mean operation on the multiple first similarity results after migration to obtain the first mean result after migration; if the first mean result after migration is greater than a fourth set threshold, determining the candidate migration model as the second largest language model after migration; if the first mean result after migration is less than or equal to the fourth set threshold, canceling the embedding operation of the current target network module on the second largest language model, migrating to the second largest language model based on the next target network module corresponding to the second sorting result, and obtaining a new candidate migration model; using the new candidate migration model as the candidate migration model, and returning to the step of judging whether the candidate migration model is the second largest language model after migration based on the data set corresponding to the target knowledge reasoning function, and re-executing.
[0099] For example, a dataset corresponding to the target knowledge reasoning function is extracted from the question-answer pair result dataset baselineSet, which is denoted as S set =[S1={Q1,A 1, A obj1},S2={Q2,A2,A obj2},...,S n ={Q n ,A n ,A o bjn}], the first knowledge reasoning result set corresponding to the target knowledge reasoning function is A obj1 、Aobj2 ,...,A objn Based on the candidate transfer model, the dataset S corresponding to the target knowledge reasoning function set Execute the knowledge reasoning operation and generate the result as the second knowledge reasoning result set C after migration set seq =[C seq1 ,C seq2 ,....,C seqn ], calculate C seqi With A obji , where i represents the position index and seq is the sequence number of the target network module.
[0100] In this embodiment, there is no restriction on the similarity algorithm, and any similarity calculation method in the field of artificial intelligence can be used, such as the cosine similarity algorithm, which is denoted as R. s et seq ={...,R(A obji ,C seqi ),...}. AVG is the average operation. If the first average result after migration is AVG(M set seq ) If it is greater than the fourth set threshold (there is no restriction on the fourth set threshold, for example, it can be 0.7), the candidate migration model is determined to be the second largest language model after migration, which means that the candidate migration model at this time has successfully migrated the target knowledge reasoning function. If the first mean value result after migration is less than or equal to the fourth set threshold, the embedding operation of the current target network module on the second largest language model is canceled, and the next target network module corresponding to the second sorting result is migrated to the second largest language model to obtain a new candidate migration model; the new candidate migration model is used as the candidate migration model, and the step of judging whether the candidate migration model is the second largest language model after migration based on the data set corresponding to the target knowledge reasoning function is returned and re-executed.
[0101] In this embodiment, a candidate migration model is used to generate a migrated second knowledge reasoning result set based on a data set corresponding to a target knowledge reasoning function, and based on multiple migrated first similarity results between the migrated second knowledge reasoning result set and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, the first mean result after migration is obtained by determining the mean result of the multiple migrated first similarity results, and the first mean result after migration is compared with a fourth set threshold to determine whether the candidate migration model is the second largest language model after migration. This method can further improve the accuracy of recognition of the second largest language model after migration.
[0102] Optionally, the current target network module corresponding to the second sorting result is migrated to the second largest language model to obtain a candidate migration model, including: determining the embedding position of the current target network module in the second largest language model; modifying the current target network module according to the data format information corresponding to the embedding position to obtain the modified current target network module; and embedding the modified current target network module into the embedding position.
[0103] In this embodiment, the entire module of the current target network module in the first language model can be embedded into the architecture of the second language model. The specific steps are as follows: Step 1, determine the embedding position. If the current target network module is used to process or enhance the representation of input data, it can be inserted immediately after the embedding layer of the second language model, that is, the embedding position is after the embedding layer of the second language model; if the current target network module is used to enhance the model's ability to understand text, it can be inserted between multiple encoders of the second language model, that is, the embedding position is between multiple encoders of the second language model; if the current target network module is used to improve the quality or diversity of generated text, it can be placed before or between multiple decoders of the second language model, that is, the embedding position is before or between multiple decoders of the second language model. Step 2, modify the current target network module to achieve input and output interface adaptation. First, determine the data format information corresponding to the embedding position, such as the data type, format, dimension and additional information (such as attention weight) of the input and output before and after the embedding layer. According to the data format information, modify the relevant components of the current target network module to meet the input and output requirements of the second language model. If the internal structure of the current target network module is not easy to modify directly, or in order to maintain its independence, a special adaptation layer or middleware can be developed to bridge the current target network module and the second largest language model. The adaptation layer or middleware is responsible for data conversion, formatting, and possible pre-processing / post-processing. Step three, embed the modified current target network module into the embedding position. Add the initialization code of the modified current target network module to the initialization function of the second largest language model, and add an instance of the modified current target network module to the class of the second largest language model or the corresponding architecture configuration file. Modify the data flow path to ensure that at the embedding position, data can flow into the modified current target network module and continue to flow to the subsequent modules in the second largest language model after processing. Finally, load the parameters of the modified current target network module in the first largest language model into the second largest language model.
[0104] In this embodiment, the embedding position of the current target network module in the second largest language model is determined; the current target network module is modified according to the data format information corresponding to the embedding position, and the modified current target network module is embedded into the embedding position. This method can solve the structural difference problem during cross-model migration, accurately migrate the current target network module to the second largest language model, and ensure the compatibility and functional accuracy of the second largest language model with the current target network module.
[0105] Optionally, the current target network module corresponding to the second sorting result is migrated to the second largest language model to obtain a candidate migration model, including: determining the target operator of the second largest language model based on the current target network module; adjusting the current target network module according to the interface information of the target operator to obtain the adjusted current target network module; establishing a mapping relationship between the parameters of the target operator and the parameters of the adjusted current target network module; migrating the adjusted current target network module to the target operator to obtain a composite operator; loading the parameters of the adjusted current target network module into the composite operator according to the mapping relationship; and adjusting the interface format of the composite operator to be compatible with the second largest language model.
[0106] Operators can be understood as the basic operation units in deep learning or mathematical computing, used to perform specific calculations (such as matrix multiplication and convolution). They can be regarded as "functional modules" or "computational steps" in the model. In neural networks, operators can correspond to specific operations at a certain layer (such as activation functions and normalization).
[0107] In this embodiment, the current target network module can be integrated into the target operator in the second largest language model to form a new composite operator. The specific steps are as follows: Step 1: Target operator selection. The current target network module is analyzed to determine its network structure, core functions, and its position and capabilities (such as feature extraction and sequence processing capabilities) in the network structure of the first largest language model. Based on this, operators with similar or complementary network structures and functions to the current target network module are selected from the operators of the second largest language model as target operators. Step 2: Adjust the current target network module. The interface information of the target operator is analyzed. The interface information is the input and output interface, including data types, dimensionality requirements, etc. Based on the interface information of the target operator, the input and output structure of the current target network module is adjusted to ensure that the current target network module and the target operator can seamlessly connect at the data level. In this embodiment, the consistency of the target operator logic before and after fusion can also be verified through simulation operation or unit testing to ensure that no logical errors are introduced during the fusion process. Step 3: Fusion reconstruction implementation. Analyze the parameters of the target operator and the adjusted current target network module, and establish a mapping relationship between the parameters of the target operator and the adjusted parameters of the current target network module. Migrate (merge) the code of the adjusted current target network module into the code of the target operator to obtain a composite operator. Based on the mapping relationship, load the parameters of the current target network module into the composite operator. Adjust the interface format of the composite operator, such as designing a unified external interface, to ensure compatibility between the composite operator and the second largest language model.
[0108] It should be noted that you can use optimization algorithms such as hyperparameter search and gradient descent to tune the parameters of composite operators. You can also optimize the memory usage and computational efficiency of composite operators by sharing memory and reducing redundant computations.
[0109] In this embodiment, by adjusting the current target network module to dynamically adapt to the target operator interface, establishing a mapping relationship between the parameters of the target operator and the adjusted parameters of the current target network module, migrating the adjusted current target network module to the target operator to obtain a composite operator, and loading the adjusted parameters of the current target network module into the composite operator according to the mapping relationship, and adjusting the format compatibility of the composite operator interface, the current target network module can be accurately migrated to the second largest language model, and the compatibility and functional accuracy of the second largest language model with the current target network module can be guaranteed.
[0110] This embodiment adjusts the interface information of the current network module to ensure that it matches the target operator structure, reducing inter-operator call conflicts. By reusing the adjusted parameters of the current target network module, retraining overhead can be avoided. After adjusting the interface format of the composite operator, the composite operator can be seamlessly integrated into the second language model, improving its scalability.
[0111] Optionally, the current target network module corresponding to the second sorting result is migrated to the second largest language model to obtain a candidate migration model, including: determining the target operator of the second largest language model based on the current target network module; spatially aligning the parameters of the target operator and the parameters of the current target network module to obtain the aligned target operator parameters and the aligned current target network module parameters; determining the weight matrix of the target operator; determining the correction direction and correction strength of the aligned current target network module parameters on the weight matrix; determining a parameter merging operation function based on the weight matrix, the correction direction and the correction strength; and executing the parameter merging operation function to merge the aligned current target network module parameters into the aligned target operator parameters.
[0112] In this embodiment, the parameter set of the target network module can be integrated and fused with the parameter set of the target operator in the second largest language model to form a unified parameter space. The specific steps are as follows: Step 1: Target operator selection. The current target network module is analyzed to determine its network structure, core functions, and its position and capabilities (such as feature extraction and sequence processing capabilities) within the network structure of the first largest language model. Based on this, operators in the second largest language model are selected that have similar network structure and functions to those of the current target network module and serve as target operators. Step 2: Spatial alignment of the parameters of the target operator with those of the current target network module. Spatial alignment can be achieved by adjusting the data type, data dimension, data range, etc. to ensure that the parameters of the target operator and the parameters of the current target network module are consistent in shape and distribution, facilitating subsequent fusion or migration. Specifically, the parameters of the current target network module and the target operator (there may be multiple parameters) are obtained, and their detailed properties, including data type, data dimension, and data range, are recorded. The parameter sets of the current target network module and the target operator are compared and analyzed to identify potential compatibility issues, such as data type mismatch and inconsistent data dimension. For any compatibility issues discovered, design and implement necessary parameter adjustment strategies or conversion algorithms, such as data type conversion, dimension and precision adjustment, and data range normalization. There are no compatibility issues between the aligned target operator parameters and the aligned current target network module parameters. Step 3: Parameter merging. Introduce the weight matrix W of the target operator. The shape of W is expressed as (input_features, output_features), abbreviated as m*n. Set the hyperparameter k to represent the degree of freedom for correction. The value of k can be much smaller than min(m,n). Introduce the low-rank matrix D with a shape of m*k and initialize it to a random small value. Introduce the low-rank matrix E with a shape of k*n. The E matrix can be set to fixed or trainable as needed to adapt to different merging scenarios and performance requirements. The low-rank matrices D and E represent the direction and strength of correction of the weight matrix W by the aligned current target network module parameters, respectively. The parameter merging operation function is W_new = W + D*E. According to the above parameter merging operation function, the merging operation is implemented in the code, that is, the parameter merging operation function is executed to merge the aligned current target network module parameters into the aligned target operator parameters.
[0113] It should be noted that for the training optimization portion, a loss function can be set, and the remaining parts of the second largest language model (except the current target network module and target operator) are set to evaluation mode. Then, only the correction direction and correction strength of the weight matrix are trained, and the weight matrix is optimized using methods such as gradient descent. In this embodiment, there is no limitation on the specific loss function, and for example, it can be mean square error, cross entropy loss, etc.
[0114] In this embodiment, the target operator is determined based on the current target network module, and parameter space alignment is performed to determine the weight matrix of the target operator. The correction direction and correction strength of the weight matrix are determined in combination with the aligned current target network module parameters. Based on the weight matrix, correction direction and correction strength, a parameter merging operation function is determined and executed. In this way, the aligned current target network module parameters are merged into the target operator, so that the current target network module can be accurately migrated to the second largest language model, and the compatibility and functional accuracy of the second largest language model with the current target network module can be guaranteed.
[0115] In this embodiment, through parameter alignment and dynamic correction of the weight matrix, the aligned current target network module parameters can be better merged into the aligned target operator parameters, thereby improving the generalization ability of the second language model. Through the parameter merging operation function, knowledge fusion between the first language model and the second language model can be achieved, avoiding redundant calculations and improving training or reasoning efficiency. The introduction of correction direction and correction strength can control the update amplitude of the weight matrix and prevent performance oscillation or degradation during the merging process.
[0116] S308 : Based on the second training set, train the migrated second largest language model to obtain a second target large language model.
[0117] In this embodiment, the target network module corresponding to the target knowledge reasoning function in the first large language model is determined based on the data set corresponding to the target knowledge reasoning function and the data set corresponding to the non-target knowledge reasoning function through multiple intervened network modules, so that the target network module can be accurately located. At least one target network module is sorted according to the first mean difference result corresponding to each target network module in at least one target network module to obtain a second sorting result. The current target network module corresponding to the second sorting result is migrated to the second large language model to obtain the migrated second large language model. This method can realize the sharing and use of the current target network module across models and the expansion of the second large language model, thereby reducing repeated training and reducing the amount of calculation, thereby improving the training efficiency of the second large language model and reducing the hardware training cost of the second large language model. At the same time, it can also reduce the research cost and development cost of the second target large language model. Based on the second training set, the migrated second large language model is trained to obtain a second target large language model, which can further optimize the performance and thus obtain a high-precision second target large language model.
[0118] Figure 4 A schematic diagram of a device for determining a target large language model provided in an embodiment of the present application is shown in FIG. Figure 4As shown, the apparatus includes: a target network module determination module 410, a cropping module 420, a first training module 430, a migration module 440, and a second training module 450;
[0119] A target network module determination module 410 is configured to determine, in a target knowledge learning scenario, a target network module corresponding to the target knowledge reasoning function in the first language model based on a data set corresponding to the target knowledge reasoning function;
[0120] A trimming module 420 is configured to trim the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function to obtain a trimmed first large language model if the target knowledge learning scenario is a knowledge trimming scenario;
[0121] A first training module 430 is configured to train the trimmed first large language model based on a first training set to obtain a first target large language model;
[0122] A migration module 440 is configured to migrate the target network module to a second language model based on a data set corresponding to the target knowledge reasoning function to obtain the migrated second language model if the target knowledge learning scenario is a knowledge migration scenario;
[0123] The second training module 450 is configured to train the migrated second large language model based on a second training set to obtain a second target large language model.
[0124] The technical solution of the embodiment of the present application is as follows: if the target knowledge learning scenario is a knowledge tailoring scenario, the target network module corresponding to the target knowledge reasoning function in the first large language model is determined by the target network module determination module based on the data set corresponding to the target knowledge reasoning function; if the target knowledge learning scenario is a knowledge tailoring scenario, the target network module in the first large language model is tailored based on the data set corresponding to the target knowledge reasoning function to obtain the tailored first large language model; if the target knowledge learning scenario is a knowledge transfer scenario, the target network module is trained on the first large language model based on the first training set to obtain the first target large language model; if the target knowledge learning scenario is a knowledge transfer scenario, the target network module is migrated to the second large language model based on the data set corresponding to the target knowledge reasoning function to obtain the migrated second large language model; if the target knowledge learning scenario is a knowledge transfer scenario, the migrated second large language model is trained on the second training set to obtain the second target large language model. In an embodiment of the present application, in a knowledge tailoring scenario, by tailoring the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function, the number of parameters and the amount of calculation can be reduced. Thus, the tailored first large language model is trained based on the first training set to obtain the first target large language model. This method can improve the training efficiency of the first large language model and reduce the hardware training cost of the first large language model, and also reduce the hardware application cost of the first target large language model. In the case of knowledge migration, by migrating the target network module to the second large language model based on the data set corresponding to the target knowledge reasoning function, the target network module can be reused across models. Thus, the migrated second large language model is trained based on the second training set to obtain the second target large language model. This method can avoid repeated training of the corresponding parameters of the target network module (or target knowledge reasoning function), reduce the amount of calculation, and thus improve the training efficiency of the second large language model and reduce the hardware training cost of the second large language model. At the same time, it can also reduce the research cost and development cost of the second target large language model.
[0125] Optionally, the target network module determination module is specifically used to: perform knowledge reasoning based on the question and answer pair data set through the first large language model to obtain a first knowledge reasoning result set; wherein, the first large language model includes multiple network modules; the question and answer pair data set and the first knowledge reasoning result set are correspondingly composed into a question and answer pair result data set; wherein, the question and answer pair result data set includes a data set corresponding to the target knowledge reasoning function and a data set corresponding to non-target knowledge reasoning functions; execute an intervention strategy on each network module in the multiple network modules to obtain multiple intervened network modules; through the multiple intervened network modules, determine the target network module corresponding to the target knowledge reasoning function in the first large language model based on the data set corresponding to the target knowledge reasoning function and the data set corresponding to non-target knowledge reasoning function.
[0126] Optionally, the target network module determination module is also used to: determine a second knowledge reasoning result set based on the data set corresponding to the target knowledge reasoning function through each of the intervened network modules; wherein the data set corresponding to the target knowledge reasoning function includes a question-answer pair data set corresponding to the target knowledge reasoning function and a first knowledge reasoning result set corresponding to the target knowledge reasoning function; determine a third knowledge reasoning result set based on the data set corresponding to the non-target knowledge reasoning function through each of the intervened network modules; wherein the data set corresponding to the non-target knowledge reasoning function includes a question-answer pair data set corresponding to the non-target knowledge reasoning function and a first knowledge reasoning result set corresponding to the non-target knowledge reasoning function; determine the similarity between the elements with the same position index in the second knowledge reasoning result set and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtain multiple first similarity results; determine the similarity between the elements with the same position index in the third knowledge reasoning result set and the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function, and obtain multiple second similarity results; determine the target network module corresponding to the target knowledge reasoning function in the first large language model based on the multiple first similarity results and the multiple second similarity results.
[0127] Optionally, the target network module determination module is also used to: perform a mean operation on the multiple first similarity results to obtain a first mean result; perform a mean operation on the multiple second similarity results to obtain a second mean result; determine a first mean difference result based on the first mean result and the second mean result; if the first mean difference result is greater than a first set threshold, then use the corresponding intervened network module as the target network module to obtain at least one target network module.
[0128] Optionally, the trimming module is specifically used to: sort the at least one target network module according to the first mean difference result corresponding to each of the at least one target network module to obtain a first sorting result; trim the first large language model based on the current target network module corresponding to the first sorting result to obtain a candidate trimmed model; and determine whether the candidate trimmed model is the trimmed first large language model based on the data set corresponding to the target knowledge reasoning function.
[0129] Optionally, the cropping module is also used to: determine a cropped second knowledge reasoning result set based on the data set corresponding to the target knowledge reasoning function through a candidate cropping model; determine a cropped third knowledge reasoning result set based on the data set corresponding to the non-target knowledge reasoning function through a candidate cropping model; determine the similarity between the elements with the same position index in the cropped second knowledge reasoning result set and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtain multiple cropped first similarity results; determine the similarity between the elements with the same position index in the cropped third knowledge reasoning result set and the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function, and obtain multiple cropped second similarity results; and determine whether the candidate cropping model is the cropped first large language model based on the multiple cropped first similarity results and the multiple cropped second similarity results.
[0130] Optionally, the trimming module is further used to: perform an average operation on the multiple trimmed first similarity results to obtain a trimmed first average result; perform an average operation on the multiple trimmed second similarity results to obtain a trimmed second average result; if the trimmed first average result is less than the second set threshold, and the trimmed second average result is greater than the third set threshold, then determine the candidate trimmed model as the trimmed first large language model; if the trimmed first average result is greater than or equal to the second set threshold, or the trimmed second average result is less than or equal to the third set threshold, then cancel the trimming operation of the current target network module on the first large language model, and trim the first large language model based on the next target network module corresponding to the first sorting result to obtain a new candidate trimmed model; use the new candidate trimmed model as the candidate trimmed model, and return to the step of determining whether the candidate trimmed model is the trimmed first large language model based on the data set corresponding to the target knowledge reasoning function, and re-execute.
[0131] Optionally, the migration module is specifically used to: sort the at least one target network module according to the first mean difference result corresponding to each of the at least one target network module to obtain a second sorting result; migrate the current target network module corresponding to the second sorting result to the second large language model to obtain a candidate migration model; and determine whether the candidate migration model is the second large language model after migration based on the data set corresponding to the target knowledge reasoning function.
[0132] Optionally, the migration module is also used to: determine the embedding position of the current target network module in the second largest language model; modify the current target network module according to the data format information corresponding to the embedding position to obtain the modified current target network module; and embed the modified current target network module into the embedding position.
[0133] Optionally, the migration module is also used to: determine the target operator of the second largest language model based on the current target network module; adjust the current target network module according to the interface information of the target operator to obtain the adjusted current target network module; establish a mapping relationship between the parameters of the target operator and the parameters of the adjusted current target network module; migrate the adjusted current target network module to the target operator to obtain a composite operator; load the adjusted parameters of the current target network module into the composite operator according to the mapping relationship; adjust the interface format of the composite operator to be compatible with the second largest language model.
[0134] Optionally, the migration module is also used to: determine the target operator of the second largest language model based on the current target network module; spatially align the parameters of the target operator and the parameters of the current target network module to obtain the aligned target operator parameters and the aligned current target network module parameters; determine the weight matrix of the target operator; determine the correction direction and correction strength of the aligned current target network module parameters on the weight matrix; determine a parameter merging operation function based on the weight matrix, the correction direction and the correction strength; execute the parameter merging operation function to merge the aligned current target network module parameters into the aligned target operator parameters.
[0135] Optionally, the migration module is also used to: determine the second knowledge reasoning result set after migration based on the data set corresponding to the target knowledge reasoning function through the candidate migration model; determine the similarity between the elements with the same position index in the second knowledge reasoning result set after migration and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtain multiple first similarity results after migration; perform an average operation on the multiple first similarity results after migration to obtain the first average result after migration; if the first average result after migration is greater than a fourth set threshold, determine the candidate migration model as the second largest language model after migration; if the first average result after migration is less than or equal to the fourth set threshold, cancel the embedding operation of the current target network module on the second largest language model, and migrate the next target network module corresponding to the second sorting result to the second largest language model to obtain a new candidate migration model; use the new candidate migration model as the candidate migration model, and return to the step of judging whether the candidate migration model is the second largest language model after migration based on the data set corresponding to the target knowledge reasoning function, and re-execute.
[0136] The target large language model determination device provided in the embodiment of the present application can execute the target large language model determination method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0137] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0138] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the random access memory (RAM) 13. The processor 11, the read-only memory (ROM) 12, and the random access memory (RAM) 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0139] Various components in the electronic device 10 are connected to an input / output (I / O) interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0140] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors for running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as determining the target large language model.
[0141] In some embodiments, the determination of the method target large language model can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via a read-only memory (ROM) 12 and / or a communication unit 19. When the computer program is loaded into the random access memory (RAM) 13 and executed by the processor 11, one or more steps of determining the method target large language model described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the determination of the method target large language model in any other appropriate manner (for example, by means of firmware).
[0142] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0143] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable target large language model determination device, so that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0144] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0146] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0147] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0148] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the method for determining a target large language model as provided in any embodiment of the present application.
[0149] The computer program product, during implementation, may be written in one or more programming languages or a combination thereof, for performing the operations of the present application, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0150] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the appended claims.
Claims
1. A method for determining a target large language model, characterized in that: include: In a target knowledge learning scenario, determining a target network module corresponding to the target knowledge reasoning function in the first language model based on a data set corresponding to the target knowledge reasoning function; If the target knowledge learning scenario is a knowledge tailoring scenario, tailoring the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function to obtain the tailored first large language model; training the tailored first large language model based on the first training set to obtain a first target large language model; If the target knowledge learning scenario is a knowledge transfer scenario, the target network module is migrated to the second largest language model based on the data set corresponding to the target knowledge reasoning function to obtain the migrated second largest language model; based on the second training set, the migrated second largest language model is trained to obtain a second target large language model.
2. The method according to claim 1, characterized in that In the target knowledge learning scenario, determining a target network module corresponding to the target knowledge reasoning function in the first language model based on a data set corresponding to the target knowledge reasoning function includes: Performing knowledge reasoning on the dataset based on the question and answer using the first language model to obtain a first knowledge reasoning result set; wherein the first language model includes multiple network modules; The question-answer pair dataset and the first knowledge reasoning result set are correspondingly combined into a question-answer pair result dataset; wherein the question-answer pair result dataset includes a dataset corresponding to the target knowledge reasoning function and a dataset corresponding to the non-target knowledge reasoning function; executing an intervention strategy on each of the plurality of network modules to obtain a plurality of intervened network modules; The target network module corresponding to the target knowledge reasoning function in the first large language model is determined through the multiple intervened network modules based on the data set corresponding to the target knowledge reasoning function and the data set not corresponding to the target knowledge reasoning function.
3. The method according to claim 2, characterized in that Determining, by using the multiple intervened network modules, a target network module corresponding to the target knowledge reasoning function in the first language model based on a data set corresponding to the target knowledge reasoning function and a data set not corresponding to the target knowledge reasoning function, including: Determining, through each of the intervened network modules, a second knowledge reasoning result set based on a data set corresponding to the target knowledge reasoning function; wherein the data set corresponding to the target knowledge reasoning function includes a question-answer pair data set corresponding to the target knowledge reasoning function and a first knowledge reasoning result set corresponding to the target knowledge reasoning function; Determining, by each of the intervened network modules, a third knowledge reasoning result set based on the data set corresponding to the non-target knowledge reasoning function; wherein the data set corresponding to the non-target knowledge reasoning function includes the question-answer pair data set corresponding to the non-target knowledge reasoning function and the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function; Determine the similarity between the second knowledge reasoning result set and the elements with the same position index in the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtain a plurality of first similarity results; Determine the similarity between the third knowledge reasoning result set and the elements with the same position index in the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function, and obtain a plurality of second similarity results; A target network module corresponding to the target knowledge reasoning function in the first large language model is determined based on the multiple first similarity results and the multiple second similarity results.
4. The method according to claim 3, characterized in that Determining a target network module corresponding to the target knowledge reasoning function in the first large language model based on the multiple first similarity results and the multiple second similarity results includes: Performing an average operation on the plurality of first similarity results to obtain a first average result; Performing an average operation on the plurality of second similarity results to obtain a second average result; determining a first mean difference result based on the first mean result and the second mean result; If the first mean difference result is greater than a first set threshold, the corresponding network module after intervention is used as a target network module to obtain at least one target network module.
5. The method according to claim 4, characterized in that If the target knowledge learning scenario is a knowledge tailoring scenario, tailoring the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function to obtain the tailored first large language model, including: Sorting the at least one target network module according to the first mean difference result corresponding to each target network module in the at least one target network module to obtain a first sorting result; trimming the first large language model based on the current target network module corresponding to the first sorting result to obtain a candidate trimmed model; Based on the data set corresponding to the target knowledge reasoning function, it is determined whether the candidate trimmed model is the trimmed first large language model.
6. The method according to claim 5, characterized in that Determining whether the candidate tailored model is the tailored first language model based on the data set corresponding to the target knowledge reasoning function includes: Determine a tailored second knowledge reasoning result set based on the data set corresponding to the target knowledge reasoning function through the candidate tailoring model; Determining a tailored third knowledge reasoning result set based on a data set corresponding to the non-target knowledge reasoning function through a candidate tailoring model; Determine the similarity between the elements with the same position index in the clipped second knowledge reasoning result set and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtain a plurality of clipped first similarity results; Determine the similarity between the elements with the same position index in the clipped third knowledge reasoning result set and the first knowledge reasoning result set corresponding to the non-target knowledge reasoning function, and obtain a plurality of clipped second similarity results; It is determined whether the candidate cropped model is the cropped first large language model based on the multiple cropped first similarity results and the multiple cropped second similarity results.
7. The method according to claim 6, characterized in that Determining whether the candidate cropped model is the cropped first large language model based on the multiple cropped first similarity results and the multiple cropped second similarity results includes: performing a mean operation on the plurality of clipped first similarity results to obtain a clipped first mean result; performing a mean operation on the plurality of cropped second similarity results to obtain a cropped second mean result; If the first mean value after trimming is less than the second set threshold, and the second mean value after trimming is greater than the third set threshold, determining the candidate trimmed model as the first large language model after trimming; If the first mean result after trimming is greater than or equal to the second set threshold, or the second mean result after trimming is less than or equal to the third set threshold, the trimming operation of the current target network module on the first large language model is canceled, and the first large language model is trimmed based on the next target network module corresponding to the first sorting result to obtain a new candidate trimmed model; the new candidate trimmed model is used as the candidate trimmed model, and the step of determining whether the candidate trimmed model is the trimmed first large language model based on the data set corresponding to the target knowledge reasoning function is returned and re-executed.
8. The method according to claim 4, characterized in that If the target knowledge learning scenario is a knowledge transfer scenario, migrating the target network module to the second largest language model based on the data set corresponding to the target knowledge reasoning function to obtain the migrated second largest language model includes: sorting the at least one target network module according to the first mean difference result corresponding to each target network module in the at least one target network module to obtain a second sorting result; Migrating the current target network module corresponding to the second sorting result to the second largest language model to obtain a candidate migration model; Based on the data set corresponding to the target knowledge reasoning function, it is determined whether the candidate migration model is the second largest language model after the migration.
9. The method according to claim 8, characterized in that Migrating the current target network module corresponding to the second sorting result to the second largest language model to obtain a candidate migration model includes: Determining an embedding position of the current target network module in the second largest language model; Modify the current target network module according to the data format information corresponding to the embedding position to obtain the modified current target network module; The modified current target network module is embedded into the embedding position.
10. The method according to claim 8, characterized in that Migrating the current target network module corresponding to the second sorting result to the second largest language model to obtain a candidate migration model includes: Determining a target operator of the second largest language model based on the current target network module; Adjusting the current target network module according to the interface information of the target operator to obtain the adjusted current target network module; Establishing a mapping relationship between the parameters of the target operator and the adjusted parameters of the current target network module; Migrating the adjusted current target network module to the target operator to obtain a composite operator; According to the mapping relationship, the adjusted parameters of the current target network module are loaded into the composite operator; The interface format of the composite operator is adjusted to be compatible with the second largest language model.
11. The method according to claim 8, characterized in that Migrating the current target network module corresponding to the second sorting result to the second largest language model to obtain a candidate migration model includes: Determining a target operator of the second largest language model based on the current target network module; Spatially aligning the parameters of the target operator and the parameters of the current target network module to obtain aligned target operator parameters and aligned current target network module parameters; Determining a weight matrix of the target operator; Determining the direction and strength of correction of the weight matrix by the aligned current target network module parameters; Determine a parameter merging operation function based on the weight matrix, the correction direction, and the correction strength; The parameter merging operation function is executed to merge the aligned current target network module parameters into the aligned target operator parameters.
12. The method according to claim 8, characterized in that Determining whether the candidate transfer model is the second largest language model after the transfer based on the data set corresponding to the target knowledge reasoning function includes: Determine a second knowledge reasoning result set after migration based on the data set corresponding to the target knowledge reasoning function by using the candidate migration model; Determine the similarity between the elements with the same position index in the second knowledge reasoning result set after migration and the first knowledge reasoning result set corresponding to the target knowledge reasoning function, and obtain a plurality of first similarity results after migration; Performing an average operation on the plurality of migrated first similarity results to obtain a migrated first average result; If the first mean value after the migration is greater than a fourth set threshold, determining the candidate migration model as the second largest language model after the migration; If the first mean result after migration is less than or equal to the fourth set threshold, the embedding operation of the current target network module on the second largest language model is canceled, and the next target network module corresponding to the second sorting result is migrated to the second largest language model to obtain a new candidate migration model; the new candidate migration model is used as the candidate migration model, and the step of determining whether the candidate migration model is the second largest language model after migration based on the data set corresponding to the target knowledge reasoning function is returned and re-executed.
13. A device for determining a target large language model, characterized in that: include: a target network module determination module, configured to determine, in a target knowledge learning scenario, a target network module corresponding to the target knowledge reasoning function in the first language model based on a data set corresponding to the target knowledge reasoning function; a tailoring module, configured to, if the target knowledge learning scenario is a knowledge tailoring scenario, tailor the target network module in the first large language model based on the data set corresponding to the target knowledge reasoning function to obtain a tailored first large language model; A first training module is configured to train the trimmed first large language model based on a first training set to obtain a first target large language model; A migration module, configured to migrate the target network module to a second language model based on a data set corresponding to the target knowledge reasoning function, if the target knowledge learning scenario is a knowledge migration scenario, to obtain the migrated second language model; The second training module is used to train the migrated second large language model based on a second training set to obtain a second target large language model.
14. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining a target large language model as described in any one of claims 1 to 12.
15. A storage medium comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to perform the method for determining a target large language model according to any one of claims 1 to 12.
16. A computer program product comprising a computer program, characterized in that When executed by a processor, the computer program implements the method for determining a target large language model according to any one of claims 1 to 12.
Citation Information
Patent Citations
Large-scale pre-training language model compression method based on hardware perception
CN116822593A
Multi-scale cross-domain index transaction attribution method and system based on knowledge enhancement
CN119558687A
Analyzing an inference of a machine learning predictor
US20250094811A1