Large language model integration method supporting semantic correction
By using intelligent fusion language model technology, which dynamically adjusts weights and parameter fusion, the problem of insufficient generalization ability in large language model ensembles is solved, and efficient semantic correction of diverse text data is achieved, improving the adaptability and accuracy of the model.
Patent Information
- Application Number
- CN202410781522.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2025-11-14
AI Technical Summary
Existing large language model ensemble methods that support semantic correction suffer from insufficient generalization ability and inflexible correction of new types of errors when dealing with highly diverse text data due to fixed weight allocation and static parameter fusion.
By employing intelligent fusion language model technology, dynamic weight allocation and model parameter fusion are combined with context-aware technology and hybrid regularization technology to optimize model training and parameter adjustment, thereby achieving accurate correction for different text types.
It improves the model's adaptability and accuracy in diverse text environments, enhances its ability to correct complex semantic errors, maximizes the advantages of each source model, and enhances the overall performance and flexibility of the model.
Smart Images

Figure CN120952144A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of language model integration technology, and more specifically, to a method for integrating large language models that supports semantic correction. Background Technology
[0002] A large language model integration method supporting semantic correction aims to achieve efficient context-aware semantic understanding and correction through intelligent fusion of language model technology, and to achieve accurate correction of various text types and semantic errors by utilizing dynamic weight allocation and model parameter fusion.
[0003] Existing large language model ensemble methods that support semantic correction suffer from insufficient generalization ability and inflexible correction of new types of errors when processing highly diverse text data due to the fixed weight allocation and static parameter fusion of large language model ensemble. Therefore, this paper proposes a large language model ensemble method that supports semantic correction. Summary of the Invention
[0004] The purpose of this invention is to provide a large language model ensemble method that supports semantic correction, in order to solve the problems mentioned in the background art, which are insufficient generalization ability and insufficient flexibility in correcting new types of errors due to the fixed weight allocation and static parameter fusion of large language model ensemble.
[0005] To achieve the above objectives, the present invention aims to provide a method for integrating large language models that supports semantic correction, comprising the following steps:
[0006] S1. Select several large language models and organize and preprocess the text data for training and testing;
[0007] S2. Train each source model and optimize the model using intelligent fusion language model technology;
[0008] S3. Use intelligent fusion language model technology to fuse the parameters of each source model, dynamically select model parameters for fusion based on the context of the input text data, and adjust the parameters of the fused model.
[0009] S4. Evaluate the semantic correction capability of the fused model on different types of text data, and adjust the fusion strategy and regularization parameters based on the evaluation results.
[0010] S5. Integrate the trained model into the target system and continuously collect user feedback and performance data.
[0011] As a further improvement to this technical solution, in S1, a large language model is selected for model integration, a model pre-trained on a large-scale and diverse dataset is selected, and the performance of the model on similar tasks is considered.
[0012] The text data mentioned above comes from public datasets, professional datasets, and text data collected independently from websites, forums, and social media through web crawlers.
[0013] The aforementioned preprocessing of text data is used to improve data quality, enhance data consistency and standardization. The specific operations involved include: removing noise information from the text, performing word segmentation according to the requirements of the model, and standardizing the text format.
[0014] As a further improvement to this technical solution, in S2, the intelligent fusion language model technology is based on the fusion of FuseLLM fusion technology, context-aware technology and hybrid regularization technology. It is used to improve the model performance in semantic correction tasks, integrate the characteristics of multiple source models, and enable the final model to understand, correct and generate semantically correct text.
[0015] The specific operational steps involved in training each source model and optimizing the model using intelligent fusion language model technology are as follows:
[0016] S2.1 Context-aware training of the selected model, initialization of the model, and generation of context embedding vectors for each input text data using the model’s built-in self-attention mechanism;
[0017] S2.2 For the initial output of the model, select the weighted cross-entropy loss function to evaluate the difference between the model output and the true label, and adjust and add task adaptation layers.
[0018] S2.3. Use L2 regularization and Dropout, setting the L2 coefficient to 0.01 and the Dropout ratio to 0.1. Monitor the training loss and validation metrics, and adjust the model training hyperparameters accordingly.
[0019] S2.4. Evaluate the model on an independent validation set, and adjust the regularization parameters and training strategy as needed based on the validation results.
[0020] As a further improvement to this technical solution, the weighted cross-entropy loss function in S2.2 can balance the weights of errors between different categories, allowing for a greater focus on one aspect of error types. The mathematical formula involved is as follows:
[0021]
[0022] Where L is the value of the loss function; C is the total number of categories; w i The weight of category i; y i p represents the actual probability distribution of the target label. i Predict the probability that the data belongs to category i for the model;
[0023] The adjustment and addition of the task adaptation layer are used to better adapt to the semantic correction task, and the mathematical model formulas involved are as follows:
[0024] output = softmax(Wx + b);
[0025] Where output is the output layer of the model; softmax is the activation function used to convert the original predicted values of the output layer into a probability distribution; W is the weight matrix; x is the feature vector of the previous layer; and b is the bias vector.
[0026] As a further improvement to this technical solution, the L2 regularization used in S2.3 is achieved by adding a term proportional to the square of the weights to the model's loss function. This is used to control model complexity and smooth the model's learning parameters, reducing model complexity and overfitting. The mathematical formulas involved are as follows:
[0027]
[0028] Where L1 is the total loss value, including cross-entropy loss and regularization term; λ is the L2 regularization coefficient, controlling the influence of the regularization term, λ = 0.01; w j These are the weight parameters of the model;
[0029] The use of Dropout is described below, as well as its application in randomizing the training process and enhancing the model's ability to handle unseen data. The mathematical model formulas involved in using Dropout are as follows:
[0030] h' = h⊙m;
[0031] Where h is the original neuron activation output; m is a randomly generated binary mask, where each element independently takes the value 0 or 1; ⊙ is element-wise multiplication, i.e., the Hadamard product.
[0032] As a further improvement to this technical solution, step S3 uses intelligent fusion language model technology to fuse the parameters of each source model, dynamically selecting model parameters for fusion based on the context of the input text. The specific steps involved are as follows:
[0033] S3.1 Extract parameters from each source model, including weight matrix and bias term;
[0034] S3.2 For the input text data, use context-aware technology to calculate and adjust the fusion weights of each model;
[0035] S3.3. Use the obtained fusion weights to fuse the parameters of different models by weighted averaging;
[0036] S3.4. Use the Adam optimizer to further train and fine-tune the fused model.
[0037] As a further improvement to this technical solution, the fusion weight in S3.3 is used to determine the contribution of each source model to the final fused model, and the fusion weight w is calculated. i The mathematical model formulas involved are as follows:
[0038]
[0039] Among them, w i For the source model M i The fusion weights are: C is the context representation of the current input text; score is a function that calculates the fusion weights of the source model M given the context C. i The correlation; N is the total number of source models;
[0040] The mathematical model formula involved in fusing parameters from different models through weighted averaging is as follows:
[0041]
[0042] Where P represents the fused model parameters; P i For the source model M i The parameters; w i For weight fusion.
[0043] As a further improvement to this technical solution, step S3.4 uses the Adam optimizer, which combines momentum and adaptive learning rate techniques, to train the neural network. The mathematical formulas involved in further training and fine-tuning the fused model using the Adam optimizer are as follows:
[0044]
[0045] Where, θ t+1 The updated parameter value; θ t The updated parameter values; η is the learning rate; This is the first-order momentum estimate after bias correction, i.e., the exponential moving average of the gradient; This is the second-order momentum estimate after bias correction, i.e., the exponential moving average of the squared gradient; ∈ is the smoothing term to prevent the denominator from being zero.
[0046] As a further improvement to this technical solution, in step S4, the semantic correction capability of the fused model is evaluated for different types of text data, and the fusion strategy and regularization parameters are adjusted based on the evaluation results. The specific steps involved are as follows:
[0047] S4.1 Construct different types of text datasets, including formal text, informal dialogues, and technical documents, and select accuracy, recall, F1 score, and scoring metrics for accuracy, recall, and semantic similarity.
[0048] S4.2 Run the fused model, collect the model's output and compare it with the correct semantic expression to identify the model's strengths and weaknesses;
[0049] S4.3 Adjust the model fusion strategy based on the evaluation results. If one of the source models performs well, increase the weight of that model and adjust the L2 regularization coefficient and Dropout rate.
[0050] S4.4 After adjustment, perform model evaluation again to verify the effect of the adjustment.
[0051] As a further improvement to this technical solution, the specific steps involved in S5, namely integrating the trained model into the target system and continuously collecting user feedback and performance data, are as follows:
[0052] S5.1 Integrate the trained model into the target system and conduct preliminary model testing, including load testing;
[0053] S5.2 Implement a monitoring system to track the model's performance and health status, including monitoring model response time, error rate, and system load metrics;
[0054] S5.3 Collect user feedback and usage data, and analyze the collected usage data and user feedback to optimize the model.
[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0056] 1. In this large language model integration method that supports semantic correction, by using context-aware technology and hybrid regularization technology in intelligent fusion language model technology to train each source model and optimize the model, the adaptability and accuracy of the model in various text environments can be improved, and the effective correction of complex and diverse semantic errors can be achieved.
[0057] 2. In this method of integrating large language models that supports semantic correction, the parameters of each source model are integrated by using FuseLLM fusion technology in intelligent fusion language model technology, which can maximize the use of the advantages of each source model and enhance the overall performance and flexibility of the model. Attached Figure Description
[0058] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation
[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] Example:
[0061] Please see Figure 1 As shown, this embodiment provides a method for integrating large language models that supports semantic correction, including the following steps:
[0062] S1. Select several large language models and organize and preprocess the text data for training and testing;
[0063] In S1, a large language model is selected for model ensemble, a model pre-trained on a large-scale and diverse dataset is selected, and the performance of the model on similar tasks is considered.
[0064] The text data mentioned above comes from public datasets, professional datasets, and text data collected independently from websites, forums, and social media through web crawlers.
[0065] The aforementioned preprocessing of text data is used to improve data quality, enhance data consistency and standardization. The specific operations involved include: removing noise information from the text, performing word segmentation according to the requirements of the model, and standardizing the text format.
[0066] S2. Train each source model and optimize the model using intelligent fusion language model technology;
[0067] In S2, the intelligent fusion language model technology is based on the fusion of FuseLLM fusion technology, context-aware technology and hybrid regularization technology. It is used to improve the model performance in semantic correction tasks, integrate the features of multiple source models, and enable the final model to understand, correct and generate semantically correct text.
[0068] The FuseLLM fusion technique is a method for merging multiple large language models (LLMs) to improve the overall model's performance and robustness by combining the advantages of different models. Its steps include: parameter extraction and standardization, context-driven weight allocation, parameter fusion, and post-processing and optimization.
[0069] The context-aware technology utilizes the model's ability to understand and leverage contextual information from input data to make more accurate predictions. This technology is particularly important in semantic correction scenarios, as correct semantic understanding often depends on the completeness and accuracy of the context, including context embedding, dynamic context analysis, and personalized adaptation.
[0070] The specific operational steps involved in training each source model and optimizing the model using intelligent fusion language model technology are as follows:
[0071] S2.1 Context-aware training of the selected model, initialization of the model, and generation of context embedding vectors for each input text data using the model’s built-in self-attention mechanism;
[0072] S2.2 For the initial output of the model, select the weighted cross-entropy loss function to evaluate the difference between the model output and the true label, and adjust and add task adaptation layers.
[0073] S2.3. Use L2 regularization and Dropout, setting the L2 coefficient to 0.01 and the Dropout ratio to 0.1. Monitor the training loss and validation metrics, and adjust the model training hyperparameters accordingly.
[0074] S2.4. Evaluate the model on an independent validation set, and adjust the regularization parameters and training strategy as needed based on the validation results.
[0075] The weighted cross-entropy loss function in S2.2 can balance the weights of errors between different categories, allowing for a greater focus on one aspect of error types. The mathematical formula involved is as follows:
[0076]
[0077] Where L is the value of the loss function; C is the total number of categories; w i The weight of category i; y i p represents the actual probability distribution of the target label. i Predict the probability that the data belongs to category i for the model;
[0078] The adjustment and addition of the task adaptation layer are used to better adapt to the semantic correction task, and the mathematical model formulas involved are as follows:
[0079] output = softmax(Wx + b);
[0080] Where output is the output layer of the model; softmax is the activation function used to convert the original predicted values of the output layer into a probability distribution; W is the weight matrix; x is the feature vector of the previous layer; and b is the bias vector.
[0081] The L2 regularization used in S2.3 is achieved by adding a term proportional to the square of the weights to the model's loss function. This is used to control model complexity and smooth the model's learning parameters, reducing model complexity and overfitting. The mathematical formulas involved are as follows:
[0082]
[0083] Where L1 is the total loss value, including cross-entropy loss and regularization term; λ is the L2 regularization coefficient, controlling the influence of the regularization term, λ = 0.01; w j These are the weight parameters of the model;
[0084] The use of Dropout is described below, as well as its application in randomizing the training process and enhancing the model's ability to handle unseen data. The mathematical model formulas involved in using Dropout are as follows:
[0085] h′=h⊙m;
[0086] Where h is the original neuron activation output; m is a randomly generated binary mask, where each element independently takes the value 0 or 1; ⊙ is element-wise multiplication, i.e., the Hadamard product;
[0087] The L2 regularization reduces overfitting by penalizing the squared value of the model weights, prompting the model to prefer simpler or smoother functions.
[0088] Dropout involves randomly "dropping" a portion of the neuron outputs during training, forcing the model to learn through the remaining connections, thereby increasing the model's dependence on different subsets of neurons.
[0089] S3. Use intelligent fusion language model technology to fuse the parameters of each source model, dynamically select model parameters for fusion based on the context of the input text data, and adjust the parameters of the fused model.
[0090] In step S3, intelligent fusion language model technology is used to fuse the parameters of each source model. The model parameters are dynamically selected for fusion based on the context of the input text. The specific steps involved are as follows:
[0091] S3.1 Extract parameters from each source model, including weight matrix and bias term;
[0092] S3.2 For the input text data, use context-aware technology to calculate and adjust the fusion weights of each model;
[0093] S3.3. Use the obtained fusion weights to fuse the parameters of different models by weighted averaging;
[0094] S3.4. Use the Adam optimizer to further train and fine-tune the fused model.
[0095] In S3.3, the fusion weight is used to determine the contribution of each source model to the final fused model. The fusion weight w is calculated. i The mathematical model formulas involved are as follows:
[0096]
[0097] Among them, w i For the source model M i The fusion weights are: C is the context representation of the current input text; score is a function that calculates the fusion weights of the source model M given the context C. i The correlation; N is the total number of source models;
[0098] The mathematical model formula involved in fusing parameters from different models through weighted averaging is as follows:
[0099]
[0100] Where P represents the fused model parameters; P i For the source model M i The parameters; W i For weight fusion.
[0101] In step S3.4, the Adam optimizer is used to combine momentum and adaptive learning rate techniques to train the neural network. The mathematical formulas involved in further training and fine-tuning the fused model using the Adam optimizer are as follows:
[0102]
[0103] Where, θ t+1 The updated parameter value; θ t The updated parameter values; η is the learning rate; This is the first-order momentum estimate after bias correction, i.e., the exponential moving average of the gradient; This is the second-order momentum estimate after bias correction, i.e., the exponential moving average of the squared gradient; ∈ is the smoothing term to prevent the denominator from being zero.
[0104] S4. Evaluate the semantic correction capability of the fused model on different types of text data, and adjust the fusion strategy and regularization parameters based on the evaluation results.
[0105] In step S4, the semantic correction capability of the fused model is evaluated for different types of text data, and the fusion strategy and regularization parameters are adjusted based on the evaluation results. The specific steps involved are as follows:
[0106] S4.1 Construct different types of text datasets, including formal text, informal dialogues, and technical documents, and select accuracy, recall, F1 score, and scoring metrics for accuracy, recall, and semantic similarity.
[0107] S4.2 Run the fused model, collect the model's output and compare it with the correct semantic expression to identify the model's strengths and weaknesses;
[0108] S4.3 Adjust the model fusion strategy based on the evaluation results. If one of the source models performs well, increase the weight of that model and adjust the L2 regularization coefficient and Dropout rate.
[0109] S4.4 After adjustment, perform model evaluation again to verify the effect of the adjustment.
[0110] S5. Integrate the trained model into the target system and continuously collect user feedback and performance data.
[0111] In step S5, the specific steps involved in integrating the trained model into the target system and continuously collecting user feedback and performance data are as follows:
[0112] S5.1 Integrate the trained model into the target system and conduct preliminary model testing, including load testing;
[0113] S5.2 Implement a monitoring system to track the model's performance and health status, including monitoring model response time, error rate, and system load metrics;
[0114] S5.3 Collect user feedback and usage data, and analyze the collected usage data and user feedback to optimize the model.
[0115] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for integrating large language models that supports semantic correction, characterized in that: Includes the following steps: S1. Select several large language models and organize and preprocess the text data for training and testing; S2. Train each source model and optimize the model using intelligent fusion language model technology; S3. Use intelligent fusion language model technology to fuse the parameters of each source model, dynamically select model parameters for fusion based on the context of the input text data, and adjust the parameters of the fused model. S4. Evaluate the semantic correction capability of the fused model on different types of text data, and adjust the fusion strategy and regularization parameters based on the evaluation results. S5. Integrate the trained model into the target system and continuously collect user feedback and performance data.
2. The method for integrating large language models supporting semantic correction according to claim 1, characterized in that: In S1, a large language model is selected for model ensemble, a model pre-trained on a large-scale and diverse dataset is selected, and the performance of the model on similar tasks is considered. The text data mentioned above comes from public datasets, professional datasets, and text data collected independently from websites, forums, and social media through web crawlers. The aforementioned preprocessing of text data is used to improve data quality, enhance data consistency and standardization. The specific operations involved include: removing noise information from the text, performing word segmentation according to the requirements of the model, and standardizing the text format.
3. The method for integrating large language models supporting semantic correction according to claim 1, characterized in that: In S2, the intelligent fusion language model technology is based on the fusion of FuseLLM fusion technology, context-aware technology and hybrid regularization technology. It is used to improve the model performance in semantic correction tasks, integrate the features of multiple source models, and enable the final model to understand, correct and generate semantically correct text. The specific operational steps involved in training each source model and optimizing the model using intelligent fusion language model technology are as follows: S2.1 Context-aware training of the selected model, initialization of the model, and generation of context embedding vectors for each input text data using the model’s built-in self-attention mechanism; S2.2 For the initial output of the model, select the weighted cross-entropy loss function to evaluate the difference between the model output and the true label, and adjust and add task adaptation layers. S2.
3. Use L2 regularization and Dropout, setting the L2 coefficient to 0.01 and the Dropout ratio to 0.
1. Monitor the training loss and validation metrics, and adjust the model training hyperparameters accordingly. S2.
4. Evaluate the model on an independent validation set, and adjust the regularization parameters and training strategy as needed based on the validation results.
4. The method for integrating large language models supporting semantic correction according to claim 3, characterized in that: The weighted cross-entropy loss function in S2.2 can balance the weights of errors between different categories, allowing for a greater focus on one aspect of error types. The mathematical formula involved is as follows: in; L is the value of the loss function; C is the total number of categories; w i The weight of category i; y i p represents the actual probability distribution of the target label. i Predict the probability that the data belongs to category i for the model; The adjustment and addition of the task adaptation layer are used to better adapt to the semantic correction task, and the mathematical model formulas involved are as follows: output = softmax(Wx + b); Where output is the output layer of the model; softmax is the activation function used to convert the original predicted values of the output layer into a probability distribution; W is the weight matrix; x is the feature vector of the previous layer; and b is the bias vector.
5. The method for integrating large language models supporting semantic correction according to claim 3, characterized in that: The L2 regularization used in S2.3 is achieved by adding a term proportional to the square of the weights to the model's loss function. This is used to control model complexity and smooth the model's learning parameters, reducing model complexity and overfitting. The mathematical formulas involved are as follows: Where L1 is the total loss value, including cross-entropy loss and regularization term; λ is the L2 regularization coefficient, controlling the influence of the regularization term, λ = 0.01; w j These are the weight parameters of the model; The use of Dropout is described below, as well as its application in randomizing the training process and enhancing the model's ability to handle unseen data. The mathematical model formulas involved in using Dropout are as follows: h′=h⊙m; Where h is the original neuron activation output; m is a randomly generated binary mask, where each element independently takes the value 0 or 1; ⊙ is element-wise multiplication, i.e., the Hadamard product.
6. The method for integrating large language models supporting semantic correction according to claim 1, characterized in that: In step S3, intelligent fusion language model technology is used to fuse the parameters of each source model. The model parameters are dynamically selected for fusion based on the context of the input text. The specific steps involved are as follows: S3.1 Extract parameters from each source model, including weight matrix and bias term; S3.2 For the input text data, use context-aware technology to calculate and adjust the fusion weights of each model; S3.
3. Use the obtained fusion weights to fuse the parameters of different models by weighted averaging; S3.
4. Use the Adam optimizer to further train and fine-tune the fused model.
7. The method for integrating large language models supporting semantic correction according to claim 1, characterized in that: In S3.3, the fusion weight is used to determine the contribution of each source model to the final fused model. The fusion weight w is calculated. i The mathematical model formulas involved are as follows: Among them, w i For the source model M i The fusion weights are: C is the context representation of the current input text; score is a function that calculates the fusion weights of the source model M given the context C. i The correlation; N is the total number of source models; The mathematical model formula involved in fusing parameters from different models through weighted averaging is as follows: Where P represents the fused model parameters; P i For the source model M i The parameters; w i For weight fusion.
8. The method for integrating large language models supporting semantic correction according to claim 1, characterized in that: In step S3.4, the Adam optimizer is used to combine momentum and adaptive learning rate techniques to train the neural network. The mathematical formulas involved in further training and fine-tuning the fused model using the Adam optimizer are as follows: Where, θ t+1 The updated parameter value; θ t The updated parameter values; η is the learning rate; This is the first-order momentum estimate after bias correction, i.e., the exponential moving average of the gradient; This is the second-order momentum estimate after bias correction, i.e., the exponential moving average of the squared gradient; ∈ is the smoothing term to prevent the denominator from being zero.
9. The method for integrating large language models supporting semantic correction according to claim 1, characterized in that: In step S4, the semantic correction capability of the fused model is evaluated for different types of text data, and the fusion strategy and regularization parameters are adjusted based on the evaluation results. The specific steps involved are as follows: S4.1 Construct different types of text datasets, including formal text, informal dialogues, and technical documents, and select accuracy, recall, F1 score, and scoring metrics for accuracy, recall, and semantic similarity. S4.2 Run the fused model, collect the model's output and compare it with the correct semantic expression to identify the model's strengths and weaknesses; S4.3 Adjust the model fusion strategy based on the evaluation results. If one of the source models performs well, increase the weight of that model and adjust the L2 regularization coefficient and Dropout rate. S4.4 After adjustment, perform model evaluation again to verify the effect of the adjustment.
10. The method for integrating large language models supporting semantic correction according to claim 1, characterized in that: In step S5, the specific steps involved in integrating the trained model into the target system and continuously collecting user feedback and performance data are as follows: S5.1 Integrate the trained model into the target system and conduct preliminary model testing, including load testing; S5.2 Implement a monitoring system to track the model's performance and health status, including monitoring model response time, error rate, and system load metrics; S5.3 Collect user feedback and usage data, and analyze the collected usage data and user feedback to optimize the model.
Citation Information
Cited By
Fine-grained text uncertainty monitoring method and system based on semantic compression regularization
CN121722918A
Explanatability fusion and recovery method after large language model training based on interpretability
CN121936572A