A semantic recognition method and system of a static-dynamic hybrid large language model, a terminal and a storage medium
Patent Information
- Application Number
- CN202610934177.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-18
AI Technical Summary
[0005]本发明的主要目的在于提供一种静动混合的大语言模型的语义识别方法、系统、终端及存储介质,旨在解决现有技术对各阶段的SFT训练数据强行进行同比例划分造成大量数据浪费,且未对跨数据源难度进行统一校准,导致每个阶段SFT训练数据分配的分配结果不够准确,造成大语言模型的语义识别结果不准确的问题
[0016] In this invention, a supervised fine-tuning data pool containing multiple data sources and the sampling quantity of each data source are obtained. The supervised fine-tuning data pool is evaluated to obtain a data evaluation result. Cross-source calibration is performed on the data evaluation result to obtain a difficulty calibration result. Spatial mapping is then performed on the difficulty calibration result and all the sampling quantities to obtain a global total difficulty. A stage allocation ratio is searched based on the global total difficulty. The global total difficulty is statically solved based on the stage allocation ratio to obtain a data solution result. A stage dataset is constructed. The stage dataset is dynamically solved based on the data solution result to obtain a target stage allocation ratio. Data is then allocated to the supervised fine-tuning data pool based on the target stage allocation ratio to obtain a data allocation result. This invention improves data utilization by using an inherent distribution as the partitioning ratio and also performs unified calibration of cross-data source difficulty, improving the accuracy of SFT training data allocation results, thereby improving semantic recognition accuracy.
Smart Images

Figure CN122779082A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semantic recognition technology, and in particular to a semantic recognition method, system, terminal, and computer-readable storage medium for a large language model that combines static and dynamic elements. Background Technology
[0002] As the scale and capabilities of large language models continue to expand, the data organization method in the SFT (Supervised Fine-Tuning) stage directly determines the final performance, generalization ability and convergence stability of the model. In recent years, course learning has been introduced into the SFT and post-training stages of LLM (Large Language Model) (after the model pre-training has completed basic learning, an additional round of targeted fine-tuning training is performed).
[0003] Currently, existing multi-stage SFT training data allocation methods based on course learning include traditional soft course learning, multi-domain weight optimization, and hierarchical dynamic sampling. However, existing methods forcibly divide the SFT training data of each stage proportionally, resulting in a large amount of data waste. Furthermore, they do not uniformly calibrate the difficulty across data sources, leading to inaccurate allocation results for SFT training data at each stage and inaccurate semantic recognition results of large language models.
[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0005] The main objective of this invention is to provide a semantic recognition method, system, terminal, and storage medium for a large language model that combines static and dynamic elements. This invention aims to solve the problems of existing technologies that forcibly divide the SFT training data at each stage in the same proportion, resulting in a large amount of data waste, and that do not uniformly calibrate the difficulty of cross-data sources, leading to inaccurate allocation results of SFT training data at each stage and thus inaccurate semantic recognition results of the large language model.
[0006] To achieve the above objectives, the present invention provides a semantic recognition method for a large language model that combines static and dynamic elements. The semantic recognition method for the large language model that combines static and dynamic elements includes the following steps: Obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and evaluate the supervised fine-tuning data pool to obtain data evaluation results; Cross-source calibration is performed on the data evaluation results to obtain difficulty calibration results. Spatial mapping is performed on the difficulty calibration results and all the sampling quantities to obtain the total global difficulty. The stage allocation is searched according to the total global difficulty, and the total global difficulty is statically solved according to the stage allocation to obtain the data solution results. Construct a stage dataset, dynamically solve the stage dataset based on the data solution results to obtain the target stage allocation, allocate data to the supervised fine-tuning data pool based on the target stage allocation to obtain the data allocation results, and train the model based on the data allocation results to obtain the target large language model; Obtain the semantic information to be identified, input the semantic information to be identified into the target large language model, and obtain the semantic recognition result.
[0007] Optionally, the semantic recognition method for a large language model with static and dynamic hybrid processing, wherein obtaining a supervised fine-tuning data pool containing multiple data sources and the number of samples from each data source, and evaluating the supervised fine-tuning data pool to obtain data evaluation results, specifically includes: Obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and obtain a set of difficulty estimators, wherein the set of difficulty estimators includes multiple types of estimators; The supervised fine-tuning data pool is evaluated for difficulty based on all the evaluators to obtain multiple difficulty evaluation results. All the difficulty evaluation results are then integrated to obtain the data evaluation result.
[0008] Optionally, the semantic recognition method for the large language model with static and dynamic hybrid processing, wherein performing cross-source calibration on the data evaluation results to obtain difficulty calibration results, and spatially mapping the difficulty calibration results and all the sample quantities to obtain the total global difficulty, specifically includes: Construct a calibration matrix, and perform cross-source calibration on the original difficulty distribution of each data source in the data evaluation results based on the calibration matrix to obtain the difficulty calibration results; Obtain the difficulty space information of the difficulty calibration result, perform spatial mapping on the difficulty calibration result and all sampling quantities based on the difficulty space information to obtain multiple spatial mapping results, and obtain the total global difficulty based on all the spatial mapping results.
[0009] Optionally, in the semantic recognition method of the large language model with static and dynamic hybridity, the step of performing cross-source calibration on the native difficulty distribution of each data source in the data evaluation results based on the calibration matrix specifically involves: ; The process of obtaining the total global difficulty based on all the spatial mapping results is as follows: ; in, For the first Difficulty distribution after calibration of each data source For the first Difficulty calibration matrix for each data source For the first The difficulty distribution of each data source For the first Each difficulty level represents the total global difficulty after calibration across all data sources. The first one set by humans Total number of samples from each data source For the first The first data source The natural proportion of each difficulty level This represents the total number of data sources.
[0010] Optionally, the semantic recognition method for the large language model with static and dynamic hybrid approach, wherein the step of searching for a stage ratio based on the total global difficulty, and statically solving the total global difficulty based on the stage ratio to obtain the data solution result, specifically includes: A phase allocation ratio is obtained by performing a ratio search based on the total global difficulty, and a phase data volume matrix is obtained by performing a matrix solution on the total global difficulty based on the phase allocation ratio. The stage data volume matrix is constrained and solved to obtain the matrix constraint solution. The stage data volume matrix is then integrated and generalized to obtain the continuous difficulty score. The data solution results are obtained based on the stage data volume matrix and the continuous difficulty score.
[0011] Optionally, the semantic recognition method for the large language model with static and dynamic hybrid approach, wherein the construction phase dataset involves dynamically solving the phase dataset based on the data solution results to obtain a target phase allocation, and allocating data to the supervised fine-tuning data pool based on the target phase allocation to obtain a data allocation result, specifically including: Construct a phase dataset, and periodically solve the phase dataset based on the data solution results to obtain multiple training feedbacks; The stage allocation is updated based on all the training feedback to obtain the target stage allocation, and the stage dataset is determined based on all the training feedback to determine whether it needs to be dynamically solved. If the stage dataset needs to be solved dynamically, the total global difficulty is re-solved according to the target stage ratio to obtain the target solution result; If the stage dataset does not require dynamic solution, then the supervised fine-tuning data pool is allocated according to the target stage ratio to obtain the data allocation result.
[0012] Optionally, in the semantic recognition method of the large language model with static and dynamic hybrid approach, the step of updating the stage allocation based on all the training feedback specifically involves: ; in, For the first The iteration of the ... In the first stage The stage ratio for each difficulty level, For the first The iteration of the ... In the first stage The stage ratio for each difficulty level, To update the step size, For the first In the first stage The gradient of stage proportions for each difficulty level. For the first The validation set loss in the next iteration.
[0013] Optionally, the semantic recognition method for the static-dynamic hybrid large language model includes a semantic recognition system comprising: The data evaluation module is used to obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and to evaluate the supervised fine-tuning data pool to obtain the data evaluation results. The data solving module is used to perform cross-source calibration on the data evaluation results to obtain difficulty calibration results, perform spatial mapping on the difficulty calibration results and all the sampling quantities to obtain the total global difficulty, search the stage ratio according to the total global difficulty, and perform static solution on the total global difficulty according to the stage ratio to obtain the data solving results. The model training module is used to construct a stage dataset, dynamically solve the stage dataset based on the data solution results to obtain the target stage allocation, allocate data to the supervised fine-tuning data pool based on the target stage allocation to obtain the data allocation results, and train the model based on the data allocation results to obtain the target large language model. The semantic recognition module is used to acquire the semantic information to be recognized, input the semantic information to be recognized into the target large language model, and obtain the semantic recognition result.
[0014] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a semantic recognition program for a large language model with static and dynamic hybridity stored in the memory and executable on the processor, wherein when the semantic recognition program for the large language model with static and dynamic hybridity is executed by the processor, it implements the steps of the semantic recognition method for the large language model with static and dynamic hybridity as described above.
[0015] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a semantic recognition program for a large language model with static and dynamic hybrid characteristics, and the semantic recognition program for the large language model with static and dynamic hybrid characteristics, when executed by a processor, implements the steps of the semantic recognition method for the large language model with static and dynamic hybrid characteristics as described above.
[0016] In this invention, a supervised fine-tuning data pool containing multiple data sources and the sampling quantity of each data source are obtained. The supervised fine-tuning data pool is evaluated to obtain a data evaluation result. Cross-source calibration is performed on the data evaluation result to obtain a difficulty calibration result. Spatial mapping is then performed on the difficulty calibration result and all the sampling quantities to obtain a global total difficulty. A stage allocation ratio is searched based on the global total difficulty. The global total difficulty is statically solved based on the stage allocation ratio to obtain a data solution result. A stage dataset is constructed. The stage dataset is dynamically solved based on the data solution result to obtain a target stage allocation ratio. Data is then allocated to the supervised fine-tuning data pool based on the target stage allocation ratio to obtain a data allocation result. This invention improves data utilization by using an inherent distribution as the partitioning ratio and also performs unified calibration of cross-data source difficulty, improving the accuracy of SFT training data allocation results, thereby improving semantic recognition accuracy. Attached Figure Description
[0017] Figure 1 This is a flowchart of a preferred embodiment of the semantic recognition method of the large language model with static and dynamic hybridity of the present invention; Figure 2 This is a schematic diagram of the overall architecture of the semantic recognition method for a large language model that combines static and dynamic elements, as described in this invention. Figure 3 This is a schematic diagram of the overall process of the semantic recognition method of the large language model with static and dynamic hybridity in this invention; Figure 4 This is a flowchart illustrating the visual transition of difficulty between stages in this invention; Figure 5 This is a structural diagram of a preferred embodiment of the semantic recognition system of the large language model with static and dynamic hybridization of the present invention; Figure 6 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0020] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0021] Currently, existing multi-stage SFT training data allocation methods based on course learning include traditional soft course learning, multi-domain weight optimization, and hierarchical dynamic sampling. For traditional soft course learning, training is divided into several stages, with each stage mixing data according to a preset difficulty ratio, for example, P1=0.7:0.2:0.1, P2=0.2:0.6:0.2, P3=0.1:0.2:0.7. The default assumption is that each data source can provide three types of data (D1, D2, and D3) in a uniform ratio, where D1, D2, and D3 refer to the data difficulty level, with three levels: D1 being easy, D2 being medium, and D3 being difficult. For multi-domain weight optimization, the optimal sampling weights for each domain are learned through a surrogate model or sampling law, without distinguishing difficulty levels or handling multi-stage courses. For hierarchical dynamic sampling, global domain weights and local difficulty weights are adjusted online, and gradient information is used to dynamically adjust task sampling weights.
[0022] However, existing methods still have shortcomings: (1) The static method assumes too much: it assumes that each data source has a balanced distribution of D1, D2 and D3, but in actual engineering, cognitive classes have almost no D3, mathematical classes are dominated by D3, and code classes are scarce with D2. Forcing a proportional division will result in a lot of data waste. (2) Four uncertainties exist simultaneously: how much to sample from each data source, how to divide each data source across stages, how much data is needed for each stage, and what the difficulty ratio of each stage is. Existing solutions lack a systematic solution mechanism. (3) Dynamic methods destroy the prior sampling quantity and the scaling law experiment is distorted: Existing dynamic methods change the actual usage of each data source during training, which makes the artificially preset total sampling quantity invalid, resulting in the drift of the curve design point and the inability to reproduce the experimental conclusion when plotting the scaling law curve. (4) Insufficient robustness of evaluation: Existing methods mostly rely on a single heuristic scoring, lack a multi-evaluator integration mechanism, and are sensitive to scoring bias; (5) The proportioning matrix depends on human experience: the existing stage proportioning matrix is usually set based on expert experience and lacks an automated search mechanism.
[0023] Therefore, this invention proposes a semantic recognition method for large language models that combines static and dynamic approaches. This addresses the problem that existing technologies forcibly divide SFT training data at each stage in the same proportion, resulting in a large amount of data waste, and fail to uniformly calibrate the difficulty across data sources, leading to inaccurate allocation results of SFT training data at each stage and thus inaccurate semantic recognition results of large language models.
[0024] The semantic recognition method of the large language model with static and dynamic hybridity described in the preferred embodiment of the present invention, such as... Figure 1 As shown, the semantic recognition method of the static-dynamic hybrid large language model includes the following steps: Step S10: Obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and evaluate the supervised fine-tuning data pool to obtain data evaluation results.
[0025] Specifically, in this embodiment of the invention, a multi-stage data volume analytical solution mechanism based on the inherent difficulty distribution of the data source is used to model the course learning data allocation as a system of linear equations, and the required data volume for each stage is obtained through matrix inversion; through an end-to-end course learning data allocation framework of "main static solution + extension module", the corresponding framework structure is as follows: Figure 2As shown, this framework includes an input layer, a difficulty evaluation layer, a difficulty calibration layer, a matching configuration layer, a solution layer, and a scheduling layer. The input layer is used to input a supervised fine-tuning data pool from multiple data sources (e.g., medical, mathematical, code, literary, and cognitive data) and the number of samples from each data source. The difficulty evaluation layer includes an extension module G (a multi-evaluator integration mechanism for difficulty evaluation) for multi-evaluator integration voting, such as LLM-as-Judge (an evaluation paradigm that automatically evaluates the output quality of another LLM using a Large Language Model (LLM)) evaluator 1, a loss function evaluator 2, and a token. Entropy (character entropy) estimator 3, Length-based estimator 4; the difficulty calibration layer includes extension module C (cross-data source difficulty calibration matrix mechanism) for cross-data source difficulty calibration matrix; the allocation configuration layer includes extension module E (automatic search of stage allocation matrix) for automatic search of stage allocation matrix; the solution layer includes core module A (static linear solution based on inherent distribution), extension module B (extended solution with physical constraints) and extension module F (continuous difficulty level generalization) for statically solving the total global difficulty; the scheduling layer includes extension module D (static-dynamic hybrid scheduling) and extension module H (scaling law experimental repeatability guarantee) for dynamically solving the stage dataset.
[0026] The specific implementation process of the end-to-end course learning data allocation framework is as follows: Figure 3 As shown, Figure 3 In step S1, after obtaining the supervised fine-tuning data pool containing multiple data sources, the sampling quantity for each data source is set; then, as... Figure 3 In step S2, multiple evaluators need to be integrated through the extension module G. Specifically, a set of difficulty evaluators is obtained, which includes multiple types of evaluators, such as LLM-as-Judge (an evaluation paradigm that automatically evaluates the output quality of another LLM using a large language model (LLM)) evaluator 1, loss evaluator 2, token entropy evaluator 3, and length-based evaluator 4. The supervised fine-tuning data pool is then evaluated for difficulty based on all of these evaluators, resulting in multiple difficulty evaluation results. These results are then integrated to obtain the data evaluation result, with the corresponding expression being: ; in, For data evaluation results, For learnable integrators, The output difficulty of the first evaluator. The output difficulty for the second evaluator, For the first The difficulty of the output of each evaluator For the evaluator weights.
[0027] Step S20: Perform cross-source calibration on the data evaluation results to obtain difficulty calibration results. Perform spatial mapping on the difficulty calibration results and all the sampling quantities to obtain the total global difficulty. Search the stage ratio according to the total global difficulty. Perform static solution on the total global difficulty according to the stage ratio to obtain the data solution results.
[0028] Specifically, after obtaining the data evaluation results, such as Figure 3 S3 in the example requires cross-data source difficulty calibration via extension module C. Specifically, a calibration matrix is constructed, and cross-source calibration is performed on the original difficulty distribution of each data source in the data evaluation results based on the calibration matrix to obtain the difficulty calibration result. The corresponding expression is: ; in, For the first Difficulty distribution after calibration of each data source For the first Difficulty calibration matrix for each data source For the first The difficulty distribution of each data source is determined; then, the difficulty spatial information of the difficulty calibration result is obtained, and spatial mapping is performed on the difficulty calibration result and all sampling quantities based on the difficulty spatial information to obtain multiple spatial mapping results. Finally, the total global difficulty is obtained based on all the spatial mapping results, and the corresponding expression is: ; in, For the first Each difficulty level represents the total global difficulty after calibration across all data sources. The first one set by humans Total number of samples from each data source For the first The first data source The natural proportion of each difficulty level This represents the total number of data sources.
[0029] After completing cross-data source difficulty calibration, such as Figure 3 S4 in the example needs to be automatically searched for stage allocation ratios through the extension module E. Specifically, the allocation ratio is searched based on the total global difficulty to obtain the stage allocation ratio, and the corresponding expression is: ; in, The optimal stage ratio is... for The stage allocation matrix, This represents the total number of training phases. To satisfy the monotonicity constraint of the course, As a comprehensive indicator, This is a matrix representing the total global difficulty; comprehensive indicators. The expression is: ; in, The loss after smoothing. To achieve convergence, the main optimization is training stability. For accuracy metrics, , and All are weights.
[0030] After that, as Figure 3 S5 in the algorithm is statically solved using core module A, extended module B, and extended module F. Specifically, firstly, core module A determines the linear equation system and matrix form based on the stage ratio. Then, it performs a matrix solution on the total global difficulty based on the linear equation system and matrix form to obtain the stage data volume matrix. The corresponding expression is: ; in, This is a matrix representing the amount of data in each stage. Stage ratio matrix The transpose of the matrix. Next, the stage data matrix is constrained using extension module B to obtain the matrix constraint solution, the corresponding expression of which is: ; in, For matrix constraint solutions, and All are regularization coefficients. To punish sudden changes in data volume between stages, To prevent excessive concentration of data in a single stage, the constraints are as follows: ; in, For the first Total amount of data required for each stage For the first The upper limit of the computing power budget for the stage. The total number of difficulty levels. For the first The difficulty level is the global total after sampling from all data sources. Then, the stage data matrix is integrally generalized using extension module F to obtain the continuous difficulty score, with the corresponding expression being: ; in, For the first The continuous difficulty ratio function of the stage. For continuous difficulty scoring, This represents the global difficulty distribution density. Finally, the data solution results are obtained based on the stage data volume matrix and the continuous difficulty score.
[0031] Step S30: Construct a stage dataset, dynamically solve the stage dataset according to the data solution results to obtain the target stage allocation, allocate data to the supervised fine-tuning data pool according to the target stage allocation to obtain the data allocation results, and train the model according to the data allocation results to obtain the target large language model.
[0032] Specifically, after obtaining the data solution results, such as Figure 3 In S6, the stage dataset is constructed to perform subsequent SFT training; afterwards, as... Figure 3 In S7, the extended module D determines whether to dynamically resolve the dataset. Specifically, based on the data resolution results, the stage dataset is periodically resolved to obtain multiple training feedbacks (i.e., every...). (Collect training feedback in each training step), and update the stage allocation based on all the training feedback to obtain the target stage allocation; then, determine whether the stage dataset needs to be dynamically solved based on all the training feedback; such as Figure 3 In step S8, if the stage dataset needs to be dynamically solved, the total global difficulty is recalculated based on the target stage ratio to obtain the target solution result, and then fed back to S5. Here, the total number of samples from each data source is always equal to the prior human calculation, and the corresponding expression is: ; in, For the first The iteration of the ... Total number of samples from each data source For the 0th iteration The total sampling amount from each data source is adjusted, and only the stage ratio and stage data volume are adjusted to ensure that the scaling law experimental design point does not drift; the total global difficulty is recalculated based on the target stage ratio, and the corresponding expression is: ; in, For the first The iteration of the ... In the first stage The stage ratio for each difficulty level, For the first The iteration of the ... In the first stage The stage ratio for each difficulty level, To update the step size, For the first In the first stage The gradient of stage proportions for each difficulty level. For the first The validation set loss of the next iteration; the corresponding inter-stage difficulty transition visualization is as follows: Figure 4 As shown, starting from the simple stage P1 with D1: 70%, D2: 20%, and D3: 10%, after a smooth transition via extension module B, the intermediate stage P2 with D1: 20%, D2: 60%, and D3: 20% is obtained. Then, after a smooth transition via feedback adjustment via extension module D, the difficult stage P3 with D1: 10%, D2: 20%, and D3: 70% is obtained. Figure 3 In step S9, if the stage dataset does not require dynamic solving, then the supervised fine-tuning data pool is allocated according to the target stage ratio to obtain the data allocation result. The scaling law experiment repeatability is guaranteed through the extension module H, strictly satisfying the target condition across all stages and all training steps. The corresponding expression is: ; in, For data source Assigned to a stage Difficulty The actual number of samples used.
[0033] Subsequently, the created large language model is trained based on the data allocation results to obtain the target large language model.
[0034] Step S40: Obtain the semantic information to be identified, input the semantic information to be identified into the target large language model, and obtain the semantic recognition result.
[0035] Specifically, after training the large language model to obtain the target large language model, the semantic information to be identified is obtained, the semantic information to be identified is input into the target large language model, and semantic recognition is performed through the target large language model to output the corresponding semantic recognition result.
[0036] The technical effects that this invention can bring are as follows: (1) Zero data waste: Using the inherent distribution as the division ratio, the data utilization rate is close to 100%; (2) Improved training stability: smooth transition between stages, and significantly reduced loss fluctuation. (3) Improved reliability of scaling law experiments: Strict conservation ensures that the curve design point does not drift; (4) Enhanced Adaptability: Without disrupting Under the premise of prior knowledge, online feedback and adjustment should also be taken into account; (5) Automated proportioning design: Eliminating reliance on expert experience; (6) Robust difficulty assessment: Multi-evaluator voting suppresses single judge bias; In summary, this invention improves data utilization by using the inherent distribution as the partitioning ratio, and also performs unified calibration on the difficulty across data sources, thereby improving the accuracy of the SFT training data allocation results and thus improving the accuracy of semantic recognition results.
[0037] Furthermore, such as Figure 5 As shown, based on the above-mentioned semantic recognition method for a large language model that combines static and dynamic elements, this invention also provides a semantic recognition system for a large language model that combines static and dynamic elements, wherein the semantic recognition system for the large language model that combines static and dynamic elements includes: The data evaluation module 51 is used to obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and to evaluate the supervised fine-tuning data pool to obtain the data evaluation result. The data solving module 52 is used to perform cross-source calibration on the data evaluation results to obtain difficulty calibration results, perform spatial mapping on the difficulty calibration results and all the sampling quantities to obtain the total global difficulty, search the stage ratio according to the total global difficulty, and perform static solution on the total global difficulty according to the stage ratio to obtain the data solving results. The model training module 53 is used to construct a stage dataset, dynamically solve the stage dataset according to the data solution results to obtain the target stage ratio, allocate data to the supervised fine-tuning data pool according to the target stage ratio to obtain the data allocation result, and train the model according to the data allocation result to obtain the target large language model. The semantic recognition module 54 is used to acquire the semantic information to be recognized, input the semantic information to be recognized into the target large language model, and obtain the semantic recognition result.
[0038] Furthermore, such as Figure 6 As shown, based on the above-mentioned semantic recognition method of a large language model with static and dynamic hybridity, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 6 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0039] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a semantic recognition program 40 for a large language model with static and dynamic content, which can be executed by the processor 10 to implement the semantic recognition method for a large language model with static and dynamic content in this application.
[0040] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the semantic recognition method of the static-dynamic hybrid large language model.
[0041] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface.
[0042] In one embodiment, when the processor 10 executes the semantic recognition program 40 of the large language model with static and dynamic hybridity in the memory 20, the following steps are performed: Obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and evaluate the supervised fine-tuning data pool to obtain data evaluation results; Cross-source calibration is performed on the data evaluation results to obtain difficulty calibration results. Spatial mapping is performed on the difficulty calibration results and all the sampling quantities to obtain the total global difficulty. The stage allocation is searched according to the total global difficulty, and the total global difficulty is statically solved according to the stage allocation to obtain the data solution results. Construct a stage dataset, dynamically solve the stage dataset based on the data solution results to obtain the target stage allocation, allocate data to the supervised fine-tuning data pool based on the target stage allocation to obtain the data allocation results, and train the model based on the data allocation results to obtain the target large language model; Obtain the semantic information to be identified, input the semantic information to be identified into the target large language model, and obtain the semantic recognition result.
[0043] The process of acquiring a supervised fine-tuning data pool containing multiple data sources and the number of samples from each data source, and then evaluating the supervised fine-tuning data pool to obtain data evaluation results, specifically includes: Obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and obtain a set of difficulty estimators, wherein the set of difficulty estimators includes multiple types of estimators; The supervised fine-tuning data pool is evaluated for difficulty based on all the evaluators to obtain multiple difficulty evaluation results. All the difficulty evaluation results are then integrated to obtain the data evaluation result.
[0044] Specifically, the process of performing cross-source calibration on the data evaluation results to obtain difficulty calibration results, and spatially mapping the difficulty calibration results and all the sample quantities to obtain the total global difficulty, includes: Construct a calibration matrix, and perform cross-source calibration on the original difficulty distribution of each data source in the data evaluation results based on the calibration matrix to obtain the difficulty calibration results; Obtain the difficulty space information of the difficulty calibration result, perform spatial mapping on the difficulty calibration result and all sampling quantities based on the difficulty space information to obtain multiple spatial mapping results, and obtain the total global difficulty based on all the spatial mapping results.
[0045] Specifically, the step of performing cross-source calibration on the native difficulty distribution of each data source in the data evaluation results based on the calibration matrix includes: ; The process of obtaining the total global difficulty based on all the spatial mapping results is as follows: ; in, For the first Difficulty distribution after calibration of each data source For the first Difficulty calibration matrix for each data source For the first The difficulty distribution of each data source For the first Each difficulty level represents the total global difficulty after calibration across all data sources. The first one set by humans Total number of samples from each data source For the first The first data source The natural proportion of each difficulty level This represents the total number of data sources.
[0046] Specifically, the step of searching for the global total difficulty based on the stage allocation, and then statically solving the global total difficulty based on the stage allocation to obtain the data solution result includes: A phase allocation ratio is obtained by performing a ratio search based on the total global difficulty, and a phase data volume matrix is obtained by performing a matrix solution on the total global difficulty based on the phase allocation ratio. The stage data volume matrix is constrained and solved to obtain the matrix constraint solution. The stage data volume matrix is then integrated and generalized to obtain the continuous difficulty score. The data solution results are obtained based on the stage data volume matrix and the continuous difficulty score.
[0047] Specifically, the construction phase dataset involves dynamically solving the phase dataset based on the data solution results to obtain the target phase allocation, and then allocating data to the supervised fine-tuning data pool according to the target phase allocation to obtain the data allocation result. This includes: Construct a phase dataset, and periodically solve the phase dataset based on the data solution results to obtain multiple training feedbacks; The stage allocation is updated based on all the training feedback to obtain the target stage allocation, and the stage dataset is determined based on all the training feedback to determine whether it needs to be dynamically solved. If the stage dataset needs to be solved dynamically, the total global difficulty is re-solved according to the target stage ratio to obtain the target solution result; If the stage dataset does not require dynamic solution, then the supervised fine-tuning data pool is allocated according to the target stage ratio to obtain the data allocation result.
[0048] Specifically, updating the stage allocation based on all the training feedback includes: ; in, For the first The iteration of the ... In the first stage The stage ratio for each difficulty level, For the first The iteration of the ... In the first stage The stage ratio for each difficulty level, To update the step size, For the first In the first stage The gradient of stage proportions for each difficulty level. For the first The validation set loss in the next iteration.
[0049] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a semantic recognition program for a large language model with static and dynamic hybrid characteristics, and the semantic recognition program for the large language model with static and dynamic hybrid characteristics, when executed by a processor, implements the steps of the semantic recognition method for the large language model with static and dynamic hybrid characteristics as described above.
[0050] In summary, this invention provides a semantic recognition method, system, terminal, and storage medium for a large language model employing a combination of static and dynamic methods. The method includes: acquiring a supervised fine-tuning data pool containing multiple data sources and the number of samples from each data source; evaluating the supervised fine-tuning data pool to obtain a data evaluation result; performing cross-source calibration on the data evaluation result to obtain a difficulty calibration result; spatially mapping the difficulty calibration result and all the sample numbers to obtain a global total difficulty; searching for a stage allocation based on the global total difficulty; statically solving the global total difficulty based on the stage allocation to obtain a data solution result; constructing a stage dataset; dynamically solving the stage dataset based on the data solution result to obtain a target stage allocation; and allocating data to the supervised fine-tuning data pool based on the target stage allocation to obtain a data allocation result. This invention improves data utilization by using an inherent distribution as the partitioning ratio and also improves the accuracy of SFT training data allocation by uniformly calibrating the difficulty across data sources.
[0051] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0052] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.
[0053] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A semantic recognition method for a large language model that combines static and dynamic elements, characterized in that, The semantic recognition method of the large language model with static and dynamic hybridity includes: Obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and evaluate the supervised fine-tuning data pool to obtain data evaluation results; Cross-source calibration is performed on the data evaluation results to obtain difficulty calibration results. Spatial mapping is performed on the difficulty calibration results and all the sampling quantities to obtain the total global difficulty. The stage allocation is searched according to the total global difficulty, and the total global difficulty is statically solved according to the stage allocation to obtain the data solution results. Construct a stage dataset, dynamically solve the stage dataset based on the data solution results to obtain the target stage allocation, allocate data to the supervised fine-tuning data pool based on the target stage allocation to obtain the data allocation results, and train the model based on the data allocation results to obtain the target large language model; Obtain the semantic information to be identified, input the semantic information to be identified into the target large language model, and obtain the semantic recognition result.
2. The semantic recognition method for a large language model with static and dynamic hybrid representation according to claim 1, characterized in that, The process of acquiring a supervised fine-tuning data pool containing multiple data sources and the number of samples from each data source, and then evaluating the supervised fine-tuning data pool to obtain data evaluation results, specifically includes: Obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and obtain a set of difficulty estimators, wherein the set of difficulty estimators includes multiple types of estimators; The supervised fine-tuning data pool is evaluated for difficulty based on all the evaluators to obtain multiple difficulty evaluation results. All the difficulty evaluation results are then integrated to obtain the data evaluation result.
3. The semantic recognition method for a large language model with static and dynamic hybrid representation according to claim 1, characterized in that, The process of performing cross-source calibration on the data evaluation results to obtain difficulty calibration results, and spatially mapping the difficulty calibration results and all the sample quantities to obtain the total global difficulty, specifically includes: Construct a calibration matrix, and perform cross-source calibration on the original difficulty distribution of each data source in the data evaluation results based on the calibration matrix to obtain the difficulty calibration results; Obtain the difficulty space information of the difficulty calibration result, perform spatial mapping on the difficulty calibration result and all sampling quantities based on the difficulty space information to obtain multiple spatial mapping results, and obtain the total global difficulty based on all the spatial mapping results.
4. The semantic recognition method for a large language model with static and dynamic hybrid representation according to claim 3, characterized in that, The cross-source calibration of the native difficulty distribution of each data source in the data evaluation results based on the calibration matrix specifically involves: ; The process of obtaining the total global difficulty based on all the spatial mapping results is as follows: ; in, For the first Difficulty distribution after calibration of each data source For the first Difficulty calibration matrix for each data source For the first The difficulty distribution of each data source For the first Each difficulty level represents the total global difficulty after calibration across all data sources. The first one set by humans Total number of samples from each data source For the first The first data source The natural proportions of each difficulty level This represents the total number of data sources.
5. The semantic recognition method for a large language model with static and dynamic hybrid representation according to claim 1, characterized in that, The step of searching for the stage allocation based on the total global difficulty, and then statically solving for the total global difficulty based on the stage allocation to obtain the data solution results, specifically includes: A phase allocation ratio is obtained by performing a ratio search based on the total global difficulty, and a phase data volume matrix is obtained by performing a matrix solution on the total global difficulty based on the phase allocation ratio. The stage data volume matrix is constrained and solved to obtain the matrix constraint solution. The stage data volume matrix is then integrated and generalized to obtain the continuous difficulty score. The data solution results are obtained based on the stage data volume matrix and the continuous difficulty score.
6. The semantic recognition method for a large language model with static and dynamic hybrid representation according to claim 1, characterized in that, The construction phase dataset is dynamically solved based on the data solution results to obtain the target phase allocation. Data is then allocated to the supervised fine-tuning data pool based on the target phase allocation to obtain the data allocation results. Specifically, this includes: Construct a phase dataset, and periodically solve the phase dataset based on the data solution results to obtain multiple training feedbacks; The stage allocation is updated based on all the training feedback to obtain the target stage allocation, and the stage dataset is determined based on all the training feedback to determine whether it needs to be dynamically solved. If the stage dataset needs to be solved dynamically, the total global difficulty is re-solved according to the target stage ratio to obtain the target solution result; If the stage dataset does not require dynamic solution, then the supervised fine-tuning data pool is allocated according to the target stage ratio to obtain the data allocation result.
7. The semantic recognition method for a large language model with static and dynamic hybrid representation according to claim 6, characterized in that, The step of updating the stage allocation based on all the training feedback specifically involves: ; in, For the first The iteration of the ... In the first stage The stage ratio for each difficulty level, For the first The iteration of the ... In the first stage The stage ratio for each difficulty level, To update the step size, For the first In the first stage The gradient of stage proportions for each difficulty level. For the first The validation set loss in the next iteration.
8. A semantic recognition system for a large language model that combines static and dynamic elements, characterized in that, The semantic recognition system of the large language model with static and dynamic hybrid features includes: The data evaluation module is used to obtain a supervised fine-tuning data pool containing multiple data sources and the number of samples for each data source, and to evaluate the supervised fine-tuning data pool to obtain the data evaluation results. The data solving module is used to perform cross-source calibration on the data evaluation results to obtain difficulty calibration results, perform spatial mapping on the difficulty calibration results and all the sampling quantities to obtain the total global difficulty, search the stage ratio according to the total global difficulty, and perform static solution on the total global difficulty according to the stage ratio to obtain the data solving results. The model training module is used to construct a stage dataset, dynamically solve the stage dataset based on the data solution results to obtain the target stage allocation, allocate data to the supervised fine-tuning data pool based on the target stage allocation to obtain the data allocation results, and train the model based on the data allocation results to obtain the target large language model. The semantic recognition module is used to acquire the semantic information to be recognized, input the semantic information to be recognized into the target large language model, and obtain the semantic recognition result.
9. A terminal, characterized in that, The terminal includes a memory, a processor, and a program stored in the memory and executable on the processor. When executed by the processor, the program implements the steps of the semantic recognition method for a large language model with static and dynamic hybridity as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program thereon, and the computer-readable storage medium stores a semantic recognition program for a large language model that combines static and dynamic elements. When the semantic recognition program for the large language model that combines static and dynamic elements is executed by a processor, it implements the steps of the semantic recognition method for a large language model that combines static and dynamic elements as described in any one of claims 1-7.