Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

101 results about "Fine-tuning" patented technology

In theoretical physics, fine-tuning is the process in which parameters of a model must be adjusted very precisely in order to fit with certain observations. Theories requiring fine-tuning are regarded as problematic in the absence of a known mechanism to explain why the parameters happen to have precisely the observed values that they return. The heuristic rule that parameters in a fundamental physical theory should not be too fine-tuned is called naturalness.

Fine tuning method and device for large language model, equipment and storage medium

The embodiment of the invention provides a fine tuning method and device for a large language model, equipment and a computer readable storage medium. According to the method, the performance of a to-be-fine-tuned large language model is tested, a sample set of prediction errors of the large language model is collected, the samples of the prediction errors are classified on the basis of real categories and error prediction categories of the samples of the prediction errors, namely, the prediction errors of the large language model are classified, and the classification accuracy of the prediction errors of the large language model is improved. Then, a trained pre-training language model is utilized to analyze reasons for generation of the prediction errors, and similar error samples with the same error condition are generated based on the reasons, so that targeted fine adjustment is carried out on the model, and the over-fitting phenomenon is eliminated. According to the method, the performance of the model on different types of errors can be better analyzed, and the overfitting problem of the model in a specific scene can be deeply known and solved, so that the performance of the model in each vertical field is remarkably improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Large model parameter fine tuning method and device for electric power multi-modal data fusion and medium

The invention relates to a large model parameter fine tuning method and device for electric power multi-modal data fusion and a medium. The method comprises the steps of obtaining electric power multi-modal data, constructing an electric power multi-modal feature vector and performing time sequence calibration, dividing the electric power multi-modal data into a plurality of sub-data sets according to a geographic position and an electric power equipment type, and distributing the sub-data sets to a plurality of computing nodes of a large model; calculating the local gradient of each calculation node, obtaining a parameter difference coefficient by calculating the gradient similarity between the adjacent calculation nodes, updating the weight coefficient of the calculation nodes, carrying out weighted aggregation on the local gradients of the plurality of calculation nodes according to the updated weight coefficient, obtaining a global gradient, and updating model parameters; and the parameter updating process is repeatedly executed until the model converges, and a large model after parameter fine tuning is obtained and used for outputting a power multi-modal data fusion result. Compared with the prior art, the method has the advantages that the problems of inconsistent power multi-modal data time sequence and unbalanced parameter aggregation are solved, and the model fine tuning efficiency and accuracy are improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Fine tuning method and system for retrieval enhancement generative model based on causal reasoning

The invention relates to a causal reasoning-based retrieval enhancement generative model fine tuning method and system, and the method comprises the following steps: A, modeling a causal relationship to reveal an influence mechanism of an input variable on related knowledge extraction and irrelevant knowledge filtering; step B, defining a knowledge revenue score (KGS), and calculating the knowledge revenue score through expert labeling or automatic evaluation to construct a causal enhancement data set (CED); c, designing a fine tuning strategy based on intervention and anti-factual reasoning, systematically performing knowledge variable intervention and anti-factual condition simulation, and enhancing the sensitivity of the model to related knowledge and the robustness of the model to irrelevant knowledge; and D, dynamically integrating the context and the knowledge to generate a high-quality response in a reasoning stage. The method is beneficial to improving the accuracy and stability of the dialogue generation model.
Owner:FUZHOU UNIV

GIS (Geographic Information System) industry large language model parameter fine tuning method considering parameter adaptability difference

The invention provides a GIS industry large language model parameter fine tuning method considering parameter adaptability difference, and relates to the field of geographic information science and deep learning, the method comprises the following steps: giving a GIS downstream task training data set and a pre-training large language model, and calculating the adaptability of each parameter of the model to the data set; calculating the adaptability of each level of the large language model according to the calculated adaptability of each parameter; based on the adaptability of each layer, distributing different trainable parameters for the model layers with different adaptability; and based on the distributed trainable parameters, performing fine tuning training on the pre-trained large language model to obtain a GIS industry large language model, and completing efficient fine tuning of the large language model parameters. According to the technical scheme, the heterogeneity of the large language model parameters in the GIS professional knowledge adaptation process is considered, and the characteristic is utilized to optimize the LoRA fine adjustment process so as to improve the performance of the large language model on GIS professional tasks.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Precipitation runoff time sequence simulation method based on pre-training-fine tuning

The invention provides a rainfall runoff time sequence simulation method based on pre-training-fine tuning, and belongs to the field of hydrological simulation. The method comprises the following steps: firstly, acquiring a consistent hydrometeorological data set for global multi-basin pre-training and target basin fine tuning, constructing an LSTM runoff prediction model, performing global multi-basin pre-training to obtain a base model parameter weight, migrating the base model parameter weight to a target basin, freezing LSTM network layer parameters, and only finely tuning a regression output layer; and testing and evaluating in the target drainage basin. According to the method, firstly, transferable hydrological response is learned on global multi-basin large samples, and then lightweight fine tuning is carried out on the target basin, so that the prediction precision and cross-regional generalization ability of a data scarce region can be effectively improved, and rapid deployment is facilitated.
Owner:DALIAN UNIV OF TECH

Providing a suitability prompt to evaluate and improve the output of a generative model without fine tuning

PCT designated stage expiredWO2025155426A1Biological modelsEngineeringData mining
The present technology provides a mechanism to obtain results of similar quality to that which can be obtained by fine-tuning a generative model from the foundational model without fine-tuning. In particular, the present technology can provide a suitability prompt to evaluate and improve the output of a generative model without fine-tuning. A suitability prompt is an engineered prompt that is provided to a generative model that prompts the generative model to evaluate a candidate response that has been generated by the generative model. Often the suitability prompt can include an indication of one or more attributes of a quality candidate response. When the generative model provides a response to the suitability prompt that indicates that the candidate response is a quality response, the candidate response can be deemed good enough to be returned to a user.
Owner:APPLE INC

Large model fine-tuning optimization method based on multi-strategy fusion

The invention discloses a large model fine tuning optimization method based on multi-strategy fusion, which comprises the following steps: designing a dynamic parameter selection mechanism, adaptively determining a parameter subset needing fine tuning according to a task demand and a model structure, and reducing unnecessary parameter updating calculation; constructing a dynamic low-rank decomposition framework, dynamically adjusting the rank of a low-rank matrix according to a model training state and data characteristics, and keeping key information while compressing a parameter scale; a self-adaptive task sensing mechanism is introduced, a fine adjustment strategy is automatically adjusted according to different task characteristics, and the adaptability of the model to various tasks is improved; and a mixed precision training method is adopted, so that the calculation complexity and the memory occupation are reduced on the premise of ensuring the model precision. According to the method, a parameter efficient fine tuning technology and a dynamic low-rank decomposition strategy are innovatively combined, and an adaptive task perception mechanism and a mixed precision training technology are introduced, so that the operand and resource requirements of model training are effectively reduced, and the fine tuning efficiency and the model performance are improved.
Owner:JIANGSU JIYUAN MEDICAL TECH CO LTD

Hybrid adaptive enhanced fine tuning method, system and device

The invention discloses a hybrid adaptive enhanced fine tuning method, system and device, relates to the technical field of artificial intelligence, and is particularly suitable for a deep reasoning task of a large language model. The problems of response level length deviation, problem difficulty level deviation, insufficient exploration efficiency, insufficient sample utilization and the like existing in an existing reinforcement learning algorithm are solved. The method comprises the steps of data preprocessing, model parameter and reference strategy initialization, multiple response generation, award calculation, advantage calculation and correction, model strategy updating, sampling probability adjustment and iterative optimization. Wherein all potential useful samples are ensured to be fully utilized by introducing the correction advantages of a length normalization factor and a difficulty normalization factor and combining a hybrid cutting mechanism and an adaptive sampling strategy update model. Deviation is effectively eliminated, the model reasoning ability and training efficiency are improved, the long reasoning task exploration ability is enhanced, and diversified task requirements are met.
Owner:BEIJING ZHONGHAIJIYUAN DIGITAL TECH DEV CO LTD

Large model intelligent question setting system and method based on fuzzy mathematics fine tuning

The invention discloses a large-model intelligent question setting system and method based on fuzzy mathematics fine tuning, and belongs to the technical field of intelligent question setting. Comprising a fuzzification processing module, a multi-dimensional reward feedback module, a rule-driven attribute reasoning module, an attribute accurate quantification module, a self-adaptive membership degree adjustment module, a strategy optimization and intra-cluster standardization module and a question generation and structured output module. According to objective fields such as post specifications, post levels, technical stack depth and team scales, hard indexes are converted into soft boundary semantics through five membership functions, and then quantitative attributes of surface test questions on difficulty, openness, prejudice and distinction are reasoned by using N rules without subjective wording. And the business result data is used for driving rewards to be updated online, and finally, daily automatic fine tuning is realized on the 32B large model, so that each subjective surface test question generated by the AI not only meets the objective requirements of posts, but also has explainable, auditory and iterable closed-loop capabilities.
Owner:HEBEI NOAH HUMAN RESOURCES DEVELOPMENT GROUP CO LTD

Forgetting learning method and device for large language model

The embodiment of the invention discloses a forgetting learning method and device for a large language model. The method comprises the steps that firstly, a first large language model and a forgetting sample set are obtained, the first large language model is initially a fine-tuning model obtained by fine-tuning a fine-tuning sample set, and the forgetting sample set is a subset of the fine-tuning sample set; then, for any first forgotten sample, a plurality of similar samples are determined based on a second large language model, and each similar sample and the first forgotten sample have similar semantics but different expressions; processing the first forgotten sample by using a first large language model to obtain a first hidden layer representation; then, training loss is determined, the training loss is negatively correlated to the distance between the first hidden layer representation and the hidden layer representation of each similar sample and is positively correlated to the distance between the first hidden layer representation and the random vector, and the hidden layer representation of each similar sample is obtained based on a fine tuning model; and then training the first large language model by using the training loss to realize forgetting learning.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Tibetan language large language model fine tuning method, device and system and storage medium

The invention discloses a Tibetan language large language model fine tuning method, device and system, and a storage medium. The method comprises the following steps: S1, obtaining a TIFD data set; s2, finely adjusting the Tibetan language large language model according to the TIFD data set; wherein low-rank increments are injected into the weight matrix of the base model through LoRA fine tuning. By adopting the technical scheme of the invention, the problems of high model fine tuning cost and low efficiency in a Tibetan language data scarcity scene are solved; and the generation accuracy of the Tibetan large language model on complex grammar structures and cultural terms is improved.
Owner:NORTHWEST UNIVERSITY FOR NATIONALITIES

Method and device for obtaining training data used for fine tuning of large model

The embodiment of the invention provides a method and device for obtaining training data used for fine adjustment of a large model, and the method comprises the steps: obtaining a first instruction for the large model, inputting the first instruction into a first large model, and obtaining a first answer; determining a complexity score for the first instruction by inputting the first instruction into the second large model, determining a response quality score for the first answer by inputting the first instruction and the first answer into the second large model, and determining a comprehensive score corresponding to the first instruction and the first answer according to the complexity score and the response quality score; and determining whether the first instruction and the first answer are used for fine tuning of the first large model according to the comprehensive score.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Brain-like scene safety assessment method combining process supervision and fast and slow thinking

The invention relates to a brain-like scene safety assessment method combining process supervision and fast and slow thinking. The method comprises the following steps: firstly, designing a brain-like scene security cognition process and a scene security level division standard; then, constructing a supervision fine tuning data set and a reinforcement learning fine tuning data set of a process supervision normal form based on the brain-like scene security cognition process and a scene security level division standard; then, performing cold start on an open-source pre-training multi-modal large language model based on the supervision fine tuning data set to obtain a supervision fine tuning model; carrying out reinforcement learning training on the supervision fine tuning model based on the reinforcement learning fine tuning data set of the process supervision normal form; and finally, inputting a to-be-evaluated scene image into the trained supervision fine tuning model to obtain a scene security level. The inference process of the model can be aligned with the thinking process of human in the macroscopic level, the flexibility of the thinking process is maintained, and the interpretability of the result of each step is improved.
Owner:SICHUAN UNIV

Small language model fine tuning method based on LLaMA Factory tool

The invention discloses a small language model fine tuning method based on an LLaMA Factory tool, and the method comprises the following steps: 1) carrying out the thinking chain distillation of an open source data set through a large model, and generating a training data set for the fine tuning of a small language model; 2) based on an LLaMA Factory tool, building a model fine tuning environment; 3) fine-tuning real-time monitoring and adjustment, in the fine-tuning process, monitoring loss of the verification set and the test set, and if the performance of the model is reduced or an over-fitting phenomenon occurs, adjusting training parameters and retraining; and 4) completing training and exporting the model: after the fine tuning process is completed, if the model reaches preset performance on the verification set, completing training, exporting the trained small language model, and deploying the small language model to practical application for reasoning. According to the method, targeted training is performed on the pre-trained small model through a fine tuning technology, specific task requirements can be efficiently met, and rapid migration of a new task can be realized only by using a small-scale task data set to adjust part of parameters of the model.
Owner:HUAZHONG UNIV OF SCI & TECH

A pre-training-fine-tuning-based precipitation runoff time series simulation method

The application provides a pre-training-fine-tuning-based precipitation runoff time series simulation method, and belongs to the field of hydrological simulation. First, a consistent hydro-meteorological data set for global multi-basin pre-training and target basin fine-tuning is obtained, an LSTM runoff prediction model is constructed, a base model parameter weight is obtained through global multi-basin pre-training, the base model parameter weight is migrated to the target basin, the LSTM network layer parameter is frozen and only the regression output layer is fine-tuned, and test evaluation is carried out in the target basin. The application learns the transferable hydrological response on a large sample of global multi-basins first, and then performs light fine-tuning in the target basin, which can effectively improve the prediction accuracy and cross-region generalization ability in the data scarce area, and is convenient for rapid deployment.
Owner:DALIAN UNIV OF TECH

Method for Quantifying the Value of Scientific Problems Embedded with a Multi-Dimensional Mixture-of-Experts Mechanism

The present application provides a method for quantifying the value of scientific questions by embedding a multi-dimensional mixture-of-experts mechanism, which relates to the technical field of data processing and includes: reading predetermined quantification dimensions; introducing a predetermined triple strategy to construct a target instruction fine-tuning dataset; performing parameter quantification based on the fine-tuning principle of the QLoRA large language model; obtaining a scientific question text, performing semantic feature analysis through a gating network, and obtaining the target weight coefficient of the scientific question text according to the text semantics; analyzing the scientific question text through a target recognition large model and combining the target weight coefficient to obtain a target quantification result. Through the present application, the technical problem in the prior art that due to relying on expert evaluation and subjective judgment, there are inconsistent evaluation criteria, resulting in low efficiency of scientific question value evaluation can be solved. By integrating multiple dimensions and combining large language models for fine-tuning, the internal evaluation of the value of scientific questions is realized, and the efficiency of scientific question value evaluation is improved.
Owner:DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

Large model compression method based on continuous layer pruning and endpoint tuning

The invention relates to a large model compression method based on continuous layer pruning and endpoint tuning, and the method comprises the steps: firstly introducing a learnable continuous interval soft mask, and building a differentiable hierarchical mask mechanism in a model in cooperation with a residual bypass; secondly, by minimizing the KL divergence between output distributions before and after pruning, the optimal pruning starting point and length are automatically learned, and adaptive selection of continuous layer segments is achieved; then, executing physical layer deletion according to the optimized interval parameters, and reconnecting the network structures before and after pruning; and finally, implementing an endpoint tuning strategy, only carrying out all-parameter fine tuning on key layers on two sides of the sheared interval, and recovering the model performance at the lowest calculation overhead. According to the method, through combination of differential interval search and end point directional optimization, accurate compression and high-performance maintenance of the depth dimension of the large model are realized, model storage occupation and reasoning delay are remarkably reduced, the model output reliability in a key task scene is guaranteed, and the method is suitable for large-scale popularization and application. The method is particularly suitable for efficient deployment of the large language model in a resource-constrained environment.
Owner:ZHEJIANG UNIV OF TECH

A manifold constraint multi-track adaptation-based large language model parameter fine-tuning method, device and medium

The application discloses a large language model parameter fine-tuning method and device based on manifold constraint multi-track adaptation and a medium. In view of the problems of unstable training and limited expression capacity of an existing low-rank adaptation technology, a plurality of parallel low-rank tracks are constructed, and a double random matrix is introduced to constrain the information flow between the tracks, so that the stability of gradient propagation is ensured. Meanwhile, the expression capacity of the adaptation module is enhanced under a limited parameter budget by dynamically fusing the outputs of the tracks through a dynamic gating mechanism. The method can realize stable, efficient and high-performance model fine-tuning by fine-tuning a small number of parameters, and reasoning has no additional overhead, and is particularly suitable for application scenarios with limited resources and high stability requirements, and has wide practical value.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

A model fine-tuning method, system, terminal and medium for parameter update control in large language model fine-tuning process

The application belongs to the technical field of model fine-tuning, and specifically discloses a model fine-tuning method, system, terminal and medium for parameter update control in a large language model fine-tuning process. A fine-tuning model is established based on a large language model to be fine-tuned, input data and corresponding expected output data are obtained after pre-processing of training data, feature encoding and feature extraction are performed on the input data through a word vector encoding layer and a backbone neural network, intermediate feature results are updated and fused layer by layer using a multi-layer encoding structure, and model output results are obtained. Training loss is calculated based on the deviation between the model output results and the expected output data, the weight parameters in the backbone neural network are evaluated for importance according to the training loss, the weight parameters with a higher contribution degree to the training loss are selected as a target parameter set, and a parameter update operation is performed on the target parameter set. The application can improve the model fine-tuning efficiency and stability, and enhance the adaptation ability of the model in specific task scenarios.
Owner:INSPUR ZHUOSHU BIG DATA IND DEV CO LTD

Electronic equipment for realizing large language model parameter fine tuning method and storage medium

The invention discloses electronic equipment for realizing a large language model parameter fine tuning method and a storage medium, and relates to the field of natural language processing. The fine tuning method comprises the steps that a pre-training model and a parameter matrix of the pre-training model are acquired, a parameter fine tuning module is arranged on a self-attention layer of the pre-training model, the parameter fine tuning module comprises LoRA modules of different task types and a router, each LoRA module comprises a general feature matrix and a specific feature matrix, and the general feature matrixes of the multiple LoRA modules are shared; obtaining a fine tuning data set; loading and freezing a parameter matrix of the pre-training model, and initializing parameters of a parameter fine tuning module; and performing fine tuning on the parameters of the parameter fine tuning module based on the fine tuning data set to obtain a fine-tuned large language model. Parameters of the MoE architecture are remarkably reduced, it is ensured that the model captures differences of various tasks to the maximum extent, and the cross-task generalization of the model is ensured.
Owner:BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

Method and device for finely adjusting pre-training model, equipment and storage medium

The invention discloses a method and device for relieving catastrophic forgetting of a language model, equipment and a storage medium, and relates to the technical field of machine learning. According to the method, linear interpolation is carried out between the pre-training model parameter set and the fine-tuning model parameter set, so that a linear interpolation point of the model can still keep a relatively low loss value when the model learns new task knowledge; then the pre-training model parameter set and the fine-tuning model parameter set are added to form a fusion model parameter set, and the linear interpolation operation only performs one-time simple arithmetical operation after fine-tuning is completed without introducing any additional trainable parameters or changing the model structure, so that when the model adapts to a new task, the model can be quickly and accurately trained. Original general knowledge and performance on old tasks can be reserved to the maximum extent, and meanwhile complexity and expenditure of a model reasoning stage are not increased.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Diffusion model post-training fine tuning method and system

The invention belongs to the technical field of model fine tuning, and discloses a post-training fine tuning method and system for a diffusion model, and the method comprises the steps: carrying out the recursive structure analysis of a pre-obtained diffusion model based on a recursive likelihood ratio optimizer, obtaining recursive parameters, and determining the generation condition of the diffusion model through multi-scale prompt information; according to recursion parameters and generation conditions of the diffusion model, parameters are injected in the recursion process of the diffusion model, and the gradient of the diffusion model is estimated in combination with a gradient estimation method; and according to the estimated gradient of the diffusion model, updating parameters of the diffusion model by utilizing a model parameter updating formula so as to realize post-training adjustment of the diffusion model. The method has the characteristics of lower variance and higher sample efficiency, and effectively reduces the variance of gradient estimation by combining zero-order, half-order and first-order gradient estimation technologies.
Owner:北京大学武汉人工智能研究院

Task processing method based on large language model fine tuning, electronic equipment and medium

The invention discloses a task processing method based on large language model fine tuning, electronic equipment and a medium. The method comprises the following steps: acquiring a training data set related to a downstream task; adding a low-rank adaptation module to all layers; setting a stratified sampling strategy: determining the relative importance of each intermediate layer according to the weight norm distribution of each layer during low-rank adaptation fine tuning of the large language model so as to set the sampling probability of each layer in the training process, enabling each intermediate layer to be uniformly sampled in a single training period through non-return sampling, and obtaining a stratified sampling result; the number of middle layers actually participating in parameter updating is reduced, and video memory occupation in the training process is greatly reduced on the premise that the fine tuning performance is not reduced; according to a stratified sampling strategy, layers needing to be unfrozen in the current training period are dynamically selected in the large language model, parameter updating is carried out on low-rank adaptation modules of the layers, and a training data set is traversed to carry out fine adjustment on the large language model; and the fine-tuned large language model is used for executing downstream tasks.
Owner:ZHEJIANG UNIV

A large model anti-forgetting fine-tuning method and device, computer equipment and medium

The application discloses a large model anti-forgetting fine-tuning method and device, computer equipment and medium. The fine-tuning method solves the problem of catastrophic forgetting by introducing an orthogonal penalty loss function. That is, by constructing an orthogonal penalty loss function, adding it to the original loss function to obtain a total loss function, and based on the total loss function, the model training learns in the direction of not forgetting the original knowledge during the model fine-tuning process. When the model learns new knowledge, the Lora fine-tuning technology is prone to cause forgetting of the original knowledge in the incremental pre-training process. At this time, the orthogonal penalty loss function will increase, and through back propagation for adjustment, the orthogonal relationship between the original parameter matrix is maintained, thereby reducing the catastrophic forgetting of the anti-forgetting large model after fine-tuning.
Owner:SHENZHEN SMARTCITY TECH DEV GRP CO LTD +1

Providing a suitability prompt to evaluate and improve the output of a generative model without fine tuning

The present technology provides a mechanism to obtain results of similar quality to that which can be obtained by fine-tuning a generative model from the foundational model without fine-tuning. In particular, the present technology can provide a suitability prompt to evaluate and improve the output of a generative model without fine-tuning. A suitability prompt is an engineered prompt that is provided to a generative model that prompts the generative model to evaluate a candidate response that has been generated by the generative model. Often the suitability prompt can include an indication of one or more attributes of a quality candidate response. When the generative model provides a response to the suitability prompt that indicates that the candidate response is a quality response, the candidate response can be deemed good enough to be returned to a user.
Owner:APPLE INC

Defense method and device for member reasoning attack in large model fine tuning, and medium

The invention discloses a defense method and device for member reasoning attacks in large model fine tuning and a medium, when the accuracy difference of a target classifier on training data and verification data is larger than a threshold value, the defense method for the member reasoning attacks in large model fine tuning is executed, and the defense method comprises the steps that before gradient descent of the target classifier, the target classifier is subjected to gradient descent; according to the cardinal number change of the label set corresponding to each label between the kth batch and the (k-1) th batch, performing data fusion on samples in the label set corresponding to each label between the kth batch and the (k-1) th batch; and / or, in the gradient descending process of the target classifier, averaging gradient updating parameters of the pth layer in the target classifier, and updating each gradient element value in the gradient matrix corresponding to the pth layer by using the average value; the gradients of other hidden layers except the pth layer in the target classifier are updated based on gradient updating parameters obtained through calculation of a gradient descent method.
Owner:ZHEJIANG UNIV +1

Rotor slip root cause traceability regulation and control method based on large model supervised fine tuning

The invention discloses a rotor slip root cause traceability regulation and control method based on a large model and supervised fine tuning. The method comprises the following steps: constructing a slip text knowledge base based on slip parameter division rules, slip field cues and organized pre-declarations; the constructed slip text knowledge base is divided into a training set and a test set, and the structure of each piece of data comprises slip reason analysis, slip regulation and control measures and supplementary description; supervised fine tuning is carried out on the large language model by using the constructed slip text data set in combination with a quantized low-rank adapter technology; and matching the collected rotor working condition parameters with a slip parameter division rule to obtain corresponding rotor working condition parameter grades, and inputting the corresponding rotor working condition parameter grades into the large language model subjected to supervised fine tuning so as to determine a root cause causing slip and formulate regulation and control measures in a targeted manner. According to the invention, a reliable, credible and accurate rotor slip root cause traceability regulation result is obtained.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Method and device for inhibiting large model vertical domain fine tuning overfitting and storage medium

The invention discloses a method and device for inhibiting large model vertical domain fine tuning overfitting and a storage medium, and belongs to the technical field of large model fine tuning and deep learning optimization. In order to solve the problem of overfitting possibly caused when LoRA is used for efficient fine tuning of parameters, a low-rank matrix decomposition technology introducing random masks is mainly adopted, and model integration is carried out in combination with multiple times of mask sampling. Through the method, the generalization ability of the model can be effectively improved, overfitting is prevented, and the expression ability of the model is kept in a downstream task even under the condition that the data size is small. Compared with a traditional method, the method has the advantages of being easy to implement, efficient and good in generalization performance.
Owner:PEKING UNIV

LLM-FEM fused steel structure damage intelligent prediction method

The invention relates to a steel structure damage intelligent prediction method based on LLM-FEM fusion, and relates to the technical field of civil engineering structure health monitoring and artificial intelligence crossing. The invention aims to improve the intelligence and interpretability level of civil engineering structure health monitoring. A steel frame structure model is established through finite element analysis software (ANSYS), various health and damage states are simulated, and response index data including displacement, stress, modal frequency and the like are obtained; secondly, structuring the data into a natural language input format, and constructing a fine tuning instruction set related to an impairment identification task; performing efficient parameter fine adjustment on the large language model (LLaMA-38B) to enable the large language model to have structural damage identification and interpretation capabilities; and finally, any structural response data can be input, and structural state prediction, damage type judgment and semantic explanation output are realized. According to the method, physical simulation data and language reasoning ability are fused, generalization, interpretation and practicability are achieved, and the method is suitable for the whole process of steel structure design, monitoring and operation and maintenance management.
Owner:HEBEI UNIV OF TECH

Efficient fine tuning method for multi-task large language model parameters

The invention discloses a multi-task large language model parameter efficient fine tuning method which comprises the following steps: acquiring data sets of different task types, and performing preprocessing and task embedding processing on the data sets; based on the pre-trained large language model, introducing a low-rank adaptive expert to construct a multi-task large language model through an expert mixing mode, a semantic perception routing mechanism and a task adaptive scaling mechanism; and training the constructed multi-task large language model through the data sets of different task types, so as to carry out parameter fine tuning on the low-rank matrix introduced when the low-rank adaptation expert is introduced and parameters in the two mechanisms, thereby obtaining the optimized multi-task large language model. According to the method, the adaptability, generalization performance and reasoning accuracy of the large language model in multi-task learning can be improved while low parameter overhead is kept.
Owner:BEIJING JIAOTONG UNIV