Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Fine-tuning" patented technology

In theoretical physics, fine-tuning is the process in which parameters of a model must be adjusted very precisely in order to fit with certain observations. Theories requiring fine-tuning are regarded as problematic in the absence of a known mechanism to explain why the parameters happen to have precisely the observed values that they return. The heuristic rule that parameters in a fundamental physical theory should not be too fine-tuned is called naturalness.

A manifold constraint multi-track adaptation-based large language model parameter fine-tuning method, device and medium

The application discloses a large language model parameter fine-tuning method and device based on manifold constraint multi-track adaptation and a medium. In view of the problems of unstable training and limited expression capacity of an existing low-rank adaptation technology, a plurality of parallel low-rank tracks are constructed, and a double random matrix is introduced to constrain the information flow between the tracks, so that the stability of gradient propagation is ensured. Meanwhile, the expression capacity of the adaptation module is enhanced under a limited parameter budget by dynamically fusing the outputs of the tracks through a dynamic gating mechanism. The method can realize stable, efficient and high-performance model fine-tuning by fine-tuning a small number of parameters, and reasoning has no additional overhead, and is particularly suitable for application scenarios with limited resources and high stability requirements, and has wide practical value.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

A model fine-tuning method, system, terminal and medium for parameter update control in large language model fine-tuning process

The application belongs to the technical field of model fine-tuning, and specifically discloses a model fine-tuning method, system, terminal and medium for parameter update control in a large language model fine-tuning process. A fine-tuning model is established based on a large language model to be fine-tuned, input data and corresponding expected output data are obtained after pre-processing of training data, feature encoding and feature extraction are performed on the input data through a word vector encoding layer and a backbone neural network, intermediate feature results are updated and fused layer by layer using a multi-layer encoding structure, and model output results are obtained. Training loss is calculated based on the deviation between the model output results and the expected output data, the weight parameters in the backbone neural network are evaluated for importance according to the training loss, the weight parameters with a higher contribution degree to the training loss are selected as a target parameter set, and a parameter update operation is performed on the target parameter set. The application can improve the model fine-tuning efficiency and stability, and enhance the adaptation ability of the model in specific task scenarios.
Owner:INSPUR ZHUOSHU BIG DATA IND DEV CO LTD

Strong-weak large model cycle fine-tuning training method for fact-rich conversation content generation

The application discloses a strong-weak large model cycle fine-tuning training method for fact-rich conversation content generation, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing one-time cleaning and cold start on an original fact-rich data set by using a strong large model; mixing the cleaned data and the original data according to a probability to form a mixed data set, and performing initial fine-tuning on a weak large model; performing performance evaluation on the fine-tuned model and calculating a performance value; when the performance value exceeds a dynamic threshold, triggering the weak large model to generate a new answer and updating a historical answer queue; performing multi-round training cycle fine-tuning based on the updated data set, and traversing multiple groups of parameter configurations; and finally saving a performance-optimal model. Through the strong-weak model decoupling cooperation and the cycle self-enhancement mechanism, the application effectively reduces the training cost, avoids overfitting, and improves the performance and generalization ability of the model in the fact-rich conversation generation.
Owner:AEROSPACE INTERNET OF THINGS TECH CO LTD

Federal parameter fine-tuning method and system for heterogeneous quantization large language model

This invention provides a method and system for fine-tuning federated parameters for heterogeneous quantized large language models. The server broadcasts data to each client; each client utilizes metadata to map its local quantization scaling factor to a unified latent space independent of the quantization bit width using a normalization coefficient related to the quantization bit width, reconstructs the global update residual, and performs zero-order gradient estimation based on low-rank subspace perturbation on the quantization scaling factor while freezing low-bit integer weights. A scalar gradient estimate is calculated through forward inference, and a seed-gradient scalar is formed by combining a random seed with the scalar gradient estimate. The server groups and aggregates the seed-gradient scalar pairs according to the quantization bit width and updates the seed sampling probability distribution using a time-aware exponential moving average mechanism. This invention eliminates precision-related biases, reduces client memory usage, and compresses the amount of uploaded data.
Owner:SHANGHAI JIAOTONG UNIV

A VLA model fine-tuning method based on structured stage and key frame supervision

A VLA model fine-tuning method based on structured stage and key frame supervision, containing five steps of basic model initialization, automatic label extraction, auxiliary architecture construction, joint training and inference execution, which can overcome the defects of structured operation supervision, long-range gain and low key frame prediction error. Zero artificial annotation, general adaptation automatically extracts stage / key frame label from demonstration gripper state, greatly improves the success rate of operation; long-term multi-key frame task gain is greater, no migration cost, no modification of basic VLA model architecture and inference process, plug and play, light and efficient, no performance loss auxiliary head is light MLP, the parameter amount of learning query token is extremely small, the training / inference speed is completely consistent with the basic model, the representation is accurate, the stage representation of long-term stable learning faithfully tracks the operation stage, the key frame prediction error is as low as 10 ‑4 -10 ‑5 orders of magnitude, and there is no error divergence in complex tasks.
Owner:ZHONGKE FIFTH CENTURY (HANGZHOU) INTELLIGENT TECHNOLOGY CO LTD +1

A sequence knowledge fusion enhanced large model fine-tuning method

The application discloses a kind of fusion sequence knowledge enhanced large model fine-tuning method.The application includes the following steps: first, by carrying out sequence question and answer test to large model, the deficiency of model is identified and iterative sequence knowledge is induced, to prepare knowledge base for data set construction;Second, based on sequence knowledge, problem generation rule is designed, and sequence knowledge question and answer SeqKQA data set is obtained by man-machine cooperation marking division;Then, on the basis of conventional evaluation index, for the characteristics of sequence relative positioning question and answer task, new untried rate and attempt correct rate are added, combined with harmonic mean F1 value, a multi-dimensional evaluation index system is constructed, the quantitative representation of the deviation of model sequence relative position understanding is realized;Finally, based on SeqKQA training set, semantic pairing strategy is designed to construct training data, and integrated fine-tuning input format is formed by matching customized CoT instruction, semantic pairing instruction fine-tuning is carried out on open source LLM, and the effect is verified through distribution in and distribution out double test sets, to improve the question and answer accuracy of model on the task.
Owner:EAST CHINA UNIV OF SCI & TECH

Large model translation fine-tuning method based on entropy change parameter importance perception

This invention relates to a large-scale model translation fine-tuning method based on the importance awareness of entropy-varying parameters, belonging to the field of efficient large-scale model fine-tuning technology. Addressing the problem of poor model performance caused by data scarcity in low-resource tasks, this invention proposes a method that first freezes all trainable parameters, feeds the input data of the low-resource task into a pre-trained large language model, performs multiple forward propagations, and collects the entropy sequence of the output distribution of each Transformer layer. Based on the entropy sequence, the importance score of each layer is calculated from two dimensions: the intensity of inter-layer information change and the stability of intra-layer response. According to the importance score, a LoRA rank is dynamically assigned to each layer. The LoRA module with the assigned differential ranks is then activated, and the model undergoes standard fine-tuning training to obtain a fine-tuned model adapted to the target low-resource task. This invention effectively improves the performance of large models in low-resource translation tasks without increasing the number of trainable parameters.
Owner:KUNMING UNIV OF SCI & TECH +4

A small-sample fine-tuning method based on regularization constraints to prevent catastrophic forgetting.

ActiveCN120764613BSmall sampleAlgorithm
This invention provides a few-sample fine-tuning method based on regularization constraints to prevent catastrophic forgetting. A preliminary customized model is obtained by training a basic Stable Diffusion model using standard LoRA on small sample data. During continued training, small sample data and auxiliary data are fused. A regularization loss term to prevent catastrophic forgetting is introduced into the loss function. This regularization loss term constrains the consistency between the current model output and the previous model output, reducing the destruction of learned features during new sample learning, thereby improving the model's generalization ability and stability under the target style. This method balances rapid model adaptation to new samples with the preservation of existing capabilities, avoiding problems such as style drift and loss of detail.
Owner:SHENZHEN MIRACLE HILL TECHNOLOGY CO LTD

A method for fine-tuning a putter and manufacture thereof

Described herein is a method for fine-tuning a putter and manufacture thereof. More specifically, an application (App) that is utilised for static and dynamic data point analysis along with algorithms to best determine how attributes of a putter need to be set up to maximise consistency of a user's putter stroke and roll of the ball. A customised putter is tailored and manufactured for each individual based on the output data of the App and a fitting system of the putter is multi-adjustable therein.
Owner:FINE TUNED COMPONENTS LTD

Construction machinery fault diagnosis method based on multi-stage perturbation fine-tuning large model

PendingCN122365067ARobustificationData set
The application discloses an engineering machinery fault diagnosis method based on multi-stage disturbance fine-tuning of a large model, which is based on fault troubleshooting data, combined with three types of interference construction strategies, and according to the data enhancement principle and standard troubleshooting logic, a preferred sample data set containing logical complex interference, colloquial interference and adversarial interference is constructed; then an engineering machinery fault diagnosis model based on a large language model is constructed, and the model is trained in stages based on the logical incomplete interference data set, the colloquial interference data set and the adversarial interference data set in turn; and in each stage, the corresponding training target is mapped to the key training parameters, and the key training parameters are dynamically adjusted according to the disturbance type of the current stage, so that the model gradually learns the fault diagnosis knowledge and anti-interference strategies contained in various types of disturbance data in the parameter space, and the accuracy and overall robustness of the fault diagnosis task are improved. The application can improve the accuracy and reliability of engineering machinery fault diagnosis.
Owner:ZHEJIANG UNIV

A method for generating spatiotemporal physics fields based on Mamba-Transformer architecture and fine-tuning of physical information.

This invention discloses a spatiotemporal physics field generation method based on the Mamba-Transformer architecture and physical information fine-tuning, belonging to the field of rapid prediction and simulation of complex physics fields. This method achieves a balance between data generalization and physical conservation through two-stage training: the first stage constructs a backbone network containing an encoder, Mamba modules, and a decoder, and obtains generalization ability through supervised training; the second stage freezes the backbone network parameters, calculates the gradient of the predicted field and the residuals of the physical equations through numerical differencing, generates feature corrections through a lightweight residual encoder, and decodes and outputs a prediction result with optimized physical consistency after fusing the original latent features. This invention solves the problems of large physical errors in purely data-driven models and poor generalization and training difficulties in traditional physical constraint methods, while possessing advantages such as high accuracy, wide adaptability, and low deployment cost, providing a new solution for rapid and reliable prediction of complex spatiotemporal physics fields.
Owner:HARBIN INST OF TECH

Data construction and dynamic resampling fine-tuning method and system for multi-dialect speech recognition

PendingCN122435922ASolve fitting deficienciesImprove recognition stabilityData setEngineering
The application provides a data construction and dynamic resampling fine-tuning method and system for multi-dialect speech recognition. Firstly, the existing speech recognition model is used to recognize and transcribe the dialect audio, clean it, and perform environment-related data enhancement on the audio based on acoustic environment simulation. Secondly, the dialect type, gender and speech speed features of the audio are extracted, and the data set is divided into multiple category buckets according to the feature combination. In the batch generation stage of model fine-tuning, exponential decay probability is used for sampling in the bucket to meet the allocated basic extraction quota, and dynamic balance scores are used to guide cross-bucket compensation. Finally, combined with the ladder data enhancement suitable for the number of historical sample extractions, the model parameter update is completed. The application solves the problem that the existing dialect recognition method cannot simultaneously consider multi-dimensional attribute balance, sample coverage and over-sampling risk control, and improves the recognition robustness and generalization ability of the model in harsh recording environments and multi-dialect cross-scenarios.
Owner:南京通达海软件有限公司