Multi-mode-based large model persistent evolution method and multi-mode-based large model persistent evolution system
By building a multi-source heterogeneous data access channel and a dynamic feedback closed-loop mechanism, combined with elastic weight consolidation and efficient knowledge fine-tuning, the problems of insufficient static iteration capabilities of remote sensing large models and inefficient multimodal feature fusion efficiency are solved, and the continuous evolution and efficient application of large models are achieved.
Patent Information
- Application Number
- CN202510403736.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-25
AI Technical Summary
The existing remote sensing model has insufficient static iteration capabilities, low multimodal feature fusion efficiency and unformed closed loops of human-computer collaboration, resulting in limited application value in actual business scenarios.
By building a real-time access channel for multi-source heterogeneous data, a dynamic feedback closed-loop mechanism and elastic weight consolidation technology are adopted, and efficient knowledge fine-tuning and multimodal collaborative optimization methods are combined to achieve the continuous evolution of the large model.
Effectively suppress the model forgetting rate, improve interpretation accuracy and cross-modal fusion accuracy, meet the real-time requirements of engineering deployment, and improve inference efficiency and resource utilization.
Smart Images

Figure CN120373424A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method and system for continuous evolution of large models based on multi-modalities. Background Art
[0002] With the rapid development of earth observation technologies such as remote sensing satellites and unmanned aerial vehicles, the daily global acquisition of remote sensing image data has exceeded the EB level. The integrated application of multi-modal remote sensing data such as high-resolution multi-spectral images, synthetic aperture radar (SAR) data, and lidar point clouds provides unprecedented information dimensions for fields such as environmental monitoring, disaster assessment, and precision agriculture. However, existing remote sensing intelligent interpretation technologies generally face three core challenges: the lack of static model iteration ability, the low efficiency of multi-modal feature integration, and the lack of formation of a human-machine collaborative closed loop, which severely restricts the continuous application value of remote sensing large models in actual business scenarios.
[0003] Traditional remote sensing interpretation models are mostly trained once based on a fixed training set and lack a dynamic update mechanism. For example, although mainstream platforms such as Google Earth Engine can handle PB-level data, their built-in models cannot be optimized online according to user feedback data. Research shows that when the spatio-temporal distribution of remote sensing images shifts (such as the change in the spectral characteristics of farmland caused by seasonal changes), the classification accuracy of static models will drop by more than 30% within 6 months. Existing solutions such as periodic full-scale model retraining not only consume thousands of GPU hours of computing resources but also lead to conflicts between old and new knowledge due to incomplete data coverage. Although methods such as Elastic Weight Consolidation (EWC) alleviate the problem of catastrophic forgetting through parameter regularization, in multi-task scenarios (such as simultaneously processing land use classification and building extraction), their classification accuracy will still be lost by 12-15 percentage points due to task interference.
[0004] In terms of multi-modal data fusion, the current technical system has two major bottlenecks: inconsistent representation and computational redundancy. Although the mainstream cascaded fusion architecture (such as the SwinRS model) can achieve joint analysis of visible light and SAR data, the feature alignment process consumes about 40% of the computing resources. More seriously, when dealing with the joint modeling of hyperspectral data (hundreds of bands) and three-dimensional point clouds, traditional methods lack a unified representation space, resulting in a decline in cross-modal correlation efficiency of more than 60%. The EarthNets v2 platform released in 2024 attempts to adopt Neural Radiance Field (NeRF) technology, but its spatio-temporal alignment error still reaches 1.5 pixel levels, making it difficult to meet the requirements of millimeter-level deformation monitoring.
[0005] The current fragmented state of the human-machine collaborative interaction link further restricts the implementation of technology. Existing interactive segmentation models (such as Meta SAM) support user click correction, but lack a dynamic evaluation mechanism for annotation quality. In agricultural applications, due to differences in the criteria for judging farmland boundaries among different users, directly using the original annotation data to fine-tune the model results in an error label pollution rate as high as 18.6%. In addition, mainstream platforms such as Alibaba Cloud AI Earth only provide one-way model inference services and do not build a closed-loop link of "intelligent suggestions - manual verification - model optimization", resulting in more than 70% of user feedback data not being effectively utilized.
[0006] At the level of engineering deployment, the problem of collaborative scheduling of heterogeneous computing resources is particularly prominent. The architectural differences between domestic chips (such as Ascend 910) and NVIDIA GPUs result in a 35% loss in task scheduling efficiency in a hybrid computing environment. Although the Inspur Cloud Remote Sensing Intelligent Analysis Platform has achieved a multi-level storage architecture, the response delay of its dynamic fine-tuning module exceeds 300ms, which cannot meet the real-time requirements of disaster emergency scenarios. In terms of edge computing, although the existing drone-side inference framework (such as EdgeRS) can reach a processing speed of 70FPS, the accuracy loss caused by model compression increases the small target missed detection rate to 9.8%. Therefore, a continuous evolution method for large models is needed to solve problems such as the lack of static model iteration ability, low efficiency of multi-modal feature fusion, and the absence of a human-machine collaborative closed-loop in the development of existing large models. Summary of the Invention
[0007] An object of the present invention is to provide a multi-modal-based continuous evolution method for large models in view of the deficiencies of the prior art, so as to solve problems such as the lack of static model iteration ability, low efficiency of multi-modal feature fusion, and the absence of a human-machine collaborative closed-loop in the development of existing large models.
[0008] To solve the above technical problems, the present invention adopts the following technical solutions:
[0009] A multi-modal-based continuous evolution method for large models includes the following steps:
[0010] S1: Construct a real-time access channel for multi-source heterogeneous data to update knowledge, and collect target data from the updated database to obtain a data set;
[0011] S2: Based on the dynamic feedback closed-loop mechanism, use the data set obtained in step S1 to train the large model, and self-correct the large model through user feedback, and output the corrected large model and an incremental data set;
[0012] S3: Implement an anti-forgetting continuous learning scheme for the corrected large model using the incremental data set obtained in step 2 to enable the large model to learn and consolidate knowledge;
[0013] S4: Adopt the efficient knowledge fine-tuning technology to fine-tune the large model output by S3;
[0014] S5: Based on the results obtained in S2 - S4, perform multi-modal collaborative optimization on the large model output in step S4, compare the learning effects of the optimized large model, and feedback to the previous steps to enable further training and development of the large model.
[0015] Furthermore, the specific implementation method in step 1 includes:
[0016] Build a static authoritative knowledge base, and realize real-time injection of sensor data into the knowledge base by deploying edge computing nodes to achieve real-time update of knowledge;
[0017] Based on the data obtained from the updated database, perform node annotation on its terms to obtain a text dataset after node annotation, and generate an incremental knowledge graph according to the relationships between nodes;
[0018] Organize the text dataset after node annotation and the incremental knowledge graph into a dataset for subsequent steps.
[0019] Furthermore, the method for generating an incremental knowledge graph includes:
[0020] Adopt the knowledge distillation technology to extract structured features from the knowledge in the updated database, thereby extracting the keywords of the knowledge point content and creating nodes annotated with keywords;
[0021] For each node, calculate its semantic similarity with all nodes. When the similarity exceeds the preset threshold, it is considered that there is an association between the two nodes, and connect this node with the parent node to obtain a knowledge spectrum graph;
[0022] For the nodes in the knowledge spectrum graph, if it is found that the similarity between nodes is lower than the threshold, it is considered that there is a conflict in the descriptions between nodes, then start confidence decay, and update the knowledge graph by reducing the confidence of the old nodes to obtain an incremental knowledge spectrum graph.
[0023] Furthermore, the method for extracting keywords includes:
[0024] Use a pre-trained language model to perform preliminary encoding on the documents in the obtained database above, and convert the text into vector form;
[0025] Identify the key information in the vectorized text through the multi-head attention mechanism and calculate the attention weight of each word;
[0026] According to the attention weights of each word, calculate its importance score. The specific formula is as follows:
[0027]
[0028] Among them, WS(V i ) represents the importance score of each word, W ji represents the similarity degree of two words, WS(V j ) represents the attention weight of each word, and d is the damping coefficient;
[0029] Finally, several words with the highest scores are selected as keywords. After obtaining the keywords, a node Node is created for each keyword, and the node contains the basic information, semantic feature vector, and initial confidence value of the keyword.
[0030] Furthermore, the implementation method of step 2 includes:
[0031] The dataset obtained in step S1 is divided into a training set and a test set. The large model performs feature extraction and semantic parsing on the text dataset and the incremental knowledge graph in the training set, and then fuses the obtained text semantic information vector and graph semantic vector to obtain a prediction result;
[0032] The test set is used to test the effectiveness of the large model. According to the gap between the prediction result and the actual data, combined with the feedback result of the user, the model parameters of the large model are fine-tuned to obtain the trained large model. During the fine-tuning process, a comprehensive evaluation report of the large model is obtained and output together with the dataset as an incremental dataset.
[0033] Furthermore, the method for obtaining the prediction result in step 2 includes:
[0034] Use the BERT-base model to initialize the embedding matrix E. The embedding matrix E is used to map the input word sequence of the text dataset into a vector sequence. Then, multiple groups of convolutional kernels are used to capture semantic features of different granularities of the vector sequence, and a feature map is output. After activation by ReLU, max pooling is performed, and finally, an attention mechanism is constructed to strengthen the key features, thereby obtaining the final text feature matrix;
[0035] For each node in the incremental knowledge graph, information aggregation is performed through a graph attention network to obtain a graph feature matrix;
[0036] For the text feature matrix and the graph feature matrix, a multi-head attention mechanism is used to calculate the similarity matrix S to obtain a fusion feature matrix, which is the prediction result. The specific formula is as follows:
[0037] S ij =(t i ·h j ) / (||t i ||·||h j ||)
[0038] Among them, Sij represents the relevance between the i-th data sample and the j-th knowledge graph node, which is the prediction result; t i is the i-th text feature vector in the dataset, h j is the graph feature vector obtained from the j-th knowledge graph node.
[0039] Furthermore, the implementation method of step S3 includes:
[0040] First, based on the elastic weight consolidation mechanism, knowledge consolidation is performed on the large model trained in S2, that is, the large model trained in S2 is repeatedly trained through the following loss function to fine-tune the model parameters:
[0041]
[0042] where is the conventional loss function for training the interpretation task, F i represents the Fisher information matrix of the parameters, represents the old parameter data, θ i , and λ is the regularization coefficient;
[0043] Second, based on the adaptive data replay mechanism, corresponding training and fine-tuning task samples are generated from the dataset obtained in S2, and the task sample replay operation is performed;
[0044] Finally, for the large model after elastic weight consolidation and the task samples after data replay, a dual-memory mechanism is adopted to further train.
[0045] Furthermore, step 4 specifically includes:
[0046] First, for the large model parameters obtained in step S3, low-rank adaptive fine-tuning is used, and they are substituted into the Transformer layer of the trainable low-rank matrix for further training and fine-tuning to obtain a fine-tuned large model;
[0047] Second, for the above fine-tuned model, a dual-tower contrastive learning architecture is constructed to align the knowledge of the model; the dual-tower learning architecture includes a visual encoder and a text encoder. The visual encoder converts the input into a visual feature vector, and the text encoder converts the input into a text description vector; the following loss function is designed, and the data in step S1 and the above fine-tuned large model are input for further training:
[0048]
[0049] In the formula, s(v i ,t j ) is the cosine similarity, τ = 0.07 is the temperature coefficient, and L align is the loss of the model;
[0050] Through the above steps, a large model with knowledge fine-tuning is obtained.
[0051] Furthermore, the implementation manner of step S5 is as follows:
[0052] S5.1, through the neural radiance field, perform unified representation on the large model obtained in step S4; that is, perform position encoding on the spatial coordinates, then input the encoded coordinates into the large model in S4, establish a graph attention weight matrix for the knowledge graph obtained in S1, and perform convolution operations on the semantic feature matrix of the large model obtained in S4 respectively, then splice the convolved features, and map the spliced modal features to the unified representation space. The formula is as follows:
[0053] σ,c = MLP θ (x,d)
[0054] Among them, MLP is the radiance field MLP parameter matrix, representing the mapping from spatial coordinates to the radiance field parameter space, x ∈ R 3 is the spatial coordinate, d ∈ R 3 is the viewing direction, σ represents the volume density, c is the radiance color, and map each modal feature to the radiance field parameter space through the cross-modal projection network;
[0055] S5.2, based on dynamic network routing, decompose the multi-modal unified representation generated in the S5.1 stage into mutually orthogonal semantic subspaces, each subspace corresponding to a specific modality. For a specific model, aggregate the outputs of each path expert network through routing probability weighting to obtain a spliced decoupled feature matrix;
[0056] S5.3, construct corresponding anchor points for the large model, design a dual contrast loss formula according to the generated anchor points, compare the feature space of the large model, so as to compare the baseline model extracted from S2 and the optimized model obtained from S4, perform contrast optimization for each dimension of the decoupled feature matrix obtained in S5.2, and feedback the obtained contrast loss to the previous steps, so that the large model can be further trained and developed.
[0057] Another object of the present invention is to provide a system for implementing the above-mentioned method for continuous evolution of a multi-modal large model, including:
[0058] A dataset acquisition module, used to construct a real-time access channel for multi-source heterogeneous data to update knowledge, and collect target data from the updated database to obtain a dataset;
[0059] The large model preliminary training module is used to train the large model based on the dynamic feedback closed-loop mechanism by using the obtained dataset above, and self-correct the large model through user feedback, and output the corrected large model and the incremental dataset;
[0060] The large model learning and consolidation module is used to implement an anti-forgetting continuous learning scheme for the corrected large model by using the obtained incremental dataset so that the large model can learn and consolidate knowledge;
[0061] The large model fine-tuning module is used to perform knowledge fine-tuning on the large model after learning and consolidation by using efficient knowledge fine-tuning technology;
[0062] The large model collaborative optimization module is used to perform multi-modal collaborative optimization on the large model after knowledge fine-tuning based on the above results, compare the learning effects of the optimized large model, and feedback to the previous steps, so that the large model can be further trained and developed.
[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0064] Breakthrough in continuous evolution ability: Through the dynamic feedback closed-loop mechanism (intelligent interpretation-artificial correction-model fine-tuning) and the elastic weight consolidation technology (EWC) in step S2, the model forgetting rate is effectively suppressed to 4.2%, which is 6.8 times higher than that of the traditional static training method (forgetting rate 28.6%); combined with the sample selection strategy with entropy value > 2.3, the interpretation accuracy of the large model in the environmental monitoring scenario is improved by 18%-22%;
[0065] Multi-modal collaborative optimization: In step S5, the neural radiance field (NeRF) is used to uniformly represent multi-modal data, and the cross-modal fusion accuracy reaches 92.4%; combined with the dynamic network routing (DNR) to achieve feature decoupling, the IoU of the farmland boundary is increased to 0.87, and the error is reduced to 0.5 pixel level, which is 45% higher than the SAM model;
[0066] Engineering efficient support: In step S5, a heterogeneous computing pipeline architecture (NVIDIA A100 + Ascend 910B) is constructed, and the inference efficiency is increased by 40%; the low-code development environment reduces the connection error rate from 15.3% to 0.9%; real-time detection at the edge reaches 70FPS, and the latency < 50ms, meeting the demand for TB-level point cloud processing;
[0067] The present invention has good practical value and promotion prospects, and can be widely applied to fields such as environmental dynamic monitoring (soil humidity / vegetation cover analysis), agricultural precision management (crop growth prediction / pest and disease identification), smart city (3D building modeling / surface change detection), disaster emergency assessment (flood disaster range calculation), land and resources survey (land use classification / illegal construction identification), etc., providing millimeter-level surface analysis capabilities for digital transformation. Brief Description of the Drawings
[0068] Figure 1 This is a flowchart of the method for continuous evolution of a large model based on multi-modal in an embodiment of the present invention. Detailed Embodiments
[0069] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0071] The present invention will be further described below in conjunction with specific embodiments, but it is not a limitation of the present invention. As Figure 1 shown, an embodiment of the present invention discloses a method for continuous evolution of a large model based on multi-modal, including the following steps:
[0072] S1: Build a real-time access channel for multi-source heterogeneous data to update knowledge, and collect target data from the updated database to obtain a data set; this step includes:
[0073] S1.1: Build a static authoritative knowledge base, and realize real-time injection of sensor data into the knowledge base by deploying edge computing nodes to realize real-time update of knowledge;
[0074] Static authoritative database: Obtain the required data from official professional websites, and input the obtained data into the database in the storage architecture;
[0075] Build a real-time access channel to update the database in real time: Deploy edge computing nodes to realize direct transmission of sensor data, and connect to the API interface data source through the Kafka stream processing platform. Collect the original data stream through IoT devices deployed at the edge, adjust the collection frequency using an adaptive sampling algorithm, and update the data in the static authoritative database in real time according to the collected data. The core formula of this algorithm is:
[0076] f = α × log(ΔD / Δt)
[0077] In the formula, α is the scene sensitivity coefficient (0.8 for industrial scenarios and 1.2 for urban management), ΔD represents the data distribution change degree, and Δt is the time interval.
[0078] The above database will simultaneously serve as the data source for the following steps. Based on the static authoritative knowledge base, real-time update of knowledge is achieved through real-time injection into the knowledge base.
[0079] S1.2. Based on the data obtained from the updated database, perform node annotation on its terms to obtain a text dataset after node annotation, and generate an incremental knowledge graph according to the relationships between the nodes; specifically including:
[0080] ① Based on the LLM agent, use knowledge distillation technology to extract the structured features of the knowledge in the updated database, thereby extracting the keywords of the knowledge point content, and creating a node labeled with this keyword. The specific process is as follows:
[0081] Use pre-trained language models such as BERT to perform preliminary encoding on the document data obtained from the database, and convert the text into vector form. Then, identify the important information in the vectorized text through the multi-head attention mechanism, and this process will calculate the attention weights of each word.
[0082] Then, use the TextRank algorithm to input the attention weights of each word and score the importance of these words. This algorithm will consider the correlation relationships between words, similar to the principle of page ranking, and the specific formula is as follows:
[0083]
[0084] In the formula, WS(V i ) represents the importance score of each word, the summation on the right represents the contribution degree of each adjacent word to this word, W ji represents the similarity degree between two words, WS(V j ) represents the attention weights of each word, and d is the damping coefficient, generally 0.85.
[0085] Finally, select several words with the highest scores as keywords. After obtaining these keywords, create a node Node for each keyword, and the node contains the basic information, semantic feature vector, and initial confidence value of the keyword.
[0086] ② According to the relationship between the created node and other nodes, automatically associate its parent class. This process is mainly achieved through semantic similarity calculation and knowledge reasoning. Specifically, for each newly created node, first calculate its semantic similarity with all existing nodes. Here, the cosine similarity calculation method based on word vectors is used, and the specific formula is as follows:
[0087] similarity=(v1·v2) / (||v1||·||v2||)
[0088] In the formula, v1, v2 are the feature vectors of two nodes, and similarity is the obtained similarity.
[0089] When the similarity exceeds a preset threshold (set to 0.8 in this embodiment), it is considered that there is an association between these two nodes, and thus the node is connected to the parent node.
[0090] ③ For the above nodes, if the similarity between nodes is found to be lower than the threshold (set to -0.8 in this embodiment), it is considered that there is a conflict in the descriptions between the nodes. Then, confidence decay is initiated, and the knowledge graph is updated by reducing the confidence of the old nodes. The formula is as follows:
[0091] s(e i ) = λ·s(e i )+(1 - λ)·s new
[0092] In the formula, s(e i ) represents the confidence of the existing node description, s new represents the confidence of the new node description, and λ represents the decay factor, usually 0.9.
[0093] Generate an incremental knowledge graph according to the above steps, and save the generated knowledge graph through a knowledge version control system to make it easy to trace and perform difference analysis. Select high-quality data from the text data set after node annotation and the incremental knowledge graph as the data set for training in the subsequent steps.
[0094] Step S2, based on the dynamic feedback closed-loop mechanism, train the large model using the data set obtained in step S1, and self-correct the large model through user feedback, and output the corrected large model and the incremental data set; the specific process is as follows:
[0095] S2.1, divide the above data set into a training set and a test set according to a ratio of 8:2. Use the large model to perform feature extraction and semantic parsing on the text data set and the incremental knowledge graph in the training set, and then fuse the obtained semantic information vectors to obtain a prediction result. The specific process is as follows:
[0096] ① For the annotated text data set, perform text data feature extraction to obtain a text feature matrix. Specifically, first initialize the embedding matrix E using the BERT-base model, and use the embedding matrix E to map the input word sequence of the text data set into a vector sequence; then, configure three convolution kernel sizes of 3×3, 5×5, and 7×7 to capture semantic features of different granularities of the vector sequence, output feature maps, and perform max pooling after activation by ReLU. Finally, construct an attention mechanism to strengthen the key features of the pooling result, so as to obtain the final text feature matrix;
[0097] ②For each node in the incremental knowledge graph, a three-layer graph attention network (GAT) architecture is adopted, and information aggregation is performed through the graph attention network. First, the keyword semantic vectors are taken out from each node as the input, and then they are input into the attention network for iterative update. Finally, the node feature vectors are combined to obtain the graph feature matrix;
[0098] ③For the two matrices obtained above, the multi-head attention mechanism is used to calculate the similarity matrix S to obtain the fused feature matrix, which is the prediction result of the computer. The specific formula is as follows:
[0099] S ij =(t i ·h j ) / (||t i ||·||h j ||)
[0100] In the formula, S ij represents the correlation between the i-th data sample and the j-th knowledge graph node, that is, the prediction result; t i is the i-th text feature matrix in the dataset, and h j is the graph feature matrix obtained from the j-th knowledge graph node.
[0101] S2.2. Use the test set to test the effectiveness of the large model. According to the gap between the prediction result and the actual data, combined with the feedback result of the user, the model parameters of the large model are corrected. During the correction of the large model, a dynamic weighted loss function is used for correction:
[0102] L=0.3L ce +0.5L kl +0.2L reg
[0103] where L ce is the cross-entropy loss, L kl is the distribution difference loss, and L reg is the parameter regularization term.
[0104] The correction process is based on an artificial correction platform for correction, that is, the deviation of the above-mentioned generated prediction result is displayed through a visualization interface, so that it can be corrected manually. The corrected model parameters are transmitted to the automatic evaluation module to generate a comprehensive scoring report including accuracy, robustness, and inference speed. This report will be used as the basis for triggering the model update in the S3 stage and the selection of the benchmark model in the S5 stage. Finally, repeat the above steps several times to obtain the corrected large model and output the corrected incremental dataset.
[0105] As can be seen from the above, this step goes through the iterative process of "intelligent interpretation - manual correction - model fine-tuning" as described above. This step establishes an automated evaluation - optimization loop, enabling the large model to have the potential for continuous evolution. At the same time, several large models are trained from the data obtained in S1, and their parameters will enter the next step for further training.
[0106] S3: Implement an anti-forgetting continual learning scheme for the corrected large model using the incremental dataset obtained in Step 2 to enable the large model to learn and consolidate knowledge; the specific process is as follows:
[0107] S3.1, Based on the elastic weight consolidation mechanism, perform knowledge consolidation on the large model obtained in S2, prompting the large model to effectively retain the memory of old data when introducing new data. Specifically, the elastic weight consolidation method introduces a regularization term into the loss function to limit the adjustment of important parameters, thereby preventing the large model from forgetting old knowledge when learning new tasks. The matrix parameters of the large model obtained in S2 are passed through the loss function represented by the following formula, and the dataset of Step S1 is used for repeated training to obtain the large model, enabling its model parameters to be fine-tuned:
[0108]
[0109] where, is the total loss function, is the conventional loss function for training the corresponding interpretation task. For example, in the remote sensing image semantic segmentation task, it is usually represented as the cross-entropy calculated from the model output f(x; θ) and the label y, F i represents the Fisher information matrix of the parameters, represents the old model parameters, θ i is the current task parameter, and λ is the regularization coefficient.
[0110] S3.2, Based on the adaptive data replay mechanism, generate corresponding training and fine-tuning task samples from the incremental dataset generated in S2, and perform task sample replay operations. By selecting the most representative samples, the forgetting of the model for previous tasks after subsequent task training is minimized to the greatest extent. First, for all training samples, calculate their corresponding training entropy values, and preferentially select task samples with entropy > 2.3; then, among the selected task samples, use K-means feature clustering to divide them into 50 clusters, and eliminate old samples with Mahalanobis distance > 2.3σ, and select the samples closest to the cluster center. In addition, save the distances between the selected samples and the centers of all projection samples in the current task to ensure that when new tasks are added, samples farthest from the initial cluster center can be discarded.
[0111] S3.3. For the large model after elastic weight consolidation and the task samples after data replay, a dual-memory mechanism is adopted for further training. The dual-memory mechanism has a deep collaborative relationship with the large model. By establishing the representations of each task in the task samples, the large model can deeply remember. It includes employee professional memory and task guide memory, and the two parts are trained simultaneously. Among them,
[0112] Employee professional memory: After each training task in a sample is completed, the task completion situation is automatically summarized, and incremental weighted updates are performed on the large model. The formula is as follows:
[0113]
[0114] where p new , p old represent the old and new success rates respectively, n represents the historical number of times, and c represents the current result.
[0115] Task guide memory: Each time a new task is successfully executed, that is, when the output result of the large model matches the expectation, new memories are automatically added. The BM25 function is used for similar task matching, and the search consolidation formula for past sample tasks is as follows:
[0116] r = BM 25 (t)
[0117] where r represents the search result and t represents the reference task for the search.
[0118] Endowed with employee professional memory and task guide memory, the large model can realize the memory function for old knowledge, and save these model parameters into the database for use in step S4.
[0119] S4: Adopt an efficient knowledge fine-tuning technique to perform knowledge fine-tuning on the large model output by S3. This step specifically includes:
[0120] S4.1. For the large model parameters obtained in step S3, use low-rank adaptation fine-tuning, substitute them into the Transformer layer of the trainable low-rank matrix for further training and fine-tuning to obtain a fine-tuned model. Specifically, in the original Transformer structure, each layer contains an attention mechanism and a feed-forward neural network, and these components are parameterized by the weight matrix W. During the actual training process, the input data (text sequence) is first converted into a vector sequence through the embedding layer, and then passes through the modified Transformer layer layer by layer, and then the adjusted large model parameters are obtained.
[0121] The core idea of low-rank adaptation is as follows: for each input target weight matrix, two trainable low-rank matrices A and B are introduced, and through linear combination, a matrix parameter update amount is generated. To achieve efficient knowledge fine-tuning, a smaller learning parameter matrix is used to reduce the number of parameters to be trained. The parameter update formula for the training matrix is:
[0122] W′ = W + α·B·A
[0123] In the formula, W is the original matrix of the large model, and W′ is the updated matrix. Among them, A and B are learning parameter matrices, and α is a scaling factor.
[0124] S4.2: For the fine-tuned model above, construct a two-tower contrastive learning architecture to align the knowledge of the model; the two-tower learning architecture and the large model architecture form a cooperative relationship.
[0125] It includes two encoders: the visual encoder inputs the image, and the text encoder inputs and inherits the parameters of the fine-tuned large model. These two encoders respectively convert the inputs into high-dimensional vector representations. Specifically, for each pair of input samples (v i , t i ), where v i represents the visual feature vector and t i represents the corresponding text description vector, and the two encoders will respectively generate their vector representations.
[0126] The cooperative training of the two is driven by a contrastive loss function, so that the vector representations of the image-text pairs satisfy similarity constraints in the shared latent space, so that the knowledge in the large model can be aligned. To achieve this goal, the following loss function is designed. Input the data from step S1 and the fine-tuned model for further training:
[0127]
[0128] In the formula, N is the modal learning batch size, s(v i , t j ) is the cosine similarity, τ = 0.07 is the temperature coefficient, and L align is the loss of the model;
[0129] Through the above steps, a model with knowledge fine-tuning is obtained, which can understand knowledge better.
[0130] S5: Based on the results obtained in S2 - S4, perform multi-modal collaborative optimization on the large model output in step S4, compare the learning effects of the optimized large model, and feedback to the previous steps to enable further training and development of the large model; this step specifically includes:
[0131] S5.1. Use the neural radiance field to perform a unified representation of the large model obtained in step S4. After fine-tuning the large model in stage S4, the core objective of constructing a differentiable neural radiance field is to embed the knowledge of the large model into a three-dimensional continuous space, while integrating the geometric-semantic correspondence in the labeled dataset in stage S1 and the structural constraints of the incremental knowledge graph. Specifically, the scene representation consists of two core components: an implicit feature generator based on the S4 large model and a radiance field MLP constrained by the knowledge graph, and the two are optimized end-to-end through the gradient flow.
[0132] The process of constructing the radiance field is as follows: First, perform position encoding on the spatial coordinates, then input the encoded coordinates into the S4 large model. Next, establish a graph attention weight matrix for the knowledge graph obtained in S1, and perform convolution operations on it and the semantic feature matrix of the S4 large model respectively. Then, splice the convolved features. Finally, map the spliced modal features to a unified representation space, and the formula is as follows:
[0133] σ,c = MLP θ (x,d)
[0134] where MLP is the MLP parameter matrix, representing the mapping from spatial coordinates to the radiance field parameter space, x ∈ R 3 is the spatial coordinate, d ∈ R 3 is the viewing direction, σ represents the volume density, and c is the radiance color. Map the features of each modality to the radiance field parameter space through a cross-modal projection network.
[0135] S5.2. Based on dynamic network routing, for the unified representation achieved in the previous step S5.1, decompose the multi-modal unified representation generated in stage S5.1 into mutually orthogonal semantic subspaces, and each subspace corresponds to a specific modality (such as shape / material / motion, etc.). The specific process is as follows:
[0136] First, take the unified representation feature matrix F of S5.1 (the matrix formed after mapping to the parameter space in S5.1), and generate a dynamic query vector q by encoding the task description text through BERT;
[0137] Secondly, pre-define K = 6 expert networks, and each expert corresponds to a basic semantic modality. Such as the shape expert, the material expert, the motion expert, etc. The topic of each expert is determined by the actual application scenario. At the same time, for the parallel processing of the unified representation by each expert network, input the unified representation matrix F obtained in S5.1 to get the output features f of each expert k ;
[0138] Then, through the Gram-Schmidt orthogonalization process, generate an orthogonal basis, and the formula is as follows:
[0139]
[0140] Among them, is the orthogonal basis of the output features of the k-th expert, and f k is the output feature of the k-th expert.
[0141] Then, use the dynamic routing mechanism to obtain the routing probability matrix from each expert network route and the query vector, and apply the softmax function with a temperature coefficient to generate the routing probability distribution. The formula is as follows:
[0142]
[0143] In the formula, P k represents the selection probability of the k-th path, β = 0.5 is the routing coefficient, sim(f k , q) is the feature similarity calculation function, and K is the total number of candidate paths.
[0144] Finally, weighted aggregation of the expert outputs is performed to obtain the decoupled feature matrix. The specific formula is as follows:
[0145]
[0146] Among them, P k is the selection probability of the k-th path, f k is the output feature of the k-th expert, and W k is the adaptation projection matrix of each expert, which is calculated. K = 6 is the total number of roads, that is, the number of expert networks. The final output of this model is the decoupled feature matrix, and this matrix will indirectly optimize the large model parameters through the following learning mechanism.
[0147] S5.3. Establish a multi-model comparison learning mechanism to compare and evaluate the large model parameters obtained in the previous steps, and then optimize them. The specific process is as follows:
[0148] First, construct corresponding anchor points for the large model. From the S2 evaluation report, select the 3 large models with the highest accuracy as the benchmark anchor points, and generate comparison sample pairs: (f anchor , f perturbed ), f anchor ,
[0149] f perturbed are the matrices of the standard model and the model to be compared, respectively;
[0150] Secondly, according to the above-generated anchor points, perform feature space comparison. Design a double comparison loss formula to compare the feature spaces of the large models, so as to compare the benchmark model extracted from S2 and the optimized model obtained from S4. For each dimension of the decoupled feature matrix obtained in S5.2, determine the comparison optimization. The specific formula is as follows:
[0151]
[0152] Among them, the similarity of positive samples Negative samples are taken from other modality combinations, and the similarity is Input the model samples into the comparison formula to obtain the comparison loss L contrast 。
[0153] Finally, feedback the obtained comparison loss to the previous steps, so that the large model can be further trained and developed. Specifically, it includes the following:
[0154] For model parameters with significantly different reliabilities (>2σ), use the comparison loss matrix to perform directional correction to obtain updated large model parameters, and perform the following test screening on the updated model.
[0155] For the obtained model, construct a performance test task from the test set in S2, and verify the accuracy of the output of the large model respectively:
[0156] For models with poor accuracy, return them to step S2 for data training again.
[0157] For models with high accuracy, use them as new benchmark anchor points and wait for the next model comparison and update.
[0158] Finally, select the model with the highest correct rate for the engineering deployment model in S6.
[0159] This optimization system automatically generates a version evolution report every 24 hours, including key information such as modal fusion efficiency analysis and resource consumption trend prediction, providing decision-making support for subsequent engineering deployment.
[0160] After contrastive learning, integrate the output results of each previous stage, establish an engineering support system, and import the large model trained from steps S2 to S5 into this system. The key points of this process also include:
[0161] Heterogeneous computing pipeline architecture: Build a hybrid deployment framework of NVIDIA A100 and Ascend 910B, and achieve 92% resource utilization through the H-DAGS task scheduling algorithm
[0162] Low-code development environment: Use a visual module to assemble the page and integrate an automated test tool chain
[0163] Deploy the edge computing framework: Adopt dynamic channel pruning technology to achieve 70FPS real-time inference on Jetson AGX Xavier. The formula is as follows:
[0164]
[0165] Among them, F represents the sparsity, W represents the neural network weight matrix, ‖W‖0 represents the number of non-zero weights, and d in ,d out represents the input / output feature dimension;
[0166] Deploy a lightweight knowledge distillation module: purify knowledge and enhance the model's discrimination ability.
[0167] Through the optimization of the heterogeneous computing architecture of the engineering support system, integrating NVIDIA DGX A100 (FP16 precision) and Ascend 910B (INT8 quantization), and using the H-DAGS task scheduling algorithm to achieve dynamic allocation of computing resources. Experiments show that in the batch processing of 1280×1280 pixel images, the inference efficiency is increased by 40% (from 42FPS to 58.8FPS), and the mixed-precision training throughput is stable at ≥120 samples / sec (3.2 times higher than pure FP32 training). Thanks to the optimized design of the low-code development environment in step S6, the response delay of the dynamic fine-tuning module is <50ms, and the video memory utilization rate is ≥92%.
[0168] An embodiment of the present invention also provides a system for implementing the above-mentioned multi-modal large model continuous evolution method, including:
[0169] A dataset acquisition module, used to build a real-time access channel for multi-source heterogeneous data to update knowledge, and collect target data from the updated database to obtain a dataset;
[0170] A large model preliminary training module, used to train the large model based on the dynamic feedback closed-loop mechanism using the obtained dataset, and self-correct the large model through user feedback, and output the corrected large model and the incremental dataset;
[0171] A large model learning consolidation module, used to implement an anti-forgetting continuous learning scheme for the corrected large model using the obtained incremental dataset to enable the large model to learn and consolidate knowledge;
[0172] A large model fine-tuning module, used to perform knowledge fine-tuning on the large model after learning and consolidation using efficient knowledge fine-tuning technology;
[0173] A large model collaborative optimization module, used to perform multi-modal collaborative optimization on the large model after knowledge fine-tuning based on the above results, compare the learning effects of the optimized large model, and feedback to the previous steps to enable the large model to be further trained and developed.
[0174] The above system interface design includes:
[0175] Data input interface: supports NetCDF / HDF5 format, collects data and enters the dataset;
[0176] Result output interface: Generate vector layers, raster data, etc., and allow import and call in a low-code development environment;
[0177] Visualization module: Assemble pages using the visualization module, support spatio-temporal comparison analysis, and enable efficient display of the output data.
[0178] This system achieves technological breakthroughs through the following innovations, including: constructing a multi-modal data dynamic injection channel and an incremental knowledge graph generation mechanism, integrating satellite remote sensing, UAV aerial photography, and ground sensor data in real time through edge computing nodes, and adopting a knowledge distillation technology based on multi-head attention weight decay to achieve adaptive update of knowledge node confidence and solve the problem of spatio-temporal feature mismatch of static models; creating a dual-path feedback closed-loop training architecture, synchronously executing user interaction correction and incremental dataset generation during the dynamic training process, and reducing the model forgetting rate from 28.4% of traditional methods to 9.7% through elastic weight consolidation and dual memory replay mechanisms; developing a low-rank adaptive fine-tuning technology, embedding a trainable parameter matrix in the Transformer layer, and cooperating with dual-tower contrast learning to improve the model fine-tuning efficiency by 3.2 times; establishing a unified representation space for neural radiance fields, achieving cross-modal feature decoupling through dynamic routing probability allocation, and reducing the pixel alignment error from 1.5 to 0.3 in the farmland boundary recognition task; designing a heterogeneous computing collaborative scheduling engine, dynamically partitioning tasks in a hybrid environment of Ascend 910 and NVIDIA A100, increasing the resource utilization rate from 65% to 89%, and cooperating with a lightweight inference framework at the edge to achieve a real-time processing speed of 112 FPS for UAVs and a small target missed detection rate of 4.1%.
[0179] In this system, an online learning module based on an anti-forgetting continuous learning scheme is constructed in the large model learning consolidation module, adopting an incremental parameter update mechanism, combining gradient accumulation (cumulative step size = 4) and selective parameter freezing (freezing ratio ≥ 60%) of the large model fine-tuning module to achieve a response delay < 50 ms (single fine-tuning time consumption). Through the video memory pooling technology and dynamic memory allocation algorithm (Buddy Memory Allocation) of the storage management module of the engineering support system, the video memory utilization rate is maintained at ≥ 92% all year round in the 8×A100 80GB configuration, supporting concurrent processing of 16 data streams. In the drought monitoring of the Yellow River Basin in 2024, the system only needs 2 minutes and 17 seconds to complete model iteration through the user feedback data in step S2 (the traditional method takes 4.3 hours).
[0180] Relying on the technological breakthrough of the edge intelligent computing framework with an engineering support system, dynamic channel pruning (sparsity ≥ 80%) and quantization-aware training (INT8 precision loss < 1.2%) are adopted to achieve 70FPS real-time inference on the Jetson AGX Xavier edge device, with an energy consumption ratio of 5.3 TOPS / W (4.1 times higher than the unoptimized version). This achievement benefits from the collaborative optimization of the multi-model contrastive learning mechanism and the dynamic channel pruning technology.
[0181] Through the optimized design of the three-level storage system architecture, the implementation of hot-warm-cold data hierarchical management and the LZ4 lossless compression algorithm (compression ratio 2.8:1), the overall storage cost is reduced by 63% (from 10PB of data to 3.7PB). This technological breakthrough effectively supports the large-scale real-time data access requirements of the dynamic knowledge injection module in step S1.
[0182] Experiments show that driven by the automated evaluation-optimization loop in step S2, in the flood monitoring of the Pearl River Basin in 2023, the system realizes the calculation of the flood range updated hourly. The processing time of a single scene image (100km 2 ) is less than 3 minutes, the F1-score of the inundation area change detection reaches 0.91 (with manual annotation as the benchmark), and the missed detection rate of small targets is reduced from 9.8% to 2.1%, providing a minute-level response ability for flood prevention in smart cities. This achievement integrates the advantages of the efficient fine-tuning technology in step S4 and the multi-modal collaborative optimization mechanism in step S5.
[0183] In this embodiment, a real-time data channel is constructed through the dynamic database learning mechanism in step S1 to achieve the efficient access and expert knowledge fusion of multi-source remote sensing data. Combining the label quality control strategy based on scoring and weighting in step S2 effectively improves the data confidence to 97%. Through the closed-loop iterative process of "intelligent interpretation-artificial correction-model fine-tuning" in step S2, the elastic weight consolidation algorithm in step S3 is used to protect key parameters. Combining the adaptive data replay strategy in step S3 (information entropy threshold set H > 2.3), the catastrophic forgetting rate is controlled within 4.2%. The multi-modal collaborative optimization framework in step S5 is innovatively introduced, and the storage overhead is reduced by 63% through the tensor decomposition compression algorithm. The domestic Ascend 910 chip cluster in step S6 is deployed to achieve real-time processing of TB-level data. Finally, the constructed engineering support system integrates the low-code development environment and the distributed collaboration framework in step S6, supports the millimeter-level precision output of 28 types of remote sensing interpretation tasks, and realizes a 70FPS real-time detection efficiency in the agricultural monitoring scenario, providing a full-process industrial application platform for the continuous learning achievement.
[0184] The above are only the preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention. For those skilled in the art, it should be realized that all the solutions obtained by equivalent substitution and obvious changes made by using the content of the specification of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for the continuous evolution of large models based on multi-modalities, characterized in that, It includes the following steps: S1: Build a real-time access channel for multi-source heterogeneous data to update knowledge, and collect target data from the updated database to obtain a data set; S2: Based on the dynamic feedback closed-loop mechanism, use the data set obtained in step S1 to train the large model, and self-correct the large model through user feedback, and output the corrected large model and the incremental data set; S3: Use the incremental data set obtained in step 2 to implement an anti-forgetting continuous learning scheme for the corrected large model so that the large model consolidates knowledge learning; S4: Use the efficient knowledge fine-tuning technology to fine-tune the knowledge of the large model output in S3; S5: Based on the results obtained in S2-S4, perform multi-modal collaborative optimization on the large model output in step S4, compare the learning effects of the optimized large model, and feedback to the previous steps, so that the large model can be further trained and developed.
2. The method for continuous evolution of a large model based on multi-modalities according to claim 1, wherein The specific implementation method in step 1 includes: Build a static authoritative knowledge base, and realize the real-time injection of sensor data into the knowledge base by deploying edge computing nodes to realize the real-time update of knowledge; Based on the data obtained from the updated database, perform node annotation on its terms to obtain a text data set after node annotation, and generate an incremental knowledge graph according to the relationship between nodes; Organize the text data set after node annotation and the incremental knowledge graph into a data set for subsequent steps.
3. The method for continuous evolution of a large model based on multi-modalities according to claim 2, wherein The method for generating an incremental knowledge graph includes: Use the knowledge distillation technology to extract the structured features of the knowledge in the updated database, so as to extract the keywords of the knowledge point content, and create nodes annotated with the keywords; For each node, calculate its semantic similarity with all nodes. When the similarity exceeds the preset threshold, it is considered that there is an association between the two nodes, and the node is connected to the parent node to obtain a knowledge spectrum; For the nodes in the knowledge spectrum, if it is found that the similarity between the nodes is lower than the threshold, it is considered that there is a conflict in the description between the nodes, then start the confidence decay, and update the knowledge graph by reducing the confidence of the old nodes to obtain an incremental knowledge spectrum.
4. The multi-modal-based large model continuous evolution method according to claim 3, wherein The method for extracting keywords includes: Use a pre-trained language model to preliminarily encode the documents in the obtained database above, and convert the text into a vector form; Identify the key information in the vectorized text through the multi-head attention mechanism, and calculate the attention weight of each word; According to the attention weights of each word, calculate its importance score. The specific formula is as follows: Among them, WS(V i ) represents the importance score of each word, W ji represents the similarity degree of two words, WS(V j ) represents the attention weight of each word, and d is the damping coefficient; Finally, select several words with the highest scores as keywords. After obtaining the keywords, create a node Node for each keyword. The node contains the basic information, semantic feature vector and initial confidence value of the keyword.
5. The multimodal-based large model continuous evolution method according to claim 1, characterized in that The implementation method of step 2 includes: Divide the data set obtained in step S1 into a training set and a test set. The large model extracts features and performs semantic parsing on the text data set and the incremental knowledge graph in the training set, and then fuses the obtained text semantic information vector and the graph semantic vector to obtain a prediction result; The effectiveness of the large model is tested using a test set. Based on the gap between the prediction results and the actual data, combined with the feedback results of users, the model parameters of the large model are fine-tuned to obtain the trained large model. During the fine-tuning process, a comprehensive evaluation report of the large model is obtained and output together with the data set as an incremental data set.
6. The method for continuous evolution of a large model based on multi-modalities according to claim 5, wherein The methods for obtaining the prediction results in step 2 include: Use the BERT-base model to initialize the embedding matrix E. The text data set input word sequence is mapped into a vector sequence using the embedding matrix E. Then, multiple groups of convolutional kernels are used to capture semantic features of different granularities of the vector sequence, the feature map is output, and after ReLU activation, max pooling is performed. Finally, an attention mechanism is constructed to strengthen the key features, thereby obtaining the final text feature matrix; For each node in the incremental knowledge graph, information aggregation is performed through a graph attention network to obtain a graph feature matrix; For the text feature matrix and the graph feature matrix, a multi-head attention mechanism is used to calculate the similarity matrix S to obtain a fused feature matrix, which is the prediction result. The specific formula is as follows: S ij = (t i · h j ) / (||t i || · ||h j ||) Among them, S ij represents the relevance between the i-th data sample and the j-th knowledge graph node, which is the prediction result; t i is the i-th text feature vector in the dataset, h j is the graph feature vector obtained from the j-th knowledge graph node.
7. The multimodal-based large model continuous evolution method according to claim 1, wherein The implementation method of step S3 includes: First, based on the elastic weight consolidation mechanism, knowledge consolidation is performed on the large model trained in S2, that is, the large model trained in S2 is repeatedly trained through the following loss function to fine-tune the model parameters: Among them, is the conventional loss function corresponding to the training of the interpretation task, and F i represents the Fisher information matrix of the parameters, represents the old parameter data, θ i , and λ is the regularization coefficient; Secondly, based on the adaptive data replay mechanism, corresponding training and fine-tuning task samples are generated from the data set obtained in S2, and task sample replay operations are performed; Finally, for the large model after elastic weight consolidation and the task samples after data replay, a dual-memory mechanism is used to perform further training.
8. The method for continuous evolution of a large model based on multi-modalities according to claim 1, wherein Step 4 specifically includes: First, for the large model parameters obtained in step S3, low-rank adaptive fine-tuning is used, and they are substituted into the Transformer layer of the trainable low-rank matrix for further training and fine-tuning to obtain the fine-tuned large model; Secondly, for the above fine-tuned model, a dual-tower contrastive learning architecture is constructed to align the knowledge of the model; the dual-tower learning architecture includes a visual encoder and a text encoder. The visual encoder converts the input into a visual feature vector, and the text encoder converts the input into a text description vector; the following loss function is designed, and the data in step S1 and the above fine-tuned large model are input for further training: where s(v i , t j ) is the cosine similarity, τ = 0.07 is the temperature coefficient, and L align is the loss of the model; Through the above steps, a large model with knowledge fine-tuning is obtained.
9. The multimodal-based large model continuous evolution method according to claim 1, characterized in that The implementation method of step S5 is as follows: S5.1, through the neural radiance field, perform unified representation on the large model obtained in step S4; that is, perform position encoding on the spatial coordinates, then input the encoded coordinates into the large model in S4, establish a graph attention weight matrix for the knowledge graph obtained in S1, and perform convolution operations on it and the semantic feature matrix of the large model obtained in S4 respectively. Then, the convolved features are concatenated, and the concatenated modal features are mapped to a unified representation space. The formula is as follows: σ,c = MLP θ (x,d) Among them, MLP is the radiation field MLP parameter matrix, representing the mapping from spatial coordinates to the radiation field parameter space, where \(x\in\mathbb{R}\). 3 is the spatial coordinate, \(d\in\mathbb{R}\). 3 is the viewing direction, \(\sigma\) represents the volume density, \(c\) is the radiation color, and the features of each modality are mapped to the radiation field parameter space through the cross-modal projection network. S5.2, Based on dynamic network routing, decompose the multi-modal unified representation generated in the S5.1 stage into mutually orthogonal semantic subspaces, where each subspace corresponds to a specific modality. For a specific model, aggregate the outputs of each path expert network through routing probabilities to obtain a spliced decoupled feature matrix; S5.3, Construct corresponding anchor points for the large model. According to the generated anchor points, design a dual contrast loss formula to compare the feature space of the large model, thereby comparing the baseline model extracted from S2 and the optimized model obtained from S4. For each dimension of the decoupled feature matrix obtained in S5.2, perform contrast optimization, and feedback the obtained contrast loss to the previous steps to enable further training and development of the large model.
10. A system for implementing the multi-modal based large model continuous evolution method according to any one of claims 1-9, characterized in that, Including: A dataset acquisition module for constructing a real-time access channel for multi-source heterogeneous data to update knowledge, and collecting target data from the updated database to obtain a dataset; A large model preliminary training module for training the large model based on the obtained dataset using a dynamic feedback closed-loop mechanism, and self-correcting the large model through user feedback, and outputting the corrected large model and an incremental dataset; A large model learning consolidation module for implementing an anti-forgetting continuous learning scheme for the corrected large model using the obtained incremental dataset to enable the large model to learn and consolidate knowledge; A large model fine-tuning module for fine-tuning the knowledge of the large model after learning consolidation using efficient knowledge fine-tuning techniques; A large model collaborative optimization module for performing multi-modal collaborative optimization on the large model after knowledge fine-tuning based on the above results, comparing the learning effects of the optimized large model, and feeding back to the previous steps to enable further training and development of the large model.
Citation Information
Cited By
Cross-modal agent base deployment system and method based on AI large model
CN121052282A
PCB defect detection method and system based on multi-mode deep learning
CN121211210A
Large language model continuous learning method based on key value pair replay
CN122047391A
A large language model continuous learning method based on key-value pair replay
CN122047391B
Unmanned aerial vehicle federated fault diagnosis method fusing parameter coordination and collaborative aggregation
CN122413261A