Method and system for continuously updating visual large model

Through the combination of visual Mamba model and bidirectional state space model, combined with position information and Prompt template, the problem of catastrophic forgetting in continuous learning is solved, the generalization ability and adaptability of the model are significantly improved, and the stability and reliability of the model in a variable environment are achieved.

CN120162731APending Publication Date: 2025-06-17SUZHOU UNIV OF SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510148635.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art results in degradation in model performance due to catastrophic forgetting in continuous learning tasks.

Method used

The visual Mamba model is used to extract feature representations, and combined with position information and Prompt templates, a two-way state space model is built to simulate the timing changes and dynamic characteristics of visual data, and feature fusion is performed through multimodal pre-trained models, and an online learning mechanism is introduced to update model parameters in real time.

Benefits of technology

It significantly enhances the model's understanding of visual data and feature expression ability, improves the generalization ability and adaptability of the model, reduces catastrophic forgetting problems, enhances the robustness of outliers and noise, and ensures stability and reliability in variable environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162731A_ABST
    Figure CN120162731A_ABST
Patent Text Reader

Abstract

The invention relates to a continuous updating method and system for a visual large model, and belongs to the technical field of artificial intelligence. Comprising the following steps: extracting visual data features by using a Mamba model to obtain a first feature representation, and embedding position information to obtain a second feature representation; constructing an image block sequence according to the second feature representation, and combining the first feature representation and the second feature representation to construct a bidirectional state space model and a knowledge state; constructing a Prompt template and introducing the Prompt template into a visual Mangbar model to obtain first enhanced feature representation; inputting the image block sequence into the bidirectional state space model, simulating the time sequence change and dynamic characteristics of the visual data, and combining the first enhanced feature representation to obtain a second enhanced feature representation; inputting the feature representation into a multi-modal pre-training model for fusion to obtain fusion data; and through an online learning mechanism, processing new visual data in real time, and updating model parameters. According to the method, the problem of model performance reduction caused by disastrous forgetting in a continuous learning task is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a method and system for continuously updating a large visual model. Background Technique

[0002] In recent years, continuous learning or incremental learning has become a research direction that has received much attention in the field of artificial intelligence. This method processes tasks containing different knowledge points through a sequential learning mode, and can retain the ability to learn new information while simulating human beings' ability to retain old knowledge. With the in-depth research, various methods have emerged to address the challenges in continuous learning, especially the catastrophic forgetting problem, that is, the phenomenon of losing the knowledge of old tasks when learning new tasks. Among them, the knowledge distillation technique is particularly prominent, especially the implementation of distillation between the prediction outputs of the current model and the old model in terms of category. On this basis, the PODNet method further innovates by performing distillation on all intermediate layer feature maps of the network backbone, thereby retaining more key information of old tasks. With the emergence of the Vision Transformer (ViT for short), it has shown great potential in tasks such as image classification and class incremental learning. DyTox improves the adaptability of the model to incremental learning tasks by replacing the traditional convolutional neural network (CNN) backbone with ViT and introducing task-related embeddings.

[0003] However, these methods have alleviated the catastrophic forgetting problem to a certain extent, but also brought about a decline in the overall performance of the model. In the field of knowledge distillation, the synthetic data of the data-free distillation method may not fully represent the distribution and characteristics of the original data, resulting in the distortion or loss of knowledge. Quantization distillation may lead to knowledge loss and accuracy degradation, especially in extreme quantization (such as binary quantization) scenarios. For ViT, its data requirement is relatively large because, compared with CNN, the inductive bias ability of the self-attention mechanism of ViT is weaker and more data is needed to automatically learn these assumptions. Summary of the Invention

[0004] To this end, the technical problem to be solved by the present invention is to overcome the decline in model performance caused by catastrophic forgetting in the prior art in continuous learning tasks.

[0005] In a first aspect, to solve the above technical problem, the present invention provides a method for continuously updating a large visual model, including:

[0006] Obtain visual data, extract features from the visual data by using a visual mamba model to obtain a first feature representation; embed position information in the first feature representation to obtain a second feature representation; construct an image patch sequence according to the second feature representation;

[0007] Construct a bidirectional state space model and a knowledge state based on the first feature representation and the second feature representation.

[0008] Construct a Prompt template according to the knowledge state, and introduce the Prompt template into the Visual Mamba model to obtain a first enhanced feature representation; wherein, the Prompt template dynamically updates the parameters in the Visual Mamba model; input the image patch sequence into the bidirectional state space model to simulate the temporal changes and dynamic characteristics of the visual data.

[0009] Obtain a second enhanced feature representation based on the temporal changes, the dynamic characteristic input, and the first enhanced feature representation.

[0010] Input the second enhanced feature representation into a multi-modal pre-trained model for feature fusion to obtain fused data.

[0011] Based on the fused data, the Visual Mamba model, and the bidirectional state space model, introduce an online learning mechanism to process new visual data in real time; update the parameters of the Visual Mamba model and the bidirectional state space model according to the new visual data.

[0012] In one embodiment of the present invention, the bidirectional state space model updates the knowledge state through a state transition equation and an observation equation.

[0013] In one embodiment of the present invention, introducing the Prompt template into the Visual Mamba model includes: dynamically generating prompt information using a prompt pool method, and the prompt information guides the Visual Mamba model to focus on the key features of the visual data; using dual prompts to prompt the visual data and text questions.

[0014] In one embodiment of the present invention, inputting the second enhanced feature representation into a multi-modal pre-trained model for feature fusion includes using the knowledge of the multi-modal pre-trained model to enhance the representation ability of the Visual Mamba model.

[0015] In one embodiment of the present invention, the online learning mechanism includes using an energy-guided discovery method to discover new categories. The energy-guided discovery method enhances the variance of the feature vectors of unseen data to obtain enhanced feature vectors. The calculation formula of the enhanced feature vectors is:

[0016]

[0017] Wherein, is the enhanced feature vector, is the mean of the unseen feature vectors, and σ uis the variance of the unseen feature vector, N(·) represents the normal distribution, K is the number of enhanced feature vectors, is the unseen data, is the estimated or predicted value of the original feature vector, and d is the dimension of the feature vector.

[0018] In one embodiment of the present invention, the energy-guided discovery method includes an energy-based contrastive loss, and the calculation method of the contrastive loss is:

[0019]

[0020] where L ec is the contrastive loss, includes seen and unseen data, g old and g new respectively represent the nodes of the known class and the newly discovered class in the online classifier, f on (x n ) represents the model output, x n represents the nth data point extracted, E(·) is the energy function; the calculation formula of the energy function is:

[0021]

[0022] where Y is the label set, |Y| is the number of all possible labels, g i (·) is a function that maps the model output to a numerical value, f(x) is the model output, i is the class number, and g represents the node.

[0023] In one embodiment of the present invention, during the process of using the Vision Mamba model to extract features from the visual data, it includes classifying, detecting, or segmenting visual tasks according to the first feature representation.

[0024] Second, to solve the above technical problems, the present invention provides a visual large model continuous update system, including:

[0025] A feature extraction module, configured to obtain visual data, use the Vision Mamba model to extract features from the visual data to obtain a first feature representation; embed position information in the first feature representation to obtain a second feature representation; construct an image patch sequence according to the second feature representation;

[0026] A model construction module, configured to construct a bidirectional state space model and a knowledge state according to the first feature representation and the second feature representation;

[0027] A feature enhancement module, configured to construct a Prompt template according to the knowledge state, introduce the Prompt template into the Visual Mamba model to obtain a first enhanced feature representation; wherein, the Prompt template dynamically updates the parameters in the Visual Mamba model; input the image patch sequence into the bidirectional state space model to simulate the temporal changes and dynamic characteristics of the visual data; based on the temporal changes, the dynamic characteristic input and the first enhanced feature representation, obtain a second enhanced feature representation;

[0028] A data fusion module, configured to input the second enhanced feature representation into a multi-modal pre-trained model for feature fusion to obtain fused data;

[0029] A real-time update module, configured to introduce an online learning mechanism based on the fused data, the Visual Mamba model and the bidirectional state space model to process new visual data in real time; update the parameters of the Visual Mamba model and the bidirectional state space model according to the new visual data.

[0030] In a third aspect, to solve the above technical problems, the present invention provides a computer program product, which includes a computer program, and when the computer program runs, the above method for continuously updating a visual large model is executed.

[0031] In a fourth aspect, to solve the above technical problems, the present invention provides an electronic device, including the above system for continuously updating a visual large model.

[0032] The above technical solutions of the present invention have the following beneficial effects compared with the prior art:

[0033] (1) For the method and system for continuously updating a visual large model of the present invention, the Visual Mamba model is used to extract feature representations, and the position information and the Prompt template are fused, significantly enhancing the model's understanding depth and feature expression ability of visual data. By introducing the bidirectional state space model, the temporal changes and dynamic characteristics of visual data can be accurately simulated, effectively capturing the temporal correlation in the image sequence. Further, the model has the ability to dynamically adjust and enhance the feature representation according to the dynamic input, and the enhanced feature representation is sent into the multi-modal pre-trained model for feature fusion, outputting more comprehensive fused data, greatly improving the generalization ability and adaptability of the model.

[0034] (2) The present invention introduces an online learning mechanism, enabling the model to process new visual data in real time and update model parameters based on the new data, flexibly coping with various scenarios such as domain increment, class increment, and task increment, enhancing the model's continuous learning ability and adaptability. Dynamically updating parameters reduces the common catastrophic forgetting problem in online learning. At the same time, the model improves its robustness to outliers and noise by continuously updating to adapt to new data, ensuring stability and reliability in a changing environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in combination with the drawings, where

[0036] Figure 1 is a flowchart of a method for continuously updating a large visual model in a preferred embodiment of the present invention;

[0037] Figure 2 is a structural diagram of a Visual Mamba model in a preferred embodiment of the present invention;

[0038] Figure 3 is a schematic structural diagram of extracting Prompts in a preferred embodiment of the present invention;

[0039] Figure 4 is a schematic diagram of a multi-modal pre-training model in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The following further illustrates the present invention in combination with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the examples given are not intended to limit the present invention.

[0041] Embodiment 1

[0042] Referring to Figure 1 as shown, an embodiment of the present invention provides a method for continuously updating a large visual model, including:

[0043] Obtain visual data, use the Visual Mamba model to extract features from the visual data to obtain a first feature representation; embed position information in the first feature representation to obtain a second feature representation; construct an image patch sequence according to the second feature representation;

[0044] Construct a bidirectional state space model and a knowledge state according to the first feature representation and the second feature representation;

[0045] Construct a Prompt template according to the knowledge state, introduce the Prompt template into the Vision Mamba model to obtain the first enhanced feature representation; wherein, the Prompt template dynamically updates the parameters in the Vision Mamba model; input the image patch sequence into the bidirectional state space model to simulate the temporal changes and dynamic characteristics of visual data;

[0046] Based on the temporal changes, dynamic characteristic input, and the first enhanced feature representation, obtain the second enhanced feature representation;

[0047] Input the second enhanced feature representation into the multi-modal pre-trained model for feature fusion to obtain the fused data;

[0048] Based on the fused data, the Vision Mamba model, and the bidirectional state space model, introduce an online learning mechanism to process new visual data in real time; update the parameters of the Vision Mamba model and the bidirectional state space model according to the new visual data.

[0049] A method for continuously updating a vision large model provided by an embodiment of the present invention uses the Vision Mamba model to extract feature representations, and combines position information and the Prompt template, effectively improving the model's understanding ability and feature expression ability for visual data. Through the bidirectional state space model, the model can simulate the temporal changes and dynamic characteristics of visual data and capture the temporal dependencies in the image sequence. In addition, the model can capture the dynamic changes of visual data according to the dynamic characteristics and the first enhanced feature representation, enhancing the adaptability to complex scenes. Inputting the second enhanced feature representation into the multi-modal pre-trained model for feature fusion to obtain more comprehensive fused data, thereby enhancing the generalization ability of the model. Introducing an online learning mechanism enables the model to process new visual data in real time and update the model parameters according to the new data, flexibly coping with various scenarios such as domain increment, class increment, and task increment, and improving the model's continuous learning ability and adaptability. Dynamically updating the parameters reduces the common catastrophic forgetting problem in online learning, that is, forgetting old tasks when learning new tasks. The model continuously updates to adapt to new data, improving the robustness to outliers and noise, and ensuring stability and reliability in a changing environment. Therefore, this method not only effectively solves the forgetting problem in continuous learning but also significantly improves the overall performance of the model.

[0050] Specifically, Vision Mamba is developed from Text Mamba, which is widely used in natural language processing (NLP) tasks. It is a new type of deep learning model that effectively processes image data and captures the global context information of images through a Bidirectional State Space Model (SSM for short) and positional embedding technology. This meets the requirements of visual tasks in a continuous learning environment, enabling continuous learning from new data while retaining previously learned knowledge. Further, Mamba introduces a selective scanning algorithm that can dynamically adjust the model's computational process according to the input data, focusing only on important parts of the input. This algorithm utilizes the parallel computing power of the GPU to accelerate the model's training and inference processes through parallel scanning operations; at the same time, Mamba reduces the data transfer between GPU memory and video memory through kernel fusion and recomputation techniques, thereby improving computational efficiency; Mamba simplifies the traditional SSM architecture, removing the attention mechanism and multi-layer perceptron (MLP) modules, making the model lighter and faster, and easier to parallelize and utilize computing resources on the GPU; moreover, Mamba achieves a computational complexity that scales linearly with the sequence length through a selective state space and hardware-aware algorithm (i.e., its computational time complexity depends only on itself, and loop iterations and other computations are constant-level, with a time complexity of O(L), while the traditional Transformer has a time complexity of O(L 2 ))), meaning that Mamba can more efficiently process large-scale long-sequence data on the GPU.

[0051] Specifically, for the structure of the Vision Mamba model (Vim model for short), reference can be made to Figure 2 . The Vision Mamba model is a visual processing model based on the state space model. It is an extension of the Text Mamba model used in natural language processing and effectively processes sequence data through a selective state space. In this embodiment, the Vision Mamba model first receives raw visual data, such as images, and then efficiently extracts the key features of the images through a series of transformation operations, including convolution, normalization, and attention mechanisms, to obtain the first feature representation. To deeply capture the spatial relationships between various parts of the image, the model further embeds positional information and constructs an image patch sequence, which enables the model to more comprehensively understand the global structure of the image, thereby enhancing the overall ability to grasp the image content. In this embodiment, the process of using the Vision Mamba model to extract features from visual data includes classifying, detecting, or segmenting visual tasks based on the first feature representation. The specific steps for the Vision Mamba model to process visual data are as follows:

[0052] S201. In the initial stage, the input image is first segmented into non-overlapping image patches of a settable fixed size (the term "patches" can also be translated as "image patches"). For example, an image of 224×224 pixels can be segmented into patches of 16×16 pixels, and these patches are converted into a series of vectors through a linear projection layer to prepare for subsequent processing.

[0053] S202. For visual processing, a flatten operation is performed on the 2D image (the flatten operation refers to converting the 2D image data into a one-dimensional array), and each patch is mapped to a feature space of a fixed dimension using a learnable linear transformation. That is, the 2D image is converted into a flattened 2D patch where (H, W) is the size (length × width) of the input image, C is the number of channels, P is the size of the image patch, and J is the number of patches into which the image is divided.

[0054] S203. Project l P linearly onto a vector of dimension D, add positional encoding, and embed positional information in the image.

[0055] Specifically, the positional encoding is usually a vector of dimension D, where D is the embedding dimension of the model. For each dimension d, the positional encoding uses sine and cosine functions, and the calculation formula is:

[0056]

[0057] where PaE (pos,d) represents the value of the positional encoding vector at position pos and dimension d, and the value of d is 2n or 2n + 1. pos is the position index in the sequence, n is the dimension index, and D is the dimension of the positional encoding.

[0058] Furthermore, the token sequence L0 after linear projection and positional embedding has the following mathematical expression:

[0059]

[0060] where t cls represents the class token of the entire patch sequence, is the Jth patch of t, J is the number of patches into which the image is divided and is an integer greater than 0, and is the learnable projection matrix.

[0061] S204. The token sequence (L l-1) Sent to the l-th layer of the Vision Mamba encoder (Vim encoder) and obtain the output L l . Normalize the output class token to obtain the normalized result y and feed it into the MLP head to get the final prediction The specific calculation formula is as follows:

[0062] L l = Vim(L l-1 ) + L l-1 ;

[0063]

[0064] where Vim(·) is the Vision Mamba Block, Norm(·) is the normalization function, and MLP(·) is the multi-layer perceptron function.

[0065] Specifically, the image input is first normalized to stabilize the training process and accelerate model convergence. Subsequently, the encoder performs two linear projections on the normalized token sequence to generate new sequences x and z. Each token in the model undergoes one-dimensional convolutional processing in the forward and backward directions, using Forward Conv1d and Backward Conv1d respectively. In this process, the embodiments of the present invention preferably use the SiLU activation function to enhance the non-linear expression ability of the model. The processed results pass through a bidirectional state space model, which helps capture long-range dependencies and compress the sequence data, while learning the dynamic characteristics of the sequence through linear transformation.

[0066] Furthermore, after processing the forward and backward sequences, a gating mechanism is used to combine the outputs of the forward and backward directions with the residual connection of the input sequence to obtain the final token sequence. This design of residual connection helps alleviate the problem of gradient disappearance, thereby maintaining the stability of the gradient signal during the training process. Subsequently, after passing through a linear layer and performing a residual connection (i.e., addition) with the original input, the final sequence token is obtained. This processed token sequence is then fed into the MLP, where the MLP is a feed-forward neural network that learns the non-linear representation of the data through multiple fully connected layers (also known as dense layers). The head of the MLP is specifically used for classification prediction of the class token. Such a design not only enhances the model's ability to process sequence information but also improves the accuracy of the classification task.

[0067] Furthermore, for an n-layer MLP, its calculation can be expressed as:

[0068] h1 = f(W1x + b1);

[0069] h2 = f(W2h1 + b2);

[0070]

[0071] where x is the input feature vector, h1, h2, and h n are the output of the hidden layers of the first, second, and nth layers respectively, is the prediction result of the output layer, W1, W2, W n are the weight matrices, b1, b2, b n are the bias vectors, and f(·) represents the activation function.

[0072] Through steps S201 - S204, the stability of the model training process is promoted, the convergence is accelerated, and the performance of image classification and other visual tasks is improved.

[0073] Specifically, the embodiment of the present invention constructs a continuous learning framework based on a bidirectional state space model, which integrates multiple strategies such as replay, regularization, and parameter isolation, aiming to effectively solve the continuous learning problem. For the bidirectional state space model, it is a mathematical model used to describe the behavior of a dynamic system, representing the evolution of the internal state of the system through a set of first-order differential equations (continuous-time system) or difference equations (discrete-time system), and using another set of equations to describe the relationship between the system state and the output. In the bidirectional state space model, the system is defined by two sets of equations, namely: the state equation and the observation equation, and the knowledge state can be updated through the state transition equation and the observation equation.

[0074] The state equation y′(t) describes the evolution of the internal state of the system over time, and its mathematical expression is:

[0075] y′(t) = Py(t) + Qu(t) + ω(t);

[0076] where y(t) is the state vector, P is the state transition matrix, Q is the control input matrix, u(t) is the control vector, and ω(t) is the process noise.

[0077] The observation equation describes how the output of the system depends on the system state and the control input, and its mathematical expression is:

[0078] z(t) = Rx(t) + Hu(t) + e(t);

[0079] where z(t) is the output vector or the observation vector, R is the output matrix, H is the direct transmission matrix, e(t) is the measurement noise, and x(t) is the system state.

[0080] The embodiment of the present invention introducing the bidirectional state space model enables the model to understand and predict the dynamic changes of visual data in a structured manner.

[0081] Specifically, introducing the Prompt template into the Visual Mamba model includes: dynamically generating prompt information using the prompt pool method to guide the Visual Mamba model to focus on the key features of visual data, and these prompt information helps the model to be more accurate when identifying and processing visual content. Further, a dual prompt strategy is adopted to effectively prompt visual data and text questions. Specifically, the system learns two non-overlapping prompt spaces, which respectively encode task-invariant and task-specific instructions, so as to guide the model to focus on the key features of the image and the semantic information of the text at the same time. This fusion strategy enables the model to more comprehensively understand visual content and consider context information when processing visual tasks to optimize the response to specific text questions.

[0082] For this method of using Prompt, its structural schematic diagram can be referred to Figure 3 . Utilize the representative features of the pre-trained model and guide the model to sequentially learn tasks by learning a set of dynamic prompts. These prompts are organized in a key-value shared memory space called the prompt pool, and a query mechanism is used to dynamically find the task-related prompt subset according to the instance input features. The prompt pool is jointly optimized with the supervised loss to ensure that the shared prompt encoding shares knowledge, realizes knowledge transfer, and at the same time maintains the plasticity of the model. The specific implementation steps are as follows:

[0083] S401. Construct a prompt pool, which is defined as Pr = {Pr1, Pr2, …, Pr M}, where is a single prompt, L p is the token length of the prompt, and D is the embedding dimension of the model.

[0084] S402. In order to dynamically select prompts suitable for different inputs, an instance-level prompt query strategy based on key-value pairs is designed in this solution. Each prompt is associated with a learnable key The set of all keys is Given the input x, the input x is encoded into the same dimension as the key through the query function q(x), and then the cosine similarity is used as the matching function γ(·) to select the top keys that match the query most. The goal of the query is to minimize the following objective:

[0085]

[0086] where, KEY x represents the subset of the top keys selected for x, s i is the index i selected from the set , is the total number of all keys in the hint pool.

[0087] S403, using the Prompt method to adopt optimization target training, the training target is to minimize the end-to-end training loss function, including classification loss and key selection loss. In each training step, After prompts, the adaptive embedding feature x p It is input into the rest of the pre-trained model fr and finally the classifier g parameterized by φ φ , the result prediction is performed. The specific calculation formula is:

[0088]

[0089] in, represents the average of the output hidden vectors corresponding to the prompt position, g φ is the classifier, L is the softmax cross entropy loss, λ is the weight to balance the two loss terms, Ε (x,y) is the expected loss function, over the entire dataset The average of the differences between the model's predicted output and the actual output. Here, x represents the input data and y represents the corresponding true label or output.

[0090] S404: After the prediction is completed, the contrast loss is calculated to update the Prompt template. The loss function aims to shorten the distance between positive sample pairs and increase the distance between negative sample pairs. The loss function is calculated as:

[0091]

[0092] Among them, que is the query vector; key + is the positive sample key vector; key i are all key vectors (including positive and negative samples); τ is the temperature parameter used to control the smoothness of the distribution.

[0093] Specifically, the second enhanced feature representation is input into the multi-modal pre-trained model to achieve feature fusion and generate fused data. This process fully utilizes the rich knowledge contained in the multi-modal pre-trained model, significantly enhancing the representation ability of the Visual Mamba model. Through cross-modal knowledge transfer, the generalization ability of the model to new tasks is enhanced, enabling it to handle more complex tasks, such as cross-modal retrieval and visual question answering. For the multi-modal pre-trained model, by combining various types of data such as text, images, and videos, it can better understand cross-modal context information, enabling the model to capture more abundant information when processing long videos or high-resolution images, thereby enhancing the understanding ability of continuous scenes and the comprehensive understanding of data. During the process of continuous learning, the multi-modal pre-trained model usually has better long-term memory ability. The model can remember more previously learned information and use this information for new, unseen tasks or data. Therefore, referring to Figure 4 , in the embodiment of the present invention, a contrastive language-image pre-trained model is preferably used as a visual encoder in the multi-modal pre-trained model to encode visual information, and an MLP is used to map visual features to the Mamba structure to utilize its efficient long-sequence learning ability.

[0094] Specifically, the core idea of the text-image multi-modal large model is to map images and text into the same vector space, enabling the model to directly calculate the similarity between images and text in the vector space without additional intermediate representations. The text-image multi-modal large model includes two main components: a text encoder and an image encoder. In existing research, the text encoder usually uses the Transformer architecture, while the image encoder can be ResNet or ViT. Due to the efficient computational performance, causal propagation, memory and speed efficiency, and wide applicability of the Mamba architecture and its Vim, they are preferably used as the encoders for text and images in the multi-modal pre-trained model. In the embodiment of the present invention, the text encoder preferably uses the Mamba architecture, and the image encoder preferably uses Visual Mamba.

[0095] Furthermore, Online Continual Learning (OCL) is a machine learning paradigm that is crucial for building intelligent systems capable of adapting to ever-changing environments. It aims to simulate the natural learning process of humans where they do not forget old knowledge when learning new knowledge. In an online learning environment, data is usually presented to the model gradually in the form of a data stream, rather than providing the entire dataset at once. Online learning does not assume prior knowledge of the task boundaries. The model can learn new data in real-time without waiting to collect a large amount of data before training and needs to learn without an explicit task-switching signal. In particular, in the field of online continual learning, Novel Category Discovery (NCD) is a key and challenging problem. Since it involves unsupervised learning, novel categories are unknown during the training phase, requiring the model to be able to identify these novel categories without explicit label guidance. The embodiments of the present invention effectively solve this difficult problem by innovatively combining the preferred Centroid Method and the Energy-Guided Discovery Method. The specific content of the Centroid Method and the Energy-Guided Discovery Method is as follows:

[0096] (1) Centroid Method

[0097] When receiving the input x obtained from the previous step, the Visual Mamba model extracts patches of x according to the given patch size and stride. Then, it uses the Visual Mamba encoder (i.e., the above-mentioned multi-layer Vim(·) block) for encoding and extracting features to obtain the feature vector h. Next, it uses an online clustering algorithm (such as KNN, KMeans++) to process the features in sequence. The specific process is as follows: Given multiple feature vectors h, these feature vectors are clustered to obtain multiple clusters, and the clustering center c of each cluster is calculated respectively. The most similar cluster c is found by mapping the clustering center to the nearest cluster centroid in the existing centroid set j , and its calculation formula is:

[0098]

[0099] where d(·,·) represents any distance metric (e.g., using cosine distance), z is the centroid of the category, and C is the set of all clustering centers obtained after clustering. If c j is less than the threshold from r, it is classified as an old category. And when the cluster c j obtained by clustering the feature vectors is assigned to a category, the category centroid is updated as follows:

[0100]

[0101] Among them, is the centroid of the updated new category, and α is the centroid learning rate.

[0102] The original patch p and the selected centroid are temporarily stored in the auxiliary memory M j . The centroid learning rate α controls the degree of influence of a single input on the centroid. If the distance between all category centroids z and their nearest clustering centroid c j is greater than the novelty detection threshold then a new category is created, and c j becomes the centroid of the new category. At the same time, the centroid is updated synchronously as the training progresses to avoid concept drift. The model recalculates each centroid using the j examples in the centroid memory M . The specific calculation formula is:

[0103]

[0104] where F φ (·) represents the output function for extracting features. It should be noted that this output function can be the above-mentioned multi-layer Vim(·) block.

[0105] (2) Energy-guided discovery method

[0106] First, the data needs to be divided into known and unknown categories, and the offline model θ off (·) = g off (g off (·)) is used to obtain the unnormalized predicted value z t of the model output (the unnormalized predicted value can be simply referred to as logits), and these logits are used to calculate the energy score. The specific calculation formula is:

[0107]

[0108] where z t is the logits calculated for the batch data B off using the offline model θ t , e t is the set of energy scores, is the i-th energy score, and y l is the set of labels of known categories.

[0109] Secondly, the unknown data needs to be divided into seen and unseen categories. To improve the clustering accuracy of unseen data, the embodiment of the present invention adopts a feature enhancement method based on variance. The specific method is:

[0110] The feature vector of the unseen data is enhanced by variance, which is mathematically expressed as: Among them, the enhanced feature vector has the following calculation formula:

[0111]

[0112] where is the enhanced feature vector, is the mean of the unseen feature vectors, σ u is the variance of the unseen feature vectors, N(·) represents the normal distribution, K is the number of enhanced feature vectors, is the unseen data, is the estimated or predicted value of the original feature vector, and d is the dimension of the feature vector, that is, the number of elements in each feature vector.

[0113] Specifically, in order to effectively absorb new knowledge while retaining old knowledge in the online learning environment, the embodiments of the present invention adopt an efficient parameter adjustment strategy and an energy-based contrast loss method to achieve continuous update and maintenance of knowledge. Among them, the energy function E(f(x); g) of the energy-based contrast loss method has the following calculation formula:

[0114]

[0115] where Y is the label set, that is, all possible output categories, |Y| is the number of all possible labels, g i (·) is a function that maps the output f(x) of the model to a numerical value, which represents the degree of matching between the model output and the specific class number i, and g represents the node.

[0116] The energy-based contrast loss L ec has the following calculation formula:

[0117]

[0118] where includes seen and unseen data, g old and g new respectively represent the nodes in the online classifiers of the known class and the newly discovered class, f on (x n ) represents the model output, x n represents the nth data point extracted from the data set .

[0119] Furthermore, the total loss function L inc combines the cross-entropy loss and the energy-based contrast loss L ec , and the specific calculation formula is:

[0120] L inc = Lce +L ec ;

[0121] where L ce is the cross-entropy loss using pseudo-labeled data.

[0122] To further illustrate the superiority of the embodiments of the present invention, comparative experiments are conducted on different backbones in the ImageNet dataset.

[0123] Specifically, the embodiments of the present invention preferably use the CIFAR-10, CIFAR-100, and ImageNet datasets as benchmarks to evaluate the robustness of the classification model. The specific data is shown in Table 1. In Table 1, Convnets represents convolutional neural networks, including ResNet-18, ResNet-50, ResNet-101, ResNet-152, ResNeXt50-32x4d, and RegNetY-4GF; Transformers represents transformers, including ViT-B / 16, ViT-L / 16, DeiT-Ti, DeiT-S, and DeiT-B; SSMs represents bidirectional state space models, including the method-Ti and the method-S of the present invention.

[0124] Furthermore, Table 1 shows different architectures of the Vision Mamba model. By reducing the number of model parameters, they effectively reduce the storage and computational requirements and reduce the risk of overfitting. Fewer parameters not only mean less computational resources are needed during training and inference, accelerating the model's learning speed and prediction ability, but also achieve an accuracy comparable to that of models with a larger number of parameters. In addition, the smaller number of parameters and optimized computational process enable the Vision Mamba model to achieve real-time or near-real-time processing speed, making it suitable for application scenarios that require fast response.

[0125] Table 1 Comparison results of different backbones on the ImageNet dataset

[0126]

[0127] The embodiments of the present invention provide a method for continuously updating a large vision model based on the Mamba framework. The Vision Mamba model is used to process visual data, and position information is embedded to construct an image patch sequence; an SSM is used to construct a continuous learning framework, and a dynamic system is defined to simulate the temporal changes of visual data; the Prompt method is combined to achieve dynamic update of model parameters; combined with a multi-modal pre-trained large model to solve complex continuous learning problems. Therefore, the present invention constructs a continuous learning framework for the Vision Mamba model and SSM, effectively solves the catastrophic forgetting problem, improves the generalization and adaptation ability of the model, introduces an online learning mechanism, enables the model to process new visual data in real time and update parameters, adapts to various incremental scenarios, and obtains significant performance improvements.

[0128] Embodiment 2

[0129] Based on the same inventive concept, this embodiment provides a visual large model continuous update system, and the principle of solving problems is similar to that of a visual large model continuous update method provided in Embodiment 1, and the repeated parts will not be elaborated.

[0130] This embodiment provides a visual large model continuous update system, including:

[0131] A feature extraction module, configured to obtain visual data, extract features from the visual data by using a visual mamba model to obtain a first feature representation; embed position information in the first feature representation to obtain a second feature representation; and construct an image patch sequence according to the second feature representation;

[0132] A model construction module, configured to construct a bidirectional state space model and a knowledge state according to the first feature representation and the second feature representation;

[0133] A feature enhancement module, configured to construct a Prompt template according to the knowledge state, introduce the Prompt template into the visual mamba model to obtain a first enhanced feature representation; wherein, the Prompt template dynamically updates the parameters in the visual mamba model; input the image patch sequence into the bidirectional state space model to simulate the temporal variation and dynamic characteristics of the visual data; and obtain a second enhanced feature representation based on the temporal variation, dynamic characteristic input, and the first enhanced feature representation;

[0134] A data fusion module, configured to input the second enhanced feature representation into a multi-modal pre-trained model for feature fusion to obtain fusion data;

[0135] A real-time update module, configured to introduce an online learning mechanism based on the fusion data, the visual mamba model, and the bidirectional state space model to process new visual data in real time; and update the parameters of the visual mamba model and the bidirectional state space model according to the new visual data.

[0136] Embodiment 3

[0137] This embodiment provides a computer program product, which includes a computer program that, when running, causes the visual large model continuous update method described in Embodiment 1 to be executed.

[0138] Embodiment 4

[0139] This embodiment provides an electronic device, including the visual large model continuous update system described in Embodiment 2.

[0140] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0141] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0144] Obviously, the above embodiments are only examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for continuously updating a large visual model, characterized in that: include: Acquire visual data, and perform feature extraction on the visual data using a visual mamba model to obtain a first feature representation; embedding position information in the first feature representation to obtain a second feature representation; constructing an image block sequence according to the second feature representation; constructing a bidirectional state space model and a knowledge state according to the first feature representation and the second feature representation; Constructing a Prompt template according to the knowledge state, and introducing the Prompt template into the visual mamba model to obtain a first enhanced feature representation; wherein the Prompt template dynamically updates the parameters in the visual mamba model; inputting the image block sequence into the bidirectional state space model to simulate the temporal changes and dynamic characteristics of the visual data; Obtaining a second enhanced feature representation based on the time series change, the dynamic characteristic input, and the first enhanced feature representation; Inputting the second enhanced feature representation into a multimodal pre-trained model for feature fusion to obtain fused data; Based on the fused data, the visual mamba model and the bidirectional state-space model, an online learning mechanism is introduced to process new visual data in real time; and the parameters of the visual mamba model and the bidirectional state-space model are updated according to the new visual data.

2. A method for continuously updating a large visual model according to claim 1, characterized in that: The bidirectional state space model includes updating the knowledge state through state transfer equations and observation equations.

3. A method for continuously updating a large visual model according to claim 1, characterized in that: Introducing the Prompt template into the Visual Mamba model includes: using a prompt pool method to dynamically generate prompt information, wherein the prompt information guides the Visual Mamba model to focus on key features of the visual data; and using double prompts to prompt the visual data and text questions.

4. A method for continuously updating a large visual model according to claim 1, characterized in that: Inputting the second enhanced feature representation into the multimodal pre-trained model for feature fusion includes using the knowledge of the multimodal pre-trained model to enhance the representation capability of the visual mamba model.

5. A method for continuously updating a large visual model according to claim 1, characterized in that: The online learning mechanism includes discovering new categories using an energy-guided discovery method, which performs variance enhancement on feature vectors of unseen data to obtain enhanced feature vectors. The calculation formula of the enhanced feature vectors is: in, is the enhanced feature vector, is the mean of the unseen eigenvectors, σ u is the variance of the unseen eigenvector, N(·) represents the normal distribution, K is the number of enhanced eigenvectors, No data found. is the estimated or predicted value of the original feature vector, and d is the dimension of the feature vector.

6. A method for continuously updating a large visual model according to claim 5, characterized in that: The energy-guided discovery method includes an energy-based contrast loss, and the contrast loss is calculated as follows: Among them, L ec is the contrast loss, Including seen and unseen data, g old and g new Represent the nodes in the online classifier of known categories and discovered new categories, respectively, on (x n ) represents the model output, x n represents the nth data point extracted, E(·) is the energy function; the calculation formula of the energy function is: Among them, Y is the label set, |Y| is the number of all possible labels, and g i (·) is a function that maps the model output to a numerical value, f(x) is the model output, i is the category number, and g represents the node.

7. A method for continuously updating a large visual model according to claim 1, characterized in that: The process of extracting features from the visual data using the visual mamba model includes classifying, detecting or segmenting visual tasks based on the first feature representation.

8. A visual large model continuous updating system, characterized in that: include: A feature extraction module, used to obtain visual data, and perform feature extraction on the visual data using a visual Mamba model to obtain a first feature representation; embedding position information in the first feature representation to obtain a second feature representation; constructing an image block sequence according to the second feature representation; A model building module, used for building a bidirectional state space model and a knowledge state according to the first feature representation and the second feature representation; A feature enhancement module, configured to construct a Prompt template according to the knowledge state, introduce the Prompt template into the visual mamba model, and obtain a first enhanced feature representation; wherein the Prompt template dynamically updates the parameters in the visual mamba model; input the image block sequence into the bidirectional state space model to simulate the temporal changes and dynamic characteristics of the visual data; and obtain a second enhanced feature representation based on the temporal changes, the dynamic characteristics input, and the first enhanced feature representation; A data fusion module, used for inputting the second enhanced feature representation into a multimodal pre-training model for feature fusion to obtain fused data; A real-time update module is used to introduce an online learning mechanism based on the fusion data, the visual mamba model and the bidirectional state-space model to process new visual data in real time; and update the parameters of the visual mamba model and the bidirectional state-space model according to the new visual data.

9. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is run, a method for continuously updating a visual large model as claimed in any one of claims 1 to 7 is executed.

10. An electronic device, characterized in that: Including a visual large model continuous updating system as described in claim 8.

Citation Information

Cited By

  • Robot operation track generation method based on structure perception and knowledge enhancement reasoning

    CN120765961A

  • Visual language model calibration method and device based on consensus perception

    CN121412692A

  • Displacement monitoring compensation method, training method and device based on image spatial-temporal characteristics

    CN121616597A