A multi-modal large model incremental joint training method and system
Patent Information
- Application Number
- CN202611077216.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-11
AI Technical Summary
[0005]因此,本发明提供了一种多模态大模型增量式联合训练方法解决现有技术中难以根据新增知识与历史知识的关联变化动态确定模型调整区域的问题
[0016] The beneficial effects of this invention are as follows: By analyzing the correlation changes between newly added multimodal training data and historical multimodal knowledge, a knowledge evolution path and knowledge evolution mode are constructed, realizing the dynamic characterization of the influence relationship of newly added knowledge and providing a basis for the accurate determination of the model adjustment region; by determining the target adjustment region based on incremental knowledge state information and dynamically generating joint training strategies based on knowledge gain changes, adaptive incremental updates of model parameters are realized, improving the pertinence of incremental training and the efficiency of joint training; by continuously updating historical multimodal knowledge state information using updated model parameters, continuous evolution and dynamic maintenance of model knowledge are realized, enhancing the continuous learning capability of large multimodal models.
Smart Images

Figure CN122734535A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to a method and system for incremental joint training of multimodal large models. Background Technology
[0002] With the development of deep learning and pre-trained models, multimodal large models can uniformly model multiple modal data such as images, text, speech, and video, and are widely used in fields such as intelligent question answering, content generation, and intelligent retrieval. As multimodal data continues to increase, incremental training technology has gradually become an important way for multimodal large models to continuously learn new knowledge. Existing technologies usually use new training data to perform local fine-tuning or joint training of the model in order to achieve continuous updating of model knowledge.
[0003] There is a complex relationship between newly added multimodal training data and historical knowledge. Different newly added knowledge has different effects on the knowledge regions inside the model. Existing incremental joint training methods usually adopt a fixed parameter update strategy and lack a mechanism to dynamically determine the model adjustment region based on the changes in the relationship between newly added knowledge and historical knowledge. This makes it difficult for the joint training strategy to adaptively adjust to different knowledge regions, thus affecting the relevance and efficiency of incremental training. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a multimodal large-scale incremental joint training method to solve the problem in the prior art that it is difficult to dynamically determine the model adjustment region based on the correlation changes between new knowledge and historical knowledge.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides an incremental joint training method for a multimodal large model, comprising: acquiring new multimodal training data to be added to the multimodal large model and preprocessing it to obtain a standardized incremental training dataset; The standardized incremental training dataset is input into the multimodal large model, and incremental multimodal features are obtained through the multimodal feature extraction network. Based on incremental multimodal features and historical multimodal knowledge state information, the correlation changes between newly added multimodal training data and historical multimodal knowledge are analyzed to obtain incremental knowledge state information; Based on the incremental knowledge state information, determine the target adjustment region in the multimodal large model corresponding to the newly added knowledge, and obtain the model adjustment region information; Based on the model adjustment region information, a joint training strategy is generated. Based on the joint training strategy, the multimodal large model is incrementally jointly trained to obtain the updated multimodal model parameters. Update the historical multimodal knowledge state information using the updated multimodal model parameters.
[0007] As a preferred embodiment of the multimodal large model incremental joint training method of the present invention, the standardized incremental training dataset includes, The newly added multimodal training data is cleaned and aligned to obtain preprocessed new multimodal training data. Data associations between different modalities are established. The preprocessed new multimodal training data and the data associations are stored uniformly to obtain a standardized incremental training dataset.
[0008] As a preferred embodiment of the incremental joint training method for multimodal large models described in this invention, the incremental multimodal features include: A multimodal large model is obtained by pre-training using image data, text data, speech data, and video data. A multimodal feature extraction network is obtained by performing feature learning on image data, text data, speech data, and video data. The multimodal feature extraction network in the multimodal large model extracts features from image data, text data, speech data, and video data in the standardized incremental training dataset to obtain visual features, text features, speech features, and video features. By using a multimodal feature extraction network in a large multimodal model to perform feature mapping and semantic association fusion, incremental multimodal features with correlations between different modalities are obtained.
[0009] As a preferred embodiment of the multimodal large model incremental joint training method of the present invention, the incremental knowledge state information includes: Based on incremental multimodal features and historical multimodal knowledge state information, newly added knowledge feature units and historical knowledge association units are determined respectively. Based on the positional changes of newly added knowledge feature units in the visual feature space, text semantic space, and cross-modal association space, establish the knowledge influence relationship between newly added knowledge feature units and historical knowledge association units; Based on the knowledge influence relationship, the propagation direction, propagation range and propagation intensity of the newly added knowledge feature unit acting on the historical knowledge association unit are calculated to determine the knowledge evolution path of the newly added knowledge feature unit affecting the historical multimodal knowledge state information, and to obtain knowledge evolution path information. Based on the knowledge evolution path information, a knowledge evolution relationship is constructed in which the newly added knowledge feature unit acts on the historical multimodal knowledge state information. Based on the association range expansion relationship, association direction migration relationship, association structure reorganization relationship and association relationship replacement relationship of the newly added knowledge feature unit in the historical multimodal knowledge state information, the corresponding knowledge evolution mode information is generated. When a new knowledge feature unit expands the scope of historical knowledge association, the knowledge evolution mode information corresponds to the knowledge expansion state; when a new knowledge feature unit changes the direction of historical knowledge association, the knowledge evolution mode information corresponds to the knowledge migration state. When a newly added knowledge feature unit forms a historical knowledge association relationship, the knowledge evolution mode information corresponds to a knowledge recombination state; when a newly added knowledge feature unit replaces a historical knowledge association relationship, the knowledge evolution mode information corresponds to a knowledge replacement state. Based on knowledge evolution path information and knowledge evolution pattern information, incremental knowledge state information is generated to indicate the direction of influence of newly added multimodal training data on historical multimodal knowledge state information.
[0010] As a preferred embodiment of the multimodal large model incremental joint training method of the present invention, the model adjustment region information includes: The knowledge evolution pattern information corresponding to the newly added knowledge features is determined based on the knowledge expansion state, knowledge migration state, knowledge reorganization state, and knowledge replacement state recorded in the incremental knowledge state information. Based on the knowledge evolution path information, determine the propagation path of the new knowledge features in the historical multimodal knowledge state information, and based on the knowledge influence direction information, determine the influence range corresponding to the new knowledge features; Based on the knowledge evolution mode information, the knowledge evolution path information, and the knowledge influence direction information, the knowledge propagation path corresponding to the newly added knowledge features is mapped to the visual knowledge region, text knowledge region, and cross-modal fusion region in the multimodal large model. The knowledge action area information is determined based on the degree of correlation between the newly added knowledge features and the visual knowledge region, text knowledge region, and cross-modal fusion region. Based on the knowledge application area information, the target adjustment area of the newly added knowledge features in the multimodal large model is determined, and the model adjustment area information is obtained.
[0011] As a preferred embodiment of the incremental joint training method for multimodal large models described in this invention, the updated multimodal model parameters include: Based on the model adjustment region information, the changes in knowledge gain of newly added knowledge features in the visual knowledge region, text knowledge region, and cross-modal fusion region are statistically analyzed. The changes in knowledge gain are then continuously sorted according to the knowledge propagation direction recorded in the knowledge evolution path information to form a knowledge gain change sequence. Based on the knowledge gain change sequence, determine the knowledge gain change difference between adjacent knowledge regions, increase or decrease the parameter update frequency of the corresponding knowledge region according to the knowledge gain change difference, and generate a joint training strategy according to the adjusted parameter update frequency. According to the joint training strategy, the parameters of the visual knowledge region, text knowledge region, and cross-modal fusion region are updated sequentially according to the parameter update frequency. After completing parameter updates in the visual knowledge region, the parameter update trajectory corresponding to the visual knowledge region is written into the joint training strategy. After completing parameter updates in the text knowledge region, the parameter update trajectory in the joint training strategy is used to correct the parameter update path corresponding to the text knowledge region. After completing parameter updates in the cross-modal fusion region, the corrected parameter update path is used to adjust the parameter update rhythm corresponding to the cross-modal fusion region to obtain the incrementally updated multimodal model parameters. Based on the incrementally updated multimodal model parameters, the updated knowledge gain states corresponding to the visual knowledge region, text knowledge region, and cross-modal fusion region are calculated using the newly added multimodal training data. The updated knowledge gain states are then correlated with the corresponding knowledge gain states before incremental training to obtain the knowledge gain evolution sequence. When the knowledge gain evolution sequence changes, the parameter update frequency corresponding to the joint training strategy is readjusted according to the changed knowledge gain evolution sequence to obtain the updated multimodal model parameters.
[0012] As a preferred embodiment of the incremental joint training method for multimodal large models described in this invention, the method for updating historical multimodal knowledge state information includes: Using the updated multimodal model parameters, extract the knowledge parameter change features corresponding to the updated multimodal model parameters to obtain the knowledge parameter change features. Reorganize the knowledge parameter change features according to the knowledge organization method corresponding to the historical multimodal knowledge state information to obtain the knowledge state update content. By utilizing the knowledge state update content, the knowledge state update content is associated with the knowledge association structure corresponding to the historical multimodal knowledge state information to obtain the updated historical knowledge state content, and the historical multimodal knowledge state information is updated to obtain the historical multimodal knowledge state information.
[0013] Secondly, this invention provides a multimodal large-model incremental joint training system, comprising: The preprocessing module acquires the new multimodal training data to be added to the multimodal large model and performs preprocessing to obtain a standardized incremental training dataset. The extraction module inputs the standardized incremental training dataset into the multimodal large model and obtains incremental multimodal features through the multimodal feature extraction network; The analysis module analyzes the correlation changes between newly added multimodal training data and historical multimodal knowledge based on incremental multimodal features and historical multimodal knowledge state information to obtain incremental knowledge state information. The adjustment module determines the target adjustment region corresponding to the newly added knowledge in the multimodal large model based on the incremental knowledge status information, and obtains the model adjustment region information. The training module generates a joint training strategy based on the model adjustment region information, and performs incremental joint training on the multimodal large model based on the joint training strategy to obtain the updated multimodal model parameters. The update module uses the updated multimodal model parameters to update the historical multimodal knowledge state information.
[0014] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the multimodal large model incremental joint training method as described in the first aspect of the present invention.
[0015] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the multimodal large model incremental joint training method as described in the first aspect of the present invention.
[0016] The beneficial effects of this invention are as follows: By analyzing the correlation changes between newly added multimodal training data and historical multimodal knowledge, a knowledge evolution path and knowledge evolution mode are constructed, realizing the dynamic characterization of the influence relationship of newly added knowledge and providing a basis for the accurate determination of the model adjustment region; by determining the target adjustment region based on incremental knowledge state information and dynamically generating joint training strategies based on knowledge gain changes, adaptive incremental updates of model parameters are realized, improving the pertinence of incremental training and the efficiency of joint training; by continuously updating historical multimodal knowledge state information using updated model parameters, continuous evolution and dynamic maintenance of model knowledge are realized, enhancing the continuous learning capability of large multimodal models. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a multimodal large-scale incremental joint training method.
[0019] Figure 2 This is a schematic diagram of a multimodal large-model incremental joint training system. Detailed Implementation
[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0021] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0022] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0023] Reference Figures 1-2 As an embodiment of the present invention, this embodiment provides a multimodal large model incremental joint training method, including the following steps: S1. Obtain the new multimodal training data to be added to the multimodal large model and preprocess it to obtain a standardized incremental training dataset.
[0024] S1.1 Perform data cleaning and modal data alignment on the newly added multimodal training data to obtain preprocessed newly added multimodal training data, establish data association relationships between different modalities, and uniformly store the preprocessed newly added multimodal training data and the data association relationships to obtain a standardized incremental training dataset.
[0025] Furthermore, the image data, text data, speech data, and video data in the newly added multimodal training data are formatted and standardized to remove data content with missing information, abnormal markers, and noise interference. Modal data association relationships are established based on the semantic correspondence between different modal data. For example, for training samples containing image data and text data, the association relationship between image data and text data is determined by the target object, scene description, and semantic tags corresponding to the text content; for training samples containing speech data and text data, the association relationship between speech data and text data is determined by matching the semantic information corresponding to the speech content with the text description information.
[0026] Specifically, for training samples containing video data, the correlation between different modalities within the video data is established through the temporal correspondence between video frame content, video description information, and speech information. Based on the established data correlation, the preprocessed new multimodal training data is organized, and the data of different modalities with correlation are stored accordingly, so that the correspondence between image data, text data, speech data, and video data is maintained, and a standardized incremental training dataset is obtained.
[0027] S2. Input the standardized incremental training dataset into the multimodal large model, and obtain incremental multimodal features through the multimodal feature extraction network.
[0028] S2.1. Pre-train using image data, text data, speech data, and video data to obtain a multimodal large model. Then, perform feature learning on image data, text data, speech data, and video data to obtain a multimodal feature extraction network.
[0029] Furthermore, during pre-training using image, text, speech, and video data, training samples of different modalities are obtained. These data are then input into their respective data processing flows to learn the feature information within each modality. For example, for training samples containing both image and text data, the multimodal model learns the mapping relationship between visual and semantic information by analyzing the correspondence between image content and text description. For training samples containing both speech and text data, the multimodal model learns the conversion relationship between sound and language information by analyzing the consistency between speech content and text description. For continuous image frames and corresponding speech information in video data, the multimodal model learns the association between video scenes and multimodal semantics by analyzing the temporal correlation between changes in video content and speech description. This enables the multimodal model to learn features and express associations across different modalities, resulting in a multimodal feature extraction network.
[0030] S2.2. The multimodal feature extraction network in the multimodal large model is used to extract features from the image data, text data, speech data and video data in the standardized incremental training dataset to obtain visual features, text features, speech features and video features.
[0031] Furthermore, different modalities of data from the standardized incremental training dataset are input into a multimodal feature extraction network. Corresponding feature information is extracted based on the feature attributes of each modality. For image data, the multimodal feature extraction network extracts target objects, spatial relationships, and visual semantic information to obtain visual features. For text data, it extracts semantic content, entity relationships, and contextual information to obtain text features. For speech data, it extracts acoustic variations, semantic expressions, and speech content information to obtain speech features. For video data, it analyzes the changing relationships between consecutive video frames and the semantic information corresponding to the video content to obtain video features. By extracting feature information corresponding to different modalities, the different modalities of data in the standardized incremental training dataset are transformed into feature representations suitable for incremental learning of a large multimodal model, resulting in visual, text, speech, and video features.
[0032] S2.3. Feature mapping and semantic association fusion are performed through the multimodal feature extraction network in the multimodal large model to obtain incremental multimodal features with correlation between different modalities.
[0033] Furthermore, when performing feature mapping and semantic association fusion through the multimodal feature extraction network in the multimodal large model, visual features, text features, speech features, and video features are mapped to a unified feature representation space. Association matching is performed based on the semantic correspondence between different modal features, establishing correspondences between them. For example, when visual features contain target object information and text features contain corresponding object description information, the semantic association between visual and text features is determined through the feature mapping process. When speech features contain event description information and video features contain corresponding event change processes, the temporal association between speech and video features is determined through the feature association fusion process. When image data, text data, speech data, and video data jointly describe the same knowledge content, the multimodal feature extraction network fuses different modal features, forming a unified association expression between different modalities. Through the above feature mapping and semantic association fusion process, associations are established between different modal features, obtaining incremental multimodal features with associations between different modalities.
[0034] S3. Based on incremental multimodal features and historical multimodal knowledge state information, analyze the correlation changes between newly added multimodal training data and historical multimodal knowledge to obtain incremental knowledge state information.
[0035] S3.1. Based on the incremental multimodal features and the historical multimodal knowledge state information, determine the newly added knowledge feature units and the historical knowledge association units respectively.
[0036] Furthermore, based on the semantic, object, and modal association information contained in visual, text, speech, and video features, the newly added knowledge content corresponding to the newly added multimodal training data is segmented to obtain new knowledge feature units. For example, when the incremental multimodal features contain target object features corresponding to image data and object description features corresponding to text data, the visual and text features that can represent the same newly added knowledge content are associated to determine the corresponding new knowledge feature units; when the incremental multimodal features contain behavioral change information in video data and event description information in speech data, the new knowledge feature units that can represent new event information are determined based on the correspondence between video content changes and speech semantics.
[0037] Specifically, by utilizing the historical knowledge association structures stored in the historical multimodal knowledge state information, association analysis is performed on the visual, textual, audio, and video knowledge content within the historical multimodal knowledge state information. Historical knowledge association units are determined based on the semantic relationships, feature connectivity relationships, and cross-modal correspondences between different historical knowledge contents. For example, when the historical multimodal knowledge state information contains object category information corresponding to a newly added knowledge feature unit, the attribute information, scene information, and textual description information associated with the object category are determined as the corresponding historical knowledge association units. When the historical multimodal knowledge state information contains event information corresponding to a newly added knowledge feature unit, the historical knowledge content associated with the event's occurrence conditions, event process, and event result is determined as the historical knowledge association units. This ensures that the newly added knowledge feature units can represent the newly added knowledge content in the new multimodal training data, and that the historical knowledge association units can represent the associated knowledge content in the historical multimodal knowledge state information, thus obtaining both the newly added knowledge feature units and the historical knowledge association units.
[0038] S3.2. Based on the positional changes of the newly added knowledge feature units in the visual feature space, text semantic space, and cross-modal association space, establish the knowledge influence relationship between the newly added knowledge feature units and the historical knowledge association units.
[0039] Furthermore, the newly added knowledge feature units are mapped to the visual feature space, text semantic space, and cross-modal association space. The spatial feature state is determined based on the changes in the feature distribution positions of the newly added knowledge feature units in different spaces. In the visual feature space, the degree of visual association between the newly added knowledge feature units and historical knowledge association units is determined by analyzing the distance changes between their corresponding visual feature vectors. In the text semantic space, the semantic association changes between the newly added knowledge feature units and historical knowledge association units are determined by analyzing the changes in the matching degree between their corresponding text semantic features. In the cross-modal association space, the cross-modal association changes between the newly added knowledge feature units and historical knowledge association units are determined by analyzing the changes in the mapping relationships between their corresponding visual features, text features, speech features, and video features and historical knowledge association units. Based on the association changes in the visual feature space, text semantic space, and cross-modal association space, the knowledge influence relationship between the newly added knowledge feature units and historical knowledge association units is established.
[0040] Specifically, a comprehensive analysis is conducted on the positional changes of newly added knowledge feature units in different spaces. Changes in feature distance, semantic association, and cross-modal matching relative to historical knowledge association units are used as the basis for determining knowledge influence relationships. For example, if a newly added knowledge feature unit only changes position in the visual feature space while remaining stable in the text semantic space and cross-modal association space, it is determined that the newly added knowledge feature unit primarily affects the visual association content in historical knowledge association units. When a newly added knowledge feature unit changes position in both the text semantic space and the cross-modal association space, it is determined that the newly added knowledge feature unit affects the semantic and cross-modal association relationships in historical knowledge association units. By analyzing the positional changes of newly added knowledge feature units in different feature spaces, the knowledge influence relationship between newly added knowledge feature units and historical knowledge association units is established.
[0041] S3.3. Based on the knowledge influence relationship, calculate the propagation direction, propagation range and propagation intensity of the newly added knowledge feature unit acting on the historical knowledge association unit, determine the knowledge evolution path of the newly added knowledge feature unit affecting the historical multimodal knowledge state information, and obtain knowledge evolution path information.
[0042] Furthermore, by analyzing the positional change trends of newly added knowledge feature units relative to historical knowledge association units in different feature spaces, the direction of influence of newly added knowledge feature units on historical knowledge association units can be determined. For example, when newly added knowledge feature units change from the visual feature space to the cross-modal association space, it is determined that newly added knowledge feature units have a propagation effect on cross-modal association relationships; when newly added knowledge feature units mainly change in the text semantic space, it is determined that newly added knowledge feature units have a propagation effect on text semantic association relationships. The propagation range is determined based on the association distance between newly added knowledge feature units and historical knowledge association units and the knowledge influence range. By analyzing the number of historical knowledge association units that newly added knowledge feature units can be associated with, the association level, and the association space range, the influence area of newly added knowledge feature units on historical multimodal knowledge state information can be determined.
[0043] Specifically, the propagation intensity is determined based on changes in the degree of association between newly added knowledge feature units and historical knowledge related units. When the degree of association between the newly added knowledge feature units and historical knowledge related units increases, the propagation intensity is determined to be enhanced; when the degree of association decreases, the propagation intensity is determined to be weakened. Based on the relationship between propagation direction, propagation scope, and propagation intensity, the propagation process of the newly added knowledge feature units influencing historical multimodal knowledge state information is determined, forming a knowledge evolution path in which the newly added knowledge feature units affect historical multimodal knowledge state information, thus obtaining knowledge evolution path information.
[0044] The expression for the direction of propagation is: ; in, For the direction of dissemination, This represents the vertical association distance between newly added knowledge feature units and historical knowledge-related units in the knowledge space, and the horizontal association distance between newly added knowledge feature units and historical knowledge-related units in the knowledge space. It is a constant.
[0045] The propagation range expression is: ; in, To expand the reach of the message, This refers to the knowledge association distance between newly added knowledge feature units and historical knowledge association units.
[0046] The expression for propagation intensity is: ; in, The intensity of transmission.
[0047] S3.4. Based on the knowledge evolution path information, construct the knowledge evolution relationship of the newly added knowledge feature unit acting on the historical multimodal knowledge state information, and generate the corresponding knowledge evolution mode information based on the association range expansion relationship, association direction migration relationship, association structure reorganization relationship and association relationship replacement relationship of the newly added knowledge feature unit in the historical multimodal knowledge state information.
[0048] Furthermore, based on the propagation direction, scope, and intensity between newly added knowledge feature units and historical knowledge association units, the process of how newly added knowledge feature units influence the changes in historical multimodal knowledge state information is described in relation to this process. This determines the role of newly added knowledge feature units in historical multimodal knowledge state information. For example, when a newly added knowledge feature unit establishes associations with multiple historical knowledge association units, and the scope of these associations continues to expand, it is determined that the newly added knowledge feature unit forms an association scope expansion relationship, indicating that it can supplement the associated content in the historical multimodal knowledge state information. When a newly added knowledge feature unit changes the original association direction between historical knowledge association units, it is determined that the newly added knowledge feature unit forms an association direction migration relationship, indicating that it changes the knowledge connection method in the historical multimodal knowledge state information. When a newly added knowledge feature unit forms a new combined association relationship with multiple historical knowledge association units, it is determined that the newly added knowledge feature unit forms an association structure reorganization relationship, indicating that it promotes the reorganization of the historical knowledge association structure. When a newly added knowledge feature unit replaces part of the associated content in a historical knowledge association unit, it is determined that the newly added knowledge feature unit forms an association relationship replacement relationship, indicating that it updates the existing association relationships in the historical multimodal knowledge state information.
[0049] Specifically, based on the knowledge change states corresponding to the association scope expansion relationship, association direction migration relationship, association structure reorganization relationship, and association relationship replacement relationship, knowledge evolution pattern information is generated to represent the specific evolutionary mode of the newly added knowledge feature unit acting on the historical multimodal knowledge state information, thereby obtaining knowledge evolution pattern information.
[0050] S3.5 When a new knowledge feature unit expands the scope of historical knowledge association, the knowledge evolution mode information corresponds to the knowledge expansion state; when a new knowledge feature unit changes the direction of historical knowledge association, the knowledge evolution mode information corresponds to the knowledge migration state.
[0051] Furthermore, when a newly added knowledge feature unit expands the scope of historical knowledge associations, the knowledge evolution pattern information corresponding to the newly added knowledge feature unit is determined to be in a knowledge expansion state based on the changes in the number and scope of associations between the newly added knowledge feature unit and historical knowledge association units. For example, when a new association is established between a newly added knowledge feature unit and a historical knowledge association unit, increasing the original scope of knowledge associations in the historical multimodal knowledge state information, it is determined that the newly added knowledge feature unit has a knowledge expansion effect on the historical multimodal knowledge state information; when the newly added knowledge feature unit only maintains an association with existing historical knowledge association units without generating new changes in the scope of associations, it is not determined to be in a knowledge expansion state. By determining whether the newly added knowledge feature unit forms a new scope of knowledge associations, the identification of the expansion of historical knowledge content by the newly added knowledge is achieved.
[0052] When a newly added knowledge feature unit changes the direction of historical knowledge association, the knowledge evolution pattern information corresponding to the newly added knowledge feature unit is determined as a knowledge migration state based on the change in the direction of the association between the newly added knowledge feature unit and the historical knowledge association units. For example, when a newly added knowledge feature unit causes the historical knowledge association units to change from the original visual association direction to a textual semantic association direction, or from a single-modal association to a cross-modal association, it is determined that the newly added knowledge feature unit has changed the knowledge association direction in the historical multimodal knowledge state information, generating a knowledge migration state; when the newly added knowledge feature unit does not change the association direction between historical knowledge association units, it is not determined to be a knowledge migration state. By judging whether the knowledge association direction has changed, the identification of situations where new knowledge changes the historical knowledge association path is achieved, generating corresponding knowledge evolution pattern information.
[0053] S3.6 When a newly added knowledge feature unit forms a historical knowledge association relationship, the knowledge evolution mode information corresponds to the knowledge recombination state; when a newly added knowledge feature unit replaces a historical knowledge association relationship, the knowledge evolution mode information corresponds to the knowledge replacement state.
[0054] Furthermore, when newly added knowledge feature units form historical knowledge associations, the knowledge evolution pattern information corresponding to the newly added knowledge feature unit is determined to be in a knowledge reorganization state based on the changes in the association structure between the newly added knowledge feature unit and multiple historical knowledge association units. For example, when a newly added knowledge feature unit establishes new associations with visual knowledge association content, textual knowledge association content, and cross-modal knowledge association content, respectively, causing multiple originally independent historical knowledge association units to form new knowledge associations, it is determined that the newly added knowledge feature unit has a knowledge structure reorganization effect on the historical multimodal knowledge state information; when a newly added knowledge feature unit only maintains an association with a single historical knowledge association unit and does not change the association relationship between multiple historical knowledge association units, it is not determined to be in a knowledge reorganization state. By analyzing whether newly added knowledge feature units form new knowledge association combinations, the identification of changes in the historical knowledge structure caused by newly added knowledge is achieved.
[0055] When a new knowledge feature unit replaces a historical knowledge association, the knowledge evolution pattern information corresponding to the new knowledge feature unit is determined to be in a knowledge substitution state based on the changes in the associated content between the new knowledge feature unit and the historical knowledge association unit. For example, if a new knowledge feature unit and a historical knowledge association unit have the same knowledge expression content, but the feature association corresponding to the new knowledge feature unit can cover the historical knowledge association, then the new knowledge feature unit is determined to replace the existing associated content in the historical knowledge association unit, generating a knowledge substitution state. If the new knowledge feature unit only supplements the historical knowledge association unit without changing the existing knowledge association, it is not determined to be a knowledge substitution state. By determining whether the new knowledge feature unit changes the historical knowledge association, the system identifies the situation where new knowledge updates existing knowledge content and generates corresponding knowledge evolution pattern information.
[0056] S3.7 Based on the knowledge evolution path information and knowledge evolution mode information, generate incremental knowledge state information that indicates the direction of influence of newly added multimodal training data on historical multimodal knowledge state information.
[0057] Furthermore, correlation analysis is conducted on the propagation direction, propagation scope, and propagation intensity in the knowledge evolution path information to determine the influence direction of newly added knowledge feature units on historical multimodal knowledge state information. Based on the knowledge evolution pattern information, the type of change caused by newly added knowledge feature units on historical multimodal knowledge state information is determined. When the knowledge evolution pattern information corresponds to a knowledge expansion state, it is determined that the newly added multimodal training data is mainly used to expand the correlation scope in historical multimodal knowledge state information; when the knowledge evolution pattern information corresponds to a knowledge transfer state, it is determined that the newly added multimodal training data mainly changes the correlation direction in historical multimodal knowledge state information; when the knowledge evolution pattern information corresponds to a knowledge transfer state, it is determined that the newly added multimodal training data mainly changes the correlation direction in historical multimodal knowledge state information; when the knowledge evolution pattern information corresponds to a knowledge transfer state, it is determined that the newly added multimodal training data mainly changes the correlation direction in historical multimodal knowledge state information. When the knowledge evolution pattern information corresponds to the knowledge recombination state, it is determined that newly added multimodal training data promotes the formation of new combination relationships in the historical knowledge association structure. When the knowledge evolution pattern information corresponds to the knowledge replacement state, it is determined that newly added multimodal training data is used to update the existing knowledge association relationships in the historical multimodal knowledge state information. The influence direction information in the knowledge evolution path information is associated with the knowledge change type in the knowledge evolution pattern information to form the state description content of the influence of newly added multimodal training data on the historical multimodal knowledge state information. This enables the incremental knowledge state information to simultaneously record the influence direction of newly added knowledge and the knowledge change mode, thereby obtaining the incremental knowledge state information.
[0058] S4. Based on the incremental knowledge state information, determine the target adjustment region in the multimodal large model corresponding to the newly added knowledge, and obtain the model adjustment region information.
[0059] S4.1 Determine the knowledge evolution mode information corresponding to the newly added knowledge features based on the knowledge expansion state, knowledge migration state, knowledge reorganization state and knowledge replacement state recorded in the incremental knowledge state information.
[0060] Furthermore, the corresponding knowledge evolution mode is determined based on the changes in the association between the newly added knowledge feature units and the historical multimodal knowledge state information. For example, when the incremental knowledge state information records that the newly added knowledge feature unit increases the number of historical knowledge associated units or expands the scope of historical knowledge association, the knowledge evolution mode information corresponding to the newly added knowledge feature is determined to be in the knowledge expansion state; when the incremental knowledge state information records that the newly added knowledge feature unit changes the connection direction between historical knowledge associated units, the knowledge evolution mode information corresponding to the newly added knowledge feature is determined to be in the knowledge migration state; when the incremental knowledge state information records that the newly added knowledge feature unit forms a new combination relationship with multiple historical knowledge associated units, the knowledge evolution mode information corresponding to the newly added knowledge feature is determined to be in the knowledge reorganization state; when the incremental knowledge state information records that the newly added knowledge feature unit replaces the existing associated content in historical knowledge associated units, the knowledge evolution mode information corresponding to the newly added knowledge feature is determined to be in the knowledge replacement state. By identifying the types of knowledge changes in the incremental knowledge state information, the knowledge evolution mode information corresponding to the newly added knowledge feature can reflect the specific changes in the historical multimodal knowledge state information caused by the newly added knowledge.
[0061] S4.2 Determine the propagation path of newly added knowledge features in historical multimodal knowledge state information based on knowledge evolution path information, and determine the scope of influence corresponding to newly added knowledge features based on knowledge influence direction information.
[0062] Furthermore, the association process between newly added knowledge feature units and historical knowledge association units is tracked to determine the propagation path of newly added knowledge features in historical multimodal knowledge state information. For example, when a newly added knowledge feature unit propagates from the visual feature space to the cross-modal association space, the propagation path corresponding to the newly added knowledge feature is determined to pass through the visual knowledge region and the cross-modal fusion region; when the newly added knowledge feature unit mainly acts on the text semantic space, the propagation path corresponding to the newly added knowledge feature is determined to pass through the text knowledge region. The influence range corresponding to the newly added knowledge feature is determined based on the knowledge influence direction information. By analyzing the degree of association between the newly added knowledge feature unit and the visual knowledge region, the text knowledge region, and the cross-modal fusion region, the knowledge region range in which the newly added knowledge feature can act is determined. For example, when a newly added knowledge feature unit is associated with both visual features and text semantic features, the influence range corresponding to the newly added knowledge feature is determined to include both the visual knowledge region and the text knowledge region; when the degree of change in the cross-modal association relationship between the newly added knowledge feature unit and the newly added knowledge feature is high, the influence range corresponding to the newly added knowledge feature is determined to include the cross-modal fusion region. By determining the propagation path and influence range, the knowledge action area corresponding to the newly added knowledge feature can be accurately located.
[0063] S4.3. Based on the knowledge evolution mode information, the knowledge evolution path information, and the knowledge influence direction information, the knowledge propagation path corresponding to the newly added knowledge features is mapped to the visual knowledge region, text knowledge region, and cross-modal fusion region in the multimodal large model. The knowledge action area information is determined based on the degree of correlation between the newly added knowledge features and the visual knowledge region, text knowledge region, and cross-modal fusion region.
[0064] Furthermore, based on the knowledge evolution pattern information, the knowledge change type corresponding to the newly added knowledge feature is determined. Then, based on the propagation direction recorded in the knowledge evolution path information, the transmission process of the newly added knowledge feature in the multimodal large model is determined. According to the propagation path of the newly added knowledge feature in the visual feature space, text semantic space, and cross-modal association space, the feature association relationship corresponding to the newly added knowledge feature is mapped to the corresponding knowledge region in the multimodal large model. For example, when the knowledge evolution pattern information is in the knowledge expansion state, and the knowledge evolution path information indicates that the newly added knowledge feature mainly propagates from the visual feature space to the cross-modal association space, the knowledge propagation path corresponding to the newly added knowledge feature is mapped to the visual knowledge region and the cross-modal fusion region; when the knowledge evolution pattern information is in the knowledge migration state, and the knowledge influence direction information indicates that the newly added knowledge feature mainly changes the text semantic association direction, the knowledge propagation path corresponding to the newly added knowledge feature is mapped to the text knowledge region.
[0065] Specifically, the knowledge application area information is determined based on the degree of correlation between newly added knowledge features and visual knowledge regions, textual knowledge regions, and cross-modal fusion regions. By analyzing the matching degree between the feature correlations corresponding to the newly added knowledge features and different knowledge regions, the knowledge regions where the newly added knowledge features primarily function are determined. For example, when the correlation between the newly added knowledge feature and the visual knowledge region is higher than that between the textual knowledge region and the cross-modal fusion region, the visual knowledge region is determined as the corresponding knowledge application area; when the newly added knowledge feature has correlations with both the textual knowledge region and the cross-modal fusion region, the textual knowledge region and the cross-modal fusion region are determined as the corresponding knowledge application areas. Through this process, a correspondence is established between the knowledge propagation path of the newly added knowledge features and the knowledge regions within the multimodal large model, thus obtaining the knowledge application area information.
[0066] S4.4. Based on the knowledge application area information, determine the target adjustment area of the newly added knowledge features in the multimodal large model, and obtain the model adjustment area information.
[0067] Furthermore, when the knowledge application area information indicates that the newly added knowledge feature is mainly associated with the visual knowledge area, the parameter area corresponding to the visual knowledge area is determined as the target adjustment area for the newly added knowledge feature. When the knowledge application area information indicates that the newly added knowledge feature is associated with both the text knowledge area and the cross-modal fusion area, the parameter areas corresponding to the text knowledge area and the cross-modal fusion area are determined as the target adjustment areas based on the propagation range and influence direction of the newly added knowledge feature in the text knowledge area and the cross-modal fusion area. The target adjustment areas are then filtered based on the knowledge evolution mode information corresponding to the newly added knowledge feature. When the newly added knowledge feature corresponds to a knowledge expansion state, the knowledge area within the newly associated range is prioritized as the target adjustment area. When the newly added knowledge feature corresponds to a knowledge migration state, the knowledge area where the association direction has changed is prioritized as the target adjustment area. When the newly added knowledge feature corresponds to a knowledge reorganization state or a knowledge replacement state, the knowledge area where the knowledge structure has changed or the knowledge association has been updated is determined as the target adjustment area. By jointly determining the parameter adjustment position corresponding to the newly added knowledge feature based on the knowledge application area information and the knowledge evolution mode information, the newly added knowledge feature can be mapped to the target adjustment area in the multimodal large model, thus obtaining the model adjustment area information.
[0068] S5. Generate a joint training strategy based on the model adjustment region information, and perform incremental joint training on the multimodal large model based on the joint training strategy to obtain the updated multimodal model parameters.
[0069] S5.1. Based on the model adjustment region information, calculate the change in knowledge gain of newly added knowledge features in the visual knowledge region, text knowledge region, and cross-modal fusion region, and continuously sort the change in knowledge gain according to the knowledge propagation direction recorded by the knowledge evolution path information to form a knowledge gain change sequence.
[0070] Furthermore, based on the model adjustment region information, the target adjustment region corresponding to the newly added knowledge feature is determined, and the changes in the knowledge state of the target adjustment region before and after incremental training are obtained. By analyzing the changes in feature expression, semantic association, and cross-modal association in different knowledge regions before and after the application of the newly added knowledge feature, the change in knowledge gain of the corresponding knowledge region is determined. For example, when the newly added knowledge feature mainly applies to the visual knowledge region, the change in knowledge gain of the visual knowledge region is determined by analyzing the changes in feature expression in the visual knowledge region; when the newly added knowledge feature applies to both the text knowledge region and the cross-modal fusion region, the change in knowledge gain of the corresponding region is determined by analyzing the changes in text semantic association and cross-modal association fusion.
[0071] Specifically, based on the knowledge propagation direction recorded in the knowledge evolution path information, the knowledge gain changes corresponding to the visual knowledge region, text knowledge region, and cross-modal fusion region are continuously sorted. During the sorting process, the knowledge gain changes are arranged from the knowledge propagation start region to the knowledge propagation end region according to the sequential relationship in the propagation path of the new knowledge features, so that the knowledge gain changes form a knowledge gain change sequence corresponding to the knowledge propagation process. For example, when the knowledge evolution path information indicates that the new knowledge features propagate from the visual knowledge region to the cross-modal fusion region, the corresponding knowledge gain changes are sorted according to the propagation relationship between the visual knowledge region, text knowledge region, and cross-modal fusion region to form a knowledge gain change sequence.
[0072] S5.2. Based on the knowledge gain change sequence, determine the knowledge gain change difference between adjacent knowledge regions, increase or decrease the parameter update frequency of the corresponding knowledge region according to the knowledge gain change difference, and generate a joint training strategy according to the adjusted parameter update frequency.
[0073] Furthermore, the larger the difference in knowledge gain, the more obvious the change in the absorption state of the new knowledge features between adjacent knowledge regions, indicating that there are differences in the degree of response of different knowledge regions to the new knowledge during the dissemination of new knowledge; when the difference in knowledge gain is small, it indicates that the knowledge absorption state between adjacent knowledge regions is relatively close, indicating that the new knowledge has a relatively stable dissemination effect between adjacent knowledge regions. The adjustment direction of parameter update frequency for the corresponding knowledge region is determined based on the difference in knowledge gain. When the change in knowledge gain for the previous knowledge region is higher than that for the next knowledge region, it indicates that the new knowledge has a higher degree of absorption in the previous knowledge region. Therefore, the parameter update frequency for the previous knowledge region is increased, enabling the multimodal large model to further enhance the expression of new knowledge in high-gain regions. When the change in knowledge gain for the next knowledge region increases, it indicates that the new knowledge has a strong influence in the propagation process to the next knowledge region. Therefore, the parameter update frequency for the next knowledge region is increased, making the parameter update process adapt to the direction of new knowledge propagation. When the change in knowledge gain for the corresponding knowledge region tends to stabilize, the parameter update frequency for the corresponding knowledge region is reduced to reduce unnecessary parameter adjustments. Based on the adjusted parameter update frequency, corresponding joint training strategies are generated for the visual knowledge region, text knowledge region, and cross-modal fusion region, enabling incremental training of different knowledge regions according to the matched parameter update frequency.
[0074] Specifically, by analyzing the difference in knowledge gain changes between adjacent knowledge regions in the knowledge gain change sequence, the parameter update frequency can be adjusted according to the actual absorption of new knowledge in different knowledge regions. The difference in knowledge gain changes is used to determine the differences in knowledge response between adjacent knowledge regions. When the gain change of new knowledge in a certain knowledge region is large, the parameter update frequency of the corresponding knowledge region is increased, enabling the multimodal large model to strengthen its learning of new knowledge. When the gain change of new knowledge in a certain knowledge region tends to stabilize, the parameter update frequency of the corresponding knowledge region is decreased, concentrating parameter updates on regions with significant knowledge changes. By correlating the difference in knowledge gain changes with the parameter update frequency, the joint training strategy can dynamically adjust the training rhythm of different knowledge regions according to the propagation process of new knowledge, improving the targeting of model parameter updates during incremental training.
[0075] The expression for the difference in knowledge gain is: ; in, This represents the difference in knowledge gain between adjacent knowledge regions. For the knowledge gain change sequence, the first... The change in knowledge gain corresponding to each knowledge region. For the first The change in knowledge gain corresponding to each knowledge region. For knowledge area indexing.
[0076] S5.3 According to the joint training strategy, the parameters of the visual knowledge region, text knowledge region and cross-modal fusion region are updated sequentially according to the parameter update frequency.
[0077] Furthermore, based on the propagation direction of newly added knowledge features along the knowledge evolution path, parameter updates are prioritized for the visual knowledge region, text knowledge region, or cross-modal fusion region corresponding to the starting position of knowledge propagation. Then, parameter updates for other knowledge regions are sequentially performed according to the parameter update frequency set in the joint training strategy. For example, when newly added knowledge features mainly propagate from the visual knowledge region to the cross-modal fusion region, parameter updates for the visual knowledge region are performed first, allowing it to absorb the newly added knowledge features. Then, parameter updates for the text knowledge region and the cross-modal fusion region are performed, enabling the newly added knowledge to be jointly updated along the knowledge propagation path. When newly added knowledge features mainly affect the text knowledge region and the cross-modal fusion region, the knowledge regions with larger changes in knowledge gain are updated first, followed by those with lower correlation, based on the parameter update frequency. By performing parameter updates for different knowledge regions according to the parameter update frequency, the visual knowledge region, text knowledge region, and cross-modal fusion region can undergo collaborative incremental training based on the degree of influence of the newly added knowledge, obtaining updated multimodal model parameters.
[0078] S5.4 After completing the parameter update in the visual knowledge region, the parameter update trajectory corresponding to the visual knowledge region is written into the joint training strategy. After completing the parameter update in the text knowledge region, the parameter update path corresponding to the text knowledge region is corrected using the parameter update trajectory in the joint training strategy. After completing the parameter update in the cross-modal fusion region, the parameter update rhythm corresponding to the cross-modal fusion region is adjusted using the corrected parameter update path to obtain the incrementally updated multimodal model parameters.
[0079] Furthermore, after the parameters in the text knowledge region are updated, the changes in the relationship between the visual knowledge region and the text knowledge region are analyzed based on the parameter update trajectory of the visual knowledge region recorded in the joint training strategy, and the parameter update path corresponding to the text knowledge region is corrected. For example, when the semantic association of newly added knowledge features in the text knowledge region is enhanced after the visual knowledge region parameters are updated, the parameter update direction of the text knowledge region is adjusted according to the visual knowledge region parameter update trajectory, so that the text knowledge region can adapt to the changes in the visual knowledge region. When the visual knowledge region parameters have a weak impact on the text knowledge region, the original parameter update path of the text knowledge region is maintained. After the parameter update in the cross-modal fusion region is completed, the parameter update rhythm corresponding to the cross-modal fusion region is adjusted according to the corrected parameter update path, so that the cross-modal fusion region can complete the collaborative update by combining the parameter changes of the visual knowledge region and the text knowledge region. Through the transmission and correction of parameter update trajectories between the visual knowledge region, the text knowledge region, and the cross-modal fusion region, the correlation between different knowledge regions during the incremental training process is maintained, and the multimodal model parameters after incremental updates are obtained.
[0080] Specifically, by transferring the parameter update trajectory of the preceding knowledge region to the subsequent knowledge region, a correlated update process is formed between the visual knowledge region, the text knowledge region, and the cross-modal fusion region. Recording the parameter update trajectory of the visual knowledge region allows the absorption of new knowledge within it to serve as a reference for adjusting the parameters of the subsequent text knowledge region. Furthermore, by using the visual knowledge region's parameter update trajectory to correct the text knowledge region's parameter update path, the text knowledge region can adapt to the knowledge changes generated by the visual knowledge region. Finally, by using the corrected parameter update path to adjust the parameter update rhythm of the cross-modal fusion region, the cross-modal fusion region can coordinate the correlated changes between the visual and text knowledge regions. This method, through the transfer of parameter update trajectories between different knowledge regions, enables the incremental training process to reflect the correlated influence between multimodal knowledge, improving the consistency of joint updates across different knowledge regions.
[0081] S5.5 Based on the incrementally updated multimodal model parameters, use the newly added multimodal training data to calculate the updated knowledge gain states corresponding to the visual knowledge region, text knowledge region, and cross-modal fusion region, respectively. Then, perform correlation analysis between the updated knowledge gain states and the corresponding knowledge gain states before incremental training to obtain the knowledge gain evolution sequence.
[0082] Specifically, newly added multimodal training data is input into the large multimodal model corresponding to the updated multimodal model parameters. The updated visual, textual, speech, and video features are obtained through a multimodal feature extraction network. The absorption and changes of the newly added knowledge features in the visual knowledge region, textual knowledge region, and cross-modal fusion region are analyzed to determine the corresponding updated knowledge gain state. The updated knowledge gain state is then correlated with the corresponding knowledge gain state before incremental training. By comparing the changes in knowledge gain in different knowledge regions before and after incremental training, the evolution trend of the newly added knowledge features in different knowledge regions is determined. For example, if the knowledge gain state increases after updating the visual knowledge region while the textual knowledge region changes less, it indicates that the newly added knowledge features are mainly absorbed in the visual knowledge region. If the knowledge gain state changes after updating the cross-modal fusion region, it indicates that the newly added knowledge features generate new knowledge changes during the cross-modal fusion process. Based on the relationship between the changes in knowledge gain state before and after updating the visual knowledge region, textual knowledge region, and cross-modal fusion region, a knowledge gain evolution sequence is formed to represent the changes in the knowledge absorption state of different knowledge regions during incremental training.
[0083] Specifically, by comparing the changes in knowledge gain states in different knowledge regions before and after incremental training, the actual absorption of new knowledge within the multimodal large model is analyzed. By calculating the updated knowledge gain states for the visual knowledge region, textual knowledge region, and cross-modal fusion region separately, the incremental learning effects of different modal knowledge regions can be independently identified. By correlating the updated knowledge gain states with those before incremental training, the knowledge gain change process can reflect the evolutionary trend of new knowledge in different knowledge regions. By generating a knowledge gain evolution sequence, subsequent joint training strategies can be adjusted according to the actual changes in the knowledge regions, achieving dynamic feedback during incremental training.
[0084] The knowledge gain state expressions for the visual region, text region, and cross-modal region are defined as follows: ; in, For the first After the first incremental training... The knowledge gain state corresponding to each knowledge region. For the first After the first incremental training... Knowledge state characteristics corresponding to each knowledge region For the first Before the first incremental training The historical knowledge state characteristics corresponding to each knowledge region.
[0085] The expression for the evolutionary change of knowledge gain is: ; in, For the first The evolutionary change of knowledge gain corresponding to each knowledge region. This represents the knowledge gain state corresponding to the previous round of incremental training.
[0086] Constructing the evolutionary sequence of knowledge gain: ; in, This is a sequence of knowledge gain evolution. This represents the evolutionary change in knowledge gain corresponding to the visual knowledge region. This represents the evolutionary change in knowledge gain corresponding to the text knowledge region. This represents the evolutionary change in knowledge gain corresponding to the cross-modal fusion region.
[0087] Based on the evolutionary sequence of knowledge gain: ; in, For the adjusted number The update frequency of parameters for each knowledge region To adjust the first The update frequency of parameters for each knowledge region This is the adjustment coefficient for changes in knowledge gain. This represents the evolutionary change in knowledge gain.
[0088] S5.6 When the knowledge gain evolution sequence changes, the parameter update frequency corresponding to the joint training strategy is readjusted according to the changed knowledge gain evolution sequence to obtain the updated multimodal model parameters.
[0089] Furthermore, the updated knowledge gain evolution sequence is obtained, and the evolution trends of knowledge gain in the visual knowledge region, text knowledge region, and cross-modal fusion region are analyzed. When the knowledge gain evolution changes, it is determined that the corresponding knowledge region is still in a state of knowledge absorption and change, and the parameter update frequency of the corresponding knowledge region is increased. When the knowledge gain evolution changes decrease or tend to stabilize, it is determined that the knowledge expression of the corresponding knowledge region tends to stabilize, and the parameter update frequency of the corresponding knowledge region is decreased. The parameter update frequencies of the visual knowledge region, text knowledge region, and cross-modal fusion region are adjusted according to the changed knowledge gain evolution sequence to make the joint training strategy match the newly added knowledge evolution state. The parameter update is then re-executed according to the adjusted joint training strategy to obtain the updated multimodal model parameters.
[0090] S6. Update the historical multimodal knowledge state information using the updated multimodal model parameters.
[0091] S6.1. Using the updated multimodal model parameters, extract the knowledge parameter change features corresponding to the updated multimodal model parameters to obtain the knowledge parameter change features. Reorganize the knowledge parameter change features according to the knowledge organization method corresponding to the historical multimodal knowledge state information to obtain the knowledge state update content.
[0092] Furthermore, based on the existing knowledge organization methods in the historical multimodal knowledge state information, the knowledge parameter change features are re-associated and rearranged to maintain a correspondence between the knowledge parameter change features and the knowledge expression forms in the historical multimodal knowledge state information. For example, when the knowledge parameter change features correspond to a visual knowledge region, they are organized according to the visual knowledge association relationships in the historical multimodal knowledge state information; when the knowledge parameter change features correspond to a cross-modal fusion region, they are organized according to the cross-modal association relationships in the historical multimodal knowledge state information. By reorganizing the knowledge parameter change features according to the knowledge organization methods corresponding to the historical multimodal knowledge state information, the updated knowledge change content can maintain a connection with the existing knowledge state, thus obtaining the updated knowledge state content.
[0093] S6.2. Using the knowledge state update content, the knowledge state update content is associated according to the knowledge association structure corresponding to the historical multimodal knowledge state information to obtain the updated historical knowledge state content, and the historical multimodal knowledge state information is updated to obtain the historical multimodal knowledge state information.
[0094] Furthermore, based on the knowledge feature type corresponding to the knowledge state update content, the knowledge state update content is associated with the corresponding historical knowledge association position. For example, when the knowledge state update content corresponds to a knowledge change in a visual knowledge region, the knowledge state update content is associated with the visual knowledge association content in the historical multimodal knowledge state information; when the knowledge state update content corresponds to a change in the association between a text knowledge region and a cross-modal fusion region, the knowledge state update content is associated with the corresponding cross-modal knowledge association position. Based on the matching between the knowledge state update content and the existing knowledge association relationships in the historical multimodal knowledge state information, the historical knowledge association structure is updated so that the new knowledge changes can supplement the existing knowledge state. When the updated knowledge status expands the scope of existing knowledge associations, the association scope in the historical multimodal knowledge status information is updated; when the updated knowledge status changes the direction of existing knowledge associations, the corresponding knowledge association relationships are adjusted; when the updated knowledge status replaces existing knowledge association content, the corresponding historical knowledge association content is updated. By associating the updated knowledge status content with the corresponding position in the historical multimodal knowledge status information, the updated historical knowledge status content is obtained, and the historical multimodal knowledge status information is updated based on the updated historical knowledge status content to obtain the historical multimodal knowledge status information.
[0095] This embodiment also provides a multimodal large model incremental joint training system, including: a preprocessing module, which acquires new multimodal training data to be added to the multimodal large model and performs preprocessing to obtain a standardized incremental training dataset; The extraction module inputs the standardized incremental training dataset into the multimodal large model and obtains incremental multimodal features through the multimodal feature extraction network; The analysis module analyzes the correlation changes between newly added multimodal training data and historical multimodal knowledge based on incremental multimodal features and historical multimodal knowledge state information to obtain incremental knowledge state information. The adjustment module determines the target adjustment region corresponding to the newly added knowledge in the multimodal large model based on the incremental knowledge status information, and obtains the model adjustment region information. The training module generates a joint training strategy based on the model adjustment region information, and performs incremental joint training on the multimodal large model based on the joint training strategy to obtain the updated multimodal model parameters. The update module uses the updated multimodal model parameters to update the historical multimodal knowledge state information.
[0096] This embodiment also provides a computer device applicable to the multimodal large model incremental joint training method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the multimodal large model incremental joint training method proposed in the above embodiment.
[0097] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0098] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the multimodal large model incremental joint training method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0099] In summary, this invention analyzes the correlation changes between newly added multimodal training data and historical multimodal knowledge, constructs knowledge evolution paths and patterns, and dynamically characterizes the influence relationship of newly added knowledge, providing a basis for accurately determining the model adjustment region. By determining the target adjustment region based on incremental knowledge state information and dynamically generating joint training strategies based on knowledge gain changes, adaptive incremental updates of model parameters are achieved, improving the targeting of incremental training and the efficiency of joint training. By continuously updating historical multimodal knowledge state information using updated model parameters, continuous evolution and dynamic maintenance of model knowledge are achieved, enhancing the continuous learning capability of large multimodal models.
[0100] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A multimodal large-scale incremental joint training method, characterized in that: This includes acquiring new multimodal training data to be added to the multimodal large model and preprocessing it to obtain a standardized incremental training dataset; The standardized incremental training dataset is input into the multimodal large model, and incremental multimodal features are obtained through the multimodal feature extraction network. Based on incremental multimodal features and historical multimodal knowledge state information, the correlation changes between newly added multimodal training data and historical multimodal knowledge are analyzed to obtain incremental knowledge state information; Based on the incremental knowledge state information, determine the target adjustment region in the multimodal large model corresponding to the newly added knowledge, and obtain the model adjustment region information; Based on the model adjustment region information, a joint training strategy is generated. Based on the joint training strategy, the multimodal large model is incrementally jointly trained to obtain the updated multimodal model parameters. Update the historical multimodal knowledge state information using the updated multimodal model parameters.
2. The multimodal large model incremental joint training method as described in claim 1, characterized in that: The standardized incremental training dataset includes, The newly added multimodal training data is cleaned and aligned to obtain preprocessed new multimodal training data. Data associations between different modalities are established. The preprocessed new multimodal training data and the data associations are stored uniformly to obtain a standardized incremental training dataset.
3. The multimodal large-model incremental joint training method as described in claim 2, characterized in that: The incremental multimodal features include, A multimodal large model is obtained by pre-training using image data, text data, speech data, and video data. A multimodal feature extraction network is obtained by performing feature learning on image data, text data, speech data, and video data. The multimodal feature extraction network in the multimodal large model extracts features from image data, text data, speech data, and video data in the standardized incremental training dataset to obtain visual features, text features, speech features, and video features. By using a multimodal feature extraction network in a large multimodal model to perform feature mapping and semantic association fusion, incremental multimodal features with correlations between different modalities are obtained.
4. The multimodal large model incremental joint training method as described in claim 3, characterized in that: The incremental knowledge state information includes Based on incremental multimodal features and historical multimodal knowledge state information, newly added knowledge feature units and historical knowledge association units are determined respectively. Based on the positional changes of newly added knowledge feature units in the visual feature space, text semantic space, and cross-modal association space, establish the knowledge influence relationship between newly added knowledge feature units and historical knowledge association units; Based on the knowledge influence relationship, the propagation direction, propagation range and propagation intensity of the newly added knowledge feature unit acting on the historical knowledge association unit are calculated to determine the knowledge evolution path of the newly added knowledge feature unit affecting the historical multimodal knowledge state information, and to obtain knowledge evolution path information. Based on the knowledge evolution path information, a knowledge evolution relationship is constructed in which the newly added knowledge feature unit acts on the historical multimodal knowledge state information. Based on the association range expansion relationship, association direction migration relationship, association structure reorganization relationship and association relationship replacement relationship of the newly added knowledge feature unit in the historical multimodal knowledge state information, the corresponding knowledge evolution mode information is generated. When a new knowledge feature unit expands the scope of historical knowledge association, the knowledge evolution mode information corresponds to the knowledge expansion state; when a new knowledge feature unit changes the direction of historical knowledge association, the knowledge evolution mode information corresponds to the knowledge migration state. When a newly added knowledge feature unit forms a historical knowledge association relationship, the knowledge evolution mode information corresponds to a knowledge recombination state; when a newly added knowledge feature unit replaces a historical knowledge association relationship, the knowledge evolution mode information corresponds to a knowledge replacement state. Based on knowledge evolution path information and knowledge evolution pattern information, incremental knowledge state information is generated to indicate the direction of influence of newly added multimodal training data on historical multimodal knowledge state information.
5. The multimodal large model incremental joint training method as described in claim 4, characterized in that: The model adjustment region information includes, The knowledge evolution pattern information corresponding to the newly added knowledge features is determined based on the knowledge expansion state, knowledge migration state, knowledge reorganization state, and knowledge replacement state recorded in the incremental knowledge state information. Based on the knowledge evolution path information, determine the propagation path of the new knowledge features in the historical multimodal knowledge state information, and based on the knowledge influence direction information, determine the influence range corresponding to the new knowledge features; Based on the knowledge evolution mode information, the knowledge evolution path information, and the knowledge influence direction information, the knowledge propagation path corresponding to the newly added knowledge features is mapped to the visual knowledge region, text knowledge region, and cross-modal fusion region in the multimodal large model. The knowledge action area information is determined based on the degree of correlation between the newly added knowledge features and the visual knowledge region, text knowledge region, and cross-modal fusion region. Based on the knowledge application area information, the target adjustment area of the newly added knowledge features in the multimodal large model is determined, and the model adjustment area information is obtained.
6. The multimodal large model incremental joint training method as described in claim 5, characterized in that: The updated multimodal model parameters include, Based on the model adjustment region information, the changes in knowledge gain of newly added knowledge features in the visual knowledge region, text knowledge region, and cross-modal fusion region are statistically analyzed. The changes in knowledge gain are then continuously sorted according to the knowledge propagation direction recorded in the knowledge evolution path information to form a knowledge gain change sequence. Based on the knowledge gain change sequence, determine the knowledge gain change difference between adjacent knowledge regions, increase or decrease the parameter update frequency of the corresponding knowledge region according to the knowledge gain change difference, and generate a joint training strategy according to the adjusted parameter update frequency. According to the joint training strategy, the parameters of the visual knowledge region, text knowledge region, and cross-modal fusion region are updated sequentially according to the parameter update frequency. After completing parameter updates in the visual knowledge region, the parameter update trajectory corresponding to the visual knowledge region is written into the joint training strategy. After completing parameter updates in the text knowledge region, the parameter update trajectory in the joint training strategy is used to correct the parameter update path corresponding to the text knowledge region. After completing parameter updates in the cross-modal fusion region, the corrected parameter update path is used to adjust the parameter update rhythm corresponding to the cross-modal fusion region to obtain the incrementally updated multimodal model parameters. Based on the incrementally updated multimodal model parameters, the updated knowledge gain states corresponding to the visual knowledge region, text knowledge region, and cross-modal fusion region are calculated using the newly added multimodal training data. The updated knowledge gain states are then correlated with the corresponding knowledge gain states before incremental training to obtain the knowledge gain evolution sequence. When the knowledge gain evolution sequence changes, the parameter update frequency corresponding to the joint training strategy is readjusted according to the changed knowledge gain evolution sequence to obtain the updated multimodal model parameters.
7. The multimodal large model incremental joint training method as described in claim 6, characterized in that: The updated historical multimodal knowledge state information includes... Using the updated multimodal model parameters, extract the knowledge parameter change features corresponding to the updated multimodal model parameters to obtain the knowledge parameter change features. Reorganize the knowledge parameter change features according to the knowledge organization method corresponding to the historical multimodal knowledge state information to obtain the knowledge state update content. By utilizing the knowledge state update content, the knowledge state update content is associated with the knowledge association structure corresponding to the historical multimodal knowledge state information to obtain the updated historical knowledge state content, and the historical multimodal knowledge state information is updated to obtain the historical multimodal knowledge state information.
8. A multimodal large model incremental joint training system, based on the multimodal large model incremental joint training method according to any one of claims 1 to 7, characterized in that: This includes a preprocessing module, which acquires new multimodal training data to be added to the multimodal large model and preprocesses it to obtain a standardized incremental training dataset; The extraction module inputs the standardized incremental training dataset into the multimodal large model and obtains incremental multimodal features through the multimodal feature extraction network; The analysis module analyzes the correlation changes between newly added multimodal training data and historical multimodal knowledge based on incremental multimodal features and historical multimodal knowledge state information to obtain incremental knowledge state information. The adjustment module determines the target adjustment region corresponding to the newly added knowledge in the multimodal large model based on the incremental knowledge status information, and obtains the model adjustment region information; The training module generates a joint training strategy based on the model adjustment region information, and performs incremental joint training on the multimodal large model based on the joint training strategy to obtain the updated multimodal model parameters. The update module uses the updated multimodal model parameters to update the historical multimodal knowledge state information.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the multimodal large model incremental joint training method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the multimodal large model incremental joint training method according to any one of claims 1 to 7.