Knowledge base construction method and device based on multi-modal large language model
Through cross-modal embedding space alignment and positive and negative sample optimization mechanism, the InstructBLIP and LLaMA-adapter models are used to generate text descriptions of multimodal data, which solves the problem of multimodal data fragmentation and achieves efficient alignment and discriminability improvement of multimodal embedding space.
Patent Information
- Application Number
- CN202511228315.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing methods fail to fully utilize the generative potential of multimodal large language models (MLLMs), resulting in sparse cross-modal semantic associations, insufficient embedding space alignment accuracy, and the separation of multimodal data and multimodal text descriptions, causing the embedding center to deviate from the true category distribution.
Through cross-modal embedding space alignment and parameter freezing training mechanism, the InstructBLIP model and LLaMA-adapter model are used to generate text descriptions of multimodal data, and the discriminability of the embedding space is improved through a comparative learning framework and positive and negative sample optimization mechanism.
It achieves semantic alignment between multimodal data and text descriptions, improves the discriminability of the embedding space, makes similar samples compact and separates heterogeneous samples, reduces the amount of computation and retains generalization ability.
Smart Images

Figure CN120744846A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of knowledge base construction, and in particular to a method and device for constructing a knowledge base based on a multimodal large language model. Background Art
[0002] As a key component of multimodal representation learning, the knowledge base provides semantic priors for multimodal embedding and center positioning. Traditional methods rely on manual annotation or structured databases, but their scalability and semantic coverage are limited. With the rise of LLMs, their generative capabilities have begun to be used to automatically build knowledge bases. For example, through prompt engineering, LLMs can generate diverse text descriptions for each category, breaking through the limitation of traditional methods that rely on a single high-quality text source. Category descriptions are generated through GPT-4, and multimodal LLMs are combined to generate data-related text to build a hybrid knowledge base to enhance the semantic richness of the embedding center. However, descriptions generated by LLMs alone may contain irrelevant semantics (such as background details), causing the embedding center to deviate from the true category distribution. For example, for the "helicopter" category, the generated text may overemphasize the "blue sky" background and ignore key features such as the rotor structure.
[0003] In summary, most existing methods separate multimodal data from multimodal text descriptions, failing to fully activate the generative potential of multimodal large language models (MLLMs), resulting in sparse cross-modal semantic associations and insufficient embedding space alignment accuracy. Summary of the Invention
[0004] This application provides a knowledge base construction method based on a multimodal large language model, aiming to build a knowledge base based on categories and MLLMs, and improve the discriminability of embedded data and positioning centers by enhancing irrelevant descriptions.
[0005] According to a first aspect of the present application, a method for constructing a knowledge base based on a multimodal large language model is provided, comprising:
[0006] Acquire multimodal data and corresponding category labels, wherein the multimodal data includes at least one of images, videos, audio, point clouds, thermal imaging, events, and text;
[0007] Based on a preset prompt word template, a text description corresponding to the multimodal data is generated through a multimodal large language model, wherein the multimodal large language model includes an InstructBLIP model and an LLaMA-adapter model. The text description corresponding to the multimodal data is generated using the following formula:
[0008]
[0009] in, represents multimodal data input, is the text description corresponding to the multimodal data, Represents the generating function of the multimodal large language model, the generating function Training is performed according to a pre-trained multimodal encoder, wherein the pre-trained multimodal encoder includes the frozen visual encoder of the InstructBLIP model and the base encoder corresponding to the cross-modal adaptation layer of the LLaMA-adapter model;
[0010] The training process of the generator function of the multimodal large language model includes:
[0011] Cross-modal embedding space alignment: Based on the pre-trained multimodal encoder, the multimodal data and its corresponding text description are input into a contrastive learning framework, and the embedding space of the multimodal data is aligned by optimizing the contrastive learning loss function;
[0012] Parameter freezing and adaptation layer training: Freeze the main parameters of the pre-trained visual encoder and text encoder, and only train the lightweight adaptation layer consisting of a cross-modal attention module or query transformer to minimize the cross-entropy loss between the text description and the true category label;
[0013] Joint optimization: The contrastive learning loss and cross entropy loss are weighted to determine the total loss function, and the adaptation layer parameters are jointly optimized through the backpropagation algorithm until the model converges.
[0014] The step of aligning the embedding space of the multimodal data by optimizing the contrastive learning loss function comprises:
[0015] Mapping the input multimodal data to an intermediate feature representation based on the visual encoder and text encoder;
[0016] The intermediate features of each modality are projected into a shared semantic space through a parameterized mapping function to form a directly comparable semantic representation.
[0017] Construct positive and negative sample pairs based on data association relationships, where positive sample pairs include semantically consistent cross-modal data, and negative sample pairs include semantically inconsistent cross-modal data;
[0018] By maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the contrastive learning loss function is optimized to achieve multimodal embedding space alignment;
[0019] Calculating the cosine similarity between the generated text description and the corresponding category label through a pre-trained text encoder;
[0020] According to the text descriptions whose cosine similarity is higher than a preset threshold, a knowledge base of the corresponding category is constructed.
[0021] As a preferred embodiment, the generating of a text description corresponding to the multimodal data by using a multimodal large language model includes:
[0022] For image, event, and thermal imaging modality data, the InstructBLIP model’s visual-language alignment module generates text descriptions through a frozen visual encoder and a lightweight Querying Transformer.
[0023] For audio, video, and point cloud modal data, the cross-modal adaptation layer of the LLaMA-adapter model is used to map non-text modal features to the language model embedding space and generate text descriptions.
[0024] As a preferred embodiment, before generating a text description corresponding to the multimodal data through a multimodal large language model based on a preset prompt word template, the multimodal data needs to be preprocessed, including:
[0025] Normalize and spatially align image, video, and thermal imaging data to ensure that the input resolution is consistent with the requirements of the multimodal large language model.
[0026] Convert the audio data into a Mel-spectrogram and perform time slicing to generate fixed-length spectrum segments to ensure input dimension consistency;
[0027] The point cloud data is subjected to dimensionality reduction and feature extraction using voxelization or the farthest point sampling (FPS) algorithm.
[0028] As a preferred embodiment, the method further comprises:
[0029] Based on a preset prompt word template, generating a negative sample description that is unrelated to the semantics of the category label through the multimodal large language model;
[0030] Performing comparative learning on the negative sample description and the multimodal data to enhance the discrimination between different categories in the embedding space;
[0031] The negative sample description and its corresponding multimodal data features are stored in the knowledge base to enhance the category boundary discrimination of the embedding space.
[0032] According to a second aspect of the present application, a knowledge base construction device based on a multimodal large language model is provided, comprising:
[0033] a multimodal data acquisition module, configured to acquire multimodal data and corresponding category labels, wherein the multimodal data includes at least one of images, videos, audio, point clouds, thermal imaging, events, and text;
[0034] A text description generation module is configured to generate a text description corresponding to the multimodal data based on a preset prompt word template using a multimodal large language model, wherein the multimodal large language model includes an InstructBLIP model and an LLaMA-adapter model, and the text description corresponding to the multimodal data is generated using the following formula:
[0035]
[0036] in, represents multimodal data input, is the text description corresponding to the multimodal data, Represents the generating function of the multimodal large language model, the generating function Training is performed according to a pre-trained multimodal encoder, wherein the pre-trained multimodal encoder includes the frozen visual encoder of the InstructBLIP model and the base encoder corresponding to the cross-modal adaptation layer of the LLaMA-adapter model;
[0037] The training process of the generator function of the multimodal large language model includes:
[0038] Cross-modal embedding space alignment: Based on the pre-trained multimodal encoder, the multimodal data and its corresponding text description are input into a contrastive learning framework, and the embedding space of the multimodal data is aligned by optimizing the contrastive learning loss function;
[0039] Parameter freezing and adaptation layer training: Freeze the main parameters of the pre-trained visual encoder and text encoder, and only train the lightweight adaptation layer consisting of a cross-modal attention module or query transformer to minimize the cross-entropy loss between the text description and the true category label;
[0040] Joint optimization: The contrastive learning loss and cross entropy loss are weighted to determine the total loss function, and the adaptation layer parameters are jointly optimized through the backpropagation algorithm until the model converges.
[0041] The step of aligning the embedding space of the multimodal data by optimizing the contrastive learning loss function comprises:
[0042] Mapping the input multimodal data to an intermediate feature representation based on the visual encoder and text encoder;
[0043] The intermediate features of each modality are projected into a shared semantic space through a parameterized mapping function to form a directly comparable semantic representation.
[0044] Construct positive and negative sample pairs based on data association relationships, where positive sample pairs include semantically consistent cross-modal data, and negative sample pairs include semantically inconsistent cross-modal data;
[0045] By maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the contrastive learning loss function is optimized to achieve multimodal embedding space alignment;
[0046] A similarity calculation unit, configured to calculate the cosine similarity between the generated text description and the corresponding category label using a pre-trained text encoder;
[0047] The knowledge base construction unit is used to construct a knowledge base of a corresponding category according to the text descriptions whose cosine similarity is higher than a preset threshold.
[0048] According to a third aspect of the present application, an electronic device, at least one processor, and a memory communicatively connected to the at least one processor are provided; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute any of the above-described methods for constructing a knowledge base based on a multimodal large language model.
[0049] Compared with the prior art, this application achieves the following beneficial effects:
[0050] (1) This application uses a contrastive learning framework to map multimodal data and their text descriptions into a shared semantic space, achieving semantic alignment between heterogeneous modalities such as images, audio, and point clouds and text, and solving the problem of semantic separation between modalities in traditional methods.
[0051] (2) This application introduces a positive-negative sample pair optimization mechanism, which significantly improves the discriminability of multimodal embedding by maximizing the similarity of positive samples and minimizing the similarity of negative samples, making similar samples more compact in the embedding space and heterogeneous samples more separated.
[0052] (3) This application freezes the main parameters of the pre-trained encoder during model training and only trains lightweight adaptation layers such as the cross-modal attention module or query transformer, which greatly reduces the amount of computation while retaining the generalization ability of the training model.
[0053] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of the present application, nor are they intended to limit the scope of the present application. Other features of the present application will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The above and other features, advantages and aspects of the embodiments of the present application will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present invention and do not constitute a limitation of the present application. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0055] Figure 1A flow chart of a method for constructing a knowledge base based on a multimodal large language model according to an embodiment of the present application is shown;
[0056] Figure 2 The following is a flowchart of the knowledge base construction process according to an embodiment of the present application;
[0057] Figure 3 A schematic diagram of description generation based on the InstructBLIP model according to an embodiment of the present application is shown;
[0058] Figure 4 A schematic diagram of description generation based on LLaMA-adapter in an embodiment of the present application is shown;
[0059] Figure 5 A block diagram of a device for constructing a knowledge base based on a multimodal large language model according to an embodiment of the present application is shown;
[0060] Figure 6 A schematic diagram of an exemplary electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0061] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0062] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0063] like Figure 1 FIG. 1 is a flow chart of a method for constructing a knowledge base based on a multimodal large language model according to the present application. The method 100 includes:
[0064] S110: Acquire multimodal data and corresponding category labels, wherein the multimodal data includes at least one of images, videos, audio, point clouds, thermal imaging, events, and text.
[0065] In some embodiments, multimodal data generally comes from public datasets or custom collection devices (such as cameras, lidar, etc.). Public datasets can directly provide standardized category labels, and data collected by custom collection devices can be labeled through manual labeling, automated pre-labeling, or cross-modal association to ensure that the labels strictly correspond to the data content. For example:
[0066] Image / Video: Captured through optical imaging devices (such as RGB cameras), image frames or continuous video streams are extracted at a preset resolution (such as 1080p);
[0067] Point cloud: Scan the scene using a 3D sensing device (such as a LiDAR) at a set sampling frequency (e.g., 10 Hz) to obtain 3D spatial coordinates and reflection intensity information.
[0068] Thermal imaging: Collecting thermal radiation distribution maps in a specific band (such as 8-14μm) through infrared sensing equipment (such as thermal imagers);
[0069] Event data: A dynamic vision sensor (DVS) is used to record pixel-level brightness change event streams.
[0070] Furthermore, to ensure that the input format is compatible with subsequent models, the multimodal data obtained above needs to be preprocessed, including:
[0071] Normalize and spatially align image, video, and thermal imaging data to ensure that the input resolution is consistent with the requirements of the multimodal large language model.
[0072] Convert the audio data into a Mel-spectrogram and perform time slicing to generate fixed-length spectrum segments to ensure input dimension consistency;
[0073] The point cloud data is subjected to dimensionality reduction and feature extraction using voxelization or the farthest point sampling (FPS) algorithm.
[0074] The impulse noise in the event data or the sensor noise in the image is smoothed using Gaussian filtering or median filtering.
[0075] S120: Based on a preset prompt word template, generate a text description corresponding to the multimodal data through a multimodal large language model.
[0076] In some embodiments, traditional methods directly align multimodal features through text embeddings of category names, but they still face two major bottlenecks:
[0077] Semantic uniformity: A single category name cannot cover the diversity within the category (e.g., "laptop" may be associated with fine-grained attributes such as "keyboard" and "display");
[0078] Modality bias: Text-driven alignment relies too much on language priors, resulting in the loss of fine-grained semantics of modalities such as vision and auditory.
[0079] To solve this problem, Figure 2 As shown, this application first establishes a prompt word template ("Generate adescription of this {modal type} that is related to the {Category}."), then uses a multimodal large language model to generate a text description corresponding to the multimodal data, and finally constructs a knowledge base of the corresponding category based on the cosine similarity between the text description and the category label.
[0080] The multimodal large language model includes the InstructBLIP model and the LLaMA-adapter model. The text description corresponding to the multimodal data is generated by the following formula:
[0081]
[0082] in, represents multimodal data input, is the text description corresponding to the multimodal data, Represents the generative function of a multimodal large language model.
[0083] In some embodiments, generating a text description corresponding to the multimodal data using a multimodal large language model includes:
[0084] For image, event, and thermal imaging modal data, the InstructBLIP model is used to generate multimodal data descriptions. Specifically, for event data and thermal imaging data, key visual features (such as edges and contours) of their paired RGB images can be extracted. These visual features and corresponding category labels are input into the InstructBLIP model. The image features and language semantics are aligned through a cross-modal attention mechanism to generate fine-grained text descriptions. The exemplary description generation process is shown in the following figure. Figure 3 shown.
[0085] For audio, video, and point cloud modal data, the LLaMA-adapter model is used to generate descriptions for multimodal data. Specifically, feature encoding is performed on the audio, video, or point cloud modal data to extract its key modal features (such as point cloud geometry, etc.). The key modal features and corresponding category labels are input into the LLaMA-adapter model to generate context-related semantic descriptions. The exemplary description generation process is as follows: Figure 4 shown.
[0086] In some embodiments, the training process of the aforementioned multimodal large language models (MLLMs) includes:
[0087] Cross-modal embedding space alignment: Based on a pre-trained multimodal encoder, the multimodal data and its corresponding text description are input into a contrastive learning framework, and the embedding space of the multimodal data is aligned by optimizing the contrastive learning loss function;
[0088] Parameter freezing and adaptation layer training: Freeze the main parameters of the pre-trained visual encoder and text encoder, and only train the lightweight adaptation layer consisting of a cross-modal attention module or query transformer to minimize the cross-entropy loss between the text description and the true category label;
[0089] Joint optimization: The contrastive learning loss and cross entropy loss are weighted to determine the total loss function, and the adaptation layer parameters are jointly optimized through the backpropagation algorithm until the model converges.
[0090] In some embodiments, aligning the embedding space of the multimodal data by optimizing a contrastive learning loss function comprises:
[0091] Mapping the input multimodal data to an intermediate feature representation based on the visual encoder and text encoder;
[0092] The intermediate features of each modality are projected into a shared semantic space through a parameterized mapping function to form a directly comparable semantic representation.
[0093] Construct positive and negative sample pairs based on data association relationships, where positive sample pairs include semantically consistent cross-modal data, and negative sample pairs include semantically inconsistent cross-modal data;
[0094] By maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the contrastive learning loss function is optimized to achieve multimodal embedding space alignment.
[0095] S130: Calculate the cosine similarity between the generated text description and the corresponding category label through a pre-trained text encoder.
[0096] In some embodiments, the generated text description (such as "a white car turns left at the intersection") and the corresponding category label (such as "car") are input into a pre-trained text encoder (such as CLIP's text encoder or BERT) to generate an embedding vector for the text description and an embedding vector for the category label, respectively.
[0097] Furthermore, the cosine similarity between the embedding vector of the text description and the embedding vector of the category label is calculated. :
[0098] ;
[0099] in, 、 Represent the embedding vector of text description and the embedding vector of category label respectively.
[0100] S140: Building a knowledge base of the corresponding category based on the text descriptions whose cosine similarity is higher than a preset threshold.
[0101] In some embodiments, a similarity threshold (eg, 0.7) is set, and screening is performed based on the similarity threshold, retaining only text descriptions above the threshold and storing them in the knowledge base according to category.
[0102] In some embodiments, to ensure that the text description is semantically consistent with the original multimodal data, the method further includes:
[0103] Extract the feature vectors of multimodal data and calculate the image-text similarity score through the visual-text matching module of CLIP , set the image-text similarity threshold (such as 0.6), and remove descriptions with scores lower than the threshold to ensure that the text description is semantically consistent with the original data.
[0104] In some embodiments, to enhance the class boundary discrimination of the embedding space, the method further includes:
[0105] Based on a preset prompt word template, a negative sample description that is unrelated to the semantics of the category label is generated through a multimodal large language model; the negative sample description is compared with the multimodal data to enhance the discrimination between different categories in the embedding space; and the negative sample description and its corresponding multimodal data features are stored in the knowledge base to enhance the category boundary discrimination of the embedding space.
[0106] According to the above embodiments of the present application, the following technical effects are achieved:
[0107] (1) This application uses a contrastive learning framework to map multimodal data and their text descriptions into a shared semantic space, achieving semantic alignment between heterogeneous modalities such as images, audio, and point clouds and text, and solving the problem of semantic separation between modalities in traditional methods.
[0108] (2) This application introduces a positive-negative sample pair optimization mechanism, which significantly improves the discriminability of multimodal embedding by maximizing the similarity of positive samples and minimizing the similarity of negative samples, making similar samples more compact in the embedding space and heterogeneous samples more separated.
[0109] (3) This application freezes the main parameters of the pre-trained encoder during model training and only trains lightweight adaptation layers such as the cross-modal attention module or query transformer, which greatly reduces the amount of computation while retaining the generalization ability of the training model.
[0110] It should be noted that, for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0111] The above is an introduction to the method embodiment. The following is a further explanation of the solution described in this application through an apparatus embodiment.
[0112] Figure 5 FIG1 is a block diagram of a knowledge base construction device based on a multimodal large language model according to an embodiment of the present application. Figure 5 The device 50 includes:
[0113] a multimodal data acquisition module 510 for acquiring multimodal data and corresponding category labels, wherein the multimodal data includes at least one of images, videos, audio, point clouds, thermal imaging, events, and text;
[0114] A text description generation module 520 is configured to generate a text description corresponding to the multimodal data using a multimodal large language model based on a preset prompt word template;
[0115] A similarity calculation unit 530 is configured to calculate the cosine similarity between the generated text description and the corresponding category label using a pre-trained text encoder;
[0116] The knowledge base construction unit 540 is configured to construct a knowledge base of a corresponding category according to the text descriptions whose cosine similarity is higher than a preset threshold.
[0117] In the technical solution of this application, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0118] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.
[0119] Figure 6A schematic block diagram of an electronic device 600 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0120] The electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a ROM 602 or a computer program loaded from a storage unit 608 into a RAM 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O interface 605 is also connected to the bus 604.
[0121] Multiple components in the electronic device 600 are connected to the I / O interface 605, including an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0122] Computing unit 601 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 601 performs the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of method 100 described above may be performed. Alternatively, in other embodiments, computing unit 601 may be configured to perform method 100 in any other suitable manner (e.g., via firmware).
[0123] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0124] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0125] In the context of this application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0126] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0127] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0128] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0129] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.
[0130] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A knowledge base construction method based on a multimodal large language model, characterized in that: include: Acquire multimodal data and corresponding category labels, wherein the multimodal data includes at least one of images, videos, audio, point clouds, thermal imaging, events, and text; Based on a preset prompt word template, a text description corresponding to the multimodal data is generated through a multimodal large language model; wherein the multimodal large language model includes an InstructBLIP model and an LLaMA-adapter model, and the text description corresponding to the multimodal data is generated using the following formula: ; in, represents multimodal data input, is the text description corresponding to the multimodal data, Represents the generating function of the multimodal large language model, the generating function Training is performed according to a pre-trained multimodal encoder, wherein the pre-trained multimodal encoder includes the frozen visual encoder of the InstructBLIP model and the base encoder corresponding to the cross-modal adaptation layer of the LLaMA-adapter model; The training process of the generator function of the multimodal large language model includes: Cross-modal embedding space alignment: Based on the pre-trained multimodal encoder, the multimodal data and its corresponding text description are input into a contrastive learning framework, and the embedding space of the multimodal data is aligned by optimizing the contrastive learning loss function; Parameter freezing and adaptation layer training: Freeze the main parameters of the pre-trained visual encoder and text encoder, and only train the lightweight adaptation layer consisting of a cross-modal attention module or query transformer to minimize the cross-entropy loss between the text description and the true category label; Joint optimization: The contrastive learning loss and cross entropy loss are weighted to determine the total loss function, and the adaptation layer parameters are jointly optimized through the backpropagation algorithm until the model converges; The step of aligning the embedding space of the multimodal data by optimizing the contrastive learning loss function comprises: Mapping the input multimodal data to an intermediate feature representation based on the visual encoder and text encoder; The intermediate features of each modality are projected into a shared semantic space through a parameterized mapping function to form a directly comparable semantic representation. Construct positive and negative sample pairs based on data association relationships, where positive sample pairs include semantically consistent cross-modal data, and negative sample pairs include semantically inconsistent cross-modal data; By maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the contrastive learning loss function is optimized to achieve multimodal embedding space alignment; the cosine similarity between the generated text description and the corresponding category label is calculated through the pre-trained text encoder; According to the text descriptions whose cosine similarity is higher than a preset threshold, a knowledge base of the corresponding category is constructed.
2. The method according to claim 1, characterized in that Generating a text description corresponding to the multimodal data by using a multimodal large language model includes: For image, event, and thermal imaging modality data, the InstructBLIP model’s visual-language alignment module generates text descriptions through a frozen visual encoder and a lightweight Querying Transformer. For audio, video, and point cloud modal data, the cross-modal adaptation layer of the LLaMA-adapter model is used to map non-text modal features to the language model embedding space and generate text descriptions.
3. The method according to claim 1, characterized in that Before generating a text description corresponding to the multimodal data through a multimodal large language model based on a preset prompt word template, the multimodal data needs to be preprocessed, including: Normalize and spatially align image, video, and thermal imaging data to ensure that the input resolution is consistent with the requirements of the multimodal large language model. Convert the audio data into a Mel-spectrogram and perform time slicing to generate fixed-length spectrum segments to ensure input dimension consistency; The point cloud data is subjected to dimensionality reduction and feature extraction using voxelization or the farthest point sampling (FPS) algorithm.
4. The method according to claim 1, wherein The method further comprises: Based on a preset prompt word template, generating a negative sample description that is unrelated to the semantics of the category label through the multimodal large language model; Performing comparative learning on the negative sample description and the multimodal data to enhance the discrimination between different categories in the embedding space; The negative sample description and its corresponding multimodal data features are stored in the knowledge base to enhance the category boundary discrimination of the embedding space.
5. A knowledge base construction device based on a multimodal large language model, characterized in that: include: a multimodal data acquisition module, configured to acquire multimodal data and corresponding category labels, wherein the multimodal data includes at least one of images, videos, audio, point clouds, thermal imaging, events, and text; A text description generation module, configured to generate a text description corresponding to the multimodal data using a multimodal large language model based on a preset prompt word template; The multimodal large language model includes the InstructBLIP model and the LLaMA-adapter model, and the text description corresponding to the multimodal data is generated by the following formula: ; in, represents multimodal data input, is the text description corresponding to the multimodal data, Represents the generating function of the multimodal large language model, the generating function Training is performed according to a pre-trained multimodal encoder, wherein the pre-trained multimodal encoder includes the frozen visual encoder of the InstructBLIP model and the base encoder corresponding to the cross-modal adaptation layer of the LLaMA-adapter model; The training process of the generator function of the multimodal large language model includes: Cross-modal embedding space alignment: Based on the pre-trained multimodal encoder, the multimodal data and its corresponding text description are input into a contrastive learning framework, and the embedding space of the multimodal data is aligned by optimizing the contrastive learning loss function; Parameter freezing and adaptation layer training: Freeze the main parameters of the pre-trained visual encoder and text encoder, and only train the lightweight adaptation layer consisting of a cross-modal attention module or query transformer to minimize the cross-entropy loss between the text description and the true category label; Joint optimization: The contrastive learning loss and cross entropy loss are weighted to determine the total loss function, and the adaptation layer parameters are jointly optimized through the backpropagation algorithm until the model converges; The step of aligning the embedding space of the multimodal data by optimizing the contrastive learning loss function comprises: Mapping the input multimodal data to an intermediate feature representation based on the visual encoder and text encoder; The intermediate features of each modality are projected into a shared semantic space through a parameterized mapping function to form a directly comparable semantic representation. Construct positive and negative sample pairs based on data association relationships, where positive sample pairs include semantically consistent cross-modal data, and negative sample pairs include semantically inconsistent cross-modal data; By maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the contrastive learning loss function is optimized to achieve multimodal embedding space alignment; A similarity calculation module is used to calculate the cosine similarity between the generated text description and the corresponding category label through a pre-trained text encoder; The knowledge base construction module is used to construct a knowledge base of a corresponding category based on the text descriptions whose cosine similarity is higher than a preset threshold.
6. An electronic device, characterized in that: The electronic device comprises: At least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Data information label processing method of large language model
CN117453921A
Domain large model multi-modal knowledge base construction method based on feature representation
CN118779469A
Embedded intelligent visual language large model knowledge base construction and application method, equipment, medium and product
CN119476463A
News picture description method based on multi-modal retrieval enhancement generation
CN120336571A
Archive knowledge base construction and retrieval method and system based on multi-modal data fusion
CN120407703A
Cited By
Audio and video associated target analysis method and system fused with multi-modal scene understanding
CN121191068A
Point cloud generation method and device based on multi-modal information, equipment and storage medium
CN122089967A
Point cloud generation method and device based on multi-modal information, equipment and storage medium
CN122089967B