Method for expanding multi-modal AI large model based on traditional database

By configuring the multimodal big model API plug-in and the citus distributed plug-in in the PostgresSQL database, combined with the multimodal AI big model, the problem that PostgresSQL does not support multiple machines and data sharding is solved, and a distributed database with AI capabilities is realized, improving database performance and expansion capabilities.

CN120336289APending Publication Date: 2025-07-18SHANGHAI DIGITAL GOVERNANCE RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510444530.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing PostgresSQL database does not support multi-machine and data sharding, and the multi-modal AI big model and database are two systems, and a distributed database solution lacks AI capabilities.

Method used

By configuring the Postgres plug-in of the multimodal big model API, combined with the citus distributed plug-in, a distributed database is built to realize the integration of the multimodal AI big model and PostgresSQL, and supports image content understanding function.

Benefits of technology

Without changing the existing architecture, expand database capabilities at low cost, improve performance, support massive data applications, reduce enterprise costs, and realize database usage of AI capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336289A_ABST
    Figure CN120336289A_ABST
Patent Text Reader

Abstract

The invention discloses a method for expanding a multi-modal AI large model based on a traditional database. The method comprises the following steps: establishing a multi-modal large model service; a Postgres plug-in which is in butt joint with the multi-mode large model API is configured; based on N servers, a Postgres cluster is established, and a distributed database is realized; the method comprises the steps that a Postgres main node is connected, an SQL file is used for creating a database table, and the database table is converted into a distributed table through a citus distributed plug-in: a trigger is activated for the database table on which AI large model operation is to be executed through the Postgres plug-in; inserting data into the database table which activates the trigger by using an SQL (Structured Query Language) statement; aiming at pictures which cannot be understood by a current multi-mode large model, requests are resent based on different problems, and targeted results of different problems are obtained. According to the method, a multi-modal large-model distributed database with an AI function is created.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of databases, and particularly to a method for expanding a multi-modal AI large model based on a traditional database. Background Art

[0002] A large model refers to a deep learning model with billions to hundreds of billions of parameters, such as the GPT (Generative Pretrained Transformer) model; while a multi-modal large model refers to a large deep learning model that performs well in processing various different modal (such as text, image, audio, etc.) data. These models usually integrate a variety of data processing technologies, such as natural language processing, computer vision, and speech recognition, to achieve more comprehensive understanding and processing. Multi-modal large models have a wide range of applications in various fields, such as visual question answering, image captioning, video understanding, etc.

[0003] PostgresSQL (abbreviated as Postgres) is an open-source relational database management system (RDBMS) that has powerful functions and features and is widely used in various scenarios such as enterprise applications, web applications, and data analysis. The PostgresSQL database only supports single-machine systems and does not support multi-machine and data sharding.

[0004] A distributed database refers to a database system that stores data in multiple physical or logical locations, which can be distributed on different computers, servers, or data centers. The distributed database system aims to solve the problems that a single database system is difficult to handle large-scale data and high-concurrency requests. By dispersing data storage and processing, the performance, availability, and scalability of the system are improved.

[0005] The traditional single-machine database PostgresSQL has strong expansion capabilities, while the commonly used distributed databases in the current market have weak expansion capabilities; and although multi-modal AI large models can provide the ability to process and understand text and pictures, they are two separate systems from the database. There is currently no distributed database that includes AI capabilities in the market. Therefore, the existing technology can no longer meet the needs of people at the present stage. Based on the current situation, it is urgent to improve the existing technology. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for expanding a multi-modal AI large model based on a traditional database to solve the problems raised in the above background art.

[0007] The present invention provides the following technical solution: A method for expanding a multi-modal AI large model based on a traditional database, including: Step S100: Set up a multi-modal large model service, configure the API standard interface of the multi-modal large model, which is used to receive requests from an external database to send to the large model for AI inference on pictures and text. The multi-modal large model includes Tongyi Qianwen multi-modal large model Qwen-VL; Preferably, in the sent AI inference request, it includes the specified multi-modal large model name, the picture link, and the question about the picture corresponding to the picture link.

[0008] Step S200: Configure the Postgres plug-in that docks with the multi-modal large model API. The Postgres plug-in is used to extend the functions of the PostgresSQL database; Preferably, the Postgres plug-in includes: a control file, an SQL file, and a c file; among them, The control file is used for Postgres to identify the Postgres plug-in and the action of loading the Postgres plug-in; The c file is used for the specific code to dock with the multi-modal large model; The SQL file is used to create a database method based on the functions in the c file when the plug-in is started, and specify this method when the database table activates the trigger.

[0009] Step S300: Based on multiple servers, build a Postgres cluster to implement a distributed database; Preferably, install and activate the Citus distributed plug-in on each server, select one server as the master node, and other servers as worker nodes, and connect the worker nodes to the master node through the connection function provided by the Citus distributed plug-in; Step S400: Dock with the Postgres master node, use the SQL file to create a database table, and convert this database table into a distributed table through the Citus distributed plug-in: Preferably, the Citus distributed plug-in automatically creates a sharded table of this table on the worker nodes, so that all insert, delete, update, and query operations performed through the master node will be automatically sharded to the worker nodes for parallel execution; Step S500: Based on the Postgres plug-in configured in Step S200, activate the trigger for the database table to be operated on by the AI large model; Step S600: Use SQL statements to insert data into the database table with the activated trigger. The data automatically processes data sharding, and performs picture understanding by accessing the multi-modal large model service, and writes the returned results into the database table; Step S700: For images that cannot be understood by the current multi-modal large model, resend requests based on different questions to obtain targeted results for different questions.

[0010] The present invention has the following beneficial effects: Based on the open-source PostgresSQL database, the present invention accesses the capabilities of the multi-modal AI large model into the database by expanding the plugin capabilities and configuring the Postgres plugin, and then combines with the citus distributed plugin to create a distributed database of multi-modal large model with AI capabilities; By adding AI capabilities on the PostgresSQL database, the present invention supports the image content understanding function, enabling enterprises to expand their database capabilities at low cost without changing their own architectures. For database users, they can achieve the ability to use AI without understanding the specific usage postures of AI, reducing the costs of enterprises and being able to incubate more business scenarios; The present invention supports horizontal expansion and data sharding through the distributed database, and can significantly improve performance in the application of massive data. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a schematic diagram of the method for expanding a multi-modal AI large model based on a traditional database according to an embodiment of the present invention; Figure 2 It is a schematic diagram of the method according to an embodiment of the present invention in the application of horizontal expansion of the architecture. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art within the scope of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0013] In the embodiments of the present invention, Citus mentioned is an open-source extension that can convert PostgresSQL into a distributed database system. It is built on PostgresSQL and enables PostgresSQL to handle large-scale data and high-concurrency requests. Citus distributes and stores data on multiple nodes through horizontal expansion, and uses parallel query and parallel computing to accelerate data processing and query operations.

[0014] The PostgresSQL extension is a set of shared codes that can add new functions to the database; for example, installing the extension through the CREATE EXTENSION command can simplify database management and facilitate migration.

[0015] Reference Figure 1 , the present invention provides the following technical solution: a method for expanding a multi-modal AI large model based on a traditional database, and the specific steps include: Step S100: Build a multi-modal large model service, configure the API standard interface of the multi-modal large model, and use it to receive requests for AI inference of pictures and texts sent by an external database to the large model.

[0016] Specifically, use an HTTP (HyperText Transfer Protocol) request to perform AI inference on pictures and texts by the large model. During inference, it is necessary to call the specified large model name, picture link, and content about the picture. Among them, the multi-modal large model includes Tongyi Qianwen multi-modal large model Qwen-Vl, and API refers to the large model inference API specification promoted by the openai company.

[0017] In this embodiment, for the multi-modal large model that provides the repository code and model files, obtain the corresponding repository code and model files of the large model, and call the model through the interface; for the large model that does not provide the repository code, generally call the model through the API interface of the large model. Send an HTTP request to the server, and the information of the request includes: the picture image url that needs to be understood by the large model and the question for asking the large model; after the large model performs inference on the picture and text, return the understanding result answer of the large model through the interface, which respectively correspond to image_url, question, and answer in the database table. That is, the plugin reads the content of the image_url and question columns in the database table, puts them into the http request, and after requesting the large model service interface, writes the interface return result back to the answer column to realize the understanding of pictures and texts; exemplarily, the external database of the present invention is, for example, a PostgresSQL relational database.

[0018] Step S200: Configure a Postgres plugin that docks with the API of the multi-modal large model. This plugin is used to expand the functions of the PostgresSQL database. The Postgres plugin includes: a control file, an SQL file, and a c file; among them, the control file is used to identify the Postgres plugin and the actions for loading the Postgres plugin; the c file is used for the specific code for docking with the multi-modal large model; the SQL file is used to configure the functions of the plugin based on the functions and types in the c file when the Postgres plugin is started, and specify the functions when the database table activates the trigger to realize the extension and implementation of custom functions.

[0019] Specifically, after writing a Postgres plug-in that interfaces with a multi-modal large model API, the functions of the PostgreSQL database are extended after compilation and installation according to the Postgres plug-in specifications.

[0020] Step S300: Based on N servers, build a PostgreSQL cluster to implement a distributed database; Install and activate the Citus distributed plug-in on each server. Select one server as the master node and the other servers as the worker nodes of this master node. Connect the worker nodes to the master node through the connection function provided by the Citus distributed plug-in.

[0021] Step S400: Interface with the PostgreSQL master node, use SQL to create a database table, and convert the created database table into a distributed database table through the Citus distributed plug-in: Specifically, the database table created by the master node includes an ID column. Shard the database table based on the ID column, and create the same database table, that is, a sharded table, on the worker nodes through the Citus distributed plug-in, so as to convert the database table into a distributed database table, enabling all insert, delete, update, and query operations performed through the master node to be automatically sharded to the worker nodes for parallel execution. When the master node inserts data, the data is evenly sent to the worker nodes of the master node through sharding, and the worker nodes will respectively call the multi-modal large model for AI inference; in this embodiment, when the user operates, they will not perceive at all that it is the worker nodes that are executing in parallel, and only need to operate on the master node as if operating a single-machine database normally.

[0022] Step S500: Based on the Postgres plug-in configured in step S200, activate the trigger for the database table to be operated on by the AI large model. When inserting data into a database table, if the Postgres plugin detects that the database table is about to perform an operation related to the AI large model, it automatically activates a trigger and executes the function of the Postgres plugin. For example, the Postgres plugin reads the inserted row data and uses fields previously agreed upon with the database table, such as image_url and question, to ask questions to the AI large model service, and writes the result answer back into the fields of the database table. Here, image_url is the picture link; question is the question requested from the large model for the picture of image_url, used to understand the content in the picture; for example: "What scene is in the picture and how many people are there?", "Is there a traffic jam in the picture?"; answer is the answer of the multimodal large model's understanding of the picture content based on the said question; for example, "In the picture is an entrance and exit of a subway station, there are about 20 people", "In the picture, the vehicles are dense and the spacing is very small, and the congestion is extremely serious."

[0023] Step S600: Use an SQL statement to insert data into the database table with the activated trigger. The data automatically processes data sharding, performs picture understanding by accessing the multimodal large model service, and writes the returned result into the database table.

[0024] In the embodiment of the present invention, the database table for expanding the large model service includes fields such as image_url, question, and answer. An HTTP request is sent to the server, and the information of the request includes: the picture imageurl that needs to be understood by the large model and the question for asking the large model. After the large model infers the picture and text, the understanding result answer of the large model is returned through the interface, corresponding to image_url, question, and answer in the database table respectively. That is, the plugin reads the content of the image_url and question columns in the database table, puts them into the http request, and after requesting the large model service interface, writes the interface return result back to the answer column to realize the understanding of the picture and text. In this embodiment, only traditional SQL is used throughout the use process of the entire database, that is, the AI capability is completed at the same time.

[0025] Step S700: For pictures that the current multimodal large model cannot understand, resend requests based on different questions proposed to obtain targeted results for different questions.

[0026] In this embodiment, multiple database tables can also be created to activate multiple triggers to identify different scenarios. For example, trigger A for identifying general objects (such as people and vehicles), and trigger B for identifying pictures in a certain vertical domain. These two plugins respectively point to the large model services in the corresponding specialized directions. Therefore, when doing specific business, it is necessary to select to activate trigger A or trigger B on the database table according to the scenario.

[0027] Generally, multi-modal large models can recognize all things, so the AI capabilities can cover various different business scenarios. However, in some fields (such as the military industry), there are very few relevant pictures on the Internet. Therefore, when it is necessary to perform business recognition on this type of pictures, the large model cannot recognize well. At this time, it is necessary to retrain the large model based on the pictures that cannot be understood, train these pictures that cannot be understood into the model, form a new vertical domain model and use it for the understanding of pictures in the vertical domain.

[0028] To avoid reducing the accuracy of the large model in recognizing previously identified objects, it is necessary to retain the base model and the vertical domain model. The base model refers to a model that has been trained and has a certain degree of generality, and can handle various types of data and tasks; the vertical domain model refers to a large language model that has been specially trained and optimized in a specific field or industry, and has the characteristics of domain specialization. Among them, for 90% of the general scenarios, by asking different questions (questions) about a large amount of picture and text data and making targeted answers, the base model (and the corresponding plugin) can be obtained after retraining, and all things can be recognized; for 10% of the vertical domain scenarios, after obtaining the vertical domain model (and the corresponding plugin) through the general multi-modal large model of the base plus the training of vertical domain samples, it can be recognized. In this embodiment, there may be multiple major categories of plugins (base, vertical domain A, vertical domain B). Different plugins are activated according to different scenarios to dock with the general or vertical domain large model. In each major category, various small category scenarios can be recognized by asking questions.

[0029] The present invention is illustrated by an embodiment of the execution steps of the present invention: for the open-source multi-modal large model (i.e., the general large model), taking the Tongyi Qianwen multi-modal large model (Qwen-VL) of Alibaba as an example, on a machine with a GPU (assuming the IP is 172.20.101.222).

[0030] First, download and obtain the repository code Qwen-VL.git and the model file Qwen-VL-Chat of Qwen-VL.

[0031] Start the multi-modal large model service, call the standard interface promoted by OpenAI, and send an http request to the server through the plugin corresponding to the large model to achieve image understanding. For example, send a request through the http_post request interface, and the request information includes the image image_url that needs to be understood by the large model. Specifically, it is image_1.jpg in the images folder of the machine with the IP address 172.20.101.222; and the question question asked to the Qwen-VL large model, specifically such as "Please describe the time, place, and content in the picture". That is, the plugin reads the content in the image_url and question columns in the table, puts it into the http request, requests the large model service interface, and then writes the interface return result back to the answer column.

[0032] Exemplarily, for some pictures in the vertical domain scenario of a specific user, when the current large model cannot accurately understand the pictures, it is necessary to retrain the vertical domain samples based on the general multi-modal large model service. For example, first prepare some scenario pictures and questions for which the large model's recognition ability is insufficient, that is, the relevant knowledge samples in the vertical domain, and form files in the form of pictures, questions, and correct answers (or results). In the repository directory of the large model, train these samples, specifically including: Such as the picture image_url; the question (question) corresponding to the picture, such as "What breed is the dog in the picture?"; the correct answer (answer) corresponding to the picture, such as "The dog in the picture is a Labrador."; and again, in the picture image_url; the question (question) corresponding to the picture, such as "Which places are in the picture?"; the correct answer (answer) corresponding to the picture, such as "The first picture is the city skyline of Chongqing, and the second picture is the skyline of Beijing." After training, output the large model file to form a vertical domain large model that can better recognize the pictures in the training samples. A corresponding plugin can be written for the vertical domain large model, such as the mllm_special plugin, to connect to the services provided by the vertical domain large model. Then, the database table selects to reactivate the general plugin, such as mllm, or the vertical domain plugin, such as mllm_special, according to the specific application scenario.

[0033] As an alternative embodiment of the present invention, this embodiment is used to illustrate the process of implementing the method of the present invention. First, a multi-modal large model service is built, and a standard openai api specification is provided for an external database. Taking the Postgres relational database as an example, an exemplary plug-in named MLLM is written, which includes a control file, an SQL file, and a c file. Exemplarily, by activating the MLLM plug-in to interact with the large model Qwen-VL, the data returned by the large model is obtained, including the image url address (image_url) in the request and the question (question) asked to the large model, and the result (answer) feedback from the large model is written back to the database. Suppose there is a database table my_table, which contains four columns: id, image_url, question, and answer, where id refers to the ID of the inserted row data, image_url refers to the image URL address, question refers to the question asked to the large model, and answer refers to the answer obtained after accessing the large model. Then, a trigger is created on the database table my_table to execute the function implemented in the MLLM plug-in. Every time a piece of data is inserted into the database table, the trigger in the plug-in will send the image and the question to the multi-modal large model and write the result back to the answer column, achieving the effect of automatically accessing the large model for image content understanding. When developers use this AI database, they only need to insert data according to traditional SQL to achieve the effect of the AI large model without having to pre-learn the knowledge of the AI large model. The MLLM file is compiled and installed in the contrib directory of the PostgresSQL database to load this plug-in in the PostgresSQL database.

[0034] In this embodiment, horizontal architecture expansion and data sharding are also supported. As Figure 2 shown, there are three servers A, B, and C, each installed with the PostgresSQL database, and the MLLM plug-in and the Citus distributed plug-in are loaded. When building a cluster, select one of the servers A as the master node, that is, the Postgres master node. Through Citus distribution, servers B and C are added as worker nodes on the basis of the master node, namely Postgres worker node 1 and Postgres worker node 2. The added worker nodes include the IP of the worker node and the PORT information of the Postgres port.

[0035] In this embodiment, the open-source Postgres database is used and started on three servers respectively. Assume the names of the three servers are pg1, pg2, and pg3; their IPs and ports correspond to 172.20.120.1:5432, 172.20.120.2:5432, and 172.20.120.3:5432 respectively. Install and activate the Citus distributed plugin to the PostgreSQL database on the three servers to turn database tables into distributed tables and use the id column for sharding. The Citus distributed plugin automatically creates the my_table table on the worker nodes and loads triggers to achieve data insertion through the master node pg1. For example, execute the SQL statements: select citus_add_node('172.20.120.2', 5432) and select citus_add_node('172.20.120.3', 5432) on pg1, which incorporates the two nodes pg2 and pg3 into pg1 as its worker nodes, and pg1 becomes the master node at this time. After creating a database table through traditional SQL on pg1, assume the database table named test1, which contains an id column. By running the SQL statement: select create_distributed_table('test1', 'id'), the same test1 database table is automatically created on pg2 and pg3 through the Citus distributed plugin; then the external entry of the entire distributed database is pg1. For example, for data insertion operations on the test1 table, pg1 performs hashing through the id column to determine whether to forward it to pg2 or pg3 for insertion and storage. For query operations on the test1 table, pg1 will escape the query SQL and forward it to pg2 and pg3, and then unify the results obtained from pg2 and pg3 to return the final result. pg1 performs hashing through the id column. Specifically, when pg1 receives a data insertion request for the test1 table, it checks the value of the id column specified in the request, then performs a hash operation (hash) on the id value to generate a hash value. Based on the generated hash value, pg1 determines whether this specific insertion operation is transferred to pg2 or pg2 through the routing rules.

[0036] In this embodiment, the distributed database evenly distributes data to the worker nodes by sharding through the id column, and the worker nodes respectively call the multi-modal large model for AI inference; for example, insert the picture data to be analyzed into the my_table table through the following SQL statement: Insert into my_table(id, image_url, question) values (1,'http: / / your_image_path', 'What is in the picture'); Then write the returned result into the answer column of the inserted row in the database, that is, obtain the content result analyzed by the AI model.

[0037] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for extending a multi-modal AI large model based on a traditional database, characterized in that the method Including: Build a multi-modal large model service and configure the API standard interface of the multi-modal large model to receive requests from an external database to send to the large model for AI inference on pictures and text. The multi-modal large model includes Tongyi Qianwen multi-modal large model Qwen-VL; Configure a Postgres plugin that docks with the API of the multi-modal large model. The Postgres plugin is used to extend the functions of the PostgresSQL database; Based on N servers, build a PostgresSQL cluster to implement a distributed database. Among them, the PostgresSQL databases on the N servers all contain the activated citus distributed plugin. Select one of the servers as the master node and add other servers as worker nodes of the master node. The worker nodes contain the worker node IP and the port PORT information in PostgresSQL. N is a natural number greater than or equal to 3. Dock with the PostgresSQL master node, use SQL to create a database table, and convert the created database table into a distributed database table through the citus distributed plugin; Based on the configured Postgres plugin, activate the trigger for the database table to be operated on by the AI large model; Use SQL statements to insert data into the database table with the activated trigger, understand the picture by accessing the multi-modal large model service, and write the returned result back to the current database table; For pictures that the current multi-modal large model cannot understand, resend requests based on different questions proposed to obtain targeted results for different questions.

2. The method for extending a multi-modal AI large model based on a traditional database according to claim 1, wherein: In the request from the external database to the large model for AI inference on pictures and text, it includes the specified multi-modal large model name, picture link, and questions about the picture corresponding to the picture link.

3. The method for expanding a multi-modal AI large model based on a traditional database according to claim 1, wherein: The Postgres plugin includes a control file, an SQL file, and a c file; wherein, The control file is used to identify the Postgres plugin and the actions for loading the Postgres plugin; The c file is used for the specific code to dock with the multi-modal large model; The SQL file is used to configure the plugin functions based on the functions and types in the c file when the plugin is started, and specify the functions when the database table activates the trigger to achieve the extension of custom functions; Interact with the multi-modal large model through the Postgres plugin to obtain the result of the http request returned by the multi-modal large model.

4. The method for extending a multi-modal AI large model based on a traditional database according to claim 1, wherein: When building the Postgres cluster, it also includes connecting the worker nodes to the master node through the connection function function of the citus distributed plugin.

5. The method for extending a multi-modal AI large model based on a traditional database according to claim 1, wherein: The database table is converted into a distributed database table, including: the database table created by the master node includes an ID column, the database table is sharded based on the ID column, and the same database table is created on the worker nodes through the Citus distributed plugin, so that all the insert, delete, update, and query operations performed through the master node are automatically sharded to the worker nodes for parallel execution.

6. The method for expanding a multi-modal AI large model based on a traditional database according to claim 5, wherein When the master node inserts data, the distributed database table evenly sends the data to the worker nodes of the master node through the sharding, and the worker nodes respectively call the multi-modal large model to perform AI inference.

7. The method for expanding a multi-modal AI large model based on a traditional database according to claim 1, wherein: During the process of activating the trigger, when each piece of data is inserted into the database table, a function of the Postgres plugin is automatically executed, including the Postgres plugin reads the inserted row data, uses the fields agreed with the database table in advance, asks questions to the AI large model service, and writes the result back into the fields of the database table.

8. The method for expanding a multi-modal AI large model based on a traditional database according to claim 1, wherein: For pictures that the current multi-modal large model cannot understand, it also includes activating multiple triggers by creating multiple database tables to identify different scenarios.

9. The method for expanding a multi-modal AI large model based on a traditional database according to claim 8, characterized in that: The multi-modal large model includes two types: a general base model and a vertical domain model. Among them, The base model is trained through a large amount of pictures and corresponding text Q&A data, and is used to identify general domain scenarios; The vertical domain model is trained through the base model and vertical domain samples, and is used to identify vertical domain scenarios. The vertical domain includes one or more domain categories, and the vertical domain samples are pictures and corresponding text Q&A data in the corresponding domain.

10. The method for expanding a multi-modal AI large model based on a traditional database according to claim 9, wherein: The vertical domain model samples at least include pictures, the questions corresponding to the pictures, and the correct answers corresponding to the pictures, and are formed into files in the form of pictures, questions, and correct answers, and are stored in the warehouse directory of the multi-modal large model for training.