Multi-modal data storage method and device based on Surreal DB

By uniformly storing and retrieving multimodal data in SurrealDB, the problem that traditional database systems cannot efficiently manage structured and unstructured data is solved, and efficient multimodal data processing and query are achieved.

CN120277067APending Publication Date: 2025-07-08特赞(上海)信息科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510321623.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional database systems cannot efficiently store and manage structured and unstructured data at the same time, resulting in complex and inefficient multimodal data processing.

Method used

SurrealDB is used as the underlying storage engine, and data of different modalities are processed and related relationships are established through vectorization, text, pictures, audio, video and other data are stored uniformly, and searched using full-text index and vector index.

Benefits of technology

It realizes unified storage and efficient retrieval of multimodal data, reduces processing complexity, and facilitates cross-modal query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277067A_ABST
    Figure CN120277067A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a Surreal DB-based multi-modal data storage method and device, and the method comprises the steps: extracting characters and / or pictures from document, webpage, audio and video data, carrying out the vectorization processing of the characters and / or pictures, and obtaining the vector representation of different modal data; calculating the similarity between the vector representations, and establishing an association relationship between different modal data based on the similarity; and uniformly storing the text data, the picture data, the document data, the webpage data, the audio data and the video data in a Surreal DB database in a full-text index and vector index mode based on the incidence relation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of data processing, and more specifically, to a multi-modal data storage method and apparatus based on SurrealDB. Background Art

[0002] Traditional database systems usually store and process structured data (such as tabular data) and unstructured data (such as text, pictures, etc.) separately, and cannot support multiple retrieval methods simultaneously, resulting in complex and inefficient management and retrieval processes, and unable to meet the complex requirements of processing cross-modal data. How to efficiently vectorize and store large-scale multi-modal data and ensure high efficiency during query is a challenge. Summary of the Invention

[0003] To solve key technical problems such as the efficiency, accuracy, and scalability of multi-modal data storage and retrieval, the embodiments described in this document provide a multi-modal data storage method, apparatus, and computer-readable storage medium storing a computer program based on SurrealDB.

[0004] According to a first aspect of the present disclosure, there is provided a multi-modal data storage method based on SurrealDB, including: extracting text and / or pictures from documents, web pages, audio, and video data, performing vectorization processing on the text and / or pictures to obtain vector representations of different modal data; calculating the similarity between the vector representations, and establishing an association relationship between different modal data based on the similarity; and storing text, pictures, documents, web pages, audio, and video data in a SurrealDB database in a unified manner based on the association relationship in the form of full-text indexing and vector indexing.

[0005] In some embodiments of the present disclosure, extracting text and / or pictures from documents, web pages, audio, and video data, and performing vectorization processing on the text and / or pictures to obtain vector representations of different modal data includes: extracting audio, frame pictures, and text from video, web page, or document data, performing vectorization processing on the extracted audio, frame pictures, and text to obtain corresponding vector representations; performing vectorization and word segmentation processing on the text to generate text vectors and text keywords; extracting text descriptions and feature vectors for pictures, and converting the text descriptions of the pictures into vectors; using a speech recognition model to convert audio into text, and converting the text corresponding to the audio into vectors.

[0006] In some embodiments of the present disclosure, extracting audio, frame pictures, and text from video, web page, or document data includes: using FFmpeg to extract subtitle text, audio, and frame pictures from video; using a web crawler tool to parse the web page source code to extract text, embedded pictures, and audio in the web page; using different document parsing libraries to extract text and pictures in the document, and extract embedded audio.

[0007] In some embodiments of the present disclosure, the text is vectorized and tokenized to generate text vectors and text keywords, including: translating non-English text data into English text through a machine translation model; using a natural language processing model to tokenize the text to obtain the tokens in the text, converting each token of the text into a vector, and extracting text keywords by the TF-IDF method.

[0008] In some embodiments of the present disclosure, text descriptions and feature vectors are extracted from pictures, and the text descriptions of the pictures are converted into vectors, including: extracting the feature vectors of the pictures through a convolutional neural network, where the feature vectors include the color, shape, and texture features of the pictures; generating text descriptions of the pictures through a picture description generation model for semantic understanding and description of the pictures, where the text descriptions include the objects and scenes in the pictures; and using a natural language processing model to convert the text descriptions into text vectors.

[0009] In some embodiments of the present disclosure, audio, frame pictures, and text are extracted from video, web pages, or document data, and the extracted audio, frame pictures, and text are vectorized to obtain corresponding vector representations, including: tokenizing and vectorizing the text extracted from video, web pages, or document data to obtain text vectors; extracting text descriptions and feature vectors from the frame pictures extracted from video, web pages, or document data to obtain picture feature vectors and text description vectors; converting the audio extracted from video, web pages, or document data into text using a speech recognition model, and converting the text corresponding to the audio into speech vectors.

[0010] In some embodiments of the present disclosure, the similarity between vector representations is calculated, and the association relationship between different modal data is established based on the similarity, including: using cosine similarity or Euclidean distance or Manhattan distance to calculate the similarity value between different modal data vectors, and if the similarity value between two vectors exceeds a preset threshold, there is an association relationship between the contents corresponding to these two vectors.

[0011] In some embodiments of the present disclosure, based on the association relationship, in the form of full-text indexing and vector indexing, text, pictures, documents, web pages, audio, and video data are uniformly stored in the SurrealDB database, including: creating a data table in SurrealDB, where the data table includes fields for full-text indexing and fields for vector indexing, and representing the association relationship by constructing foreign key relationships, tags, or metadata in the data table; inserting the data extracted from documents, web pages, audio, and video and their corresponding vector representations into the SurrealDB data table for full-text indexing of text data and vector indexing of the vectors of picture, audio, and video modal data.

[0012] According to a second aspect of the present disclosure, there is provided a multi-modal data storage device based on SurrealDB. The device includes at least one processor; and at least one memory storing a computer program. When the computer program is executed by the at least one processor, the device is caused to: extract text and / or pictures from documents, web pages, audio, and video data, perform vectorization processing on the text and / or pictures to obtain vector representations of different modal data; calculate the similarity between the vector representations, and establish an association relationship between different modal data based on the similarity; and based on the association relationship, store text, pictures, documents, web pages, audio, and video data in a SurrealDB database in a full-text index and vector index manner.

[0013] In some embodiments of the present disclosure, when the computer program is executed by the at least one processor, the device is caused to extract text and / or pictures from documents, web pages, audio, and video data, perform vectorization processing on the text and / or pictures to obtain vector representations of different modal data by the following operations: extract audio, frame pictures, and text from video, web page, or document data, perform vectorization processing on the extracted audio, frame pictures, and text to obtain corresponding vector representations; perform vectorization and tokenization processing on the text to generate text vectors and text keywords; extract text descriptions and feature vectors for pictures, and convert the text descriptions of the pictures into vectors; use a speech recognition model to convert audio into text, and convert the text corresponding to the audio into a vector.

[0014] In some embodiments of the present disclosure, when the computer program is executed by the at least one processor, the device is caused to perform vectorization and tokenization processing on the text to generate text vectors and text keywords by the following operations: translate non-English text data into English text through a machine translation model; use a natural language processing model to perform tokenization processing on the text to obtain tokens in the text, convert each token of the text into a vector, and extract text keywords by the TF-IDF method.

[0015] In some embodiments of the present disclosure, when the computer program is executed by the at least one processor, the device is caused to extract text descriptions and feature vectors for pictures, and convert the text descriptions of the pictures into vectors by the following operations: extract feature vectors of pictures through a convolutional neural network, where the feature vectors include color, shape, and texture features of the pictures; perform semantic understanding and description of the pictures through a picture description generation model to generate text descriptions of the pictures, where the text descriptions include objects and scenes in the pictures; and use a natural language processing model to convert the text descriptions into text vectors.

[0016] In some embodiments of the present disclosure, when the computer program is executed by at least one processor, it causes the device to extract audio, frame pictures, and text from video, web pages, or document data through the following operations: use FFmpeg to extract subtitle text, audio, and frame pictures in the video; use a web crawler tool to parse the web page source code to extract text, embedded pictures, and audio in the web page; use different document parsing libraries to extract text and pictures in the document, and extract embedded audio.

[0017] In some embodiments of the present disclosure, when the computer program is executed by at least one processor, it causes the device to perform vectorization processing on the extracted audio, frame pictures, and text to obtain corresponding vector representations: perform word segmentation and vectorization processing on the text extracted from video, web page, or document data to obtain text vectors; extract text descriptions and feature vectors from the frame pictures extracted from video, web page, or document data to obtain picture feature vectors and text description vectors; use a speech recognition model to convert the audio extracted from video, web page, or document data into text, and convert the text corresponding to the audio into speech vectors.

[0018] In some embodiments of the present disclosure, when the computer program is executed by at least one processor, it causes the device to calculate the similarity between vector representations and establish an association relationship between different modality data based on the similarity: use cosine similarity or Euclidean distance or Manhattan distance to calculate the similarity value between different modality data vectors, and if the similarity value between two vectors exceeds a preset threshold, there is an association relationship between the contents corresponding to these two vectors.

[0019] In some embodiments of the present disclosure, when the computer program is executed by at least one processor, it causes the device to store text, pictures, documents, web pages, audio, and video data in a unified manner in the SurrealDB database based on the association relationship in the form of full-text indexing and vector indexing: create a data table in SurrealDB, the data table includes fields for full-text indexing and fields for vector indexing, and represent the association relationship by constructing foreign key relationships, tags, or metadata in the data table; insert the data extracted from documents, web pages, audio, and video and their corresponding vector representations into the SurrealDB data table for full-text indexing of text data and vector indexing of vectors of picture, audio, and video modality data.

[0020] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for storing multi-modal data based on SurrealDB according to the first aspect of the present disclosure.

[0021] The SurrealDB-based multimodal data storage method and apparatus according to embodiments of the present disclosure extract text and / or pictures from documents, web pages, audio, and video data, perform vectorization processing on the text and pictures, convert them into a unified vector form, and establish content association relationships between different data items by calculating the similarity between vectors, which can effectively realize the fusion and unified storage of multimodal data, avoid fragmentation during the storage and query of different data modalities, achieve unified management of different modality data, reduce the complexity of multimodal data processing, and facilitate cross-modal retrieval and query. Description of the Drawings

[0022] To briefly describe the technical solutions of the embodiments of the present disclosure more clearly, the drawings of the embodiments will be briefly described below. It should be understood that the following-described drawings only relate to some embodiments of the present disclosure and do not limit the present disclosure, where:

[0023] Figure 1 Shows an exemplary flowchart of a SurrealDB-based multimodal data storage method 100 according to an embodiment of the present disclosure;

[0024] Figure 2 Is a schematic block diagram of a SurrealDB-based multimodal data storage apparatus 200 according to an embodiment of the present disclosure.

[0025] It should be noted that the elements in the drawings are schematic and not drawn to scale. Detailed Embodiments

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of the present disclosure without creative efforts also belong to the scope of protection of the present disclosure.

[0027] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the subject matter of the present disclosure belongs. Further, it will be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal form unless otherwise clearly defined herein.

[0028] SurrealDB is a high-performance and scalable database system that supports the storage of various data types, including relational data and non-relational data. In the embodiments of the present disclosure, by using SurrealDB as the underlying storage engine, structured data (such as tabular data) and unstructured data (such as text, pictures, etc.) are stored in SurrealDB in a unified format. Through vector indexing, efficient retrieval based on similarity can be achieved, and through full-text indexing (such as inverted indexing), keyword-based retrieval is supported, thereby returning highly relevant text, pictures, or multi-modal combined data.

[0029] Figure 1 is a schematic flowchart of a multi-modal data storage method 100 based on SurrealDB according to an embodiment of the present disclosure. Referring to Figure 1 as shown, at Figure 1 block S102, text and / or pictures are extracted from document, web page, audio, and video data, and the text and / or pictures are vectorized to obtain vector representations of different modal data.

[0030] The multi-modal information contained in different types of data (documents, web pages, audio, video) needs to be extracted through corresponding technologies. After different modal data is extracted, it is converted into a vector representation through a vectorization model. This is to enable unified storage and similarity-based retrieval. The goal of vectorization is to convert different modal data into a numerical representation to support cross-modal queries and similarity calculations.

[0031] In some embodiments of the present disclosure, audio, frame pictures, and text are extracted from video, web page, or document data, and the extracted audio, frame pictures, and text are vectorized to obtain corresponding vector representations.

[0032] For example, FFmpeg is used to extract subtitle text, audio, and frame pictures from a video. A web crawler tool, such as BeautifulSoup, is used to parse the web page source code to extract text, embedded pictures, and audio from the web page. Different document parsing libraries (PDF, Word, or parsing tools for other formats) are used to extract text and pictures from a document and extract embedded audio.

[0033] Vectorize and tokenize the text to generate text vectors and text keywords. In some embodiments of the present disclosure, for multilingual input, the system will automatically translate non-English text into English. The translated text can use a unified English vectorization method to ensure the consistency of subsequent retrieval. For data containing non-English text, the non-English text data is translated into English through a machine translation model (such as using the translation capabilities of Google Translate, DeepL, or GPT). Use natural language processing models (such as pre-trained language models like BERT, GPT, Sentence-BERT, etc.) to convert the original text and the English text into text vectors.

[0034] Specifically, first, perform some basic preprocessing operations on the original text, such as removing punctuation marks, converting to lowercase, removing stop words, etc. Load a pre-trained model (such as BERT, DistilBERT, etc.) to generate text embeddings. Input the English text or the translated English text into the model to obtain the embedding vector of the text. The vector representation of the text can capture the semantic information of the text, enabling the text content to be compared and retrieved in a high-dimensional space. An SQL example is as follows:

[0035] CREATE text CONTENT{

[0036] data:$data,-- Original text content

[0037] vector:$vector,-- Vector representation of the original text

[0038] en_data:$en_data,-- Text automatically translated into English

[0039] en_vector:$en_vector-- Vector representation of the English text

[0040] }

[0041] Extract text descriptions and feature vectors from images, and convert the text descriptions into text vectors. Image data usually contains rich visual information, and traditional image processing methods such as edge detection or color analysis are difficult to deeply understand the content of the images. To convert image data into a form suitable for storage and retrieval, extract the semantic information and feature vectors of the images through computer vision technology. According to an embodiment of the present disclosure, extract the feature vectors of the images through multiple convolutional layers and pooling layers of a convolutional neural network (such as ResNet, VGG). The feature vectors contain visual features such as the color, shape, and texture of the images.

[0042] Semantically understand and describe the image through an image description generation model (CLIP or Llama-LLaVA), and generate a text description of the image. The natural language description includes objects, scenes, activities, etc. in the image. An attention mechanism can be introduced into the image description generation model to enable the model to focus on different regions of the image, thereby generating a more accurate description. Further, referring to the steps of converting the original text into a text vector, a natural language processing model is used to convert the text description of the image into a text vector, which is convenient for retrieval based on the similarity of the description text. The SQL example is as follows:

[0043] CREATE image CONTENT{

[0044] url:$url,-- Image URL

[0045] prompt:$prompt,-- Generated description text

[0046] vector:$vector,-- Image feature vector

[0047] prompt_vector:$prompt_vector-- Vector representation of the description text

[0048] }

[0049] Use a speech recognition model to convert the audio into text and convert the text corresponding to the audio into a vector. Similarly, perform word segmentation and vectorization on the text extracted from video, web page, or document data to obtain text vectors; extract text descriptions and feature vectors from the frame images extracted from video, web page, or document data to obtain image feature vectors and text description vectors; use a speech recognition model to convert the audio extracted from video, web page, or document data into text and convert the text corresponding to the audio into a speech vector.

[0050] Refer to Figure 1 As shown, subsequently, in box S104, calculate the similarity between the vector representations and establish an association relationship between different modality data based on the similarity.

[0051] Use cosine similarity or Euclidean distance or Manhattan distance to calculate the similarity value between different modality data vectors. If the similarity value between two vectors exceeds a preset threshold, there is an association relationship between the contents corresponding to these two vectors.

[0052] For example, the similarity between the vector of the picture description text and the picture feature vector can measure the consistency between the picture content and the text content, and is used for the hybrid retrieval of pictures and text. Similarly, the similarity between the text vector and the picture description vector is used for text retrieval, and the similarity between the picture feature vectors is used for picture retrieval. If the similarity value between two vectors exceeds a preset threshold, there is an association relationship between the contents corresponding to these two vectors. Specifically, if the description of the text has a high similarity with the picture feature vector, the system will establish a text-picture association. For example, a one-to-one or many-to-one relationship is established between the picture description and the picture to ensure that the relevant text description can be accurately returned when the picture is retrieved. For the similarity calculation between different texts, a text-text association can be established. This is particularly important in the clustering and similarity retrieval of similar contents. If the feature vectors of two pictures are highly similar, a picture-picture association is established. If the text vector and the picture feature vector are highly similar, a text-picture association is established.

[0053] Through the association relationship, a more precise index structure can be established, including inverted index, vector index, and graph structure index. This index structure can help the system more efficiently match and return relevant content during retrieval. The design of the index structure needs to support the joint query and fast similarity calculation of different types of data (text, picture).

[0054] For example, for text content, a traditional full-text index (such as an inverted index) is established to support efficient keyword search. The inverted index is a classic text retrieval method, suitable for storing the relationship between content and modality. It associates the feature vector of each modality with the identifier of the associated content (such as video ID, content ID, etc.).

[0055] For vectorized content (text, picture, speech), a vector space model (such as K-D tree, LSH, i.e., locality-sensitive hashing) can be used to index the content vectors. A hybrid index structure combining vector index and full-text index can provide more flexible and efficient query capabilities in multi-modal data storage and query. In this way, users can choose semantic-based vector queries or traditional keyword-based queries according to their own needs.

[0056] Finally, in block S106, based on the association relationship, in the way of full-text index and vector index, text, picture, document, web page, audio, and video data are uniformly stored in the SurrealDB database.

[0057] In some embodiments of the present disclosure, a data table is created in SurrealDB. The data table includes fields for full-text index and fields for vector index, and represents the association relationship by constructing foreign key relationships, tags, or metadata in the data table.

[0058] For example, create a textCONTENT table through SQL, with fields including: data: stores the original text data, vector: stores the vector representation of the original text, en_data: stores the text data after English translation, en_vector: stores the vector representation of the English text.

[0059] sql

[0060] CREATE textCONTENT {

[0061] data: $data, -- The original text

[0062] vector: $vector, -- The text vector

[0063] en_data: $en_data, -- The text after English translation

[0064] en_vector: $en_vector -- The vector representation of the English text

[0065] }

[0066] Create an imageCONTENT table through SQL, with fields including: url: the URL address of the image, prompt: the natural language description of the image, vector: the image feature vector, prompt_vector: the vector of the image description text.

[0067] sql

[0068] CREATE imageCONTENT {

[0069] url: $url, -- The URL address of the image

[0070] prompt: $prompt, -- The description text of the image

[0071] vector: $vector, -- The image feature vector

[0072] prompt_vector: $prompt_vector -- The vector of the image description text

[0073] }

[0074] Insert the data extracted from documents, web pages, audio, and video, as well as the corresponding vector representations, into the SurrealDB data table for full-text indexing of text data and vector indexing of the vectors of image, audio, and video modal data.

[0075] Figure 2Schematic block diagram of a multi-modal data storage device 200 based on SurrealDB according to an embodiment of the present disclosure. As Figure 2 shown, the device 200 may include a processor 210 and a memory 220 storing a computer program. When the computer program is executed by the processor 210, the device 200 is enabled to execute the steps of the multi-modal data storage method 100 based on SurrealDB as Figure 1 shown. In one example, the device 200 may be a computer device or a cloud computing node. The device 200 can extract text and / or pictures from documents, web pages, audio, and video data, perform vectorization processing on the text and / or pictures to obtain vector representations of different modal data; calculate the similarity between the vector representations, and establish an association relationship between different modal data based on the similarity; and based on the association relationship, uniformly store text, pictures, documents, web pages, audio, and video data in a SurrealDB database in the form of full-text indexing and vector indexing.

[0076] In some embodiments of the present disclosure, the device 200 can be implemented through the following several modules: a data extraction module: extracting text, pictures, and audio information from documents, web pages, audio, and video. A vectorization processing module: calling corresponding models (such as BERT, ResNet, VGGish, etc.) to perform vectorization processing on the extracted data. A similarity calculation module: calculating the similarity between data according to the vectorized representation and establishing an association relationship. A database module: using SurrealDB to store, manage, and query multi-modal data.

[0077] In an embodiment of the present disclosure, the processor 210 may be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. The memory 220 may be any type of memory implemented using data storage technology, including but not limited to random access memory, read-only memory, semiconductor-based memory, flash memory, disk memory, etc.

[0078] In addition, in an embodiment of the present disclosure, the device 200 may also include an input device 230, such as a keyboard, a mouse, etc., for inputting original text, pictures, voices, documents, web pages, or video data. Additionally, the device 200 may further include an output device 240, such as a display, etc., for outputting the SurrealDB database storage architecture.

[0079] In other embodiments of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, can implement the steps of the multi-modal data storage method 100 based on SurrealDB as Figure 1 shown.

[0080] In summary, according to the SurrealDB-based multimodal data storage method and apparatus of the embodiments of the present disclosure, by extracting text and / or pictures from documents, web pages, audio, and video data, vectorizing the text and pictures, converting them into a unified vector form, and calculating the similarity between vectors, the content association relationship between different data items can be established, effectively realizing the fusion and unified storage of multimodal data, avoiding the fragmentation during the storage and query of different data modalities, achieving the unified management of different modality data, reducing the complexity of multimodal data processing, and facilitating cross-modal retrieval and query.

[0081] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses and methods according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, the segment of the program, or the part of the instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the block may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0082] Unless the context clearly indicates otherwise, the singular forms of words used in this specification and the appended claims include the plural, and vice versa. Thus, when referring to the singular, the corresponding plural is usually included. Similarly, the terms "comprising" and "including" are to be construed as inclusive rather than exclusive. Likewise, the term "including" and "or" should be interpreted as inclusive, unless expressly prohibited by the context. Where the term "example" is used in this specification, especially when it is placed after a list of terms, the "example" is merely exemplary and illustrative, and should not be considered exclusive or extensive.

[0083] Further aspects and scopes of adaptability become apparent from the description provided herein. It should be understood that the various aspects of the present application may be implemented alone or in combination with one or more other aspects. It should also be understood that the description herein and the specific embodiments are for illustrative purposes only and are not intended to limit the scope of the present application.

[0084] The above has described several embodiments of the present disclosure in detail. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of the present disclosure without departing from the spirit and scope of the present disclosure. The protection scope of the present disclosure is defined by the appended claims.

Claims

1. A multi-modal data storage method based on SurrealDB, characterized in that, Including: Extracting text and / or pictures from documents, web pages, audio, and video data, and performing vectorization processing on the text and / or pictures to obtain vector representations of different modality data; Calculating the similarity between the vector representations, and establishing an association relationship between different modality data based on the similarity; And Based on the association relationship, storing text, pictures, documents, web pages, audio, and video data in a SurrealDB database in the manner of full-text indexing and vector indexing.

2. The multimodal data storage method based on SurrealDB according to claim 1, wherein The extracting text and / or pictures from documents, web pages, audio, and video data, and performing vectorization processing on the text and / or pictures to obtain vector representations of different modality data includes: Extracting audio, frame pictures, and text from video, web page, or document data, and performing vectorization processing on the extracted audio, frame pictures, and text to obtain corresponding vector representations; Performing vectorization and word segmentation processing on the text to generate text vectors and text keywords; Extracting text descriptions and feature vectors from pictures, and converting the text descriptions of the pictures into vectors; Using a speech recognition model to convert audio into text, and converting the text corresponding to the audio into a vector.

3. The multimodal data storage method based on SurrealDB according to claim 2, characterized in that, The extracting audio, frame pictures, and text from video, web page, or document data includes: Using FFmpeg to extract subtitle text, audio, and frame pictures from video; Using a web crawler tool to parse the web page source code to extract text, embedded pictures, and audio from the web page; Using different document parsing libraries to extract text and pictures from documents, and extracting embedded audio.

4. The multimodal data storage method based on SurrealDB according to claim 2, wherein, The performing vectorization and word segmentation processing on the text to generate text vectors and text keywords includes: Translating non-English text data into English text through a machine translation model; Adopting a natural language processing model to perform word segmentation processing on the text, obtaining the word segments in the text, converting each word segment of the text into a vector, and extracting text keywords through the TF-IDF method.

5. The multimodal data storage method based on SurrealDB according to claim 2, wherein The extracting text descriptions and feature vectors from pictures, and converting the text descriptions of the pictures into vectors includes: Extracting the feature vectors of pictures through a convolutional neural network, where the feature vectors include the color, shape, and texture features of the pictures; Performing semantic understanding and description on the pictures through a picture description generation model to generate text descriptions of the pictures, where the text descriptions include the objects and scenes in the pictures; and Adopting a natural language processing model to convert the text descriptions into text vectors.

6. The multimodal data storage method based on SurrealDB according to claim 2, wherein The performing vectorization processing on the extracted audio, frame pictures, and text to obtain corresponding vector representations includes: Performing word segmentation and vectorization processing on the text extracted from video, web page, or document data to obtain text vectors; Extracting text descriptions and feature vectors from the frame pictures extracted from video, web page, or document data to obtain picture feature vectors and text description vectors; Using a speech recognition model to convert the audio extracted from video, web page, or document data into text, and converting the text corresponding to the audio into a speech vector.

7. The multimodal data storage method based on SurrealDB according to claim 1, wherein The calculating the similarity between the vector representations, and establishing an association relationship between different modality data based on the similarity includes: Use cosine similarity, Euclidean distance, or Manhattan distance to calculate the similarity value between different modal data vectors. If the similarity value between two vectors exceeds a preset threshold, there is an association relationship between the contents corresponding to these two vectors.

8. The multimodal data storage method based on SurrealDB according to claim 7, wherein, Based on the association relationship, storing text, pictures, documents, web pages, audio, and video data in the SurrealDB database in the form of full-text indexing and vector indexing includes: Create a data table in SurrealDB. The data table includes fields for full-text indexing and fields for vector indexing, and represents the association relationship by constructing foreign key relationships, tags, or metadata in the data table; Insert the data extracted from documents, web pages, audio, and video and their corresponding vector representations into the SurrealDB data table, so as to perform full-text indexing on text data and vector indexing on the vectors of picture, audio, and video modal data.

9. A multimodal data storage device based on SurrealDB, characterized in that, The device includes: At least one processor; and At least one memory storing a computer program; Wherein, when the computer program is executed by the at least one processor, the device is caused to execute the steps of the SurrealDB-based multi-modal data storage method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the SurrealDB-based multi-modal data storage method according to any one of claims 1 to 8.