Cross-modal data processing method and device based on Surreal DB

Through the cross-modal data processing method of SurrealDB, textualization and vectorization processing combined with full-text index and vector indexing, the efficiency and accuracy of cross-modal data storage and retrieval are solved, and efficient unified management and query of multimodal data are realized.

CN120277256APending Publication Date: 2025-07-08特赞(上海)信息科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510321606.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional search systems and database management systems are unable to effectively process cross-modal data, resulting in dispersed data storage, inefficient retrieval performance, insufficient semantic understanding capabilities and poor system scalability.

Method used

SurrealDB is used as the database management system, and data of different modalities are processed through textualization and vectorization, and a hybrid storage method of full-text index and vector index is constructed. Combined with natural language processing and computer vision technology, the unified storage and efficient retrieval of multimodal data is achieved.

Benefits of technology

It improves the efficiency and accuracy of multimodal data retrieval, realizes unified management and efficient query of different modal data, and supports high concurrency, large-scale and high-availability application requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277256A_ABST
    Figure CN120277256A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a Surreal DB-based cross-modal data processing method and device, and the method comprises the steps: carrying out textualization processing and vectorization processing on data of various different modalities to obtain a text and a vector, and carrying out word segmentation and keyword extraction on the text to obtain a text keyword; constructing a full-text index based on text keywords, constructing a vector index based on vectors, and storing data of different modalities in a data model based on Surreal DB in a mixed storage mode of the full-text index and the vector index; receiving query data of a user, performing textualization processing and vectorization processing on the query data to obtain a query text and a query vector, and performing word segmentation and keyword extraction on the query text to obtain query text keywords; and on the basis of the query text keyword and the query vector, performing vector retrieval and full-text retrieval in the Surreal DB-based data model, and outputting a matched multi-modal data retrieval result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of data processing, and more particularly, to a cross-modal data processing method and apparatus based on SurrealDB. Background Art

[0002] As the multi-modal and heterogeneous nature of data becomes increasingly prominent, traditional search systems and database management systems can no longer meet the complex requirements for processing cross-modal data. Data not only includes structured information but also involves unstructured or semi-structured information. Therefore, traditional single-modal data processing methods face numerous challenges. Especially when dealing with cross-modal search, the main problems include the dispersion of data storage, inefficient retrieval performance, lack of good semantic understanding ability in the presentation of single search results, and poor system scalability.

[0003] To address these challenges, SurrealDB emerged as a new type of database management system, providing a high-performance, flexible, and easy-to-use database solution for modern applications, especially showing its unique advantages in processing multi-modal and heterogeneous data. It can not only effectively process structured, document-type, and graph data but also meet the application requirements of high concurrency, large scale, and high availability through efficient query and real-time synchronization functions. With its distributed architecture and flexible scalability, SurrealDB has become an ideal choice for modern applications to process multi-modal data and achieve efficient cross-modal search.

[0004] In the scenario of multi-modal data retrieval, different types of data require different processing methods. How to improve the system query efficiency and the accuracy of retrieval results is a key problem that urgently needs to be solved. Summary of the Invention

[0005] Embodiments described herein provide a cross-modal data processing method, apparatus, and computer-readable storage medium storing a computer program based on SurrealDB.

[0006] According to a first aspect of the present disclosure, there is provided a cross-modal data processing method based on SurrealDB, including: performing text processing and vectorization processing on data of multiple different modalities to obtain text and vectors, performing word segmentation and keyword extraction on the text to obtain text keywords; constructing a full-text index based on the text keywords, constructing a vector index based on the vectors, and storing the data of different modalities in a data model based on SurrealDB in a mixed storage manner of the full-text index and the vector index; receiving query data from a user, performing text processing and vectorization processing on the query data to obtain a query text and a query vector, performing word segmentation and keyword extraction on the query text to obtain query text keywords; and performing vector retrieval and full-text retrieval in the data model based on SurrealDB based on the query text keywords and the query vector, and outputting a retrieved result of matching multi-modal data.

[0007] In some embodiments of the present disclosure, performing text processing and vectorization processing on data of multiple different modalities to obtain text and vectors, and performing word segmentation and keyword extraction on the text to obtain text keywords includes: translating non-English text data into English text through a machine translation model, and converting the English text into a text vector by using a natural language processing model; extracting a feature vector of an image through a convolutional neural network, where the feature vector includes color, shape, and texture features of the image, performing semantic understanding and description on the picture through an image description generation model to generate a text description of the picture, and converting the text description into a text vector by using a natural language processing model; converting an audio into text by using a speech recognition model, and converting the text corresponding to the audio into a vector; extracting audio, frame images, and text from a video, a document, or a web page, performing vectorization processing on the extracted audio, frame images, and text to obtain corresponding vectors; and performing word segmentation and keyword extraction on the text to obtain text keywords.

[0008] In some embodiments of the present disclosure, converting an audio into text by using a speech recognition model and converting the text corresponding to the audio into a vector includes: performing preprocessing of denoising, normalization, and framing on a speech signal, and converting the preprocessed speech signal into text information based on a speech recognition model; converting the text information into a vector by using a natural language processing model.

[0009] In some embodiments of the present disclosure, audio, frame images, and text are extracted from videos, documents, or web pages, and the extracted audio, frame images, and text are vectorized to obtain corresponding vectors, including: using FFmpeg to extract subtitle text, audio, and frame images from videos; using a web crawler tool to parse the web page source code to extract text, embedded pictures, and audio in the web page; using different document parsing libraries to extract text and pictures from documents and extract embedded audio; and respectively vectorizing the subtitle text, text in the web page, text in the document, frame images, embedded pictures, and audio data to obtain corresponding vectors.

[0010] In some embodiments of the present disclosure, a full-text index is constructed based on text keywords, a vector index is constructed based on vectors, and different modality data is stored in a data model based on SurrealDB in a hybrid storage manner of full-text index and vector index, including: calculating the similarity value between vectors corresponding to different modality data, if the similarity value between two vectors exceeds a preset threshold, there is an association relationship between the contents corresponding to these two vectors, and obtaining the association information of different modality data; constructing a full-text index based on text keywords and a vector index based on vectors; and storing different modality data, full-text index, vector index, and association information in a data model based on SurrealDB.

[0011] In some embodiments of the present disclosure, query data from a user is received, the query data is texturized and vectorized to obtain a query text and a query vector, and the query text is tokenized and keyword-extracted to obtain query text keywords, including: if the query data is text, using a natural language processing model to convert the text into a text vector; if the query data is a picture, extracting the feature vector of the picture through a convolutional neural network, performing semantic understanding and description on the picture through an image description generation model to generate a text description of the picture, and using a natural language processing model to convert the text description into a text vector to obtain the query text and query vector corresponding to the picture; if the query data is audio, using a speech recognition model to convert the audio into a query text and converting the text corresponding to the audio into a query vector; if the query data is mixed data of a video, document, or web page, extracting audio, frame images, and text from the video, document, or web page, and vectorizing the extracted audio, frame images, and text to obtain corresponding query vectors; and performing tokenization processing and keyword extraction on the query text to obtain query text keywords.

[0012] In some embodiments of the present disclosure, based on the query text keywords and the query vector, vector retrieval and full-text retrieval are performed in the data model based on SurrealDB, and the output matching multimodal data retrieval results include: for the query text keywords, the keyword search results matching the keywords are obtained from the SurrealDB database through full-text indexing; the similarity between the query vector and the vectors in the SurrealDB database is calculated, and the vector search results are returned according to the similarity calculation results; for hybrid data retrieval, text retrieval and vector retrieval are performed in parallel to obtain the keyword search results and the vector retrieval results; a retrieval result fusion algorithm based on ranking is used to re-rank the keyword search results and the vector search results, and the preliminary retrieval results are output; and based on the preliminary retrieval results, associated information is queried to obtain associated data.

[0013] In some embodiments of the present disclosure, using a retrieval result fusion algorithm based on ranking to re-rank the keyword search results and the vector search results, the output preliminary retrieval results include: assigning weights to the rankings of each search result and summing them up weighted, and calculating the fusion score based on the following formula:

[0014]

[0015] where rankj(i) is the ranking given to candidate item i by the jth retrieval, k is a constant used to adjust the weight influence between different ranking sources, and i is the candidate item; the keyword search results and the vector search results are re-ranked according to the fusion score, and the preliminary retrieval results are output according to the ranking results.

[0016] According to the second aspect of the present disclosure, a cross-modal data processing device based on SurrealDB is provided. The device includes at least one processor; and at least one memory storing a computer program. When the computer program is executed by at least one processor, the device: performs text processing and vector processing on various different modal data to obtain text and vectors, and performs word segmentation and keyword extraction on the text to obtain text keywords; constructs a full-text index based on the text keywords, constructs a vector index based on the vectors, and stores the different modal data in a hybrid storage manner of the full-text index and the vector index in the data model based on SurrealDB; receives the user's query data, performs text processing and vector processing on the query data to obtain the query text and the query vector, and performs word segmentation and keyword extraction on the query text to obtain the query text keywords; and performs vector retrieval and full-text retrieval in the data model based on SurrealDB based on the query text keywords and the query vector, and outputs the matching multimodal data retrieval results.

[0017] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the cross-modal data processing method based on SurrealDB according to the first aspect of the present disclosure.

[0018] The cross-modal data processing method and apparatus based on SurrealDB according to the embodiments of the present disclosure provide more efficient and accurate multi-modal data retrieval and processing capabilities through unified data storage, vectorization processing, hybrid storage and retrieval, and similarity search of multi-modal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. It should be understood that the following described drawings only relate to some embodiments of the present disclosure and do not limit the present disclosure, where:

[0020] Figure 1 FIG. 12 shows an exemplary flowchart of a cross-modal data processing method 100 based on SurrealDB according to an embodiment of the present disclosure;

[0021] Figure 2 FIG. 16 is a schematic block diagram of a cross-modal data processing apparatus 200 based on SurrealDB according to an embodiment of the present disclosure.

[0022] It should be noted that the elements in the drawings are schematic and not drawn to scale. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of the present disclosure without creative efforts shall also fall within the scope of protection of the present disclosure.

[0024] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the subject matter of the present disclosure belongs. Further, it will be understood that terms such as those defined in commonly used dictionaries shall be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal form unless expressly defined herein.

[0025] SurrealDB supports multiple data models, including relational data models, document data models, and graph data models. It supports SQL queries and real-time data synchronization, allowing multiple systems and applications to update and share data instantaneously. The embodiments of the present disclosure aim to use SurrealDB as the underlying storage engine, combined with a unified multimodal data model and a hybrid indexing mechanism, to achieve unified storage and efficient retrieval of multimodal data.

[0026] Figure 1 An exemplary flowchart of a cross-modal data processing method 100 based on SurrealDB according to an embodiment of the present disclosure is shown.

[0027] Referring to Figure 1 as shown, at Figure 1 block S102, text processing and vectorization processing are performed on data of multiple different modalities to obtain text and vectors, and the text is tokenized and keyword extraction is performed on the text to obtain text keywords.

[0028] In practical applications, enterprises often need to manage structured data (such as product information and user data in tabular form) and unstructured data (such as text, images, audio, etc.). SurrealDB provides unified data storage management and can efficiently process structured and unstructured data simultaneously. Among them, text data can store the original text information through the TEXT field. Images can be stored as binary data (such as base64 encoding) or directly store the feature vectors of the images. Vector data supports vector data storage through a dedicated VECTOR type, such as text vectors, image feature vectors, etc. Through the design of these fields and data structures, SurrealDB supports storing data of different modalities in the same table, thus avoiding the problem of scattered storage of different data sources in traditional databases.

[0029] Vectorization processing is to convert unstructured data into a numerical form (i.e., feature vectors), and text data is one of the most common types of unstructured data. In some embodiments of the present disclosure, for data containing non-English text, the non-English text data is translated into English through a machine translation model (such as using the translation capabilities of Google Translate, DeepL, or GPT). In this way, it can ensure that multilingual texts can be uniformly processed, and all data can be compared in the same vector space. A natural language processing model (such as a pre-trained language model like BERT, GPT, Sentence-BERT, etc.) is used to convert the original text and the English text into text vectors.

[0030] Specifically, first, perform some basic preprocessing operations on the original text, such as removing punctuation marks, converting to lowercase, removing stop words, etc. Load a pre-trained model (such as BERT, DistilBERT, etc.) to generate text embeddings. Input the English text or the translated English text into the model to obtain the embedding vector of the text. The vector representation of the text can capture the semantic information of the text, enabling the text content to be compared and retrieved in a high-dimensional space.

[0031] Image data usually has a high-dimensional and complex structure, and computer vision techniques are needed to extract features. According to an embodiment of the present disclosure, the feature vector of the image is extracted through multiple convolutional layers and pooling layers of a convolutional neural network (such as ResNet, VGG), and the feature vector contains visual features such as the color, shape, and texture of the image.

[0032] Perform semantic understanding and description of the picture through an image description generation model (CLIP or Ollama LLaVA) to generate a natural language description of the picture, and the natural language description includes objects, scenes, activities, etc. in the image. An attention mechanism can be introduced into the image description generation model to enable the model to focus on different regions of the image, thereby generating a more accurate description. Further, referring to the steps of converting the original text into a text vector, a natural language processing model is used to convert the natural language description into a text vector. By text vectorizing the description of the image, the semantic information of the image can be retrieved in the same vector space as the text data.

[0033] Audio data needs to be texturized first and then vectorized. In some embodiments of the present disclosure, a speech recognition model can be used to convert the audio into text and convert the text corresponding to the audio into a vector.

[0034] First, perform preprocessing of denoising, normalization, and framing on the speech signal. Removing background noise can more accurately extract the effective information of the speech signal. Then, based on the speech recognition model, convert the preprocessed speech signal into text information. Referring to the processing method of the text, a natural language processing model is used to convert the text information into a vector.

[0035] For multimodal data (e.g., a combination of text, images, audio, etc.), it is necessary to separately vectorize the data of each modality and then perform fusion. Extract audio, frame images, and text from videos, documents, or web pages, and vectorize the extracted audio, frame images, and text to obtain corresponding vectors. During the processing, use FFmpeg to extract subtitle text, audio, and frame images from videos. Use a web crawler tool to parse the web page source code to extract text, embedded pictures, and audio from the web page. Use different document parsing libraries to extract text and pictures from documents and extract embedded audio. Vectorize the subtitle text, text in the web page, text in the document, frame images, embedded pictures, and audio data separately to obtain corresponding vectors.

[0036] Subsequently, in block S104, a full-text index is constructed based on text keywords, a vector index is constructed based on vectors, and data of different modalities are stored in a data model based on SurrealDB in a hybrid storage manner of full-text index and vector index.

[0037] To achieve unified management and retrieval of multimodal data, embodiments of the present disclosure install and configure SurrealDB. In SurrealDB, a data model is designed to store text and feature vectors. The structure of the data model may include the following parts: a text field for storing raw text data for full-text indexing. A vector field for storing vector representations of text or other data modalities for vector indexing. Whether it is text content, picture descriptions, videos, or feature vectors of images, they can all be uniformly stored in this data model. This data model enables cross-modal data to be processed and queried in a single system, eliminating the problem of scattered storage and processing of various data types in traditional systems.

[0038] Since during subsequent data retrieval, it is necessary to match relevant content based on the similarity between the query data (such as text or images) and the data stored in the database. By calculating metrics such as cosine similarity, Euclidean distance, or Manhattan distance between text vectors and image feature vectors, the system can determine the correlation between different modality data in the database.

[0039] For example, the similarity between the vector of image description text and the image feature vector can measure the consistency between the image content and the text content and is used for hybrid retrieval of images and text. Similarly, the similarity between text vectors and picture description vectors is used for text retrieval, and the similarity between image feature vectors and image feature vectors is used for image retrieval.

[0040] By calculating the similarity between text and images, the correlation between different contents can be established. If the similarity value between two vectors exceeds a preset threshold, there is a correlation between the contents corresponding to these two vectors. Specifically, if the description of the text has a high similarity with the feature vector of the image, the system will establish a text-image correlation. For example, a one-to-one or many-to-one relationship is established between the image description and the image to ensure that relevant text descriptions can be accurately retrieved when the image is searched. For the calculation of similarity between different texts, a text-text correlation can be established. This is particularly important in the clustering and similarity retrieval of similar contents. If the feature vectors of two images are highly similar, an image-image correlation is established.

[0041] Build a full-text index based on text keywords and a vector index based on vectors. Store data of different modalities, the full-text index, the vector index, and the correlation information in a data model based on SurrealDB.

[0042] For example, for text content, build a traditional full-text index (such as an inverted index) to support efficient keyword search. The inverted index is a classic text retrieval method suitable for storing the relationship between content and modality. It associates the feature vector of each modality with the identifier of the associated content (such as video ID, content ID, etc.).

[0043] For vectorized content (text, image, speech), a vector space model (such as K-D tree, LSH, i.e., Locality-Sensitive Hashing) can be used to index the content vectors. A hybrid index structure that combines the vector index and the full-text index can provide more flexible and efficient query capabilities in multi-modal data storage and query. Store data of different modalities in a data model based on SurrealDB in a hybrid storage manner of the full-text index and the vector index. In this way, users can choose semantic-based vector queries or traditional keyword-based queries according to their needs.

[0044] For example, create a documents table that contains fields such as text, text vector, image vector, audio vector, etc.

[0045] CREATE TABLE documents(

[0046] id UUID PRIMARY KEY,--Unique identifier

[0047] text TEXT,--Original text

[0048] text_vector ARRAY <real>,--Text vector (e.g., vector generated by BERT)

[0049] image_vector ARRAY <real>,-- Image vector

[0050] audio_vector ARRAY <real>,--Audio vector

[0051] _text_index TEXT INDEX TEXT, -- Create a full - text index for the text field

[0052] _vector_index VECTOR INDEX VECTOR -- Create a vector index for the vector field);

[0053] Create a full - text index for the text field to support efficient text search. Create vector indexes for fields such as text_vector, image_vector, and audio_vector. Vector indexes are usually searched through approximate nearest neighbor (ANN) algorithms to support similarity queries.

[0054] In block S106, receive the user's query data, perform text processing and vectorization on the query data to obtain the query text and query vector, and perform word segmentation and keyword extraction on the query text to obtain the query text keywords.

[0055] If the query data is text, use a natural language processing model to convert the text into a text vector. If the query data is an image, extract the feature vector of the image through a convolutional neural network, perform semantic understanding and description on the image through an image caption generation model, and generate a text description of the image. Use a natural language processing model to convert the text description into a text vector to obtain the query text and query vector corresponding to the image.

[0056] If the query data is audio, use a speech recognition model to convert the audio into query text and convert the text corresponding to the audio into a query vector. If the query data is video, extract audio, frame images, and text from the video, perform text processing and vectorization on the extracted audio and frame images to obtain the query text and query vector corresponding to the video;

[0057] If the query data is mixed data of video, document, or web page, extract audio, frame images, and text from the video, document, or web page, and perform vectorization on the extracted audio, frame images, and text to obtain the corresponding query vectors.

[0058] In some embodiments of the present disclosure, pre - process the input text, including removing stop words, punctuation marks and other irrelevant information, and word segmentation, to ensure that the meaning of each word can be accurately recognized in the subsequent retrieval process.

[0059] For the queried image data, the feature vectors of the images are extracted through a convolutional neural network, and the image is semantically understood and described through an image caption generation model to generate a natural language description of the image. For example, a pre-trained convolutional neural network (such as ResNet, VGG, Inception) is used to extract the deep features of the image. By extracting the local and global features of the image, the feature vectors can reflect the overall visual content of the image, including information such as shape, color, and texture.

[0060] In some cases, in addition to retrieving through image features, generating semantic descriptions (image captions) of images also helps to enhance the flexibility of retrieval. Using technologies such as Image Captioning, descriptive text is generated based on the features of the image, and the generated text description can be used as a supplement to the image content to help understand and retrieve the image. These descriptive texts can be transformed into text vectors through vectorization methods (such as BERT, GPT).

[0061] If the query data contains mixed information of text and images, the text and images are vectorized separately. In addition, the mixed data may also contain voice and video content. During the processing, the vectorization of the text and the textification and vectorization of the images, audio, and video are performed in parallel to obtain the query text and query vectors.

[0062] Finally, in box S108, based on the query text keywords and query vectors, vector retrieval and full-text retrieval are performed in the data model based on SurrealDB, and the retrieved results of multi-modal data that match are output.

[0063] Traditional full-text retrieval methods rely on inverted indexes and can quickly locate relevant texts through keyword matching, which is suitable for scenarios where specific keywords or phrases are queried. Vector retrieval performs semantic retrieval by calculating the similarity between vectors (such as cosine similarity or Euclidean distance) to further accurately recommend the content that best matches the query intent.

[0064] Once the feature vectors of the image or the vectors of its text description are extracted, vector similarity calculation methods such as cosine similarity and Euclidean distance can be used to compare with the image feature vectors stored in the SurrealDB database to find the most similar images. The query results are sorted according to the similarity scores, and the most relevant images are ranked at the front.

[0065] The SurrealDB database stores structured data, unstructured data, or semi-structured data in advance in the form of full-text indexing and vector indexing. Structured data includes tabular data in relational databases, usually numerical or textual information, with a predefined schema (e.g., JSON, XML). Unstructured data includes images, videos, audio, text data, log files, etc. These data do not follow a fixed structure and require specific parsing methods to extract valuable information. Semi-structured data includes tag data, log files, or web pages, which have some structured elements but do not fully conform to the format of traditional databases.

[0066] In the case of text queries, an inverted index is used to obtain text search results that match keywords from the SurrealDB database. The text keywords include the keywords of the original text and the keywords of the picture text description. The similarity between the text vector and the feature vector and the vectors in the SurrealDB database is calculated, and vector search results are returned according to the similarity calculation results. To improve the retrieval efficiency, for the retrieval of mixed data including text, pictures, audio, and video, text retrieval and vector retrieval are performed in parallel, and the keyword retrieval results and the image retrieval results are merged. The full-text retrieval and vector retrieval operate in parallel, and the retrieval results are quickly returned. Especially in the case of massive data, the overall retrieval time can be accelerated. For the query results, further filtering can be performed, such as filtering based on the quality of the image (clarity, size, etc.), category labels, time, etc., to improve the accuracy and relevance of the retrieval results.

[0067] To effectively integrate the results of full-text retrieval and vector retrieval, a weighted merging method can be used to optimize the sorting. The similarity scores of full-text retrieval and vector retrieval are weighted, and the weights are adjusted according to the query requirements, data types, and retrieval effects. The results are sorted in descending order based on the combined scores, so that the most relevant results are ranked first, improving the relevance and accuracy of the retrieval.

[0068] In some embodiments of the present disclosure, a Ranked Retrieval Fusion (RRF) algorithm based on ranking is used to merge the results of text retrieval and image retrieval into a unified result set. The text and image results are weighted and merged according to the similarity scores of each result, and finally sorted by the weighted scores.

[0069] Among them, the RRF (Ranked Retrieval Fusion) algorithm is a method for processing the ranking of multi-modal retrieval results. It ensures the comprehensive relevance of the results by weighted fusion of the rankings of different retrieval results. The steps of the RRF algorithm include: ranking each result according to different retrievals (such as text retrieval, image retrieval). Combining the rankings of different retrieval results, usually using a weighted method to fuse the ranking results of the two retrievals to ensure that the results of both retrieval channels can be reflected. Assigning weights to the rankings of each search result and summing them up weighted, and calculating the fusion score based on the following formula:

[0070] where rank j (i) is the ranking given to candidate i by the j-th retrieval (the smaller the value, the more forward), k is a constant used to adjust the weight influence between different ranking sources, usually a constant greater than zero (for example, 1 or 10), and i is the candidate, representing a web page, document, or other item. Sort the keyword search results and semantic search results according to the fusion score, and output the preliminary retrieval results according to the sorting results.

[0071] Furthermore, for other data related to the search results, an associated query can be performed to provide additional information that the user may be interested in and obtain associated data. These associated data may include similar documents, tags, or other data that helps to deeply understand the query topic. Therefore, when performing a text query, additional data queries can be based on certain fields of the main retrieval results (such as user ID, tag, entity ID) to further enrich the retrieval results.

[0072] For example, if a news article is retrieved, data such as related images, videos, or social media discussions can be further queried. For pictures or texts, relevant tag, category information, user comments, or user interaction data can be further extracted to form the final query result display. After the user browses the preliminary query results, their behavior data (such as clicks, skips, view details, etc.) can be collected for analysis to optimize the query process. For example, if the user clicks on a certain result, it means that the content is relevant, and this result can be used to further optimize subsequent retrievals. Feed back the content clicked by the user to the query system for re-ranking or adding new keywords. If the user skips some results, it can be considered that these results are less relevant, thereby reducing the appearance of similar results in subsequent retrievals. The user can further refine the query through filtering conditions (such as time range, content type, relevance ranking, etc.), and the retrieval strategy can be adjusted according to the user's refinement conditions and feedback.

[0073] Figure 2 It is a schematic block diagram of a cross-modal data processing device 200 based on SurrealDB according to an embodiment of the present disclosure. As Figure 2 shown, the device 200 may include a processor 210 and a memory 220 storing a computer program. When the computer program is executed by the processor 210, the device 200 is enabled to execute the steps of a cross-modal data processing method 100 based on SurrealDB as Figure 1 shown. In one example, the device 200 may be a computer device or a cloud computing node. The device 200 can perform text processing and vectorization processing on data of multiple different modalities to obtain text and vectors, perform word segmentation and keyword extraction on the text to obtain text keywords; construct a full-text index based on the text keywords, construct a vector index based on the vectors, and store the data of different modalities in a hybrid storage manner of the full-text index and the vector index in a data model based on SurrealDB; receive query data from a user, perform text processing and vectorization processing on the query data to obtain a query text and a query vector, perform word segmentation and keyword extraction on the query text to obtain query text keywords; and perform vector retrieval and full-text retrieval in the data model based on SurrealDB based on the query text keywords and the query vector, and output a matching multi-modal data retrieval result.

[0074] In some embodiments of the present disclosure, the device 200 can translate non-English text data into English text through a machine translation model, and convert the English text into a text vector by using a natural language processing model; extract a feature vector of an image through a convolutional neural network, where the feature vector includes the color, shape, and texture features of the image, perform semantic understanding and description on the picture through an image description generation model, generate a text description of the picture, and convert the text description into a text vector by using a natural language processing model; convert an audio into text by using a speech recognition model, and convert the text corresponding to the audio into a vector; extract audio, frame images, and text from a video, a document, or a web page, perform vectorization processing on the extracted audio, frame images, and text to obtain corresponding vectors; and perform word segmentation and keyword extraction on the text to obtain text keywords.

[0075] In some embodiments of the present disclosure, the device 200 can calculate a similarity value between vectors corresponding to different-modal data. If the similarity value between two vectors exceeds a preset threshold, there is an association relationship between the contents corresponding to the two vectors, and obtain association information of different-modal data; construct a full-text index based on the text keywords, construct a vector index based on the vectors; and store the different-modal data, the full-text index, the vector index, and the association information in a data model based on SurrealDB.

[0076] In some embodiments of the present disclosure, the apparatus 200 may obtain keyword search results that match the query text keywords from the SurrealDB database through full-text indexing; calculate the similarity between the query vector and the vectors in the SurrealDB database, and return vector search results according to the similarity calculation results; for hybrid data retrieval, perform text retrieval and vector retrieval in parallel to obtain keyword search results and vector retrieval results; use a retrieval result fusion algorithm based on ranking to reorder the keyword search results and vector search results, and output preliminary retrieval results; and perform associated information query based on the preliminary retrieval results to obtain associated data.

[0077] In an embodiment of the present disclosure, the processor 210 may be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. The memory 220 may be any type of memory implemented using data storage technology, including but not limited to random access memory, read-only memory, semiconductor-based memory, flash memory, disk memory, etc.

[0078] In addition, in an embodiment of the present disclosure, the apparatus 200 may also include an input device 230, such as a keyboard, a mouse, etc., for inputting user query data, where the query data is text, picture, audio, video or hybrid data. Additionally, the apparatus 200 may further include an output device 240, such as a display, etc., for outputting multimodal data search results.

[0079] In other embodiments of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, can implement the steps of the cross-modal data processing method 100 based on SurrealDB as Figure 1 shown.

[0080] In summary, according to the cross-modal data processing method and apparatus based on SurrealDB in the embodiments of the present disclosure, by providing a unified storage mechanism and a unified multimodal vectorization processing method through SurrealDB, text data can be converted into text vectors, and data such as images and audio can be converted into feature vectors, realizing the unified processing of different modal data. This method improves the efficiency and consistency of data processing and avoids fragmentation between different modal processing methods in the traditional method. By adopting a hybrid storage method of full-text indexing and vector indexing, text and vector searches can be performed simultaneously in the same system, realizing efficient retrieval of cross-modal data and having better query performance.

[0081] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses and methods according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0082] Unless the context clearly dictates otherwise herein, the singular forms of words used in this specification and the appended claims also include the plural, and vice versa. Thus, when reference is made to the singular, the corresponding plural is generally included. Similarly, the terms "comprising" and "including" are to be construed as inclusive rather than exclusive. Likewise, the term "or" should be interpreted as inclusive, unless expressly prohibited by the context of this application. Where the term "exemplary" is used herein, particularly when it is followed by a list of terms, the "exemplary" is merely illustrative and explanatory and should not be regarded as exclusive or extensive.

[0083] Further aspects and scope of adaptability become apparent from the description provided herein. It should be understood that the various aspects of the present application may be implemented individually or in combination with one or more other aspects. It should also be understood that the description herein and the specific embodiments are for illustrative purposes only and are not intended to limit the scope of the present application.

[0084] The above has described in detail several embodiments of the present disclosure. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of the present disclosure without departing from the spirit and scope of the present disclosure. The scope of protection of the present disclosure is defined by the appended claims.< / real> < / real> < / real>

Claims

1. A cross-modal data processing method based on SurrealDB, characterized in that Including: Performing text processing and vectorization processing on data of multiple different modalities to obtain text and vectors, and performing word segmentation and keyword extraction on the text to obtain text keywords; Constructing a full-text index based on the text keywords, constructing a vector index based on the vectors, and storing data of different modalities in a data model based on SurrealDB in a hybrid storage manner of full-text index and vector index; Receiving query data from a user, performing text processing and vectorization processing on the query data to obtain a query text and a query vector, and performing word segmentation and keyword extraction on the query text to obtain query text keywords; and Performing vector retrieval and full-text retrieval in the data model based on SurrealDB based on the query text keywords and the query vector, and outputting a retrieved result of multimodal data that matches.

2. The cross-modal data processing method based on SurrealDB according to claim 1, characterized in that The performing text processing and vectorization processing on data of multiple different modalities to obtain text and vectors, and performing word segmentation and keyword extraction on the text to obtain text keywords includes: Translating non-English text data into English text through a machine translation model, and converting the English text into a text vector by using a natural language processing model; Extracting a feature vector of an image through a convolutional neural network, where the feature vector includes color, shape, and texture features of the image, performing semantic understanding and description on the picture through an image description generation model to generate a text description of the picture, and converting the text description into a text vector by using a natural language processing model; Using a speech recognition model to convert audio into text and converting the text corresponding to the audio into a vector; Extracting audio, frame images, and text from a video, a document, or a web page, and performing vectorization processing on the extracted audio, frame images, and text to obtain corresponding vectors; and Performing word segmentation and keyword extraction on the text to obtain text keywords.

3. The multimodal data storage method based on SurrealDB according to claim 2, wherein, The using a speech recognition model to convert audio into text and converting the text corresponding to the audio into a vector includes: Performing preprocessing of denoising, normalizing, and framing on a speech signal, and converting the preprocessed speech signal into text information based on a speech recognition model; Converting the text information into a vector by using a natural language processing model.

4. The multimodal data storage method based on SurrealDB according to claim 2, wherein, The extracting audio, frame images, and text from a video, a document, or a web page, and performing vectorization processing on the extracted audio, frame images, and text to obtain corresponding vectors includes: Using FFmpeg to extract subtitle text, audio, and frame images in a video; Using a web crawler tool to parse the web page source code to extract text, embedded pictures, and audio in the web page; Using different document parsing libraries to extract text and pictures in a document and extract embedded audio; Performing vectorization processing on the subtitle text, text in the web page, text in the document, frame images, embedded pictures, and audio data respectively to obtain corresponding vectors.

5. The cross-modal data processing method based on SurrealDB according to claim 1, wherein The constructing a full-text index based on the text keywords, constructing a vector index based on the vectors, and storing data of different modalities in a data model based on SurrealDB in a hybrid storage manner of full-text index and vector index includes: Calculate the similarity values between the vectors corresponding to different modal data. If the similarity value between two vectors exceeds a preset threshold, there is an association relationship between the contents corresponding to these two vectors, and obtain the association information of different modal data; Construct a full-text index based on the text keywords and construct a vector index based on the vectors; and Store the different modal data, full-text index, vector index, and association information in a data model based on SurrealDB.

6. The cross-modal data processing method based on SurrealDB according to claim 1, wherein Receive the query data of the user, perform text processing and vectorization processing on the query data to obtain a query text and a query vector, and perform word segmentation and keyword extraction on the query text to obtain query text keywords including: If the query data is text, use a natural language processing model to convert the text into a text vector; If the query data is a picture, extract the feature vector of the picture through a convolutional neural network, perform semantic understanding and description on the picture through an image description generation model, generate a text description of the picture, and use a natural language processing model to convert the text description into a text vector to obtain the query text and query vector corresponding to the picture; If the query data is audio, use a speech recognition model to convert the audio into a query text and convert the text corresponding to the audio into a query vector; If the query data is mixed data of video, document, or web page, extract audio, frame images, and text from the video, document, or web page, and perform vectorization processing on the extracted audio, frame images, and text to obtain corresponding query vectors; and Perform word segmentation processing and keyword extraction on the query text to obtain query text keywords.

7. The cross-modal data processing method based on SurrealDB according to claim 1, wherein Based on the query text keywords and query vector, perform vector retrieval and full-text retrieval in the data model based on SurrealDB, and output the retrieved results of multi-modal data including: For the query text keywords, obtain the keyword search results that match the keywords from the SurrealDB database through the full-text index; Calculate the similarity between the query vector and the vectors in the SurrealDB database, and return the vector search results according to the similarity calculation results; For mixed data retrieval, perform text retrieval and vector retrieval in parallel to obtain keyword search results and vector retrieval results; Use a retrieval result fusion algorithm based on ranking to re-rank the keyword search results and the vector search results, and output preliminary retrieval results; and Query the association information based on the preliminary retrieval results to obtain associated data.

8. The cross-modal data processing method based on SurrealDB according to claim 7, wherein Use a retrieval result fusion algorithm based on ranking to re-rank the keyword search results and the vector search results, and output preliminary retrieval results including: Assign weights to the rankings of each search result and sum them up weighted, and calculate the fusion score based on the following formula: where rank j (i) is the ranking given to candidate i by the j-th retrieval, k is a constant used to adjust the weight influence between different ranking sources, and i is the candidate; Re-rank the keyword search results and the vector search results according to the fusion score, and output the preliminary retrieval results according to the ranking results.

9. A cross-modal data processing device based on SurrealDB, characterized in that, The device includes: At least one processor; and At least one memory storing a computer program; Wherein, when the computer program is executed by the at least one processor, the device is caused to perform the steps of the SurrealDB-based cross-modal data processing method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the SurrealDB-based cross-modal data processing method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Multi-modal picture data processing method and device, equipment and storage medium

    CN121030026A

  • Multi-modal data hybrid retrieval method and device, equipment and storage medium

    CN121478992A