Cross-modal data search method and device based on Surreal DB
Through the mixed sorting strategy of SurrealDB database combining full-text search and vector search, the problem of insufficient efficiency and accuracy in cross-modal data search is solved, and efficient and accurate multimodal data retrieval is achieved.
Patent Information
- Application Number
- CN202510321596.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional data search engines are unable to effectively process cross-modal searches for multimodal data, and lack semantic understanding capabilities, resulting in poor search efficiency and scalability.
The SurrealDB database is used to combine the mixed sorting strategy of full-text search and vector search, and the multimodal data is converted into text and vectors through natural language processing and convolutional neural network, and the search is performed using reverse order index and vector similarity calculation, and the results sorting are optimized in combination with the RRF algorithm.
It realizes efficient and accurate multimodal information retrieval, improves the processing capability and query efficiency of cross-modal data, and meets the high requirements of search efficiency and accuracy.
Smart Images

Figure CN120277255A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of information retrieval, and more particularly, to a cross-modal data search method and apparatus based on SurrealDB. Background Art
[0002] In the context of Internet big data, as the multimodality and heterogeneity of data become more and more prominent, traditional data search engines can no longer meet the complex requirements of processing cross-modal data search.
[0003] Multimodal data includes text, images, videos, audio, web pages, documents, etc. When processing cross-modal data search, the presentation of single search results lacks good semantic understanding ability, cannot fully exploit the potential of different data sources, cannot provide more in-depth retrieval results, and has poor search efficiency and scalability. Summary of the Invention
[0004] Embodiments described herein provide a cross-modal data search method, apparatus, and computer-readable storage medium storing a computer program based on SurrealDB.
[0005] According to a first aspect of the present disclosure, there is provided a cross-modal data search method based on SurrealDB, including: receiving multimodal data input by a user at a data query interface, the multimodal data including text, pictures, audio, video, web pages, and document data; converting the audio, video, web page, and document data into text and pictures, extracting keywords from the text to obtain text prompt words, and performing vectorization processing on the text and pictures to obtain query vectors; performing full-text retrieval on the text prompt words based on the SurrealDB database to obtain text search results, and performing vector retrieval on the query vectors to obtain vector search results; reordering the text search results and the vector search results using a ranking-based retrieval result fusion algorithm to output preliminary retrieval results; and performing an association query on the preliminary retrieval results to obtain associated data.
[0006] In some embodiments of the present disclosure, audio, video, web page, and document data are converted into text and pictures. Keyword extraction is performed on the text to obtain text prompt words, and vectorization processing is performed on the text and pictures to obtain query vectors, including: if the query data is text, a natural language processing model is used to convert the text into a text vector, word segmentation processing and keyword extraction are performed on the text to obtain text prompt words; if the query data is a picture, a convolutional neural network is used to extract the feature vector of the picture, semantic understanding and description are performed on the picture through an image description generation model to generate a text description of the picture, and a natural language processing model is used to convert the text description of the picture into a picture prompt word vector; if the query data is audio, a speech recognition model is used to convert the audio into text, and the text corresponding to the audio is converted into a vector; if the query data is video, web page, or document data, audio, frame pictures, and text are extracted from the video, web page, or document data video, vectorization processing is performed on the extracted audio, frame pictures, and text to obtain corresponding vectors, word segmentation processing and keyword extraction are performed on the text to obtain text prompt words.
[0007] In some embodiments of the present disclosure, the SurrealDB database stores multimodal data including text, pictures, audio, video, web pages, and documents in the form of full-text indexing and vector indexing.
[0008] In some embodiments of the present disclosure, full-text retrieval is performed on the text prompt words based on the SurrealDB database to obtain text search results, and vector retrieval is performed on the query vectors to obtain vector search results, including: using an inverted index to obtain text search results matching the text prompt words from the SurrealDB database; calculating the similarity between the query vectors and the vectors in the SurrealDB database, and returning vector search results according to the similarity calculation results.
[0009] In some embodiments of the present disclosure, full-text retrieval is performed on the text prompt words based on the SurrealDB database to obtain text search results, and vector retrieval is performed on the query vectors to obtain vector search results, which further includes: for the retrieval of mixed data including text and pictures, full-text retrieval and vector retrieval are performed in parallel, and the text retrieval results and vector retrieval results are merged.
[0010] In some embodiments of the present disclosure, a retrieval result fusion algorithm based on ranking is used to re-rank the text search results and vector search results, and preliminary retrieval results are output, including: assigning weights to the rankings of the text retrieval results and vector retrieval results and performing weighted summation, and calculating a fusion score based on the following formula:
[0011]
[0012] Among them, rankj(i) is the ranking given to candidate item i by the j-th retrieval, k is a constant used to adjust the weight influence between different ranking sources, and i is the candidate item; the text search results and vector search results are re-ranked according to the fusion score, and the preliminary retrieval results are output according to the ranking results.
[0013] In some embodiments of the present disclosure, an association query is performed on the preliminary retrieval results to obtain associated data, including: performing a secondary query in the SurrealDB database according to the preliminary retrieval results to obtain associated data; or generating an associated data request based on the preliminary retrieval results and mining external data related to the preliminary retrieval results from relevant literature, news, and databases.
[0014] In some embodiments of the present disclosure, the method further includes: pre-identifying frequently accessed data and loading the frequently accessed data into the cache; and regularly reloading the updated data into the cache according to the update period of the data.
[0015] According to a second aspect of the present disclosure, there is provided a cross-modal data search device based on SurrealDB. The device includes at least one processor; and at least one memory storing a computer program. When the computer program is executed by the at least one processor, the device is caused to: receive multi-modal data input by a user at a data query interface, the multi-modal data including text, pictures, audio, video, web pages, and document data; convert the audio, video, web pages, and document data into text and pictures, extract keywords from the text to obtain text prompt words, and perform vectorization processing on the text and pictures to obtain query vectors; perform full-text retrieval on the text prompt words based on the SurrealDB database to obtain text search results, perform vector retrieval on the query vectors to obtain vector search results; use a ranking-based retrieval result fusion algorithm to re-rank the text search results and vector search results, output preliminary retrieval results; and perform an association query on the preliminary retrieval results to obtain associated data.
[0016] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the cross-modal data search method based on SurrealDB according to the first aspect of the present disclosure.
[0017] The cross-modal data search method and device based on SurrealDB according to the embodiments of the present disclosure, combining a hybrid ranking strategy of full-text retrieval and vector retrieval, can achieve efficient and accurate multi-modal information retrieval. By vectorizing unstructured data and calculating vector similarity for retrieval, it can ensure fast and accurate processing of various types of queries. The application of the RRF algorithm further optimizes the relevance ranking of the retrieval, meeting the high requirements for search efficiency and accuracy in practical applications. Through the powerful storage engine and indexing mechanism provided by SurrealDB, it can support cross-modal data retrieval and intelligent recommendation systems, greatly improving the processing ability and query efficiency of multi-modal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To briefly describe the technical solutions of the embodiments of the present disclosure more clearly, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the following described drawings only relate to some embodiments of the present disclosure and do not limit the present disclosure, where:
[0019] Figure 1 FIG. shows an exemplary flowchart of a cross-modal data search method 100 based on SurrealDB according to an embodiment of the present disclosure;
[0020] Figure 2 FIG. is a schematic block diagram of a cross-modal data search device 200 based on SurrealDB according to an embodiment of the present disclosure.
[0021] It should be noted that the elements in the drawings are schematic and not drawn to scale. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of the present disclosure without creative efforts also belong to the scope of protection of the present disclosure.
[0023] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the subject matter of the present disclosure belongs. Further, it will be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal form unless clearly defined herein otherwise.
[0024] SurrealDB is a multimodal database that supports multiple data models such as relational data, document data, and graph data. It also supports SQL queries, GraphQL queries, graph queries, etc., aiming to simplify the process of data storage and query. In a distributed environment, data can be distributed across multiple nodes, providing high availability and fault tolerance.
[0025] The embodiments of this disclosure aim to implement a cross-modal data search method based on SurrealDB, which can efficiently retrieve multimodal data and dynamically combine search results of different data types according to query requirements. By combining vector indexing with traditional full-text indexing, it allows content-based text queries and vector queries based on semantic similarity to be performed in parallel, ensuring both retrieval efficiency and semantic understanding of multimodal data. Figure 1 FIG. shows an exemplary flowchart of a cross-modal data search method 100 based on SurrealDB according to an embodiment of the present disclosure.
[0026] Refer to Figure 1 As shown, at Figure 1 block S102, receive the multimodal data input by the user at the data query interface. The multimodal data includes text, pictures, audio, video, web pages, and document data.
[0027] The user inputs a query content through the data query interface. This content can be data such as text (e.g., keywords, articles, emails, social media content, etc.), pictures, audio, video, web pages, documents, etc.
[0028] To achieve unified management and retrieval of multimodal data, SurrealDB is used as the underlying storage engine. Whether it is text content, picture descriptions, or feature vectors of images, they can all be stored in the SurrealDB database uniformly, and it can support both traditional SQL queries and vector-based similarity searches. In this way, the system can efficiently manage different types of data in the same database and improve query efficiency.
[0029] In block S104, convert audio, video, web page, and document data into text and pictures, extract keywords from the text to obtain text prompts, and perform vectorization processing on the text and pictures to obtain query vectors.
[0030] If the query data is text, a natural language processing model is used to convert the text into a text vector, perform word segmentation on the text and extract keywords to obtain text prompts. In some embodiments of the present disclosure, the input text is preprocessed, including removing stop words, punctuation marks and other irrelevant information, performing word segmentation, and using keyword extraction techniques (such as TF-IDF, TextRank, BERT, etc.) to extract important keywords from the segmented text. The extracted keywords can help the system understand the theme and core content of the text, facilitating subsequent retrieval. Then a natural language processing model (such as Word2Vec, GloVe, BERT, etc.) is used to convert the text into a text vector.
[0031] If the query data is an image, a convolutional neural network is used to extract the feature vector of the image, and an image caption generation model is used to perform semantic understanding and description of the image to generate a text description of the image, and a natural language processing model is used to convert the text description of the image into an image prompt vector. For example, a pre-trained convolutional neural network (such as ResNet, VGG, Inception) is used to extract the deep features of the image. By extracting the local and global features of the image, the feature vector can reflect the overall visual content of the image, including information such as shape, color, and texture. In addition to retrieving through the image feature vector, generating a semantic description (image caption) of the image also helps to enhance the flexibility of the retrieval. Using techniques such as Image Captioning, descriptive text is generated based on the features of the image, and the generated text description can be used as a supplement to the image content to help the system understand and retrieve the image. A natural language processing model (such as BERT, GPT) is used to convert the text description of the image into an image prompt vector.
[0032] If the query data is audio, a speech recognition model is used to convert the audio into text and convert the text corresponding to the audio into a vector. If the query data is video, web page or document data, audio, frame pictures and text are extracted from the video, web page or document data video, and the extracted audio, frame pictures and text are vectorized to obtain the corresponding vectors, and word segmentation and keyword extraction are performed on the text to obtain text prompts.
[0033] Subsequently, in block S106, a full-text search is performed on the text prompts based on the SurrealDB database to obtain text search results, and a vector search is performed on the query vectors to obtain vector search results.
[0034] Traditional full-text search methods rely on inverted indexes and can quickly locate relevant texts through keyword matching, which is suitable for scenarios where specific keywords or phrases are queried. Vector search, on the other hand, performs semantic search by calculating the similarity between vectors (such as cosine similarity or Euclidean distance) to further accurately recommend the content that best matches the query intent.
[0035] In some embodiments of the present disclosure, a reverse index is used to obtain text search results that match the text prompt from the SurrealDB database. A vector similarity calculation method, such as cosine similarity, Euclidean distance, etc., is used to calculate the similarity between the query vector and the vectors in the SurrealDB database, and vector search results are returned according to the similarity calculation results..
[0036] To improve the retrieval efficiency, for the retrieval of mixed data including text and images, full-text retrieval and vector retrieval are performed in parallel, and the text retrieval results and image retrieval results are merged. The full-text retrieval and vector retrieval operate in parallel to quickly return the retrieval results. Especially in the case of massive data, the overall retrieval time can be accelerated. For the query results, further filtering can be performed, such as filtering based on the quality (clarity, size, etc.), category labels, time, etc. of the images to improve the accuracy and relevance of the retrieval results.
[0037] Refer to Figure 1 As shown, subsequently in block S108, a ranked retrieval result fusion algorithm is used to re-rank the text search results and vector search results, and preliminary retrieval results are output.
[0038] To effectively integrate the results of full-text retrieval and vector retrieval, a weighted merging method can be used to optimize the ranking. The similarity scores of full-text retrieval and vector retrieval are weighted, and the weights are adjusted according to the query requirements, data types, and retrieval effects. The results are sorted in descending order based on the combined scores, so that the most relevant results are ranked first, improving the relevance and accuracy of the retrieval.
[0039] In some embodiments of the present disclosure, a ranked retrieval result fusion algorithm (RankedRetrieval Fusion, RRF) is used to merge the results of text retrieval and image retrieval into a unified result set. The text and image results are weighted and merged according to the similarity scores of each result, and finally sorted according to the weighted scores.
[0040] Among them, the RRF (Ranked Retrieval Fusion) algorithm is a method for processing the ranking of multi-modal retrieval results. It ensures the comprehensive relevance of the results by weighted fusion of the rankings of different retrieval results. The steps of the RRF algorithm include: ranking each result according to different retrievals (such as text retrieval, image retrieval). Merge the rankings of different retrieval results. Usually, a weighted method is used to fuse the ranking results of the two retrievals to ensure that the results of both retrieval channels can be reflected. Assign weights to the rankings of each search result and sum them up weighted, and calculate the fusion score based on the following formula:
[0041] where rank j (i) is the ranking given to candidate i by the j-th retrieval (the smaller the value, the higher the ranking), k is a constant used to adjust the weight influence between different ranking sources, usually a constant greater than zero (e.g., 1 or 10), and i is the candidate, representing a web page, document, or other item. The text search results and semantic search results are sorted according to the fusion score, and the preliminary retrieval results are output according to the sorting results.
[0042] Finally, in block S110, an associated query is performed on the preliminary retrieval results to obtain associated data.
[0043] After the preliminary retrieval results are obtained, associated queries can be performed based on these results to further mine data related to the user's query and provide additional information that the user may be interested in. In some embodiments of the present disclosure, a secondary query can be performed in the SurrealDB database according to the preliminary retrieval results to obtain associated data; or an associated data request can be generated based on the preliminary retrieval results to mine external data related to the preliminary retrieval results from relevant literature, news, and databases.
[0044] These associated data may include similar documents, tags, or other data that helps to gain a deeper understanding of the query topic. Therefore, when performing a text query, additional data queries can be performed based on certain fields of the main retrieval results (such as user ID, tags, entity ID) to further enrich the retrieval results.
[0045] For example, for images or text, relevant tags, category information, user comments, or user interaction data, etc., can be further extracted to form the final query result display. After the user browses the preliminary query results, their behavioral data (such as clicks, skips, view details, etc.) can be collected for analysis to optimize the query process. For example, if the user clicks on a certain result, it means that the content is relevant, and this result can be used to further optimize subsequent retrievals. The content clicked by the user is fed back to the query system for re-sorting or adding new keywords. If the user skips certain results, it can be considered that these results are less relevant, thereby reducing the occurrence of similar results in subsequent retrievals. The user can further refine the query through filtering conditions (such as time range, content type, relevance sorting, etc.), and the retrieval strategy can be adjusted according to the user's refinement conditions and feedback.
[0046] In some embodiments of the present disclosure, in order to effectively improve data access speed and system performance, a caching mechanism is adopted to reduce the number of requests to the database or remote server, thereby reducing latency and enhancing the user experience. First, it is necessary to pre-identify which data is frequently accessed, and this data should be preferentially loaded into the cache. Identifying frequently accessed data can collect statistical information on data access through log analysis or monitoring tools to identify hot data that is frequently accessed. A time window can be set, such as frequently accessed data within the last day, week, or month. A sliding window algorithm can be used for dynamic monitoring and updating of the identification of frequently accessed data. If the update frequency of certain data is very low and it will not change for a long time after each update, then this data can be considered "cold data" and does not need to be frequently loaded into the cache. To ensure that the data in the cache can be updated in a timely manner, a cache update time interval can be set, and according to the data update cycle, the updated data can be reloaded into the cache regularly.
[0047] Figure 2 FIG. 4 is a schematic block diagram of a cross-modal data search device 200 based on SurrealDB according to an embodiment of the present disclosure. As Figure 2 shown, the device 200 may include a processor 210 and a memory 220 storing a computer program. When the computer program is executed by the processor 210, the device 200 can execute the steps of the cross-modal data search method 100 based on SurrealDB as Figure 1 shown. In one example, the device 200 may be a computer device or a cloud computing node. The device 200 can receive multi-modal data input by a user at a data query interface, where the multi-modal data includes text, pictures, audio, video, web pages, and document data; convert the audio, video, web page, and document data into text and pictures, extract keywords from the text to obtain text prompts, and perform vectorization processing on the text and pictures to obtain query vectors; perform full-text retrieval on the text prompts based on the SurrealDB database to obtain text search results, and perform vector retrieval on the query vectors to obtain vector search results; use a retrieval result fusion algorithm based on ranking to re-rank the text search results and vector search results, and output preliminary retrieval results; and perform an association query on the preliminary retrieval results to obtain associated data.
[0048] In some embodiments of the present disclosure, the device 200 can achieve an efficient hybrid search process through parallel processing and result set merging.
[0049] The text search process includes: Text vectorization: Using a pre-trained language model (such as BERT) to convert text into vector representations. Word segmentation processing: Performing effective word segmentation processing before text retrieval to extract key information. Full-text retrieval and vector retrieval in parallel: Text data undergoes both traditional full-text retrieval and is carried out in parallel with vector retrieval to improve retrieval accuracy. Result fusion and sorting: Combining the retrieval results of text and vectors to optimize the final sorting. Data association query: According to the preliminary retrieval results, further query relevant data (such as images, videos, etc.).
[0050] The image search process includes: Image feature extraction: Extracting image features through a convolutional neural network (CNN) or other visual models. Descriptive text generation: Using an image description generation model (such as CLIP) to generate text descriptions as the linguistic expression of the image. Vector similarity calculation: Calculating the similarity between the image and the query image or text description and performing matching. Result sorting and filtering: Sorting and filtering out irrelevant content according to the results of similarity calculation.
[0051] The hybrid search process includes: Parallel processing: Simultaneously executing the query tasks of text and images to reduce the response time. Result set merging: Merging the query results of text and images to form a complete search result. RRF algorithm sorting: Using the RRF algorithm to sort the results of hybrid retrieval to improve the relevance and accuracy of retrieval. Associated data acquisition: Further obtaining relevant data according to the preliminary query results.
[0052] In an embodiment of the present disclosure, the processor 210 may be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. The memory 220 may be any type of memory implemented using data storage technology, including but not limited to random access memory, read-only memory, semiconductor-based memory, flash memory, disk memory, etc.
[0053] In addition, in an embodiment of the present disclosure, the device 200 may also include an input device 230, such as a keyboard, a mouse, etc., for inputting multi-modal data of user queries. Additionally, the device 200 may further include an output device 240, such as a display, etc., for outputting text, images, videos, audio, documents, web pages, or mixed data related to the query content.
[0054] In other embodiments of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, can implement the steps of the cross-modal data search method 100 based on SurrealDB as Figure 1 shown.
[0055] In summary, the cross-modal data search method and apparatus based on SurrealDB according to the embodiments of the present disclosure, combining the hybrid sorting strategy of full-text retrieval and vector retrieval, can achieve efficient and accurate multi-modal information retrieval. By vectorizing unstructured data and calculating vector similarity for retrieval, it can ensure fast and accurate processing of various types of queries. The application of the RRF algorithm further optimizes the relevance sorting of the retrieval, meeting the high requirements for search efficiency and accuracy in practical applications. Through the powerful storage engine and indexing mechanism provided by SurrealDB, it can support cross-modal data retrieval and intelligent recommendation systems, greatly improving the processing ability and query efficiency of multi-modal data.
[0056] In addition, this solution has high application value. For example, for an e-commerce platform, users can query products through images or text, and the system can understand the semantics of the products and provide accurate recommendations. Based on image-text matching, personalized similar product recommendations are provided. Combining the user's historical search records, customized search results are provided.
[0057] For a media platform, by analyzing the text and images of news content, relevant news and reports are provided to users. Through information such as video titles, tags, and descriptions, a large number of pictures, videos, and audio resources are managed. Combining content analysis, relevant videos are recommended to users to improve the resource retrieval and usage efficiency.
[0058] For an enterprise knowledge base, through semantic understanding and vector retrieval technologies, it helps users quickly find the required documents. Through the correlation analysis of text and image data, an enterprise internal knowledge graph is constructed to enhance knowledge sharing and application. Combining multi-modal data, accurate intelligent question-answering services are provided to optimize the user query experience.
[0059] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the apparatus and method according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of an instruction, and the module, program segment, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0060] Unless the context clearly indicates otherwise, the singular forms of words used in this specification and the appended claims include the plural, and vice versa. Thus, when reference is made to the singular, it generally includes the plural of the corresponding term. Similarly, the words "comprising" and "including" will be interpreted as inclusive rather than exclusive. Likewise, the term "including" and "or" should be interpreted as inclusive, unless such an interpretation is expressly prohibited herein. Where the term "exemplary" is used in this specification, particularly when it is followed by a list of terms, such "exemplary" is merely illustrative and explanatory and should not be considered exclusive or exhaustive.
[0061] Further aspects and scopes of adaptability become apparent from the description provided herein. It should be understood that the various aspects of the present application can be implemented alone or in combination with one or more other aspects. It should also be understood that the description herein and the specific embodiments are for illustrative purposes only and are not intended to limit the scope of the present application.
[0062] The foregoing has described in detail several embodiments of the present disclosure. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of the present disclosure without departing from the spirit and scope of the present disclosure. The scope of protection of the present disclosure is defined by the appended claims.
Claims
1. A cross-modal data search method based on SurrealDB, characterized in that, including: Receiving multimodal data input by a user at a data query interface, where the multimodal data includes text, pictures, audio, video, web pages, and document data; Converting the audio, video, web page, and document data into text and pictures, extracting keyword prompts from the text, and performing vectorization processing on the text and pictures to obtain query vectors; Performing full-text retrieval on the text prompts based on the SurrealDB database to obtain text search results, and performing vector retrieval on the query vectors to obtain vector search results; Using a retrieval result fusion algorithm based on ranking to re-rank the text search results and the vector search results, and outputting preliminary retrieval results; and Performing an association query on the preliminary retrieval results to obtain associated data.
2. The cross-modal data search method based on SurrealDB according to claim 1, wherein The converting the audio, video, web page, and document data into text and pictures, extracting keyword prompts from the text, and performing vectorization processing on the text and pictures to obtain query vectors includes: If the query data is text, using a natural language processing model to convert the text into a text vector, performing word segmentation processing and keyword extraction on the text to obtain text prompts; If the query data is a picture, extracting a feature vector of the picture through a convolutional neural network, performing semantic understanding and description of the picture through an image caption generation model, generating a text description of the picture, and using a natural language processing model to convert the text description of the picture into a picture prompt vector; If the query data is audio, using a speech recognition model to convert the audio into text and converting the text corresponding to the audio into a vector; If the query data is video, web page, or document data, extracting audio, frame pictures, and text from the video, web page, or document data video, performing vectorization processing on the extracted audio, frame pictures, and text to obtain corresponding vectors, and performing word segmentation processing and keyword extraction on the text to obtain text prompts.
3. The cross-modal data search method based on SurrealDB according to claim 1, wherein The SurrealDB database stores multimodal data including text, pictures, audio, video, web pages, and documents in a full-text index and vector index manner.
4. The cross-modal data search method based on SurrealDB according to claim 3, characterized in that, The performing full-text retrieval on the text prompts based on the SurrealDB database to obtain text search results, and performing vector retrieval on the query vectors to obtain vector search results includes: Using an inverted index to obtain text search results matching the text prompts from the SurrealDB database; Calculating the similarity between the query vector and the vectors in the SurrealDB database, and returning vector search results according to the similarity calculation result.
5. The cross-modal data search method based on SurrealDB according to claim 4, wherein The performing full-text retrieval on the text prompts based on the SurrealDB database to obtain text search results, and performing vector retrieval on the query vectors to obtain vector search results further includes: For the retrieval of mixed data including text and pictures, performing full-text retrieval and vector retrieval in parallel, and merging the text retrieval results and the vector retrieval results.
6. The cross-modal data search method based on SurrealDB according to claim 5, wherein, The using a retrieval result fusion algorithm based on ranking to re-rank the text search results and the vector search results, and outputting preliminary retrieval results includes: Assign weights to the rankings of the text retrieval results and vector retrieval results and sum them up weighted, and calculate the fusion score based on the following formula: where rank j (i) is the ranking given to candidate i by the j-th retrieval, k is a constant used to adjust the weight influence between different ranking sources, and i is the candidate; Re - rank the text search results and the vector search results according to the fusion score, and output preliminary retrieval results according to the ranking results.
7. The cross-modal data search method based on SurrealDB according to claim 1, wherein Perform an association query on the preliminary retrieval results, and the obtained associated data includes: Perform a secondary query in the SurrealDB database according to the preliminary retrieval results to obtain associated data; or Generate an associated data request based on the preliminary retrieval results, and mine external data related to the preliminary retrieval results from relevant literature, news, and databases.
8. The cross-modal data search method based on SurrealDB according to claim 1, characterized in that, The method further includes: Pre - identify frequently accessed data and load the frequently accessed data into the cache; and Regularly reload the updated data into the cache according to the update period of the data.
9. A cross-modal data search device based on SurrealDB, characterized in that, The device includes: At least one processor; and At least one memory storing a computer program; Wherein, when the computer program is executed by the at least one processor, the device executes the steps of the cross - modal data search method based on SurrealDB according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the cross - modal data search method based on SurrealDB according to any one of claims 1 to 8.
Citation Information
Cited By
Multi-modal retrieval method and device, equipment, medium and program product
CN120929632A
Multimodal retrieval methods, devices, equipment, media, and program products
CN120929632B
Data retrieval method and device, electronic equipment and storage medium
CN121029948A
Information processing device, information processing method, and program
JP7864923B1