Multi-modal material retrieval method, electronic equipment, storage medium and program product

By transforming the query requirement information in the multimodal material retrieval method into a preset representation space and combining it with meta-information matching, the problem of low efficiency in traditional single-modal retrieval is solved, and the flexibility and accuracy of cross-modal retrieval are realized, thereby improving the efficiency and accuracy of material retrieval.

CN121456166APending Publication Date: 2026-02-03ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511519151.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Traditional material retrieval methods are based on a single modality, which leads to low efficiency and insufficient relevance of search results in large-scale multimodal material databases. It is difficult to obtain materials that match the creative intent in a timely manner, which affects the efficiency of content production and the quality of works.

Method used

By receiving query requirements described in multimodal terms, rewriting them into structured query information, and transforming them into a preset representation space, multi-dimensional matching and retrieval are performed in conjunction with the meta-information of the materials. Cross-modal retrieval is achieved using a multimodal transformation model, thereby improving retrieval efficiency and accuracy.

Benefits of technology

It achieves flexibility and accuracy in cross-modal material retrieval, significantly improves the efficiency and accuracy of multimodal material retrieval results, and meets the multi-dimensional retrieval needs of creators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456166A_ABST
    Figure CN121456166A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal material retrieval method, electronic equipment, a storage medium and a program product. The multi-modal material retrieval method comprises the following steps: receiving query demand information described by at least one modal; the query demand information is rewritten, structured query information is obtained, and the structured query information comprises query conditions described according to all modalities of the query demand information and meta-information of a target material expected to be queried; converting the query condition into a preset representation space, and obtaining a conversion material which corresponds to the multi-modal material library and is converted into the preset representation space; based on the similarity between the converted query condition and the converted material, retrieving from a multi-modal material library to obtain a first batch of candidate materials; based on the meta-information of the target material and pre-extracted meta-information of a multi-modal material library, retrieving from the multi-modal material library to obtain a second batch of candidate materials; and determining a target material from the first batch of candidate materials and the second batch of candidate materials and returning the target material.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of material retrieval technology, and in particular to a multimodal material retrieval method, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] In the process of digital content creation and media production, creators often need to select appropriate materials based on their creative intentions, including but not limited to data in various modalities such as references, photos / icons, video clips, and audio files, in order to support the conception and presentation of the content.

[0003] Traditional material retrieval methods are mostly based on a single modality, such as searching text materials by keywords or searching images and videos by image. As material databases continue to expand and content types become increasingly diverse, creators face problems such as low efficiency, insufficient relevance of search results, and difficulty in obtaining materials that match their creative intentions in a timely manner. These problems, to some extent, restrict the improvement of content production efficiency and the quality of works. Summary of the Invention

[0004] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a multimodal material retrieval method is proposed, comprising: Receive query request information described in at least one modality; The query requirement information is rewritten to obtain structured query information, which includes: query conditions described according to each modality of the query requirement information, and meta-information of the target material to be retrieved; The query conditions are transformed into a preset representation space, and the transformed materials corresponding to the multimodal material library are obtained and transformed into the preset representation space. Based on the similarity between the transformed query conditions and the transformed materials, the first batch of candidate materials is retrieved from the multimodal material library; Based on the metadata of the target material and the pre-extracted metadata of the multimodal material library, a second batch of candidate materials is retrieved from the multimodal material library; The target material is determined from the first batch of candidate materials and the second batch of candidate materials and returned.

[0005] According to a second aspect of the embodiments of this specification, an electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; Wherein, when the processor executes the executable instructions, it is used to implement the method described in the first aspect.

[0006] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in the first aspect.

[0007] According to a fourth aspect of the embodiments of this specification, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0008] As can be seen from the above embodiments, this specification receives query request information described in at least one modality and rewrites it into structured query information, providing a clear and standardized retrieval basis for subsequent accurate retrieval and improving the accuracy of query request parsing; by uniformly converting query conditions and material content in the multimodal material library to a preset representation space, it realizes cross-modal material content similarity comparison, which not only supports the retrieval of materials in the same modality, but also enables cross-modal retrieval, making the query method more flexible; by retrieving the first batch of candidate materials based on content similarity and retrieving the second batch of candidate materials based on meta-information matching, and finally determining the target material from the two types of candidate materials, it realizes multi-dimensional screening, significantly improving the retrieval efficiency and result accuracy of multimodal materials.

[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of a multimodal conversion model training process provided in an exemplary embodiment.

[0011] Figure 2 This is a flowchart of a multimodal material retrieval method provided in an exemplary embodiment.

[0012] Figure 3 This is an exemplary embodiment of a diagram illustrating how a multimodal rewriting model can be used to rewrite query requirement information.

[0013] Figure 4 This is an exemplary embodiment of a method for transforming query conditions into a preset representation space using a multimodal transformation model.

[0014] Figure 5 This is a schematic diagram of the structure of a device provided in an exemplary embodiment. Detailed Implementation

[0015] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0016] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0017] To address creators' need for efficient multimodal data retrieval during digital content creation and media production, this specification provides a multimodal material retrieval method. This method unifies query conditions and materials from different modalities into a preset representation space, supporting not only retrieval of materials within the same modality but also cross-modal retrieval, thus making the query method more flexible. Furthermore, by combining materials with metadata for multi-dimensional matching retrieval, it effectively improves the efficiency and accuracy of multimodal material retrieval.

[0018] Among them, the conversion of query conditions and materials of different modalities can be realized through a multimodal conversion model; in order to improve the conversion accuracy of the model, the multimodal conversion model can be trained before the multimodal material retrieval method is implemented.

[0019] Please see Figure 1 This paper illustrates the training process of a multimodal conversion model, which can be executed by electronic devices with sufficient computing power. These devices include, but are not limited to, servers (such as physical servers and virtual servers), high-performance computing clusters, cloud computing platforms, and edge computing nodes with distributed computing capabilities. The training process of the multimodal conversion model includes: (1) Obtain multimodal samples, which include positive samples and negative samples. Positive samples include at least two modal samples that express the same semantics, and negative samples include at least two modal samples that express different semantics.

[0020] The phrase "at least two modal samples" refers to a single training sample containing data of two or more different modalities, such as text, image, audio, or video modalities. For example, a positive sample could be "a text describing a musical scene" and "a corresponding background music audio"; or "a landscape photograph" and "a corresponding text description." Although these modal samples differ in their presentation, their semantic content is consistent, and therefore they are considered positive samples.

[0021] Correspondingly, negative samples are combinations of semantically inconsistent modal samples, such as "text describing cats" and "pictures of dogs", or "audio news about a sporting event" and "images of city scenery".

[0022] Positive samples serve to guide the model in learning semantic alignment between different modalities, while negative samples serve to enhance the model's ability to distinguish between different semantics, thereby forming clearer semantic boundaries in the representation space.

[0023] (2) Input positive samples into the multimodal conversion model to be trained, so that the multimodal conversion model to be trained can convert each modal sample in the positive samples into the preset representation space; and / or, input negative samples into the multimodal conversion model to be trained, so that the multimodal conversion model to be trained can convert each modal sample in the negative samples into the preset representation space.

[0024] In one alternative implementation, the multimodal transformation model can include neural networks for different modalities, such as a text encoder for the text modality, a convolutional neural network or visual Transformer for the image modality, and an acoustic feature extraction network for the audio modality. Neural networks for different modalities can independently extract features specific to their respective modalities, and then map them to the same representation space through a unified projection layer to enable cross-modal alignment.

[0025] In the embodiments of this specification, the "preset representation space" can have various specific forms to carry the representation of different modal data under a unified semantics.

[0026] In one optional implementation, the preset representation space is a preset vector space. Samples of different modalities (such as images, audio, video, and text) can be mapped to fixed-dimensional vector representations after being processed by corresponding feature extraction networks. For example, text modal samples can be converted into semantic vectors through language models (such as BERT or Transformer encoders); image modal samples can have visual feature vectors extracted through convolutional neural networks (CNNs) or visual Transformers (ViT); audio modal samples can have spectral features extracted through acoustic models and further encoded into vectors. After these vectors are projected onto a unified vector space, cross-modal similarity can be calculated by methods such as cosine similarity and Euclidean distance, thereby achieving unified retrieval based on vector metrics.

[0027] In another optional implementation, the preset representation space is a semantic space, and the corresponding modal samples are represented in text form. Data from different modalities can first be converted into descriptive natural language text, and then unified into the semantic space. For example, multimodal transformation models include image descriptor sub-models, speech recognition sub-models, keyframe extraction and caption generation sub-models, etc. For image samples, corresponding text descriptions can be generated through the image descriptor sub-model; for audio samples, text related to their content can be obtained through the speech recognition sub-model; for video samples, the video can be converted into scene descriptive text through the keyframe extraction and caption generation sub-model. After the above processing, each modality is represented in text form, and then similarity comparison can be directly performed at the semantic level, such as calculating semantic similarity using text embedding models, or performing semantic retrieval based on large-scale language models.

[0028] Using a vector space allows for more efficient support of large-scale retrieval scenarios, making it suitable for situations requiring rapid similarity calculations. Using a semantic space provides a more intuitive way to describe cross-modal data using natural language, making it suitable for applications requiring human-computer interaction or high interpretability. Therefore, different implementation methods can be flexibly selected based on actual application needs, and all fall within the scope of protection covered by the embodiments in this specification.

[0029] (3) The multimodal conversion model is trained with the optimization objective of maximizing the distance between at least two modal samples expressing different semantics in the preset representation space and / or minimizing the distance between at least two modal samples expressing the same semantics in the preset representation space.

[0030] In this embodiment, the purpose of training is to bring semantically consistent samples from different modalities closer together and to distance semantically inconsistent samples from each other. For example, a contrastive learning framework can be used, employing similarity metrics such as cosine similarity or Euclidean distance, and utilizing contrastive loss, triplet loss, or InfoNCE loss functions to optimize model parameters.

[0031] Through the training process described above, the multimodal transformation model can achieve semantic alignment in a unified representation space. This allows samples from different modalities to appear close in the feature space when they are semantically consistent, and to appear separate in the feature space when they are semantically different. This not only improves the accuracy of cross-modal retrieval but also lays the foundation for subsequent multimodal analysis and generation tasks.

[0032] In some embodiments, the multimodal material retrieval method provided in this application can perform material conversion processing on the multimodal material library before execution, so as to achieve efficient and unified matching and comparison in the subsequent retrieval process.

[0033] After the multimodal conversion model is trained, it can be used to convert materials from the multimodal material library to a preset representation space, resulting in converted materials. For example, materials from the multimodal material library can be converted to a preset vector space to obtain converted materials represented in vector form; or, materials from the multimodal material library can be converted to a preset semantic space to obtain converted materials represented in text form. In this way, materials from different modalities are comparable in the same space, effectively avoiding problems such as scale inconsistency and semantic gaps that exist when directly comparing across modalities.

[0034] In one possible implementation, a multimodal content library is a collection that stores and manages multiple modalities of content. The types of content included may include, but are not limited to, single-modal content such as text, images, audio, and video, or multimodal combined content such as text-image or video-audio. For example, a video with subtitles contains both video and text modalities, while a document with illustrations contains both text and image modalities. After pre-conversion using a multimodal conversion model, a converted content library corresponding to the multimodal content library can be obtained. This converted content library stores converted content that corresponds one-to-one with the original content.

[0035] In another possible implementation, the multimodal content library can be further divided into multiple sub-libraries, each corresponding to a different modality type, such as a text sub-library, an image sub-library, an audio sub-library, and a video sub-library. In this case, a multimodal transformation model can be used to transform the content in each sub-library, resulting in transformed content sets mapped to a preset representation space. This sub-library-modal approach facilitates differentiated transformation optimization for different modalities, improving retrieval efficiency and accuracy during the subsequent query phase.

[0036] For example, the materials in the multimodal material library may be added, deleted, or updated over time, such as: new user-uploaded video materials, expired audio materials being removed, and existing image materials having their resolution versions updated.

[0037] When new content is added to the multimodal content library, an incremental update process can be triggered. This involves inputting the new content into the multimodal conversion model, generating corresponding converted content, writing this converted content into the conversion content library, and simultaneously establishing an index mapping relationship with the original content. This process can be executed asynchronously in the background, thus avoiding impact on online search performance.

[0038] When materials in the multimodal material library are deleted or removed, the corresponding conversion materials in the conversion material library can be deleted simultaneously to avoid returning invalid results during subsequent searches.

[0039] When existing materials are updated, such as through image resolution upgrades or video re-editing, the updated materials can be re-input into the multimodal conversion model to generate new conversion materials, replacing the old versions in the conversion material library. Simultaneously, the system can maintain version control information to allow for retrospective or verification of the materials' historical state when needed.

[0040] In other embodiments, before executing the multimodal material retrieval method, metadata can be extracted from various types of materials in the multimodal material library, and a metadata set can be formed based on the extraction results. This metadata set can then be used to constrain and filter materials during subsequent retrieval processes, thereby improving the accuracy and efficiency of the retrieval.

[0041] Meta-information is a structured description of the content attributes and external attributes of a material. It can be extracted from the information inherent in the material or automatically generated through tool-based processing methods.

[0042] Meta-information categories include, but are not limited to: ① For text materials, titles, text length, keywords, and semantic topic tags can be extracted. ② For image materials, resolution, color histogram, main object categories, shooting time, and format type (e.g., JPG / PNG) can be extracted. ③ For video materials, video titles, descriptions, frame rates, duration, resolution, encoding formats, keyframe screenshots, and shot segmentation information can be extracted. ④ For audio materials, audio titles, descriptions, duration, sampling rates, bit rates, voiceprint features, and transcribed text content can be extracted. ⑤ For composite modal materials (such as videos with subtitles or documents with illustrations), cross-modal correlation information can also be extracted, such as the alignment relationship between video and subtitles, and the pairing relationship between document text and illustrations.

[0043] In one possible implementation, multimodal analysis tools, such as image recognition tools, speech recognition tools, video analysis tools, and natural language processing tools, can be invoked to analyze the material and automatically generate corresponding metadata. For example, OCR (Optical Character Recognition) tools can be used to recognize text in images, and ASR (Automatic Speech Recognition) tools can be used to recognize speech content in audio.

[0044] In another possible implementation, the inherent attribute information of the material, such as file name, file header information, and tag information in the storage system, can be used directly as part of the metadata.

[0045] Another possible implementation method is to combine manual annotation with higher-precision metadata for some key materials, such as manually entered tags, summaries, and content classifications.

[0046] The metadata extracted through the above steps can be stored and managed according to the material ID or index key, forming a unified metadata set. This metadata set can be maintained using a relational database, a non-relational database, or a dedicated metadata management system.

[0047] In one alternative implementation, the metadata collection can be further structured into an index, such as: ① a keyword-based inverted index for quickly locating materials containing specified keywords; ② a tree index or hash index based on numerical ranges for accelerating the filtering of numerical metadata such as duration and resolution; and ③ a tag-based multidimensional index for quickly retrieving materials under a specific category.

[0048] For example, a multimodal content library is a collection that stores and manages content of various modalities, and corresponds to a unified set of metadata. Alternatively, a multimodal content library can be further divided into multiple content sub-libraries, each corresponding to a different modality type, and each content sub-library corresponding to a pre-extracted subset of metadata.

[0049] By extracting meta-information from the multimodal material library before retrieval, the multimodal material retrieval method of this specification can add meta-information constraints on the basis of similarity retrieval, thereby improving the accuracy of retrieval results.

[0050] Please see Figure 2 This document illustrates a flowchart of a multimodal content retrieval method provided in an embodiment of this specification. This method can be executed by an electronic device, including but not limited to servers (such as physical servers, virtual servers, etc.), smartphones / mobile phones, tablet computers, personal digital assistants (PDAs), laptops, desktop computers, or other types of devices. The multimodal content retrieval method includes: In S200, query request information described in at least one modality is received.

[0051] In one implementation, a user can input query information via an electronic device. This information can be described using any modality such as text, image, audio, or video, or a combination of two or more modalities. For example, a user could input the text "I want to find a video clip of raindrops hitting glass for no more than 2 seconds"; or input the text "high-resolution images of a castle" and simultaneously upload one or more reference images containing a castle.

[0052] In another implementation, electronic devices can also obtain query information through voice input and natural language dialogue interfaces, and automatically parse it into processable modal input, thereby improving the convenience of user interaction.

[0053] In S202, the query requirement information is rewritten to obtain structured query information, which includes: query conditions described according to each modality of the query requirement information, and meta-information of the target material to be retrieved.

[0054] This step rewrites the multimodal query requirements input by the user, converting unstructured natural language descriptions, reference images, etc., into structured data containing query conditions and metadata for each modality, thereby improving the standardization and executability of the retrieval.

[0055] In one example, if the query requirement includes the text "I want to find a video clip of raindrops hitting glass for no more than 2 seconds", the corresponding structured query information would be: {"text":"snapdrops hitting glass","metadata":{"duration":"2s","type":"video"}}.

[0056] In another example, if the query requirement includes the input text "high-resolution images of a castle" and one or more reference images containing a castle, the corresponding structured query information would be: {"text":"images containing a castle","image":{...},"metadata":{"resolution":">=1080","type":"image"}}.

[0057] For example, please refer to Figure 3 Electronic devices can construct rewriting suggestions based on query requirements, metadata from a multimodal content library, and preset rewriting rules. Specifically: ① Query requirements reflect the user's natural search intent and may be described in single or combined modalities such as text, images, audio, or video. These often suffer from non-standardized expressions and incomplete information, potentially leading to inaccurate results if used directly for retrieval. ② Metadata from the multimodal content library provides structured attributes such as title, description, resolution, duration, and format type, serving as a reference for rewriting. This ensures the generated structured query information matches the search fields in the content library, improving the feasibility of subsequent matching. ③ Preset rewriting rules constrain the format and content of the rewrite. For example, they stipulate that query conditions should include the target modality and that constraints should be converted to standardized expressions, such as rewriting "high resolution" as "resolution≥1080p," ensuring the rewritten output meets uniform standards for easy use by the subsequent retrieval module.

[0058] The rewriting prompts not only include query requirements, metadata from the multimodal material library, and preset rewriting rules, but can also further include the following: ① Rewrite example: For example, “Input: a raindrop video of no more than 2 seconds; Output: {'text':'raindrop video','metadata':{'duration':'≤2s','type':'video'}}”, to help the model learn the rewrite format.

[0059] ② Contextual information: such as the user's historical search habits or recent usage scenarios, to help rewrite the model to generate results that are closer to the user's needs.

[0060] ③ Error correction rules: These are used to handle situations where there is ambiguity or missing information in the query, such as "the vaguely expressed 'short film' should be standardized to a specific duration range".

[0061] Next, the electronic device inputs the rewriting prompts into the pre-trained multimodal rewriting model. The multimodal rewriting model then rewrites the query information based on metadata from a multimodal resource library and pre-defined rewriting rules, outputting structured query information. This multimodal rewriting model can be a large language model with multimodal processing capabilities. It can understand and parse input information from different modalities, such as a combination of text and images, and generate output that conforms to the target modality and structured requirements during the rewriting process. This model already possesses cross-modal semantic understanding capabilities during the pre-training stage. By introducing rewriting rules and metadata corpora into the prompts, the model gains rewriting adaptation capabilities for retrieval scenarios.

[0062] Through the rewriting process described above, the solution provided in this embodiment can rewrite user-input natural language or multimodal descriptions into structured and standardized query information, thereby significantly improving the accuracy and robustness of retrieval. On one hand, the rewritten structured query information is consistent with the metadata of the multimedia resource library, avoiding retrieval failures or insufficient recall due to differences in expression. On the other hand, the multimodal rewriting model can intuitively reflect the user's query conditions, which is beneficial for improving query accuracy. Ultimately, users can obtain more suitable search results without mastering complex search syntax, achieving a dual improvement in search efficiency and accuracy.

[0063] In S204, the query conditions are converted to a preset representation space, and the converted materials corresponding to the multimodal material library are obtained and converted to the preset representation space.

[0064] In this step, "transforming query conditions into a preset representation space" means standardizing the query conditions so that descriptions of different modalities can be measured and compared using the same representational form. This representation space, as a cross-modal common representation domain, ensures that similarity calculations between different modalities can be performed under a unified measurement system during subsequent retrieval processes, thereby achieving cross-modal retrieval.

[0065] For example, please refer to Figure 4 Electronic devices can input query conditions into a trained multimodal transformation model, which will then transform the query conditions into a preset representation space to obtain the transformed query conditions.

[0066] In one alternative implementation, the preset representation space can be a unified vector space. Query conditions can be transformed into this vector space to obtain transformed query conditions represented in vector form. Furthermore, the materials in the multimodal material library are also pre-transformed into the unified vector space in a similar manner to obtain transformed materials represented in vector form.

[0067] In another alternative implementation, the preset representation space can also be a semantic space. Query conditions can be transformed into this semantic space to obtain transformed query conditions represented in text form. Furthermore, the materials in the multimodal material library are also pre-converted into a unified semantic space in a similar manner, resulting in transformed materials represented in text form.

[0068] In S206, the first batch of candidate materials is retrieved from the multimodal material library based on the similarity between the transformed query conditions and the transformed materials.

[0069] This step filters materials that are similar to the query criteria from a content perspective. The similarity can be calculated based on cosine similarity or Euclidean distance in vector space, or semantic matching model based on semantic space. This embodiment does not impose any restrictions on this.

[0070] Furthermore, in practical applications, the preset representation space is not limited to a single form; a hybrid space scheme can also be adopted, that is, to construct both a vector space and a semantic space simultaneously: in the first stage, the vector space is used for rapid recall to obtain candidate materials; in the second stage, the semantic space is used for fine ranking to obtain a more accurate first batch of candidate vectors, thereby improving accuracy and user experience.

[0071] In S208, based on the metadata of the target material and the pre-extracted metadata of the multimodal material library, the second batch of candidate materials is retrieved from the multimodal material library.

[0072] For example, electronic devices can match the metadata of target materials with the metadata of a multimodal material library, specifically including but not limited to: ① Field matching: directly matching the same fields in the target metadata with the metadata in the material library, such as "duration=2s" or "resolution≥1080p"; ② Range filtering: when the user's metadata is a constraint, range filtering can be used, for example, if the query requirement is "video duration not exceeding 2 seconds", then only materials in the material library whose duration field is ≤2 seconds will be filtered; ③ Combination filtering: when the user specifies multiple metadata, such as "video type + duration + resolution", multiple fields can be combined and filtered to obtain candidate materials that meet all constraints; ④ Weight matching: when the user does not specify all metadata, weights can be assigned to different metadata fields, for example, prioritizing type requirements, and then considering resolution and duration.

[0073] By introducing constraints from target material metadata during the query process, the first batch of candidate materials can be supplemented and optimized. On the one hand, metadata retrieval can quickly eliminate irrelevant materials that do not meet the basic attribute requirements, thereby narrowing the candidate set and reducing subsequent computation. On the other hand, metadata retrieval has strong deterministic and structured characteristics, which can complement content-based similarity retrieval. The combination of the two improves the overall accuracy and robustness of the retrieval.

[0074] In S210, the target material is determined from the first batch of candidate materials and the second batch of candidate materials and returned.

[0075] In the embodiments of this specification, after the electronic device completes the retrieval of the first batch of candidate materials and the second batch of candidate materials, it can further process the two batches of candidate materials to determine the final target material.

[0076] In one possible implementation, the electronic device can perform deduplication on the first and second batches of candidate materials based on the unique identifier information of the materials. The unique identifier information may include, but is not limited to, the material's index number in the multimodal material library, file path, content hash value, or other unique identifiers. The purpose of deduplication is to eliminate duplicate materials in the two batches of candidate materials, ensuring computational efficiency in subsequent sorting and processing stages, while avoiding the impact of duplicate materials on the search results.

[0077] After deduplication, the electronic device can rank the deduplicated candidate materials based on the similarity between the transformed query conditions and the transformed materials corresponding to the deduplicated candidate materials. Similarity can be obtained through multimodal similarity calculation methods, such as calculating cosine similarity or Euclidean distance in vector space, or calculating text similarity scores in semantic space. The ranked candidate materials form a sequence from high similarity to low similarity, which the electronic device can directly use as the target material and return. In this scheme, ranking ensures that materials most relevant to the user's query conditions are displayed first, thereby improving the relevance of search results and user satisfaction.

[0078] In another possible implementation, in practical applications, the similarity of candidate materials may vary significantly, especially in cross-modal retrieval or when query conditions are not precisely expressed. In such cases, the similarity of some candidate materials may be below a preset threshold. Directly using these low-similarity materials may result in search results that do not meet user expectations, thus degrading the search experience. Therefore, this embodiment proposes a multimodal generation model adjustment mechanism to optimize candidate materials with similarity below a preset threshold. The preset threshold can be set specifically according to the actual application scenario.

[0079] For example, if there are candidate materials in the first and second batches with a similarity lower than a preset threshold, the electronic device will input the converted material corresponding to the low-similarity candidate material, the converted query conditions, and the similarity between the two as input into the trained multimodal generation model. The multimodal generation model possesses cross-modal understanding and generation capabilities, and can adjust or enhance the material based on the query conditions and material content, thereby improving the matching degree between the material and the query conditions. For instance, for a search request with the query condition "a 2-second video clip of raindrops hitting glass," if the original content of a video clip is "raindrops hitting a window, slightly longer than 2 seconds, and with a blurred background," the multimodal generation model can make the generated material more closely match the query requirements by editing the video clip, enhancing the visual effect of the raindrops, or adjusting the clip length.

[0080] During the adjustment process, multimodal generative models can utilize various generative strategies and techniques. For example: ① Image / video modality adjustment: using Generative Adversarial Networks (GANs) or diffusion models to make subtle modifications to visual content, making key elements more prominent or matching query conditions; ② Audio modality adjustment: using audio enhancement or cropping techniques to extract, reduce noise, or enhance key sound features of audio materials; ③ Text modality adjustment: rewriting or summarizing text materials to make their content more accurately match query requirements; ④ Maintaining cross-modal consistency: in composite modal materials (such as videos with subtitles), the generative model needs to ensure semantic consistency between modalities to avoid mismatches between video and subtitle descriptions.

[0081] After the above adjustments, the electronic device will obtain the adjusted materials, whose similarity will typically be significantly higher than the original low-similarity materials. Subsequently, the electronic device will merge these adjusted materials with candidate materials whose original similarity is not lower than a preset threshold, and finally determine the merged material set as the target material for subsequent display or further processing. This embodiment optimizes materials with low original similarity but potential relevance, bringing their matching degree to a usable level, thereby further optimizing the search results.

[0082] In the embodiments of this specification, the multimodal material library can be organized and managed in different ways to support efficient multimodal material retrieval.

[0083] In some possible implementations, the multimodal content library can be constructed as a unified collection of content for storing and managing various modalities such as text, images, audio, and video. During the retrieval process, electronic devices can calculate the similarity between the transformed query conditions and the corresponding transformed content in the multimodal content library to obtain the first batch of candidate content matching the query conditions; and they can compare the metadata of the target content with the metadata in the unified metadata set corresponding to the multimodal content library to obtain the second batch of candidate content. This unified content library approach is suitable for scenarios with a relatively medium scale and mixed storage of content modalities, and it facilitates direct similarity measurement and sorting within a unified representation space, improving the efficiency of cross-modal retrieval.

[0084] In other possible approaches, considering the potentially large differences in the number of materials of different modalities within a large-scale material library, or the user's search requirements explicitly specifying a target modality, a multimodal material library includes multiple sub-libraries to improve search efficiency and accuracy. These sub-libraries correspond to different modality types, such as image sub-libraries, video sub-libraries, audio sub-libraries, and text sub-libraries. By independently managing and indexing these sub-libraries, electronic devices can quickly locate material sets of specific modalities, thereby reducing interference from irrelevant modal materials in the search process.

[0085] In this context, structured query information typically includes the modality type of the target material the user expects to retrieve. When performing a retrieval operation, the electronic device can first determine the target material sub-library corresponding to the target modality type from multiple material sub-libraries based on this modality type information. This operation can be achieved by querying the sub-library index or matching sub-library attributes. For example, when a user expects to retrieve video material, the electronic device only performs subsequent similarity calculations and metadata matching in the video sub-library, without having to traverse the image or audio sub-libraries, thus significantly improving retrieval efficiency.

[0086] After acquiring the target material sub-library, the electronic device can further perform a dual-path search operation: ① Content-based similarity retrieval: The electronic device inputs the converted query conditions into a trained multimodal conversion model to obtain a query vector / semantic representation converted to a preset representation space. Similarity is then calculated between this vector and the converted materials in the target material sub-library. Similarity calculation can employ vector metrics such as cosine similarity, Euclidean distance, and Manhattan distance, or text similarity and semantic matching methods in the semantic space. By evaluating the similarity between each material in the target material sub-library and the query conditions, a first batch of candidate materials can be obtained. These candidate materials are highly relevant to the user's query conditions in terms of content features.

[0087] ② Meta-information-based constraint retrieval: Electronic devices can also perform retrieval by combining the meta-information of the target material with pre-extracted meta-information from the target material sub-library. Meta-information includes the structured attributes of the material, such as file type, resolution, duration, sampling rate, author, and publication time. This retrieval step can be implemented through field matching, range filtering, combined filtering, or weight matching to obtain a second batch of candidate materials from the target material sub-library. The purpose of this step is to ensure that the candidate materials are not only relevant to the query conditions in terms of content, but also meet the attribute requirements specified by the user, thereby improving the accuracy of the retrieval results.

[0088] Through the aforementioned dual-path retrieval method, electronic devices can combine content similarity with metadata constraints to achieve multi-dimensional filtering of materials in the target material sub-library. Finally, the first and second batches of candidate materials can be further deduplicated, sorted, or adjusted for low-similarity materials to form the final target material set for subsequent use or display.

[0089] By dividing the multimodal content library into sub-libraries and combining them with target modality information, electronic devices can quickly locate target content sub-libraries that meet user needs, avoiding traversing irrelevant modality content and improving retrieval efficiency; ensuring the matching degree of candidate content in terms of content and attributes, improving the accuracy and relevance of retrieval results; supporting efficient management of large-scale content libraries, and further enhancing the system's scalability and retrieval performance by independently indexing and transforming sub-libraries.

[0090] For example, if the structured query information does not contain the modality type of the target material, the electronic device can determine the target material sub-library in at least one of the following ways: (1) Multiple material sub-libraries are identified as target material sub-libraries. By traversing all material sub-libraries, it is ensured that no material that may be related to the query conditions is missed during the retrieval operation; this is suitable when the user's query conditions are very vague or do not involve modal constraints, such as inputting "a clip of raindrops hitting the window", and the user does not specify whether they want to obtain video, image or animation material; electronic devices can access each material sub-library in a loop, and perform similarity and attribute matching calculations between the query conditions and the transformed materials and meta-information in each material sub-library to obtain a candidate material set; this ensures the comprehensiveness of the retrieval, but the amount of computation is relatively large, and it is suitable for small and medium-sized material libraries or scenarios with high requirements for result coverage.

[0091] (2) Based on the query conditions, the predicted modality type of the target material is inferred, and the target material sub-library that matches the predicted modality type is determined from multiple material sub-libraries. In this approach, the electronic device analyzes the query conditions input by the user, infers the most likely modality type through natural language processing, semantic understanding, or rule matching, and selects the target material sub-library that matches the predicted modality type for retrieval. Keywords or semantic features in the query conditions, such as "video clips," "high-definition pictures," and "background music," are used to determine the modality of the material that the user wants to retrieve. The query conditions can be parsed through a pre-trained text classification model or multimodal understanding model to output modality prediction labels, such as video, image, audio, or text. Then, the corresponding sub-library is selected from multiple material sub-libraries for subsequent retrieval. Compared with comprehensive retrieval, this improves retrieval efficiency, reduces interference from irrelevant modal materials, and improves the relevance of candidate materials. When there is uncertainty in the prediction, the system can use two or more material sub-libraries with higher probabilities as target material sub-libraries at the same time to balance efficiency and comprehensiveness.

[0092] (3) Based on the user's historical usage information, infer the predicted modality type of the target material, and determine the target material sub-library that matches the predicted modality type from multiple material sub-libraries.

[0093] Historical usage information includes at least one of the following: ① Modal types of materials whose search frequency exceeds a preset threshold during the user's historical search process. In this implementation, the electronic device can statistically analyze the modal distribution of materials searched or used by the user in the past, select the modal types with higher frequency of occurrence as predicted modal types, and make personalized recommendations based on user habits, thereby improving the consistency between search results and user preferences.

[0094] ② Modal type of materials retrieved by users under the same or similar query conditions. In this implementation, the electronic device can analyze the semantic features of historical query conditions, compare them with the current query conditions, select the modality corresponding to the past search results as the predicted modal type, provide more accurate modal prediction for specific types of queries, and improve the search hit rate.

[0095] ③ The modal type of the material corresponding to the positive interaction behavior generated by the user during the search results interaction process. Positive interaction behavior includes clicking, downloading, saving, forwarding, or browsing the searched material. Positive interaction behavior (such as clicking, downloading, saving, forwarding, or browsing) indicates the user's approval of the material and can be used as a basis for modal preference. In this implementation method, electronic devices can statistically analyze the modal distribution of the material corresponding to positive interaction behavior and select high-frequency modalities as predicted modal types. By incorporating actual user operation behavior, dynamic personalized recommendations can be achieved, making the predicted target modal type more in line with the user's actual needs.

[0096] Through these three methods, electronic devices can flexibly determine the target content sub-library, ensuring that the search operation is performed on the most relevant content set even if the user does not explicitly specify the modality type. This not only improves search efficiency and reduces interference from irrelevant modalities, but also enhances the system's adaptability to user preferences, thereby improving the accuracy, intelligence, and user satisfaction of multimodal content retrieval.

[0097] For example, if the structured query information does not contain the modality type of the target material, the electronic device can also output supplementary information about the modality type of the target material to prompt the user to further input the modality type of the target material to be retrieved.

[0098] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.

[0099] Figure 5 This is a schematic structural diagram of a device provided in an exemplary embodiment. For example... Figure 5 As shown, device 500 mainly consists of a communication interface 502, a user interface 504, a processor 506, and a data storage 508. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 510. The communication interface 502 enables device 500 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 502 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 502 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 502 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 502 may also include multiple physical communication interfaces, such as Wi-Fi interfaces, Bluetooth interfaces, and wide-area wireless interfaces.

[0100] User interface 504 includes receiving user input and providing output to the user. Therefore, user interface 504 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 504 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 504 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, device 500 may support remote access from other devices via communication interface 502 or another physical interface (not shown). User interface 504 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 504 may also be configured as a display device for rendering or displaying text fragments.

[0101] Processor 506 may contain one or more general-purpose processors and / or special-purpose processors.

[0102] Data storage 508 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 506. Data storage 508 may include removable and non-removable components.

[0103] Processor 506 is capable of executing program instructions 518 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 508 to perform the various functions described herein. Data storage 508 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 500, enable device 500 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 518 by processor 506 may result in processor 506 using data 512.

[0104] For example, program instructions 518 may include an operating system 522 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 500 and one or more applications 520 (e.g., a browser, social application, or game application). Similarly, data 512 may include operating system data 516 and application data 514. Operating system data 516 is primarily accessible to the operating system 522, while application data 514 is primarily accessible to one or more applications 520. Application data 514 may reside in a file system visible or hidden from the user of device 500.

[0105] Application 520 can communicate with operating system 522 through one or more application programming interfaces (APIs). These APIs help application 520 read and / or write application data 514, transmit or receive information via communication interface 502, receive or display information on user interface 504, etc.

[0106] In some terminology, application 520 may be simply referred to as "app". Furthermore, application 520 can be downloaded to device 500 through one or more online app stores or app markets. However, applications can also be installed on device 500 in other ways, such as through a web browser or a physical interface on device 500 (e.g., a USB port).

[0107] In some embodiments, the multimodal material retrieval device can be applied to, for example... Figure 5 The device shown implements the technical solution of this specification. The multimodal material retrieval device may include: The receiving module is used to receive query request information described in at least one modality.

[0108] The rewriting module is used to rewrite the query requirement information to obtain structured query information. The structured query information includes: query conditions described according to each modality of the query requirement information, and meta-information of the target material to be retrieved.

[0109] The conversion module is used to convert query conditions to a preset representation space, and to obtain the converted materials corresponding to the multimodal material library that have been converted to the preset representation space.

[0110] The retrieval module is used to retrieve the first batch of candidate materials from the multimodal material library based on the similarity between the transformed query conditions and the transformed materials.

[0111] The retrieval module is also used to retrieve a second batch of candidate materials from the multimodal material library based on the metadata of the target material and the pre-extracted metadata of the multimodal material library.

[0112] The retrieval module is also used to identify and return target materials from the first batch of candidate materials and the second batch of candidate materials.

[0113] In one implementation, the retrieval module is specifically used to construct rewriting prompts based on query requirement information, metadata of the multimodal material library, and preset rewriting rules; input the rewriting prompts into the trained multimodal rewriting model, so that the multimodal rewriting model rewrites the query requirement information based on the metadata of the multimodal material library and preset rewriting rules, and outputs structured query information.

[0114] In one implementation, the conversion module is specifically used to input query conditions into a pre-trained multimodal conversion model, so that the multimodal conversion model can convert the query conditions to a preset representation space to obtain the converted query conditions. The multimodal conversion model is trained using multimodal samples, and the optimization objectives of the multimodal conversion model during the training process include: maximizing the distance between at least two modal samples expressing different semantics in the preset representation space, and / or minimizing the distance between at least two modal samples expressing the same semantics in the preset representation space.

[0115] In one implementation, the conversion module is specifically used to convert the query conditions to a preset vector space to obtain the converted query conditions represented in vector form; or, to convert the query conditions to a preset semantic space to obtain the converted query conditions represented in text form.

[0116] In one implementation, the structured query information also includes: the modal type of the target material to be queried; the multimodal material library includes multiple material sub-libraries, and different material sub-libraries correspond to different modal types; The device also includes a target material sub-library determination module, which is used to determine the target material sub-library that matches the modal type of the target material from multiple material sub-libraries.

[0117] The retrieval module is specifically used to retrieve the first batch of candidate materials from the target material sub-library based on the transformed query conditions and the material content that has been pre-transformed into a preset representation space corresponding to the target material sub-library; and to retrieve the second batch of candidate materials from the target material sub-library based on the metadata of the target material and the pre-extracted metadata of the target material sub-library.

[0118] In one implementation, the target material sub-library determination module is further configured to determine the target material sub-library by at least one of the following methods if the structured query information does not contain the modality type of the target material: ① Designate multiple material sub-libraries as the target material sub-library; ② Based on the query conditions, infer the predicted modality type of the target material, and determine the target material sub-library that matches the predicted modality type from multiple material sub-libraries; ③ Based on the user's historical usage information, infer the predicted modality type of the target material, and determine the target material sub-library that matches the predicted modality type from multiple material sub-libraries; wherein, the historical usage information includes at least one of the following: ① the modality type of the material whose search frequency is greater than a preset threshold in the user's historical search process; ② the modality type of the material retrieved by the user under the same or similar query conditions; ③ the modality type of the material corresponding to the positive interaction behavior generated by the user during the interaction with the search results, and the positive interaction behavior includes clicking, downloading, collecting, forwarding or browsing the retrieved material.

[0119] In one implementation, the retrieval module is specifically used to deduplicate the first batch of candidate materials and the second batch of candidate materials according to the material identifier, sort the deduplicated candidate materials according to similarity, and determine the sorted candidate materials as target materials; or, if there are candidate materials in the first batch of candidate materials and the second batch of candidate materials with similarity lower than a preset threshold, the corresponding converted material, the converted query conditions, and the similarity between the two are input into the trained multimodal generation model to instruct the multimodal generation model to adjust the materials with the goal of improving the similarity between the two, thereby obtaining the adjusted materials, and the adjusted materials and candidate materials with similarity not lower than the preset threshold are determined as target materials.

[0120] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0121] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.

[0122] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0123] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0124] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.

[0125] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.

[0126] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.

[0127] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0128] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0129] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0130] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.

Claims

1. A multimodal material retrieval method, comprising: Receive query request information described in at least one modality; The query requirement information is rewritten to obtain structured query information, which includes: query conditions described according to each modality of the query requirement information, and meta-information of the target material to be retrieved; The query conditions are transformed into a preset representation space, and the transformed materials corresponding to the multimodal material library are obtained and transformed into the preset representation space. Based on the similarity between the transformed query conditions and the transformed materials, the first batch of candidate materials is retrieved from the multimodal material library; Based on the metadata of the target material and the pre-extracted metadata of the multimodal material library, a second batch of candidate materials is retrieved from the multimodal material library; The target material is determined from the first batch of candidate materials and the second batch of candidate materials and returned.

2. The method according to claim 1, wherein rewriting the query requirement information to obtain structured query information includes: Based on the query requirement information, the metadata of the multimodal material library, and the preset rewriting rules, a rewriting prompt is constructed; The rewriting prompts are input into the trained multimodal rewriting model, which then rewrites the query request information based on the metadata of the multimodal material library and preset rewriting rules, and outputs the structured query information.

3. The method according to claim 1, wherein converting the query conditions to a preset representation space includes: The query conditions are input into a trained multimodal transformation model, which then transforms the query conditions into a preset representation space to obtain the transformed query conditions. The multimodal conversion model is trained using multimodal samples. The optimization objectives of the multimodal conversion model during training include: maximizing the distance between at least two modal samples expressing different semantics in the preset representation space, and / or minimizing the distance between at least two modal samples expressing the same semantics in the preset representation space.

4. The method according to claim 1, wherein converting the query conditions to a preset representation space comprises: The query conditions are transformed into a preset vector space to obtain the transformed query conditions represented in vector form; Alternatively, the query conditions can be converted to a preset semantic space to obtain the converted query conditions in text form.

5. The method according to claim 1, wherein the structured query information further includes: The modal type of the target material to be retrieved; The multimodal material library includes multiple material sub-libraries, and different material sub-libraries correspond to different modal types; The method further includes: From the plurality of material sub-libraries, determine the target material sub-library that matches the modal type of the target material; Based on the transformed query conditions and the pre-converted material content corresponding to the multimodal material library in the preset representation space, the first batch of candidate materials is retrieved from the multimodal material library, including: Based on the transformed query conditions and the material content that has been pre-transformed into the preset representation space corresponding to the target material sub-library, the first batch of candidate materials is retrieved from the target material sub-library; The second batch of candidate materials is retrieved from the multimodal material library based on the metadata of the target material and the pre-extracted metadata of the multimodal material library, including: Based on the metadata of the target material and the pre-extracted metadata of the target material sub-library, a second batch of candidate materials is retrieved from the target material sub-library.

6. The method according to claim 5, further comprising: If the structured query information does not contain the modality type of the target material, the target material sub-library is determined by at least one of the following methods: All of the aforementioned multiple material sub-libraries are identified as the target material sub-library; Alternatively, the predicted modality type of the target material can be inferred based on the query conditions, and a target material sub-library that matches the predicted modality type can be determined from the multiple material sub-libraries; Alternatively, the predicted modality type of the target material can be inferred based on the user's historical usage information, and a target material sub-library that matches the predicted modality type can be determined from the multiple material sub-libraries; The historical usage information includes at least one of the following: Modal types of materials whose search frequency during the user's historical search process exceeds a preset threshold; The modal types of materials retrieved by users under the same or similar query conditions; The modal type of the material corresponding to the positive interaction behavior generated by the user during the interaction with the search results, wherein the positive interaction behavior includes clicking, downloading, collecting, forwarding or browsing the searched material.

7. The method according to claim 1, wherein determining the target material from the first batch of candidate materials and the second batch of candidate materials comprises: The first batch of candidate materials and the second batch of candidate materials are deduplicated according to the material identifier. The deduplicated candidate materials are then sorted according to the similarity. The sorted candidate materials are then determined as the target materials. Alternatively, if there are candidate materials in the first batch of candidate materials and the second batch of candidate materials with a similarity lower than a preset threshold, the converted material corresponding to the candidate material, the converted query conditions, and the similarity between the two are input into the trained multimodal generation model to instruct the multimodal generation model to adjust the material with the goal of improving the similarity between the two, thereby obtaining the adjusted material. The adjusted material and the candidate materials with a similarity not lower than the preset threshold are determined as the target material.

8. An electronic device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1-7 by executing the executable instructions.

9. A computer-readable storage medium, characterized in that, It stores computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-7.

Citation Information

Cited By

  • AI-based advertisement material multi-mode intelligent retrieval method and system

    CN121901481A

  • An AI-based advertisement material multi-modal intelligent retrieval method and system

    CN121901481B