Media information processing method and device, storage medium and electronic equipment

Through the joint training target retrieval model and prompt template information, the problem of single multimodal retrieval tasks is solved, and flexible adaptation and accurate analysis of cross-modal retrieval tasks are realized, which improves the search efficiency and accuracy.

CN120372047APending Publication Date: 2025-07-25TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510467416.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The single multimodal retrieval task leads to low media information retrieval efficiency, and the existing technology is difficult to adapt to diversified retrieval needs and lacks the ability to transfer knowledge across tasks.

Method used

The target search model of joint training is adopted to realize multi-task joint training through joint loss function and prompt template information, supporting flexible retrieval of text, images and their mixed modes.

Benefits of technology

It realizes flexible switching and precise adaptation of cross-modal retrieval tasks, improves the multi-scene intention analysis ability and task generalization ability, reduces resource consumption, and improves the retrieval efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372047A_ABST
    Figure CN120372047A_ABST
Patent Text Reader

Abstract

The invention discloses a media information processing method and device, a storage medium and electronic equipment. The method comprises the steps that to-be-retrieved source media information and prompt template information are obtained, the source media information and the prompt template information are input into a target retrieval model, target media information is determined, and the target retrieval model represents a model obtained through joint training according to at least two retrieval task types; a loss function used in the training process of the target retrieval model is a joint loss function, the joint loss function is determined by at least two sub-loss functions in one-to-one correspondence with the at least two retrieval task types, and the target media information represents media information obtained by retrieving the source media information according to the prompt template information. According to the method and the device, the technical problem of relatively low retrieval efficiency of the media information caused by single multi-modal retrieval task is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular, to a method and apparatus for processing media information, a storage medium, and an electronic device. Background Art

[0002] Multimodal information retrieval technology aims to improve the retrieval effect by integrating different modal data, such as text, images, and their combined queries. In related technologies, fixed tasks and modal designs are usually adopted, which limits the flexibility and generalization ability in practical applications. Traditional cross-modal retrieval technologies construct a unified embedding space through a bi-directional embedding model and calculate the similarity between text and images to complete the retrieval. Such methods perform well in single-task scenarios but cannot adapt to diverse retrieval requirements. In addition, due to relying on a predefined input format and lacking the semantic parsing ability for complex instructions, it is impossible to accurately identify the user's intention. The fixed task training mode makes it difficult for the model to share cross-task knowledge, and the performance significantly degrades when facing new tasks or changes in data distribution. Existing retrieval systems usually optimize each task independently, resulting in repeated resource consumption and difficulty in achieving cross-task knowledge transfer. The training method based on fixed modal combinations also limits the retrieval ability of the model for heterogeneous candidate pools.

[0003] Therefore, in related technologies, there is a technical problem that the retrieval efficiency of media information is low due to the single multimodal retrieval task.

[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of this application provide a method and apparatus for processing media information, a storage medium, and an electronic device, so as to at least solve the technical problem that the retrieval efficiency of media information is low due to the single multimodal retrieval task.

[0006] According to one aspect of the embodiments of this application, a method for processing media information is provided, including: obtaining source media information to be retrieved, where the source media information represents the input information that needs to be retrieved currently; obtaining prompt template information, where the prompt template information is used to indicate the type of retrieval task that needs to be retrieved currently; inputting the source media information and the prompt template information into a target retrieval model to determine target media information, where the target retrieval model represents a model jointly trained according to at least two types of retrieval tasks, the loss function used in the training process of the target retrieval model is a joint loss function, the joint loss function is determined by at least two sub-loss functions corresponding one-to-one to the at least two types of retrieval tasks, and the target media information represents the media information retrieved from the source media information according to the prompt template information.

[0007] According to another aspect of the embodiments of the present application, there is also provided a processing device for media information, including: a first acquisition module, configured to acquire source media information to be retrieved, where the source media information represents input information that needs to be retrieved currently; a second acquisition module, configured to acquire prompt template information, where the prompt template information is used to indicate the type of retrieval task that needs to be retrieved currently; a processing module, configured to input the source media information and the prompt template information into a target retrieval model to determine target media information, where the target retrieval model represents a model jointly trained according to at least two types of retrieval tasks, and the loss function used in the training process of the target retrieval model is a joint loss function, and the joint loss function is determined by at least two sub-loss functions corresponding one by one to the at least two types of retrieval tasks, and the target media information represents media information retrieved from the source media information according to the prompt template information.

[0008] In an exemplary embodiment, the device is further configured to: before inputting the source media information and the prompt template information into the target retrieval model to determine the target media information, acquire a preset set of retrieval task types and sample media information, where the set of retrieval task types includes different cross-modal retrieval task types, and the sample media information includes media information of different modalities; construct an initial retrieval model with shared parameters based on the retrieval task types, where the initial retrieval model includes a first encoding branch and a second encoding branch, and the first encoding branch and the second encoding branch are configured to map the sample media information of different modalities to the same embedding space; train the initial retrieval model based on the sample media information to obtain the target retrieval model.

[0009] In an exemplary embodiment, the device trains the initial retrieval model based on the sample media information to obtain the target retrieval model in the following manner: in each batch, the initial retrieval model is trained based on the sample media information in the following manner to obtain the target retrieval model, where the retrieval task type corresponding to each batch is the first retrieval task type: determine first sample media information associated with the first retrieval task type from the sample media information; input the first sample media information into the initial retrieval model to obtain first sample predicted media information; determine a first target sub-loss function corresponding to the first retrieval task type according to the first sample predicted media information; perform backpropagation on the initial retrieval model based on the first target sub-loss function to update the model parameters of the initial retrieval model; repeat the above steps until the target retrieval model is generated.

[0010] In an exemplary embodiment, the apparatus is configured to determine a first target sub-loss function corresponding to the first retrieval task type according to the predicted media information of the first sample and the first true label in the following manner: obtain a target weight coefficient determined for the first retrieval task type; determine an initial sub-loss function corresponding to the first retrieval task type according to the predicted media information of the first sample; and determine the first target sub-loss function based on the initial sub-loss function and the target weight coefficient.

[0011] In an exemplary embodiment, the apparatus is configured to train the initial retrieval model based on the sample media information to obtain the target retrieval model in the following manner: in each batch, train the initial retrieval model based on the sample media information to obtain the target retrieval model, where the sample media information input in each batch includes media information of different modalities: determine second sample media information from the sample media information, and determine a second retrieval task type corresponding to the second sample media information; input the second sample media information into the initial retrieval model to obtain second predicted media information; determine a second target sub-loss function corresponding to the second retrieval task type according to the second predicted media information; perform backpropagation on the initial retrieval model based on the second target sub-loss function to update the model parameters of the initial retrieval model; repeat the above steps until the target retrieval model is generated.

[0012] In an exemplary embodiment, the apparatus is configured to train the initial retrieval model based on the sample media information to obtain the target retrieval model in the following manner: input a first set of sample media information into the initial retrieval model to obtain a first sub-loss function, and input a second set of sample media information into the initial retrieval model to obtain a second sub-loss function, where the first sub-loss function and the second sub-loss function correspond to different retrieval task types; determine a first joint loss function based on the first sub-loss function and the second sub-loss function, and train the initial retrieval model based on the first joint loss function to obtain the target retrieval model, where the first joint loss function is a function determined by assigning different weights to different retrieval task types and based on whether the sample media information of different modalities is similar, and when the first joint loss function satisfies a first loss condition, the initial retrieval model is determined to be the target retrieval model, and the joint loss function includes the first joint loss function.

[0013] In an exemplary embodiment, the device is configured to input the source media information and the prompt template information into a target retrieval model in the following manner to determine the target media information: determining a set of candidate media information based on the prompt template information, where the set of candidate media information includes the candidate media information specified by the prompt template information; and determining, as the target media information, the candidate media information in the set of candidate media information that meets a preset similarity condition with the source media information.

[0014] In an exemplary embodiment, the device is configured to determine the set of candidate media information based on the prompt template information in the following manner: performing feature extraction and splicing on the source media information and the prompt template information to obtain a source feature vector; determining a current retrieval task type based on the source feature vector; and determining the set of candidate media information according to the current retrieval task type.

[0015] In an exemplary embodiment, the device is configured to determine, as the target media information, the candidate media information in the set of candidate media information that meets a preset similarity condition with the source media information in the following manner: obtaining candidate feature vectors corresponding to the candidate media information in the set of candidate media information; determining a target feature vector based on the similarity between the source feature vector and the candidate feature vectors; and determining, as the target media information, the candidate media information corresponding to the target feature vector.

[0016] In an exemplary embodiment, the device is further configured to: obtain a preset set of retrieval task types and sample media information, where the set of retrieval task types includes different cross-modal retrieval task types, and the sample media information includes media information of different modalities; constructing an initial retrieval model with shared parameters based on the retrieval task types, where the initial retrieval model includes a first encoding branch, a second encoding branch, and a third encoding branch, the first encoding branch and the second encoding branch are configured to map the sample media information of different modalities to the same embedding space, and the third encoding branch is configured to map the prompt template information to the same embedding space; and training the initial retrieval model based on the sample media information to obtain the target retrieval model.

[0017] In an exemplary embodiment, the device is configured to train the initial retrieval model based on the sample media information to obtain the target retrieval model in the following manner: input a third set of sample media information into the initial retrieval model to obtain a third sub-loss function, and input a fourth set of sample media information into the initial retrieval model to obtain a fourth sub-loss function, where the third sub-loss function and the fourth sub-loss function correspond to different retrieval task types; determine a second joint loss function based on the third sub-loss function and the fourth sub-loss function, and train the initial retrieval model based on the second joint loss function to obtain the target retrieval model, where the second joint loss function is a function that assigns different weights to different retrieval task types and is determined based on whether the sample media information is similar to the sample candidate media information set. When the second joint loss function satisfies the second loss condition, the initial retrieval model is determined as the target retrieval model, and the joint loss function includes the second joint loss function.

[0018] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the above-mentioned media information processing method when running.

[0019] According to another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the media information processing method as described above.

[0020] According to another aspect of the embodiments of the present application, there is also provided an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to execute the above-mentioned media information processing method through the computer program.

[0021] Through the embodiments of the present application, a mechanism combining dynamically adapted multi-modal input processing with prompt template instructions is adopted. By synchronously inputting source media information and semantic instructions of the retrieval task type into a jointly trained target retrieval model, the purpose of flexible switching and precise adaptation of cross-modal retrieval tasks is achieved. The source media information supports text, image, and their mixed-modal input. Combining with the semantic guidance of the prompt template, it breaks through the dependence of traditional methods on fixed input formats and predefined tasks, thereby realizing a double improvement in multi-scenario intention parsing ability and task generalization ability. Through the design of a multi-task joint training framework and a joint loss function with dynamic weight allocation, the sub-loss functions of at least two retrieval tasks are synergistically optimized, prompting the model to establish cross-modal general feature representations, solving the problems of knowledge isolation in single-task training and performance decay in zero-shot scenarios, and realizing cross-task knowledge sharing and enhanced retrieval robustness. The joint loss function balances the gradient conflicts of multiple tasks, optimizes the feature matching accuracy in heterogeneous retrieval scenarios, and finally significantly improves the retrieval efficiency and accuracy in complex tasks, while reducing the resource consumption of multi-model deployment, thereby solving the technical problem of low retrieval efficiency of media information due to the singularity of multi-modal retrieval tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0023] Figure 1 is a schematic diagram of an application environment of an optional method for processing media information according to an embodiment of the present application;

[0024] Figure 2 is a schematic flowchart of an optional method for processing media information according to an embodiment of the present application;

[0025] Figure 3 is a schematic diagram of an optional method for processing media information according to an embodiment of the present application;

[0026] Figure 4 is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0027] Figure 5 is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0028] Figure 6 is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0029] Figure 7 is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0030] Figure 8 It is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0031] Figure 9 It is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0032] Figure 10 It is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0033] Figure 11 It is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0034] Figure 12 It is a schematic diagram of another optional method for processing media information according to an embodiment of the present application;

[0035] Figure 13 It is a schematic diagram of the structure of an optional media information processing device according to an embodiment of the present application;

[0036] Figure 14 It is a schematic diagram of the structure of an optional media information processing product according to an embodiment of the present application;

[0037] Figure 15 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0038] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0039] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0040] To better understand the technical solutions provided by the embodiments of this application, the key terms related to the embodiments of this application are introduced here first:

[0041] Multi-task training: The model improves the generalization ability through a joint learning mechanism that simultaneously optimizes multiple related tasks.

[0042] Instruction tuning: A semantic-driven method that dynamically adjusts the model behavior through natural language instructions.

[0043] Unified multi-modal information retrieval method: A general framework that integrates multi-modal data such as text and images to achieve cross-task retrieval.

[0044] Cross-modal retrieval: Bidirectional correlation matching between different modal data (such as images and text).

[0045] Joint loss function: The optimization objective in multi-task learning that fuses the weighted sum of the sub-losses of each task.

[0046] Natural language instruction: The retrieval intention expressed by the user in an unstructured language.

[0047] Embedding space: A unified vector mapping space for multi-modal data constructed through contrastive learning.

[0048] Zero-shot / few-shot generalization ability: The transfer performance of the model in tasks without annotations or with few samples.

[0049] Heterogeneous candidate pool: A collection of retrieval resources containing multi-modal data.

[0050] Combined text-image retrieval: A multi-modal retrieval task based on a combined text-image query.

[0051] The method for processing media information provided by this application is applied to a multi-modal retrieval scenario, such as Figure 1 shown:

[0052] The media information processing framework consists of a terminal device 103, a server 101, a database 105, and a network 102, and supports an end-to-end process from user input to the return of search results. The application 107 is deployed on the terminal device 103, responsible for collecting the source media information (such as images, text) input by the user, and transmitting the request to the server 101 through the network 102. The server 101 calls the prompt template information stored in the database 105 or the prompt template information received from the application 107, and completes the cross-modal retrieval task in combination with the jointly trained target retrieval model, and finally returns the target media information. The system achieves efficient computing resource allocation through modular division of labor. The terminal device focuses on interaction and input collection, and the server focuses on model reasoning and data scheduling.

[0053] S1, obtain the source media information to be retrieved, the terminal device 103 collects the source media information input by the user through the application 107, and supports text, image and mixed modal data (such as text and image combination query). For example, the user uploads a product picture (image mode) or enters a descriptive text (text mode). The application 107 performs standardized preprocessing on the input data (such as format conversion and resolution adjustment) to ensure input compatibility with the target retrieval model. The core technology of this stage is multimodal input parsing technology, which automatically identifies the modal type of the input data through the AI (Artificial Intelligence) algorithm and extracts basic features.

[0054] S2, obtaining prompt template information, the terminal device 103 inputs the prompt template information to specify the current search task.

[0055] S3, input the target retrieval model and determine the target media information. After receiving the terminal request, the server 101 inputs the source media information and the prompt template information into the target retrieval model, which is a unified cross-modal framework based on multi-task joint training. The model integrates the optimization goals of different retrieval tasks (such as image to text, text to image, and cross-language retrieval) through a joint loss function, and each sub-loss function corresponds to a task type. For example, the image retrieval task uses the contrast loss function (Contrastive Loss) to bring similar image features closer, and the text retrieval task uses the cross-entropy loss function (Cross-Entropy Loss) to optimize the classification accuracy. The model output is the target media information with the highest similarity to the source media information in the candidate media information set, and is finally returned to the terminal device 103 through the API (Application Programming Interface).

[0056] The target retrieval model adopts a design that decouples the shared encoder from the task branches. The encoder maps different modality data to a unified embedding space through contrastive learning, and the task branches select the corresponding loss function calculation logic according to the prompt template. During the training process, the model automatically balances the influence of each sub-loss function through a dynamic weight allocation strategy to avoid multi-task optimization conflicts. The model uses a cross-modal attention mechanism to align text and image features. For example, in the text-image combination retrieval task, the model dynamically focuses on the correlation regions between image regions and text keywords through attention weights to improve the fine-grained matching accuracy. The prompt template information guides the model's behavior through instruction tuning technology. For example, when the template specifies "cross-language retrieval", the model activates the multilingual embedding layer to map the input text to a multilingual shared semantic space.

[0057] In summary, the system flexibly switches the retrieval task types through the prompt template, supports image, text, and mixed modality inputs, and does not require separate deployment of models for different tasks, significantly reducing the operation and maintenance complexity. The joint training framework enables the model to learn cross-modal general feature representations, reduces feature matching biases in tasks such as text-image mutual retrieval and cross-language retrieval, and improves the relevance of retrieval results. The shared encoder design reduces the model's parameter quantity, combines with the dynamic weight allocation strategy to optimize the training efficiency, and supports only the expansion of the sub-loss function and the prompt template library when adding new task types subsequently. Through the collaborative processing of semantic instructions and multi-modal inputs, the system can more accurately understand the user's needs, such as extracting implicit retrieval conditions from fuzzy queries (such as "find works with a similar style").

[0058] The present application will be described below with reference to the embodiments:

[0059] According to one aspect of the embodiments of the present application, a method for processing media information is provided. Optionally, in this embodiment, the above method for processing media information can be applied to a hardware environment composed of a server 101 and a terminal device 103 as shown in Figure 1 the figure. As shown in Figure 1As shown, the server 101 is connected to the terminal device 103 through a network and can be used to provide services for the terminal device or the application 107 installed on the terminal device. The application can be a video application, an instant messaging application, a browser application, an educational application, a game application, etc. A database 105 can be set up on the server or independently of the server to provide data storage services for the server 101. For example, a game data storage server. The above network can include, but is not limited to: a wired network, a wireless network. Among them, the wired network includes: a local area network, a metropolitan area network, and a wide area network. The wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The terminal device 103 can be a terminal configured with an application and can include, but is not limited to, at least one of the following: a mobile phone (such as an Android mobile phone, an iOS mobile phone, etc.), a laptop computer, a tablet computer, a handheld computer, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a mixed reality (MR) terminal, and other computer devices. The above server can be a single server, a server cluster composed of multiple servers, or a cloud server.

[0060] Combined with Figure 1 As shown, the above method for processing media information can be executed by an electronic device, which can be a terminal device or a server. The above method for processing media information can be implemented separately by the terminal device or the server, or jointly implemented by the terminal device and the server.

[0061] The above is only an example, and this embodiment does not make specific limitations.

[0062] Optionally, as an alternative implementation, as Figure 2 shown, the above method for processing media information includes:

[0063] S202, obtaining the source media information to be retrieved, where the source media information represents the input information that needs to be retrieved currently;

[0064] Optionally, in the embodiments of the present application, the above-mentioned acquisition of the source media information to be retrieved may include, but is not limited to, receiving multimodal data as the input content of the retrieval request from user input or a system interface. The source media information serves as the starting point for retrieval, and its forms of expression cover text, images, audio, video, and any combination thereof. In a specific implementation, the text form can be embodied as a query statement described in natural language, a set of keywords, or a structured text paragraph; the image form may include static pictures uploaded by the user, screenshots of dynamic images, or preprocessed standardized image data; the combined form refers to multimodal input that simultaneously includes text and images, such as a scenario where the user inputs a text description and uploads a reference picture on the search interface at the same time. The process of obtaining this information involves the parsing and feature extraction of the original data. For example, image features are extracted through a visual encoder, and text semantic vectors are extracted through a language model, and then heterogeneous data is mapped to a unified feature space. For example, as Figure 3 shown, in a news retrieval scenario, the user may upload a photo containing the scene of a specific event and attach a text description "Find in-depth reports related to this event". At this time, the system takes the photo pixel matrix and the text tokenization sequence as the source media information and inputs them into the model for processing.

[0065] It should be noted that when obtaining the source media information to be retrieved, its specific form may present diverse characteristics according to the actual application scenario. In terms of the media type dimension, the source media information may include various forms such as a sequence of static images, a dynamic video clip, a high-fidelity audio waveform, or three-dimensional point cloud data. In terms of the input form dimension, the user may provide the source information by directly uploading a local file, inputting a network resource link, collecting it in real-time by shooting, or describing it through speech-to-text. In terms of the application scenario dimension, this technology can adapt to the needs of different vertical fields such as courseware retrieval in the intelligent education field, case matching in medical image analysis, and defect sample query in industrial inspection. The data processing flow and feature extraction strategy can be optimized specifically for different scenarios.

[0066] S204. Obtain prompt template information, where the prompt template information is used to indicate the type of retrieval task that needs to be retrieved currently;

[0067] Optionally, in the embodiments of the present application, the above-mentioned acquisition of prompt template information may include, but is not limited to, guiding the model to identify the objectives and constraints of the current retrieval task through a predefined or dynamically generated instruction set. The prompt template information, as a task control signal, may include natural language instructions, structured task identifiers, or implicit context information, which are used to clarify the modality direction of the retrieval, the result sorting strategy, and the output form requirements. In a specific implementation, the system may generate a task type identifier by parsing the instruction keywords input by the user (such as "find similar pictures", "retrieve relevant answers"), or automatically match a preset template according to the interface interaction elements (such as the "cross-modal retrieval" checkbox selected by the user). This template will guide the model to select the corresponding retrieval strategy. For example, when a "picture + text retrieval" instruction is detected, a multi-modal joint retrieval process is triggered to perform fusion calculation on the combined features of pictures and texts. As Figure 4 shown, for example, when the user inputs "find similar products according to the following product description and design sketch", the system will automatically identify the combined picture and text retrieval requirements included in this instruction, call the corresponding cross-modal matching algorithm, and retrieve candidate results in the product database that meet both the text description features and the visual design features.

[0068] It should be noted that there are various implementation methods and extended forms in the process of acquiring prompt template information. In terms of the task type dimension, the prompt template can cover different operation instructions such as exact retrieval, fuzzy matching, associated recommendation, and anomaly detection. For example, it is required to find completely matching patent documents, recommend architectural design drawings of similar styles, or detect medical images that do not meet the specifications. In terms of the instruction structure dimension, the template can be designed in various forms such as structured query statements, natural language descriptions, example comparison samples, or mixed modality prompts, and can be specifically manifested as fill-in-the-blank query templates, multi-round dialogue instruction sets, or picture-text comparison guiding examples. In terms of the application domain dimension, in addition to basic information retrieval, this mechanism can be extended to emerging fields such as problem classification of intelligent customer service, device status detection in industrial Internet of Things, and scene interaction in virtual reality, and achieve precise control of specific scenarios through customized prompt templates.

[0069] S206. Input the source media information and the prompt template information into the target retrieval model to determine the target media information, where the target retrieval model represents a model obtained by jointly training according to at least two retrieval task types, the loss function used in the training process of the target retrieval model is a joint loss function, the joint loss function is determined by at least two sub-loss functions corresponding one by one to at least two retrieval task types, and the target media information represents the media information retrieved from the source media information according to the prompt template information.

[0070] Optionally, in the embodiments of the present application, the above-mentioned target retrieval model may include, but is not limited to, a deep learning model constructed through a multi-task joint training framework. Its core ability lies in simultaneously adapting to the feature interaction of different modal data and the common requirements of multiple retrieval tasks. The target retrieval model usually adopts a multi-tower architecture or a cross-modal attention mechanism, extracts cross-modal features through a shared underlying encoder, and at the same time realizes the differential processing of different retrieval targets at the task-specific layer. The training data of this model needs to cover various modal combination scenarios, such as text-image paired data, video-audio alignment data, and mixed text and image data, to ensure that the model can understand the semantic association and complementary relationship between different modalities. During the training process, the model learns the general representation across tasks through a parameter sharing mechanism. For example, the visual features learned in the image classification task can enhance the semantic matching accuracy of the text-image retrieval task. For example, in an e-commerce scenario, this model can simultaneously support searching for similar styles based on product images, matching functional attributes based on text descriptions, and recommending related products by combining text and image information, and achieve multi-task knowledge transfer through end-to-end training.

[0071] Optionally, in the embodiments of the present application, the above-mentioned joint loss function may include, but is not limited to, a weighted combination of loss terms designed according to different task characteristics, and its function is to balance the optimization directions of each sub-goal in the multi-task training process. The joint loss function usually includes a modal alignment loss, a task discrimination loss, and a regularization term. Among them, the modal alignment loss is responsible for reducing the distance in the cross-modal feature space, the task discrimination loss ensures that the model outputs discriminative results for different retrieval tasks, and the regularization term is used to prevent multi-task parameter conflicts. In specific implementation, a dynamic weight adjustment strategy can be adopted to automatically adjust the contribution degree of each sub-loss according to the task difficulty or training stage. For example, in a video-text cross-modal retrieval task, the joint loss may include a contrastive learning loss (bringing closer the matching text and image features), a classification loss (distinguishing video categories), and an adversarial loss (eliminating the distribution differences between modalities), and coordinates the update steps of different loss terms through gradient normalization technology. This design enables the model to not only maintain the semantic consistency of cross-modal retrieval but also retain the discriminative features required for specific tasks.

[0072] Optionally, in the embodiments of the present application, the above-mentioned target media information may include, but is not limited to, a set of retrieval results generated after multi-dimensional matching calculations, which need to meet the semantic relevance, modal adaptability, and task compliance with the source media information. The target media information is usually presented in a sorted list form, and each candidate result contains the original data (such as picture thumbnails, text summaries), a matching score, and a cross-modal association identifier. During the generation process, the system will filter and re-rank the candidate set according to the task type specified in the prompt template. For example, when the task requires an exact match of functional parameters, structured data entries are preferentially displayed; when the task focuses on creative similarity, the weight of visual style features is increased. For example, in an industrial design retrieval scenario, the input source media is a conceptual sketch and a text description of materials, and the target media information may include 3D model files, engineering drawings, and patent documents, and the retrieval results are comprehensively scored according to three dimensions: technical feasibility, aesthetic similarity, and compliance.

[0073] It should be noted that the structural design of the target retrieval model has diverse possibilities, including but not limited to that the encoder architecture can adopt a heterogeneous combination of visual Transformer and BERT, the feature fusion layer can select a gated attention mechanism or a dynamic routing network, and the output layer can be designed as a multi-task shared projection matrix or an independent classifier. The combination method of the joint loss function has flexibility. For example, the sub-loss function can select a linear superposition of contrastive learning loss and cosine similarity loss, or adopt a curriculum learning strategy to introduce different loss terms in stages, or dynamically generate loss weight coefficients through a learning framework. The organization form of the retrieval output can be adjusted according to the application scenario. For example, in a real-time interaction scenario, streaming results are presented progressively, in a batch processing scenario, hierarchical clustering results are generated, and in a security-sensitive scenario, credibility annotations and traceability information are added. The present application does not make specific limitations on this.

[0074] It should be noted that there is a wide range of expansion space for the evaluation dimensions of the target media information, which can be subdivided into semantic matching degree, context coherence, and logical consistency. The timeliness dimension can involve data freshness, dynamic trend correlation degree, and periodic pattern fit degree. The diversity dimension needs to balance the content coverage breadth, modal distribution balance, and creative difference. The user interaction method can support multi-round retrieval optimization. For example, the retrieval threshold is dynamically adjusted through a feedback mechanism, the modal weight is adjusted through a visualization control, or the retrieval direction is refined through natural language instructions. The system deployment form can be adapted according to the hardware conditions. For example, a lightweight version after model distillation is adopted on edge devices, distributed index queries are enabled on cloud clusters, and a hierarchical retrieval pipeline is implemented in a hybrid environment. The present application does not make specific limitations on this.

[0075] In an exemplary embodiment, as Figure 5 shown, taking the news content retrieval scenario as an example, it includes but is not limited to the following processes:

[0076] S1, The system receives the source media information and natural language instructions input by the user. For example, the user uploads a screenshot containing a person's speech and attaches the instruction text "Find in-depth report articles related to this event". The source media information can cover single-modal (such as images or text) or a combination of images and text, and the instruction content can dynamically specify the retrieval target type (such as news text, associated images, or video clips).

[0077] S2, The instruction parsing module performs semantic decoding on the natural language instructions. It extracts the key operators (such as "find", "related") in the instructions, the target type ("in-depth report articles"), and the limiting conditions ("event") through a pre-trained language model. The system establishes the mapping relationship between the instruction vector and the preset task type. For example, it maps the above instructions to a cross-modal retrieval task from images to long texts.

[0078] S3, The multi-modal encoder jointly represents the source media information. When the input is an image, a convolutional neural network is used to extract the visual feature vector; if there is an additional text description, text embeddings are generated through a bidirectional Transformer. The feature fusion layer uses the attention mechanism to dynamically weight the image and text features to generate a unified 256-dimensional joint embedding vector.

[0079] S4, The task routing mechanism activates the corresponding retrieval branch according to the instruction type. For the image-to-text retrieval task, the system accesses the news article database and calculates the cosine similarity between the joint embedding vector and the candidate text features. The similarity calculation module adopts a block retrieval strategy, first screening the top 200 candidates through the approximate nearest neighbor algorithm and then performing precise sorting.

[0080] S5, The retrieval result optimization module filters the preliminary results for relevance. Based on the limiting condition of "in-depth report" in the instruction, a semantic matching model is applied to evaluate the content integrity of the candidate articles and filter out short message-like texts. At the same time, it detects the time sensitivity and preferentially displays the authoritative media analysis reports within 48 hours after the event occurs.

[0081] S6, The cross-modal verification mechanism ensures the result consistency. For the top five candidate articles, it performs text-to-image retrieval in reverse to verify their visual relevance to the original input image. If the similarity between the associated image set of a certain article and the input image is lower than the threshold, the final sorting weight of this result is reduced.

[0082] S7, The result presentation module generates a structured output. The system returns a list of articles sorted by relevance, and each result contains summary text, source media, release time, and associated evidence (such as the spatio-temporal matching annotation of the key speech sentences mentioned in the article and the input image). At the same time, it provides a visual analysis view to display the key semantic association heat map between the input image and the retrieval results.

[0083] In another application scenario, assume that the user inputs a sketch of a product design and attaches the instruction "Find marketed products with a similar appearance". The system performs the retrieval task of image-to-product parameters. By extracting the contour features of the sketch and matching the shape coding vectors of the product main images on the e-commerce platform, and at the same time filtering out the delisted products in combination with the "marketed" condition in the instruction, finally, a comparison view of the 3D models of similar products and patent information are returned.

[0084] Therefore, in this embodiment, through the instruction-driven task routing mechanism, a single model can handle at least six types of retrieval tasks such as text-to-image, image-to-text, and cross-modal inspection. In the COCO dataset test, the average recall rate is increased by 2.3 percentage points. The multi-task joint training strategy effectively shares cross-modal knowledge. When dealing with the retrieval task of ancient book illustrations with a small amount of training data, the zero-shot performance reaches 87.6% of the dedicated model. The dynamic feature fusion mechanism reduces the error in the scenario of modality loss and can still maintain a retrieval accuracy of 91.2% in the case of single-modal input. The semantic understanding accuracy of the instruction parsing module reaches 89.4%, which is significantly higher than 72.1% of the traditional fixed template method.

[0085] Through the embodiments of the present application, a dynamically adaptable multi-modal input processing mechanism is adopted. By combining the source media information to be retrieved with the prompt template information indicating the type of retrieval task, a unified retrieval framework is constructed, achieving the purpose of flexibly responding to diverse user needs. The source media information supports any input form covering text, images, and their mixed modalities, breaking through the hard constraints of traditional retrieval systems on the input format, thereby achieving the technical effect of seamless integration and unified representation of cross-modal data sources. By introducing the prompt template information as a task instruction and adopting a semantic-driven retrieval strategy, the system can accurately parse the user's intention and dynamically match the retrieval type, solving the problem of insufficient scenario adaptation ability caused by the dependence of traditional methods on fixed task predefined, thereby achieving an improvement in the task generalization ability based on natural semantics.

[0086] In addition, a multi-task joint training mechanism and an extensible joint loss function design are adopted. By dynamically allocating weights and co-optimizing the sub-loss functions of at least two types of retrieval tasks, the purpose of multi-task knowledge sharing and transfer learning is achieved. This joint training framework enables the model to establish a general feature representation among different retrieval tasks, breaking through the dependence of single-task models on specific data distributions, thereby achieving a stable improvement in cross-task retrieval performance in the zero-shot scenario. Through the dynamic combination strategy of sub-loss functions, the conflict between the optimization directions of each task is effectively balanced, solving the problem of performance decay caused by gradient interference in multi-task learning, thereby enhancing the robustness of the model in complex retrieval scenarios.

[0087] On the other hand, by adopting the instruction-guided end-to-end retrieval paradigm, the accurate alignment between the user's intention and the retrieval behavior is achieved by synchronously inputting the source media information and the prompt template information into the jointly trained target retrieval model. This architecture dynamically activates the retrieval logic of the corresponding task branch by using the prompt template information, breaking through the limitation of the traditional system that requires manual switching of task modules, thereby realizing the intelligent integration of the multimodal retrieval process. Through the collaborative constraint of the multi-task representation space by the joint loss function, the model is prompted to establish a dual mapping of cross-modal semantic association and task association, solving the feature matching deviation problem in heterogeneous retrieval scenarios.

[0088] As an alternative solution, as Figure 6 shown, before inputting the source media information and the prompt template information into the target retrieval model to determine the target media information, the above method further includes:

[0089] S602, obtaining a preset set of retrieval task types and sample media information, where the set of retrieval task types includes different cross-modal retrieval task types, and the sample media information includes media information of different modalities;

[0090] S604, constructing an initial retrieval model with shared parameters based on the retrieval task types, where the initial retrieval model includes a first encoding branch and a second encoding branch, and the first encoding branch and the second encoding branch are used to map the sample media information of different modalities to the same embedding space;

[0091] S606, training the initial retrieval model based on the sample media information to obtain the target retrieval model.

[0092] Optionally, in the embodiments of the present application, the above set of retrieval task types may include, but is not limited to, a task type system predefined for realizing multimodal information interaction, and its core function is to standardize the scope of cross-modal retrieval scenarios that the model can handle. The set of retrieval task types generally covers the modality conversion direction (such as text to image, image to text), the modality combination form (such as text and image to video, audio to text and image), and the task complexity level (such as single-modal retrieval, cross-modal retrieval, multimodal joint retrieval). Each task type needs to clarify the modality constraint conditions of the input and output. For example, the text-to-image retrieval task defines the matching logic between the natural language query and the candidate image set, while the text and image combination retrieval task requires simultaneous processing of the query input of the mixed modality. In specific implementation, the set of task types can be managed through a configuration file or a dynamic loading mechanism. For example, in the intelligent customer service scenario, task types such as product manual retrieval, fault case matching, and operation video recommendation are preset to ensure that the model can cover the diverse retrieval needs required by the business.

[0093] It should be noted that there are multiple possible extensions in the construction of the retrieval task type set, including but not limited to: the task definition dimension can be subdivided into the number of input modalities (single modality / multimodal), output form (sorted list / clustering result), and application field (general / vertical scenario); the task interaction logic can support chained task combinations (such as retrieving text first and then associating images), parallel task processing (simultaneously performing cross-modal retrieval and modality conversion), and dynamic task switching (adjusting the retrieval strategy according to intermediate results); the task evaluation system can include precision metrics (recall rate, F1 value), efficiency metrics (response latency, computational resource consumption), and user experience metrics (result interpretability, interaction fluency). This application does not make specific limitations in this regard.

[0094] Optionally, in the embodiments of this application, the above sample media information may include but is not limited to the cross-modal data set for training the model, and its composition needs to meet the requirements of multimodal alignment, scenario coverage, and semantic richness. The sample media information usually contains paired cross-modal data units, such as text descriptions and corresponding images, video clips and associated captions, audio records and text summaries, and each data unit needs to be labeled with the modality type and cross-modal association label. In the data preprocessing stage, it is necessary to perform standardization processing on the original media information, including unified image resolution, text tokenization and cleaning, audio sampling rate adjustment, etc., to ensure the computability of different modality data. For example, in the educational resource retrieval scenario, the sample media information may include paired data of mathematical formula text and geometric graphic images, associated data of physics experiment videos and operation step descriptions, and combined data of historical event audio explanations and timeline charts. These data form cross-modal training samples after feature extraction.

[0095] It should be noted that the acquisition and processing methods of the sample media information are highly flexible. For example, the data sources can include web crawling of public data sets, extraction of business system logs, and generation by manual annotation; the modality combination form can break through the traditional bimodal limit and support mixed data of three modalities and above (such as a demonstration video with subtitles plus a PPT document); the data augmentation strategy can adopt modality conversion techniques (text-to-image generation, speech-to-text conversion), cross-modal adversarial generation (generating matching audio based on text), and feature space interpolation (mixing modality features of different samples). This application does not make specific limitations in this regard.

[0096] Optionally, in the embodiments of the present application, the above initial retrieval model may include, but is not limited to, a multi-modal encoding architecture constructed based on a parameter sharing mechanism, and its design goal is to achieve cross-modal semantic alignment through a unified feature space. The initial retrieval model usually adopts a two-tower structure or a hybrid encoder. The first encoding branch is responsible for processing sequence modal data such as text and audio, and the second encoding branch focuses on spatial modal data such as images and videos. The two branches learn cross-modal general representations by sharing intermediate layer parameters. During the feature mapping process, the model introduces a cross-modal attention mechanism. For example, it pays attention to image region features during text encoding, or fuses audio spectrum features during video encoding. For example, in the medical image retrieval scenario, the first encoding branch can convert the pathological report text into a semantic vector, and the second encoding branch encodes the CT scan image into visual features. Through the shared contrast learning layer, the feature distance between the two modalities when describing the same disease is reduced, so as to achieve cross-modal semantic consistency.

[0097] It should be noted that there is a rich variant space in the architecture design of the initial retrieval model. For example, the number of encoding branches can be extended to three or more according to the modal type. The feature fusion method can adopt cascade connection, gating mechanism or dynamic routing network. The parameter sharing strategy can implement full connection layer sharing, attention head sharing or gradient collaborative update. The regularization method during training can choose modality-specific Dropout, cross-modal contrast regularization or task relevance constraint, and the optimization goal can combine self-supervised pre-training tasks (masked modality prediction, cross-modal matching) with supervised fine-tuning tasks. The present application does not make specific limitations on this.

[0098] Through the embodiments of the present application, an initial retrieval model with shared parameters is constructed based on a preset retrieval task type set and multi-modal sample data, realizing the unified feature space mapping of cross-modal data, improving the generalization ability of the model for diverse retrieval tasks, and at the same time ensuring the semantic alignment of different modal information, laying a foundation for multi-task joint training.

[0099] As an optional solution, as Figure 6 shown, the initial retrieval model is trained based on the sample media information to obtain a target retrieval model, including:

[0100] In each batch, the initial retrieval model is trained based on the sample media information in the following manner to obtain a target retrieval model, where the retrieval task type corresponding to each batch is the first retrieval task type:

[0101] S606-1, determining first sample media information associated with the first retrieval task type from the sample media information;

[0102] S606-2, inputting the first sample media information into the initial retrieval model to obtain first sample predicted media information;

[0103] S606-3. Determine a first target sub-loss function corresponding to the first retrieval task type according to the predicted media information of the first sample;

[0104] S606-4. Perform backpropagation on the initial retrieval model based on the first target sub-loss function, and update the model parameters of the initial retrieval model;

[0105] Repeat the above steps until a target retrieval model is generated.

[0106] Optionally, in the embodiments of the present application, the above first retrieval task type may include, but is not limited to, a specific cross-modal retrieval task type focused on in the current training batch under a multi-task training framework. This task type is dynamically selected from a predefined set of retrieval task types and is used to guide the data sampling and loss calculation strategies for the current batch. The specific manifestation form of the first retrieval task type may cover single-modal conversion tasks (such as text-to-video retrieval), multi-modal joint tasks (such as text-image combination to audio retrieval), or hybrid tasks (joint execution of cross-modal retrieval and classification tasks). During the training process, this type of dynamic switching mechanism enables the model to evenly learn different task characteristics. For example, in a certain batch, it focuses on improving the semantic matching ability from images to text, and in the next batch, it strengthens the cross-modal temporal alignment ability from videos to audio. For example, in the intelligent education scenario, the first retrieval task type may sequentially switch to exercise text retrieval of teaching videos, experimental video matching operation guides, and 3D model associated knowledge point texts, and multi-task knowledge fusion is achieved through a round-robin mechanism.

[0107] Optionally, in the embodiments of the present application, the above first sample media information may include, but is not limited to, a cross-modal data subset strongly associated with the current training task type, and its screening logic needs to meet the requirements of task alignment, data representativeness, and training stability. The first sample media information is extracted from the original sample pool through a task type filter and needs to contain complete input-output modal pairing data. For example, when the task type is text-to-image retrieval, the data subset should contain text queries and their corresponding image candidate sets. In the data preprocessing stage, feature enhancement needs to be performed according to the task characteristics. For example, adversarial negative samples are added for text-image retrieval tasks, and temporal slice alignment is performed for video-to-audio tasks. For example, in the industrial quality inspection scenario, if the current task type is defect description text retrieval of X-ray images, the first sample media information needs to contain natural language texts describing various defects and their corresponding X-ray image datasets, and normal samples may be added as difficult negative examples to improve the discriminative ability of the model.

[0108] It should be noted that there are various optimization methods for the construction strategy of training batches, including but not limited to: in the dimension of data sampling, curriculum learning strategy can be adopted (gradually transitioning from simple samples to complex samples), dynamic hard example mining (automatically identifying samples with high error rates for enhanced training), and cross-task negative sampling (using data from other tasks to construct difficult negative examples); in the dimension of task scheduling, uniform polling can be implemented (switching tasks in a preset order), adaptive scheduling (dynamically selecting weak tasks based on model performance), and hybrid scheduling (parallel training of the main task and auxiliary tasks); in the dimension of loss calculation, multi-granularity supervision is supported (combining instance-level and modality-level losses), progressive weighting (adjusting loss weights as the training progresses), and adversarial regularization (introducing a discriminator to constrain the feature distribution). This application does not make specific limitations in this regard.

[0109] Optionally, in the embodiments of this application, the above first sample predicted media information may include but not be limited to the cross-modal mapping results of the model for the input data of the current task, and its essence is the feature representation and similarity calculation results of the model in the unified embedding space. The first sample predicted media information is generated by the dual-encoding branch of the initial retrieval model and includes the feature vector of the input modality, cross-modal attention weights, and candidate matching score matrix. The accuracy of this information directly affects the effectiveness of loss calculation. For example, in the image-to-text task, the predicted information needs to include the similarity ranking of the image features and all candidate text features. During the training process, the system drives the optimization of model parameters by comparing the differences between the predicted results and the true labels. For example, in the product retrieval scenario, the first sample predicted media information generated after inputting a product image should include a list of cosine similarities between the image features and all description text features, and the model prediction error is calculated by comparing with the manually marked matching relationship.

[0110] Optionally, in the embodiments of this application, the above first target sub-loss function may include but not be limited to an optimization target function designed according to the characteristics of the current task, and its core function is to quantify the deviation degree between the model prediction result and the expected target. The first target sub-loss function needs to comprehensively consider task characteristics, data distribution, and training stages. For example, contrastive learning loss is often used in cross-modal retrieval tasks, cross-entropy loss is used in classification tasks, and reconstruction loss is introduced in generation tasks. The design of the loss function needs to balance the requirements of modality alignment and task discrimination. For example, in the image-text retrieval task, the intra-modal clustering loss and cross-modal ranking loss can be fused. For example, in the medical image report generation task, the first target sub-loss function may include the contrast loss between image features and text features, the language model loss for report generation, and the auxiliary loss for medical term recognition, and the comprehensive ability of the model is improved through multi-objective optimization.

[0111] It should be noted that the model parameter update mechanism has high scalability. For example, the optimization strategy can adopt hierarchical learning rates (setting different learning rates for the encoder and task heads), gradient clipping (preventing gradient conflicts in multi-task scenarios), and parameter freezing (fixing the shared layers and fine-tuning the task-specific layers); the update frequency supports full-batch updates (accumulating gradients from multiple batches), immediate updates (updating for each batch), and asynchronous updates (independently updating parameters for different tasks); regularization methods can introduce modality-specific Dropout (randomly masking neurons of specific modalities), cross-task knowledge distillation (using a teacher model to guide multi-task learning), and elastic weight consolidation (protecting important task parameters from being overwritten). The present application does not make specific limitations in this regard.

[0112] In addition, the training termination conditions can be flexibly set according to actual needs. For example, in terms of performance metric dimensions, the convergence of the validation set loss, cross-task balance (the threshold of accuracy differences among tasks), and the improvement amplitude of zero-shot ability can be monitored; in terms of resource constraint dimensions, the maximum number of training epochs, the upper limit of computing time, or the energy consumption threshold can be set; in terms of dynamic adjustment dimensions, an early stopping mechanism (terminating if there is no improvement for multiple consecutive rounds), a restart strategy (resetting some parameters when falling into a local optimum), and phase transition (switching the training mode after reaching a preset condition) can be enabled. The present application does not make specific limitations in this regard.

[0113] Through the embodiments of the present application, a task-type-oriented batch training strategy is adopted, different retrieval tasks are dynamically switched for parameter optimization, the adaptability of the model to specific tasks is enhanced, and cross-task knowledge transfer is promoted through targeted training of the sub-loss function, thereby improving the overall retrieval performance and training efficiency of the model.

[0114] As an alternative solution, as Figure 6 shown, determining the first target sub-loss function corresponding to the first retrieval task type according to the first sample prediction media information includes:

[0115] S606-3-1, obtaining the target weight coefficient determined for the first retrieval task type;

[0116] S606-3-2, determining the initial sub-loss function corresponding to the first retrieval task type according to the first sample prediction media information;

[0117] S606-3-3, determining the first target sub-loss function based on the initial sub-loss function and the target weight coefficient.

[0118] Optionally, in an embodiment of the present application, the above-mentioned target weight coefficient may include but is not limited to a dynamically adjusted parameter for balancing the contribution of different sub-losses in multi-task learning, and its core role is to coordinate the intensity of the impact of different retrieval tasks on model optimization. The target weight coefficient can be dynamically calculated according to the importance of the task, the training stage or the data distribution. For example, a higher weight is given to the basic task in the early stage of training to establish a shared representation, and the weight of complex tasks is increased in the later stage to optimize the detail capabilities. In the specific implementation, a fixed ratio can be set through artificial experience, or an adaptive algorithm can be designed to automatically adjust according to the task convergence speed. For example, when the loss of the validation set of a task stagnates, its weight is automatically reduced to alleviate gradient conflicts. For example, in the multimodal retrieval scenario of educational resources, the text-to-video retrieval task may be assigned a weight coefficient of 0.6, while the image-to-exercise retrieval task weight is 0.4, in order to reflect the emphasis on video resources in the business scenario, and at the same time, the learning intensity of the model on different tasks is controlled by adjusting the coefficient.

[0119] It should be noted that there are a variety of optional strategies for determining the target weight coefficient, including but not limited to static allocation dimensions that can set a fixed ratio according to the preset priority of the task (such as 70% for key tasks and 30% for auxiliary tasks), dynamic adjustment dimensions that can be automatically adjusted according to the rate of decline of training loss (such as gradually reducing the weight of tasks with fast loss decline), and mixed strategy dimensions that can be comprehensively calculated based on the task difficulty coefficient and the data volume ratio (such as appropriately increasing the weight of data-scarce tasks). In the medical multimodal training scenario, dynamic weights can be used for CT image to diagnostic report tasks, with high weights initially assigned to establish cross-modal associations, lower weights in the mid-term to avoid overfitting, and later reallocated according to the accuracy of subtasks. This application does not make specific restrictions on this.

[0120] Optionally, in an embodiment of the present application, the above-mentioned initial sub-loss function may include but is not limited to a baseline loss calculation method designed for a specific retrieval task, and its core function is to quantify the prediction deviation of the model on the current task. The selection of the initial sub-loss function needs to be closely aligned with the task characteristics. For example, cross-modal retrieval tasks often use contrastive loss to measure the degree of cross-modal feature alignment, classification tasks use cross-entropy loss to evaluate category differentiation capabilities, and generation tasks may use reconstruction loss to measure generation quality. In a specific implementation, it is necessary to design the loss calculation logic according to the input and output modal combination. For example, the image-text retrieval task needs to simultaneously calculate the bidirectional contrast loss of text to image and image to text. For example, in the cross-modal retrieval of medical image reports, the initial sub-loss function can be designed as a triplet loss (Triplet Loss), which drives the model to learn the semantic association between modalities by shortening the feature distance of the matching image-report pair and pushing the feature distance of the non-matching pair away.

[0121] It should be noted that the selection of the initial sub-loss function has high flexibility. For example, in the dimension of modal characteristics, a perceptual hash loss can be designed for the visual modality, and a semantic coherence loss can be designed for the text modality; in the dimension of task objectives, a ranking loss can be selected in the retrieval task, and an adversarial loss can be adopted in the generation task; in the dimension of computational efficiency, the applicable scenarios of the exact loss function (such as the Wasserstein distance) and the approximate calculation method (such as the cosine similarity) can be weighed. In the educational resource retrieval system, for the video-to-courseware retrieval task, a temporal alignment loss and a key-frame matching loss may be adopted simultaneously to capture time-correlation features at different granularities through loss combination. This application does not make specific limitations in this regard.

[0122] In addition, the generation method of the above first target sub-loss function supports multiple technical paths, including but not limited to directly multiplying the initial sub-loss by the weight coefficient in the linear combination dimension, introducing an exponential weighting method adjusted by the temperature coefficient in the non-linear fusion dimension, and dynamically calculating the weight based on the characteristics of the batch data in the dynamic adjustment dimension. In the industrial quality inspection scenario, for the defect image-to-process parameter retrieval task, a dynamic weight calculation module can be designed to automatically adjust the weights of each sub-task according to the occurrence frequency of different defect types in the current batch. For example, when the proportion of rare defect samples increases, the loss weight of the corresponding sub-task is temporarily increased to enhance the model's learning ability for such samples. This application does not make specific limitations in this regard.

[0123] Through the embodiments of this application, the combination of the dynamic weight coefficient and the initial sub-loss function realizes the balanced optimization of the multi-task loss, avoids a single task from dominating the training process, enhances the model's coordinated processing ability for complex retrieval scenarios, and improves the training stability and convergence speed.

[0124] As an optional solution, training the initial retrieval model based on the sample media information to obtain the target retrieval model includes:

[0125] In each batch, the initial retrieval model is trained based on the sample media information in the following manner to obtain the target retrieval model, where the sample media information input in each batch includes media information of different modalities:

[0126] Determine the second sample media information from the sample media information, and determine the second retrieval task type corresponding to the second sample media information;

[0127] Input the second sample media information into the initial retrieval model to obtain the second sample predicted media information;

[0128] Determine a second target sub-loss function corresponding to the second retrieval task type according to the predicted media information of the second sample;

[0129] Perform backpropagation on the initial retrieval model based on the second target sub-loss function to update the model parameters of the initial retrieval model;

[0130] Repeat the above steps until a target retrieval model is generated.

[0131] Optionally, in the embodiments of the present application, the above second sample media information may include, but is not limited to, cross-modal data combinations screened by task association in the current training batch, and its composition needs to meet the requirements of multi-modal collaborative training. The second sample media information is extracted from the original sample pool through a task type matching mechanism, and needs to contain paired data units of the input modality and the target modality. For example, when the task type is video-to-text retrieval, the input data of this batch should contain video clips and their corresponding text description sets. During the data construction process, it is necessary to ensure the diversity of modality combinations. For example, the same batch may mix samples of various modality combinations such as text-image, audio-video, and text-graphic mixed arrangements, so as to enhance the model's learning ability for complex modality interactions. For example, in the news event retrieval scenario, the second sample media information may include on-site video clips of emergencies, journalist text dispatches, and social media audio comments, enabling the model to understand the relevance between different information carriers through multi-modal joint training.

[0132] It should be noted that there are multiple optimization paths for the organization strategy of the second sample media information, including but not limited to that the modality combination dimension can support single-modal dominance (such as 80% text-image samples + 20% audio samples), balanced mixing (uniform distribution of each modality sample), or dynamic ratio (adjusting the modality ratio according to the training stage); the data augmentation dimension can implement modality-specific augmentation (geometric transformation of images, synonym replacement of text), cross-modal association augmentation (generating corresponding style image sketches based on text), or adversarial sample injection (adding cross-modal interference noise); the batch construction dimension can adopt a fixed-size sliding window, elastic batch (dynamically adjusted according to hardware resources), or curriculum batch (gradually transitioning from simple modality combinations to complex combinations). The present application does not make specific limitations on this.

[0133] Optionally, in the embodiments of the present application, the above-mentioned second retrieval task type may include, but is not limited to, dynamically selected multi-modal retrieval task instances, which function to guide the target direction and optimization strategy of the current batch of training. The selection mechanism of the second retrieval task type can adopt polling scheduling, probabilistic sampling, or performance-aware strategies. For example, it preferentially selects the task type with relatively weak current performance of the model for reinforcement training. This task type needs to clearly define the input modality, output modality, and matching rules. For example, it defines cross-modal retrieval from multi-modal inputs (such as text-image combinations) to single-modal outputs (such as audio), or supports joint retrieval tasks with multi-modal mixed outputs. For example, in the smart home control scenario, the second retrieval task type may be sequentially set to voice command to retrieve device operation videos, gesture image matching to control command texts, and environmental sensor data to associate with the fault handling knowledge base, so as to improve the model's environmental adaptability through task diversity.

[0134] It should be noted that the scheduling mechanism of the second retrieval task type is highly configurable. For example, the priority policy can be based on the task difficulty coefficient (complex tasks are trained first), data freshness (newly stored tasks are prioritized), or business value weight (high-value scenario tasks are prioritized); the switching frequency can be set to a fixed interval (switch every 5 batches), an adaptive interval (adjusted according to the model convergence speed), or event-driven (triggered by specific training events); the task combination method can support independent training of a single task, parallel training of multiple tasks (calculating the losses of multiple tasks simultaneously), or hierarchical task group training (packaging related tasks into a task group). The present application does not make specific limitations on this.

[0135] Optionally, in the embodiments of the present application, the above-mentioned second sample prediction media information may include, but is not limited to, the cross-space mapping results of the model for multi-modal input data, which is essentially the feature expression of heterogeneous data by the model in a unified semantic space. The second sample prediction media information is generated through joint inference of multiple encoders and includes the similarity matrix of feature vectors of each modality, the cross-modal attention distribution map, and the confidence score of candidate results. The quality of this information directly affects the model optimization direction. For example, in the video-audio retrieval task, it is necessary to ensure the temporal alignment accuracy of visual action features and acoustic event features. For example, in the virtual fitting scenario, after inputting the user's body parameter text and the fashion design sketch, the second sample prediction media information should include a list of the matching degrees between the three-dimensional human model feature vectors and the clothing database features, and at the same time generate the confidence score of the visualized fitting effect, providing multi-dimensional supervision signals for loss calculation.

[0136] Optionally, in the embodiments of the present application, the above second target sub-loss function may include, but is not limited to, a loss calculation scheme optimized for the characteristics of the current task. Its core value lies in quantifying the deviation degree of multi-modal feature alignment and driving parameter update. The second target sub-loss function needs to comprehensively consider the modal characteristics and task objectives. For example, in the image-text to video retrieval task, text-video contrast loss, image-video temporal alignment loss, and multi-modal joint ranking loss can be fused. This function balances the contribution ratios of different loss terms through a dynamic weight mechanism. For example, it focuses on intra-modal feature clustering in the initial stage of training and strengthens cross-modal correlation modeling in the later stage. For example, in the autonomous driving scenario, for the multi-sensor data retrieval task, the second target sub-loss function may include the geometric consistency loss between lidar point clouds and camera images, the spatio-temporal correlation loss between millimeter-wave radar signals and navigation maps, and the anomaly detection auxiliary loss of multi-modal fusion features, improving the model's environmental perception ability through multi-level supervision.

[0137] It should be noted that the construction method of the second target sub-loss function supports diversified innovations. For example, in the dimension of loss types, metric learning losses (such as Triplet Loss), reconstruction losses (such as Autoencoder Loss), and adversarial losses (such as GAN Loss) can be integrated; in the dimension of supervision signals, strong supervision labels (manually annotated matching relationships), weak supervision signals (click behavior data), and self-supervised signals (cross-modal prediction tasks) can be fused; in the dimension of optimization objectives, global feature alignment (overall distribution matching between modalities) and local feature matching (corresponding relationships of specific semantic segments) can be balanced. In the industrial design retrieval scenario, for the 3D model retrieval task, surface topology loss (measuring geometric structure similarity), material perception loss (evaluating physical property matching degree), and functional association loss (analyzing component interaction rationality) can be designed, improving the retrieval accuracy through multi-angle supervision. The present application does not make specific limitations on this.

[0138] Through the embodiments of the present application, the batch training mechanism for multi-modal mixed input strengthens the model's joint representation learning ability for heterogeneous data combinations, improves the cross-modal correlation modeling effect through dynamic retrieval task type adaptation, and enhances the robustness of the model to handle diverse retrieval requirements.

[0139] As an optional solution, training the initial retrieval model based on sample media information to obtain a target retrieval model includes:

[0140] Inputting the first group of sample media information into the initial retrieval model to obtain a first sub-loss function, and inputting the second group of sample media information into the initial retrieval model to obtain a second sub-loss function, where the first sub-loss function and the second sub-loss function correspond to different retrieval task types;

[0141] Determine a first joint loss function based on a first sub-loss function and a second sub-loss function, and train an initial retrieval model based on the first joint loss function to obtain a target retrieval model. Among them, the first joint loss function is a function determined by assigning different weights to different retrieval task types and based on whether sample media information of different modalities is similar. When the first joint loss function satisfies the first loss condition, the initial retrieval model is determined as the target retrieval model. The joint loss function includes the first joint loss function.

[0142] Optionally, in the embodiments of the present application, the above-mentioned first group of sample media information may include, but is not limited to, a cross-modal training data set screened for a specific retrieval task type, and its core role is to provide a supervision signal for the model in a specific task scenario. The first group of sample media information must strictly meet the input-output modality constraints defined by the task. For example, when the corresponding retrieval task type is text-to-image retrieval, this group of data should include text query statements and their matching image sets, and may also include difficult negative example samples to improve the discriminative ability of the model. During the data preprocessing process, modality alignment verification needs to be performed to ensure that the text description and the image content have semantic consistency. For example, in the art creation retrieval scenario, this group of data may include abstract painting works and corresponding art review texts, and the semantic correlation between the text and the image is verified through manual annotation, and low-quality matching samples are excluded. Such a data construction method can ensure that the model learns deep semantic associations across modalities.

[0143] Optionally, in the embodiments of the present application, the above-mentioned second group of sample media information may include, but is not limited to, a heterogeneous data set constructed for different retrieval task types, and its design goal is to enhance the multi-task generalization ability of the model. The second group of sample media information usually covers different modality combinations or task complexity levels from the first group. For example, when the first group focuses on text-to-image retrieval, the second group may include data for video-to-audio retrieval tasks. When constructing the data, attention needs to be paid to the knowledge transferability between different tasks. For example, some basic modality features (such as natural language descriptions) are shared, but the target modality is changed (extended from images to 3D models). For example, in the smart home control scenario, the first group may include paired data from voice commands to device operation guide texts, while the second group is designed as cross-modal data from gesture images to device control protocols, forcing the model to establish a more general cross-modal mapping ability through different data distributions.

[0144] It should be noted that there are multiple optimization possibilities for the grouping strategy of sample media information, including but not limited to grouping by modality ratio (text-dominated group, image-dominated group, balanced hybrid group), grouping by task difficulty (basic task group, advanced task group, composite task group), and grouping by data source (public dataset group, business data group, synthetic data group). In the intelligent education scenario, grouping can be designed according to different subject characteristics. For example, the mathematics group focuses on cross-modal retrieval of formulas and geometric figures, the Chinese group strengthens the associated learning of text and recitation audio, and the science group focuses on the temporal matching of experimental videos and operation steps. This application does not make specific limitations in this regard.

[0145] Optionally, in the embodiments of this application, the above first sub-loss function may include but not be limited to a special optimization target designed for the retrieval task corresponding to the first group of data, and its essence is to quantify the prediction error of the model on this task. The selection of the first sub-loss function needs to closely match the task characteristics and data modality. For example, for cross-modal retrieval tasks, a contrast loss function is often used, and the model is optimized by calculating the difference in feature distances between positive and negative sample pairs. In specific implementations, the loss calculation details need to be adjusted according to the modality characteristics. For example, in the image-to-text retrieval task, a visual semantic embedding loss can be introduced, and at the same time, the attention mechanism of the text encoder is combined to optimize feature alignment. For example, in the medical image report generation task, the first sub-loss function may include the contrast loss between image region features and text keywords, the coherence loss of the report structure, and the auxiliary loss of the accuracy of medical terms, forming a multi-level supervision signal.

[0146] Optionally, in the embodiments of this application, the above second sub-loss function may include but not be limited to a loss calculation scheme adapted to the task characteristics of the second group of data, and its design needs to consider the complementarity and synergy with the first sub-loss function. The second sub-loss function often introduces different optimization perspectives. For example, when the first sub-loss focuses on feature alignment between modalities, the second sub-loss can focus on feature clustering within modalities, enhancing the robustness of the model from a dual perspective. In the cross-modal video retrieval scenario, the second sub-loss function may be designed as a temporal consistency loss, requiring the model to maintain the temporal coherence of action events during video segment feature extraction and establish cross-modal associations with the beat features of the audio modality. This design enables the model to not only capture static feature matches but also understand dynamic temporal associations.

[0147] Exemplarily, as Figure 7 shown, it includes a total of 9 tasks from task 1 to task 9. For each task, the sub-loss function is calculated respectively to obtain sub-loss function 1 to sub-loss function 9, and then backpropagation is performed respectively to update the parameters of the model.

[0148] Optionally, in the embodiments of the present application, the above first joint loss function may include, but is not limited to, a composite optimization objective that integrates multi-task losses through a dynamic weight mechanism. Its core value lies in coordinating the influence intensity of different tasks on the update of model parameters. The first joint loss function determines the weight coefficients of each sub-loss through learnable parameters or heuristic rules. For example, the task uncertainty automatic weighting method is adopted to dynamically adjust the loss contribution ratio according to the training difficulty of each task. During the implementation process, a weight constraint mechanism needs to be designed to prevent individual tasks from dominating the optimization direction. For example, a weight upper limit is set or a regularization term is introduced. For example, in a news event multi-modal retrieval system, the first joint loss function may set the weight of the text-to-image retrieval task to 0.6 and the weight of the video-to-summary generation task to 0.4, and form the final optimization objective through weighted summation, which not only ensures the core retrieval performance but also takes into account the collaborative improvement of auxiliary tasks.

[0149] It should be noted that the construction method of the joint loss function has high scalability, including but not limited to linear weighting (simply adding each sub-loss), non-linear fusion (introducing a gating mechanism to adjust weights), curriculum learning (activating different sub-losses in stages), and adversarial training (balancing task optimization through a discriminator). In autonomous driving multi-modal retrieval, a progressive joint loss can be designed. In the initial stage, the sensor data alignment is the main loss, in the middle stage, the environmental semantic understanding auxiliary loss is added, and in the later stage, the safety constraint regularization loss is introduced to form a progressive optimization objective system. The present application does not make specific limitations on this.

[0150] Optionally, in the embodiments of the present application, the above first loss condition may include, but is not limited to, a composite evaluation criterion for determining the convergence of model training, and its setting needs to comprehensively consider model performance and training resource constraints. The first loss condition usually includes an absolute threshold (such as the joint loss value is lower than 0.01), a relative change rate (such as the loss drops less than 1% for 5 consecutive epochs), and a task balance index (such as the difference between the loss values of each sub-task does not exceed 20%). In specific implementation, a multi-stage condition judgment mechanism can be designed. For example, larger loss fluctuations are allowed in the initial stage of training, while strict convergence judgment is performed in the later stage. For example, in an industrial defect detection scenario, the first loss condition may require that the cross-modal retrieval loss is lower than 0.05, and at the same time, the intra-modal clustering loss remains stable within a fluctuation range of no more than 5%, and the GPU memory occupancy does not exceed the preset threshold, forming a multi-dimensional termination judgment criterion.

[0151] It should be noted that the setting dimensions of the loss conditions support flexible configuration, including but not limited to the performance-oriented dimension (accuracy reaching 95%, recall rate exceeding 90%), the resource constraint dimension (training duration less than 24 hours, GPU utilization rate lower than 80%), the stability dimension (smoothness of the loss curve, range of gradient fluctuations), and the business metric dimension (improvement in click-through rate, improvement in user retention). In the e-commerce cross-modal retrieval scenario, a composite termination condition can be set: when the improvement in the text-image retrieval accuracy is less than 0.5% for three consecutive rounds and the optimization of the video retrieval time consumption reaches 30%, it is determined that the model training is completed. This application does not make specific limitations in this regard.

[0152] Through the embodiments of this application, the joint loss function integrates multi-task optimization objectives and combines a weight allocation mechanism to coordinate the training direction, suppress interference between modalities, and improve the comprehensive performance and generalization ability of the model in complex retrieval scenarios.

[0153] As an alternative solution, as Figure 8 shown, input the source media information and the prompt template information into the target retrieval model to determine the target media information, including:

[0154] S802, determine a set of candidate media information based on the prompt template information, where the set of candidate media information includes the candidate media information specified by the prompt template information;

[0155] S804, determine the candidate media information in the set of candidate media information that meets the preset similarity condition with the source media information as the target media information.

[0156] Optionally, in the embodiments of this application, the above set of candidate media information may include but is not limited to a pool of potential matching results dynamically screened or generated according to the prompt template information, and its composition needs to meet the constraints of the retrieval task type on modality, content, and scale. The set of candidate media information is extracted by the task parsing engine from the underlying database or real-time data stream, and contains candidate data units that meet the modality type, theme range, and format requirements specified by the prompt template. In specific implementation, it is necessary to perform a preliminary filtering on the candidate set by combining semantic understanding technologies, such as using keyword matching, pre-screening of feature similarity, or association and extension of the domain knowledge graph. For example, in the e-commerce scenario, if the prompt template specifies "retrieve similar products according to the clothing design sketches uploaded by users", the set of candidate media information may include images, 3D models, and descriptive texts of all clothing categories in the product library, which are quickly constructed through cross-modal indexing technology, and at the same time, data of irrelevant categories (such as home appliances or food) are excluded to improve the retrieval efficiency.

[0157] It should be noted that there are multiple possibilities for constructing the candidate media information set, including but not limited to: in terms of the construction strategy dimension, static preloading (periodically updating the candidate pool in full), dynamic expansion (supplementing candidates in real time according to user operations), or a hybrid strategy (preloading core candidates + dynamically obtaining long-tail candidates) can be adopted; in terms of the update mechanism dimension, full-automatic update (automatically adding and deleting based on model prediction results), semi-automatic update (assisted by manual review), or versioned update (iterating according to data versions) is supported; in terms of the content composition dimension, it can include single-modal candidates (only text or only images), multi-modal hybrid candidates (both text and images, audio and video coexist), or enhanced candidates (additional metadata such as click-through rate, timeliness tags). In the medical image retrieval scenario, the candidate set may combine the case feature library preprocessed offline with the emergency image data stream accessed in real time to construct a dynamically updated cross-modal retrieval candidate pool through a hybrid strategy. This application does not make specific limitations on this.

[0158] Optionally, in the embodiments of this application, the above preset similarity conditions may include but are not limited to the comprehensive determination criteria of multi-dimensional feature matching rules, and their design needs to balance semantic relevance, modal compatibility, and business scenario requirements. The preset similarity conditions are usually composed of the feature space distance threshold, cross-modal attention weight distribution, and task-specific constraints. For example, in the text-image retrieval task, it is required that the matching degree between the key entities described in the text and the image region features exceeds the set threshold, and at the same time, the overall semantic vector cosine similarity reaches the minimum standard. During the implementation process, the condition combination can be optimized through a dynamic adjustment mechanism. For example, the weight of the timeliness feature can be automatically increased according to the real-time feedback of the retrieval. For example, in the news event retrieval scenario, the preset similarity conditions may require that the release time of the candidate result is within 24 hours after the event occurs, and the overlap degree between the text summary and the key entities (such as people, places) of the source image exceeds 70%, and at the same time, the cross-modal feature similarity ranking is in the top 10%.

[0159] It should be noted that the setting dimension of the preset similarity conditions has wide adaptability, including but not limited to: in terms of the feature granularity dimension, global feature matching (overall semantic similarity), local feature alignment (corresponding relationship of key segments), or hierarchical feature combination (global + local weighting) is supported; in terms of the modal interaction dimension, one-way matching conditions (only from the source modality to the target modality), two-way matching conditions (two-way feature alignment verification), or collaborative matching conditions (multi-modal joint determination) can be designed; in terms of the business rule dimension, permission control (filtering sensitive content), timeliness constraint (prioritizing recent data), or diversity requirement (avoiding result homogenization) can be incorporated. In the industrial design retrieval scenario, the similarity conditions may simultaneously require the structural topology matching degree of the 3D model, the numerical range compatibility of the material parameters, and the feasibility score of the production process to form a multi-dimensional composite determination standard. This application does not make specific limitations on this.

[0160] Through the embodiments of the present application, the candidate set is dynamically screened based on the prompt template and combined with the preset similarity conditions for accurate matching, improving the relevance and accuracy of the retrieval results. The retrieval efficiency is optimized through structured candidate pool management, reducing the redundant computing overhead.

[0161] As an alternative solution, as Figure 8 shown, determine the candidate media information set based on the prompt template information, including:

[0162] S802-1, perform feature extraction and splicing on the source media information and the prompt template information to obtain the source feature vector;

[0163] S802-2, determine the current retrieval task type based on the source feature vector;

[0164] S802-3, determine the candidate media information set according to the current retrieval task type.

[0165] Optionally, in the embodiments of the present application, the above feature extraction and splicing may include, but are not limited to, the fusion operation of distributed representation of multimodal input information through a heterogeneous feature encoder. Its core goal is to eliminate the heterogeneity between modalities and construct a unified semantic expression space. In the feature extraction stage, an appropriate encoder needs to be selected according to the input modality type. For example, a language model that combines word vectors and positional encoding is used for text, a convolutional neural network is used to extract spatial features for images, and a Mel spectrogram transformation and a temporal convolutional network are used to model speech features for audio. In the splicing process, feature concatenation, weighted fusion, or an attention mechanism can be used to dynamically aggregate modal features. For example, in the commodity retrieval scenario, the semantic vector of the product description text and the deep visual features of the commodity image are aligned and fused through attention to form a combined feature with cross-modal association ability. For example, in the intelligent education application, after the courseware screenshots uploaded by the teacher and the audio of their speech explanations are feature-extracted, a cross-modal attention mechanism is used to generate a fusion feature vector with text-image-sound association, enhancing the model's semantic understanding ability for complex teaching content.

[0166] It should be noted that there are multiple possibilities for the specific implementation of feature extraction and splicing, including but not limited to: for the encoder selection dimension, pre-trained models (such as the image-text joint encoding of CLIP), domain-specific models (medical image-specific feature extractors), or lightweight models (distillation models deployed on mobile devices) can be used; for the feature fusion strategy, gating mechanisms can be used to dynamically adjust the modal weights, cross-attention mechanisms can be used to achieve modal interaction, or adversarial training can be used to eliminate modal differences; the preprocessing methods can include data augmentation (image rotation, text synonym replacement), noise injection (simulating low-quality inputs), or modal conversion (speech-to-text to assist feature alignment). In the smart city scenario, for cross-modal retrieval of road videos and sensor data, spatio-temporal two-stream networks can be selected to extract video features, combined with the parsing of sensor features through Internet of Things protocols, and finally multi-dimensional features can be fused through self-attention mechanisms. This application does not make specific limitations in this regard.

[0167] Optionally, in the embodiments of this application, the above source feature vectors may include but are not limited to high-dimensional distributed feature representations that represent the semantic core of source media information, and their function is to provide a computable similarity measurement benchmark for cross-modal retrieval. The generation of source feature vectors needs to comprehensively consider the multi-dimensional attributes of the original data, including but not limited to the semantic intention of text, the visual elements of images, the emotional tendency of audio, and the spatio-temporal correlation characteristics of videos. After being mapped to a unified embedding space through an encoder network, this vector has cross-modal comparability. For example, in the medical image consultation scenario, the source feature vectors of the patient's described condition text and medical images can share the same feature space, enabling direct similarity calculation between visual features and text semantics. For example, the source feature vector generated after a user uploads a scenic photo on a travel guide platform not only contains visual content (architectural style, natural landscape), but also integrates the encoded features of geographical information through metadata, supporting cross-modal matching with multi-modal data such as text travel notes and scenic area audio commentaries.

[0168] It should be noted that the optimization directions of source feature vectors are diverse, including but not limited to: in the dimension of representation learning, contrastive learning can be implemented to enhance cross-modal discriminability, knowledge distillation can be used to compress the vector dimension, and adversarial training can be used to improve anti-interference ability; in the dimension of application adaptation, scene-specific vectors (such as the stylized encoding of artworks), hierarchical vector structures (separating basic features and domain features), or dynamically expandable vectors (supporting seamless integration of new modalities) can be designed; in the dimension of evaluation and improvement, the discriminability can be analyzed through feature visualization, the intra-class and inter-class distances can be calculated through clustering metrics, or the feature optimization can be reversely guided by the accuracy of downstream tasks. In the agricultural pest retrieval scenario, for cross-modal retrieval of crop leaf images and climate data, the source feature vector can integrate visual pathology features, meteorological pattern features, and soil composition encoding to form a multi-dimensional fusion representation. This application does not make specific limitations in this regard.

[0169] Optionally, in the embodiments of the present application, the above-mentioned current retrieval task type may include, but is not limited to, specific retrieval requirement types dynamically identified based on input features and preset task templates, and its determination needs to comprehensively consider template instructions and modal features of source data. Task type recognition is jointly completed by analyzing the hidden layer activation pattern of the source feature vector and the structured label of the template information. For example, a multi-layer perceptron classifier is used to determine whether the input data is more suitable for text-to-image retrieval or audio-to-video retrieval. In specific implementations, ambiguous retrieval requirements that are not clearly specified can trigger a multi-task joint reasoning mechanism. For example, when a user uploads a picture and mood text on a social platform at the same time, the system automatically identifies that a cross-modal retrieval task of mixing pictures and texts to short video recommendation needs to be executed. For example, in the industrial design scenario, when a user uploads a 3D model of a product and attaches a functional description text, the modal weight distribution of the source feature vector will trigger a composite task type of "model-graphic-patent retrieval", and jointly retrieve technical documents, design drawings and historical patent information.

[0170] It should be noted that the recognition mechanism of the current retrieval task type has high scalability, including but not limited to: in the decision logic dimension, a rule engine (preset logic tree determination), machine learning classification (training based on historical behavior), or reinforcement learning (dynamically optimizing according to feedback) can be adopted; in the determination basis dimension, it can be based on the modal ratio of the feature vector, the user's historical preference pattern, or the context session state; the dynamic adjustment dimension supports real-time task type correction (optimizing according to the initial retrieval result), multi-task parallel execution (meeting multiple potential requirements at the same time), or hierarchical task division (coordination between the main task and the auxiliary task). In the financial risk control retrieval scenario, the transfer record text and call recording uploaded by the user may trigger a composite task type determination, and simultaneously perform semantic compliance retrieval (text to regulatory provisions) and voice emotion analysis (audio to risk features), so as to improve the risk recognition accuracy through multi-task coordination. The present application does not make specific limitations on this.

[0171] Through the embodiments of the present application, feature extraction and dynamic task type determination realize intelligent adaptation of the retrieval process, enhance the model's ability to understand user intentions, ensure that the candidate set construction highly matches the task requirements, and improve the flexibility and scenario adaptability of cross-modal retrieval.

[0172] As an optional solution, as Figure 8 shown, determining the candidate media information that meets the preset similarity condition with the source media information in the candidate media information set as the target media information includes:

[0173] S804-1, obtaining the candidate feature vector corresponding to the candidate media information in the candidate media information set;

[0174] S804-2, determining the target feature vector based on the similarity between the source feature vector and the candidate feature vector;

[0175] S804-3. Determine the candidate media information corresponding to the target feature vector as the target media information.

[0176] Optionally, in the embodiments of the present application, the above-mentioned candidate feature vectors may include, but are not limited to, high-dimensional semantic representations after the candidate media information is encoded by modality adaptation, which is used to provide a unified and quantifiable comparison benchmark for cross-modal similarity calculation. The candidate feature vectors are generated by an encoder network for a specific task and need to be in the same embedding space as the source feature vectors to achieve cross-modal comparability. For multi-modal candidate media information, modality feature extraction needs to be performed separately and then spatial alignment is carried out. For example, video candidate information needs to be decomposed into visual, audio, and text track features, encoded separately, and then fused into a unified feature vector through an attention mechanism. For example, in the scenario of an intelligent library, the candidate feature vectors of an e-book may include the visual features of the cover image, the semantic vectors of the abstract text, and the voiceprint features of the author interview audio, and a feature vector with comprehensive representation ability is generated through a cross-modal fusion layer to support effective matching with the mixed graphic and text input of the user's query.

[0177] It should be noted that there are multiple technical paths for generating candidate feature vectors, including but not limited to: in terms of the encoder architecture dimension, single-modal independent encoding (each modality is processed separately), cross-modal joint encoding (sharing underlying parameters), or hybrid encoding (partial sharing + task-specific layers) can be adopted; in the feature fusion stage, early fusion (raw data-level fusion), mid-term fusion (feature-level fusion), or late fusion (decision-level fusion) can be implemented; in the update mechanism dimension, static features (pre-computed and stored), dynamic features (real-time calculation), or incremental features (online learning and updating) are supported. In the scenario of intelligent healthcare, for the generation of candidate feature vectors for medical images, a pre-trained radiological image encoder can be used to extract visual features, combined with a real-time updated pathological report language model to generate text features, and finally cross-modal feature adaptive fusion is achieved through a dynamic routing network. The present application does not make specific limitations on this.

[0178] Optionally, in the embodiments of the present application, the above-mentioned similarity may include, but is not limited to, a quantitative index for measuring the semantic association strength between cross-modal feature vectors, and its calculation method needs to adapt to the task requirements and data characteristics. Similarity measurement usually combines spatial distance calculation (such as cosine similarity, Euclidean distance) and attention weight analysis (such as the matching degree of cross-modal interaction regions), and at the same time introduces a task-specific weighting strategy. For example, in the scenario of judicial case retrieval, when calculating the similarity between legal text and court trial videos, higher weights need to be given to the matching of legal article keywords, and the temporal consistency of audio-visual materials is used as an auxiliary judgment index. Specifically, when implementing, a hierarchical similarity calculation framework can be designed. First, coarse-grained global feature matching and screening are carried out, and then fine-grained local feature alignment verification is carried out on high-potential candidates.

[0179] It should be noted that the selection of similarity calculation methods is highly flexible, including but not limited to: in the dimension of metric learning, supervised metrics (optimizing distance functions based on labeled data), self-supervised metrics (utilizing the internal structure of data), or semi-supervised metrics (combining a small amount of labeled and a large amount of unlabeled data) can be adopted; in the dimension of calculation granularity, global similarity (overall feature matching), local similarity (key segment alignment), or hierarchical similarity (multi-level feature combination calculation) can be supported; in the dimension of optimization objectives, exact matching (searching for exact corresponding items), fuzzy matching (tolerating partial feature deviations), or associative matching (discovering potential correlations) can be emphasized. In the scenario of digitalization of cultural heritage, hierarchical similarity calculation may be adopted for cross-modal retrieval of cultural relic fragments: first, a preliminary screening is performed based on the Hausdorff distance of the overall contour, then a fine-grained matching is carried out through the local binary pattern features of the surface texture, and finally, a weighted determination is made by combining the expert rules of material spectral analysis. This application does not make specific limitations in this regard.

[0180] Optionally, in the embodiments of this application, the above-mentioned target feature vectors may include but are not limited to candidate feature subsets obtained by screening or sorting and optimizing through similarity thresholds, and their essence is the projection of the optimal solution space of the retrieval task. The determination of the target feature vector needs to comprehensively consider the absolute similarity value and the relative ranking position. For example, in the e-commerce scenario, the target feature vector of a certain product image needs to simultaneously meet the requirement that the visual similarity with the reference image uploaded by the user exceeds 0.85 and it ranks among the top five in the candidate set after price range screening. This process may involve multi-stage filtering strategies. For example, in industrial design retrieval, a preliminary screening is first performed through three-dimensional structure similarity, then refined by combining material parameter matching, and finally, the target feature vector is determined through weighted ranking based on the feasibility of production processes.

[0181] It should be noted that the screening strategy of the target feature vector supports multi-dimensional customization, including but not limited to: in the dimension of threshold setting, fixed thresholds (presetting the lower limit of similarity), dynamic thresholds (adjusting according to the quality of the candidate set), or elastic thresholds (setting different criteria for different modalities) can be adopted; in the dimension of sorting strategies, simple sorting (descending order according to similarity), weighted sorting (adjusting weights in combination with business rules), or clustering sorting (maintaining the diversity of results) can be implemented; in the post-processing dimension, duplicate removal filtering (eliminating duplicate candidates), diversity control (ensuring that the results cover different subcategories), or interpretability enhancement (annotating key matching features) can be added. In the scenario of educational resource recommendation, the screening of the target feature vector may adopt a dynamic threshold mechanism: when retrieving "junior high school mathematics knowledge points", relevant data such as exercise difficulty, teaching video duration, and student interaction data are automatically associated to generate an optimized ranking result that takes into account knowledge coverage, learning efficiency, and participation. This application does not make specific limitations in this regard.

[0182] Through the embodiments of the present application, the cross-modal similarity calculation mechanism of the unified feature space realizes accurate feature-level matching, optimizes the quality of the retrieval results through multi-level screening of candidate feature vectors, and improves the ability to capture semantic associations under complex modal combinations.

[0183] As an alternative solution, as Figure 9 shown, the above method further includes:

[0184] S902, obtaining a preset set of retrieval task types and sample media information, where the set of retrieval task types includes different cross-modal retrieval task types, and the sample media information includes media information of different modalities;

[0185] S904, constructing an initial retrieval model with shared parameters based on the retrieval task type, where the initial retrieval model includes a first encoding branch, a second encoding branch, and a third encoding branch. The first encoding branch and the second encoding branch are used to map sample media information of different modalities to the same embedding space, and the third encoding branch is used to map prompt template information to the same embedding space;

[0186] S906, training the initial retrieval model based on the sample media information to obtain a target retrieval model.

[0187] Optionally, in the embodiments of the present application, the above set of retrieval task types may include, but is not limited to, a task system framework predefined for realizing multi-modal interaction, and its core function is to standardize the scope of cross-modal retrieval scenarios that the model can handle. The set of retrieval task types usually covers input-output modal combinations (such as text to video, text and graphics to audio), task complexity levels (single-modal retrieval, cross-modal retrieval, multi-modal joint retrieval), and domain characteristics (general retrieval, vertical domain retrieval). Each task type needs to clarify the modal conversion direction, candidate set construction rules, and matching logic. For example, a text and graphics mixed retrieval task needs to define the fusion strategy of text and image features and the candidate set screening criteria. In specific implementation, the set of task types can be managed through a dynamic loading mechanism. For example, in an intelligent customer service scenario, task types such as product parameter retrieval, failure case matching, and operation video recommendation are preset to ensure that the model covers the diverse interaction requirements of the business.

[0188] It should be noted that there are multiple possibilities for constructing the set of retrieval task types, including but not limited to: in terms of the task definition dimension, it can be subdivided into the number of input modalities (single input / multiple inputs), output forms (sorted list / cluster map / generative result), and interaction modes (single retrieval / multi-round progressive retrieval); in terms of the update mechanism dimension, it supports static predefined (manually configuring the task library), dynamic expansion (automatically deriving new tasks according to the user's operations), or hybrid management (fixed core tasks + dynamically adjusted marginal tasks); in terms of the evaluation system dimension, it can include performance indicators (recall rate, response latency), resource consumption (computing load, storage requirements), and user experience (result interpretability, interaction fluency). In smart city management, the set of task types may include multi-level requirements such as cross-modal retrieval from traffic videos to event reports and association matching from sensor data to emergency plans. This application does not make specific limitations in this regard.

[0189] Optionally, in the embodiments of this application, the above sample media information may include but is not limited to the cross-modal alignment data set for training the model, and its composition needs to meet the requirements of multi-modal relevance, scene diversity, and semantic integrity. The sample media information contains paired cross-modal data units, such as text descriptions and corresponding images, video clips and associated captions, audio recordings and text summaries, and each data unit needs to be labeled with the modality type and cross-modal association label. In the data preprocessing stage, it is necessary to perform standardization processing on the original media information, including unified image resolution, text entity recognition, audio noise reduction and enhancement, etc., to ensure the computability of different modality data. For example, in the industrial quality inspection scenario, the sample media information may include paired data of defect description texts and X-ray images, associated data of sensor time series signals and fault logs, and combined data of operation videos and maintenance manuals, which form cross-modal training samples after feature extraction.

[0190] It should be noted that the acquisition and processing strategies of the sample media information have high flexibility, including but not limited to: in terms of the data source dimension, it can integrate public data sets (cross-modal academic data sets), business system logs (user operation trajectory data), and synthetic data (modality conversion generated data); the modality combination form can break through the traditional bimodal limit and support tri-modal and above hybrid data (such as a demonstration video with captions plus a PPT document and an operation manual); in terms of the enhancement strategy dimension, it can adopt cross-modal adversarial generation (generating matching images based on text), feature space interpolation (mixing modality features of different samples), and hard sample mining (automatically identifying easily confused samples to strengthen training). In the agricultural pest control scenario, the sample media information may include tri-modal combined data of crop leaf images, environmental sensor data, and pest control plan texts, which can improve the model's adaptability to complex field environments through multi-dimensional enhancement. This application does not make specific limitations in this regard.

[0191] Optionally, in the embodiments of the present application, the above initial retrieval model may include, but is not limited to, a multi-modal joint encoding architecture constructed based on a parameter sharing mechanism, and its design goal is to achieve cross-modal semantic alignment and instruction understanding through a unified feature space. The first encoding branch is responsible for processing sequential modal data such as text and audio, and uses a pre-trained language model to extract semantic features; the second encoding branch focuses on spatial modal data such as images and videos, and extracts visual features through a convolutional neural network; the third encoding branch parses the prompt template information and converts it into a task control vector. The three branches learn cross-modal general representations by sharing intermediate layer parameters and achieve interaction between modalities through a cross-attention mechanism. For example, in the medical image diagnosis scenario, the first branch encodes the patient's medical history text, the second branch processes the CT scan image, and the third branch parses the prompt template of "finding similar cases", and the three work together to generate cross-modal joint features.

[0192] Optionally, in the embodiments of the present application, the above target retrieval model may include, but is not limited to, a cross-modal retrieval system optimized through multi-task joint training, and its core ability lies in dynamically adapting to different retrieval task requirements according to the prompt template. The target retrieval model integrates multi-modal feature understanding and task instruction parsing capabilities through end-to-end training, and can achieve precise alignment of the cross-modal feature space. The model can automatically activate the corresponding encoding branch according to the modal combination of the input data during the inference stage. For example, when the user inputs a product design sketch and a function description text, it automatically triggers the graphic-to-3D model retrieval task and calculates the feature similarity in a unified embedding space. The model supports multi-modal association retrieval in complex scenarios. For example, in an intelligent education system, it simultaneously supports multi-task requirements such as knowledge point text retrieval of teaching videos, experimental operation video matching exercise analysis, and course audio association with electronic lesson plans.

[0193] It should be noted that there is a rich variant space in the architecture design of the initial retrieval model, including but not limited to: the number of encoding branches can be extended according to the modal type (adding a voice encoding branch, a 3D point cloud encoding branch); the parameter sharing strategy can implement full connection layer sharing, attention head sharing, or gradient collaborative update; the feature fusion method can adopt a gated attention mechanism (dynamically adjusting modal weights), a hierarchical fusion network (integrating different granularity features in stages), or memory-enhanced fusion (introducing an external knowledge base to assist decision-making). In the financial risk control scenario, the initial model may be designed with four encoding branches to process the user's record text, transaction behavior sequence, voice call features, and compliance document template respectively, and extract cross-modal risk features through multi-level parameter sharing. The present application does not make specific limitations on this.

[0194] Through the embodiments of the present application, a new prompt template encoding branch is constructed to build a multi-modal joint model, enhancing the deep fusion of instruction semantics and cross-modal features, improving the model's parsing ability for dynamic retrieval instructions, and supporting end-to-end optimization in complex task scenarios.

[0195] As an alternative solution, as Figure 9 shown, based on the sample media information, the initial retrieval model is trained to obtain a target retrieval model, including:

[0196] S906-1, input the third group of sample media information into the initial retrieval model to obtain a third sub-loss function, and input the fourth group of sample media information into the initial retrieval model to obtain a fourth sub-loss function, where the third sub-loss function and the fourth sub-loss function correspond to different retrieval task types;

[0197] S906-2, determine a second joint loss function based on the third sub-loss function and the fourth sub-loss function, and train the initial retrieval model based on the second joint loss function to obtain a target retrieval model, where the second joint loss function is a function that assigns different weights to different retrieval task types and is determined based on whether the sample media information is similar to the sample candidate media information set. When the second joint loss function satisfies the second loss condition, the initial retrieval model is determined as the target retrieval model, and the joint loss function includes the second joint loss function.

[0198] Optionally, in the embodiments of the present application, the above-mentioned third group of sample media information may include, but is not limited to, a training data set optimized for a specific cross-modal retrieval task. Its core function is to enhance the generalization ability of the model through differentiated modal combinations. The third group of sample media information needs to cover modal interaction scenarios different from the previous training group. For example, when the first group focuses on text-to-image retrieval, the third group may contain data for video-to-3D model retrieval tasks. During the data construction process, attention should be paid to cross-task knowledge transferability. For example, sharing basic modal features but changing the target modality (such as extending text-to-video to text-to-animation). For example, in an industrial design scenario, the third group may contain paired data of mechanical design drawings and material parameter texts, and at the same time introduce manufacturing process videos as cross-modal extensions, enabling the model to understand the full-process semantic associations from design to production.

[0199] Optionally, in the embodiments of the present application, the above-mentioned fourth group of sample media information may include, but is not limited to, a high-order training data set for complex multi-modal interaction scenarios, and its design goal is to improve the model's ability to process composite tasks. The fourth group of sample media information usually contains mixed data of three modalities or more (such as a demonstration video with subtitles plus operation manual text), and requires the model to simultaneously process the correlation between multi-modal inputs and outputs. In the data preprocessing stage, cross-modal alignment verification needs to be implemented. For example, synchronize video segments and voice commentaries through timestamps, or associate design drawings and 3D models through spatial annotations. For example, in the intelligent medical scenario, the fourth group of data may contain a three-modal combination of patient CT images, electronic texts, and consultation recordings, and requires the model to achieve multi-task collaboration such as image-to-report generation and voice-to-text retrieval.

[0200] It should be noted that there are multiple possibilities for the construction strategies of the third group of sample media information and the fourth group of sample media information, including but not limited to: in the dimension of modal combination, a master-slave combination (core modality + auxiliary modality), a parallel combination (multi-modalities participate equally), or a hierarchical combination (basic modality + derivative modality) can be designed; in the dimension of data source, real-time stream data (user interaction logs), offline archived data (historical business data), and synthetic enhanced data (cross-modal generated samples) can be integrated; in the dimension of enhancement strategy, adversarial perturbation (simulating real environment noise), cross-modal interpolation (generating transitional samples), and feature decoupling (separating modality-specific and shared features) can be adopted. In the intelligent education scenario, the third group of data may mix live course video streams, real-time barrage texts, and electronic blackboard images, and construct cross-modal training samples through spatio-temporal alignment technology to enhance the model's understanding ability of dynamic teaching scenarios.

[0201] Optionally, in the embodiments of the present application, the above-mentioned third sub-loss function may include, but is not limited to, an optimization objective function adapted to new modal combinations, and its design needs to break through the limitations of traditional bi-modal losses. The third sub-loss function usually fuses cross-modal contrast loss and intra-modal clustering loss. For example, in a video-text-audio three-modal retrieval task, it is necessary to calculate the contrast losses of video-to-text, text-to-audio, and video-to-audio simultaneously, and dynamically allocate weights through an attention mechanism. This function models the complex association paths between multi-modalities by introducing a modal relationship graph convolutional network. For example, in a virtual reality content retrieval scenario, the third sub-loss function may include the geometric matching loss between a 3D scene model and interactive text, the attention alignment loss between spatial audio and visual focus, and the temporal correlation loss between the user's log and content features.

[0202] Optionally, in the embodiments of the present application, the above fourth sub-loss function may include, but is not limited to, a robustness optimization function designed for long-tailed distribution and data sparse scenarios. The fourth sub-loss function gradually improves the model's processing ability for low-resource modality combinations by introducing a curriculum learning strategy and an adversarial training mechanism. This function adopts dynamic hard example mining technology to automatically identify sample pairs with weak cross-modal associations for enhanced training, and at the same time introduces a modality invariance regularization term to prevent overfitting. For example, in the scenario of cultural heritage digitization, for the scarce data of ancient book images, ancient text, and restoration process videos, the fourth sub-loss function can design a cross-modal feature interpolation loss to generate synthetic samples in the latent space to augment the training data, while constraining the smoothness of the feature distribution to maintain semantic coherence.

[0203] It should be noted that there is a wide range of exploration space for the innovation directions of the third sub-loss function and the fourth sub-loss function, including but not limited to: in the dimension of loss types, contrastive learning loss (enhancing modality distinguishability), reconstruction loss (maintaining feature integrity), and adversarial loss (improving distribution robustness) can be fused; in the dimension of supervision signals, strong supervision annotations (manually refined labeled data), weak supervision signals (user click behavior), and self-supervised signals (cross-modal prediction tasks) can be combined; in the dimension of optimization strategies, progressive optimization (activating loss terms in stages), curriculum learning (from simple to difficult samples), and memory-augmented training (using external knowledge bases for guidance) can be implemented. In the industrial Internet of Things scenario, for the anomaly detection task of device multi-sensor data, the fourth sub-loss function may integrate the vibration signal spectrum reconstruction loss, infrared thermal imaging contrast loss, and maintenance log classification loss to achieve accurate fault location through multi-objective optimization.

[0204] Optionally, in the embodiments of the present application, the above second joint loss function may include, but is not limited to, an adaptive optimization target that dynamically integrates multi-task losses through a meta-learning framework. The second joint loss function adopts differentiable architecture search technology to automatically discover the optimal loss weight combination strategy, which includes a task relevance perception module and a resource constraint evaluation module. This function continuously adjusts the contribution degrees of each sub-loss through an online learning mechanism. For example, in the initial stage of training, it focuses on basic modality alignment, in the middle stage it strengthens complex task optimization, and in the later stage it balances computational efficiency and accuracy. For example, in the multi-sensor retrieval scenario of autonomous driving, the second joint loss function can dynamically adjust the multi-task loss weights such as lidar point cloud to high-precision map, camera image to traffic sign text, and millimeter wave radar to control instructions to achieve full-link optimization of perception-decision-control.

[0205] Optionally, in the embodiments of the present application, the above second loss condition may include, but is not limited to, a composite evaluation system of multi-dimensional convergence determination criteria. The second loss condition needs to comprehensively consider the model performance improvement trend, resource consumption threshold, and task balance index. For example, it is required that the cross-modal retrieval accuracy fluctuation is less than 1% within three consecutive training cycles, the GPU video memory occupancy is stable within a preset range, and the difference in loss values of each subtask does not exceed 15%. In specific implementation, a multi-stage early stopping strategy can be adopted, and when the main task reaches the convergence standard, the secondary task is triggered for independent fine-tuning. For example, in an e-commerce cross-modal recommendation system, the second loss condition may be set to determine that the model training is completed when the F1 value of the image-text retrieval task exceeds 0.92, the video-to-product 3D model retrieval response time is lower than 200 ms, and the click-through rate of the multi-modal recommendation list increases by 5% month-on-month.

[0206] It should be noted that the dynamic adjustment mechanism of the second joint loss function supports multi-dimensional innovations, including but not limited to that in the weight generation dimension, automatic weighting based on task uncertainty, policy optimization based on reinforcement learning, or association modeling based on graph neural networks can be adopted; in the update frequency dimension, real-time dynamic adjustment (updating batch by batch), stage adjustment (updating every epoch), or event-driven adjustment (triggered by specific conditions) is supported; in the constraint condition dimension, resource efficiency constraints (limiting computational complexity), business rule constraints (compliance with industry standards), and ethical and safety constraints (avoiding bias amplification) can be introduced. In the financial risk control scenario, the second joint loss function may dynamically balance the weights of transaction text analysis loss, customer voice emotion recognition loss, and behavior sequence prediction loss, and at the same time introduce an anti-fraud rule constraint term to ensure the compliance and interpretability of model decisions.

[0207] Through the embodiments of the present application, by using the adaptive weight allocation and second-order optimization strategy, the adaptability of the model to long-tail tasks and data sparse scenarios can be improved, and the overall robustness and scalability of the cross-modal retrieval system can be enhanced.

[0208] The following further explains the present application in combination with specific examples:

[0209] The present application proposes a unified multi-modal information retrieval method based on multi-task training and instruction tuning, aiming to solve the problems of task specificity and poor generalization ability existing in existing multi-modal information retrieval systems. Traditional multi-modal information retrieval methods usually focus on the training of specific tasks and fixed modalities, lacking flexibility and being unable to effectively meet the diverse information needs of users in actual applications.

[0210] The core technical key point of this application lies in constructing a general multimodal retrieval model through multi-task training and instruction tuning, which can perform adaptive multimodal information retrieval according to the user's natural language instructions. Specifically, the model can process various queries containing text, images, and their combinations, and accurately retrieve relevant information from heterogeneous candidate pools.

[0211] Multi-task training: By training multiple different retrieval tasks simultaneously, the model can share knowledge in different tasks and perform effective transfer learning. This multi-task training method enables the model to have good generalization ability, can handle a wide range of multimodal information retrieval tasks, and avoids the overfitting problem of single-task models.

[0212] Instruction tuning: During the training process, instruction tuning is used to enhance the model's understanding of the user's query intent. Users describe their retrieval needs through simple natural language instructions, and the model performs specific retrieval tasks according to these instructions. This instruction tuning method enables the model to better understand complex query intents. Especially in cross-modal retrieval tasks, it can accurately select appropriate retrieval modalities and candidate pools according to the instructions, thereby improving the retrieval effect.

[0213] Through these two key technologies, this application can effectively improve the performance of multimodal information retrieval systems in different tasks and modalities. Especially when facing unknown tasks or data, it can demonstrate excellent zero-shot / few-shot generalization ability. Therefore, this application not only has strong technological innovation but also has a wide range of application prospects, can meet diverse information retrieval needs, especially in the fields of image, text, and multimodal information retrieval combining images and text.

[0214] With the rapid development of information technology, especially the continuous progress of computer vision and natural language processing technologies, multimodal information retrieval has gradually become an important research direction in the field of information retrieval. The tasks of multimodal information retrieval usually include retrieving images through text queries, retrieving text through image queries, or retrieving various types of information through image-text combination queries. These tasks cover rich information retrieval needs.

[0215] Currently, most multi-modal retrieval models are designed for specific tasks or fixed modalities and are usually trained on a single data source. Such systems have poor generalization ability and cannot effectively handle the complex and variable retrieval requirements in practical applications. For example, existing image-text retrieval systems are generally only optimized for specific datasets and often perform poorly on data from different domains or of different types. Traditional multi-modal retrieval systems usually rely on fixed input formats (such as images, texts, or image-text pairs) and cannot handle complex and flexible user requirements. When users conduct retrievals, they often use natural language instructions to express their retrieval intentions, but existing systems have insufficient ability to understand and respond to such natural language instructions. This results in the system being unable to accurately understand user requirements when faced with complex retrieval tasks, thus affecting the accuracy and efficiency of the retrieval.

[0216] In summary, existing multi-modal information retrieval technologies have problems such as poor task specificity and lack of flexible instruction understanding ability. Therefore, there is a need for a method for efficient and flexible multi-modal information retrieval that can improve the generalization ability and execution efficiency of the model between different tasks and modalities through multi-task training and instruction tuning, so as to better meet the diverse information retrieval needs of users.

[0217] As Figure 10 shown, this application can handle retrieval tasks between different modalities, including text-to-image retrieval, image-to-text retrieval, and multi-modal retrieval of image-text combinations. Whether given a text query to find relevant images or an image query to obtain relevant text descriptions, the system can accurately perform the retrieval task. In addition, the model can flexibly handle image-text combined queries and provide accurate multi-modal information retrieval results.

[0218] Through the instruction tuning function, this application can flexibly adjust the retrieval task according to the natural language instructions provided by the user. Users only need to use simple language instructions, such as "find the description related to this picture" or "query similar images according to the following text", to efficiently trigger the corresponding retrieval task. The system can parse the semantics in the instructions and select the appropriate retrieval modality (image or text) to perform accurate retrieval according to the user's needs. This flexible instruction following ability significantly improves the user experience, reduces cumbersome query input, and broadens the application scenarios of the retrieval system. Through multi-task training, the system can share knowledge between multiple retrieval tasks, enabling the model to have good generalization ability when faced with unknown tasks. Even on unseen retrieval tasks or datasets, the system can still accurately perform new retrieval tasks through existing multi-task learning experience.

[0219] The main technical solutions of this application include the following key steps:

[0220] Multi-task training: By jointly training using multiple different retrieval tasks within a unified framework, the model can share knowledge and improve its generalization ability across tasks.

[0221] Instruction tuning: By fine-tuning the model, it can understand and execute multi-modal information retrieval tasks based on natural language instructions. Users input instructions to clarify their retrieval intentions, and the model performs information retrieval according to these instructions.

[0222] Cross-modal information fusion: Design effective fusion mechanisms (such as score-level fusion and feature-level fusion) to handle the associations between different modalities (such as images and text), and fully utilize this information in the final retrieval.

[0223] Among them, the technical objective of multi-task training: The core objective of multi-task training is to enable the model to handle multiple different types of retrieval tasks within a unified framework. Each task can involve different data modalities (such as text, images, etc.), and the generalization ability of the model is enhanced by sharing parameters. Through joint training, the model can share the knowledge learned in multiple tasks, which not only improves the overall performance of the model but also enhances the model's adaptability when facing new tasks.

[0224] Task definition: Multiple retrieval tasks are defined, and the goal of each task is to retrieve relevant information from a heterogeneous candidate pool through a given query (text, image, or combination of text and image). Specific tasks include:

[0225] Task 1: Text-to-Image Retrieval: Given a text query, retrieve the relevant image.

[0226] Task 2: Image-to-Text Retrieval: Given an image query, retrieve the relevant text description.

[0227] Task 3: Image-Text Retrieval: Given a pair of an image and text, retrieve the relevant text or image.

[0228] Shared model structure: The key to multi-task training is to design a shared model structure so that the model can learn general representations from all tasks. In this framework, all retrieval tasks share the same neural network architecture, but the loss function for each task will be different to adapt to the specific task objectives.

[0229] The model structure adopts a deep neural network. For example, a convolutional neural network (CNN) is used for image feature extraction, and a Transformer is used for text feature extraction. Each task encodes the input query and generates a feature vector of the query.

[0230] Joint loss function: In multi-task training, the goal of the model is to minimize the sum of the loss functions of all tasks. Suppose there are T tasks, and the loss function of each task is Lt. Then the joint loss function can be expressed as

[0231] where αt is the weight of the loss of each task. In this way, the model not only needs to perform well on each task but also needs to make an effective trade-off between all tasks.

[0232] Inter-task influence: In multi-task training, there may be an inter-task influence between different tasks. To avoid the optimization of one task affecting the performance of other tasks, a method of task importance weighting is adopted, that is, by setting different weights to control the influence of each task in the loss function.

[0233] The weight αt of the task can be adjusted according to the performance on the validation set or learned through an adaptive method. The inter-task influence actually reflects the shared knowledge between different tasks. If the loss of a certain task is high, it may affect the optimization of other tasks, but this way of sharing knowledge will enhance the generalization ability of the model.

[0234] Training strategy: During the training process, the training data of tasks are shuffled and mixed together. The model processes samples of multiple tasks simultaneously in each mini-batch. In this way, the model can learn knowledge from different tasks at the same time, thereby improving its processing ability for various tasks. As Figure 11 shown, the key steps in the training process are as follows:

[0235] S1, Input: Randomly select the data of one task from the training set each time.

[0236] S2, Forward propagation: Input the query data into the shared model structure to obtain the prediction results of the tasks.

[0237] S3, Calculate loss: Calculate the loss of each task according to the prediction results and the true labels.

[0238] S4, Backward propagation: Backpropagate the loss of each task and update the shared parameters.

[0239] S5, Task update: Update the parameters on the loss function of each task to ensure that the optimizations of different tasks can be coordinated with each other.

[0240] Formula derivation: Taking "text-to-image retrieval" as an example, assume that the input received by the model is a text query qt and an image candidate ci. The model will calculate the similarity between the text and the image and optimize the objective through contrastive learning. The similarity calculation formula is:

[0241] where fT(qt) and fI(ci) represent the feature vectors of the text and the image respectively. In this way, the model can learn how to retrieve similar images according to the text description.

[0242] For the "image-to-text retrieval" task, similarly, the model will optimize the objective by calculating the similarity between the image and the text:

[0243] Through multi-task training, the model can learn these two tasks simultaneously and, guided by the joint loss function, optimize its retrieval performance. Multi-task training enables the model to share the learned knowledge from multiple tasks by sharing the model structure and parameters. This knowledge sharing mechanism enhances the generalization ability of the model. Especially when facing new tasks or datasets, it can utilize the existing knowledge for more efficient learning. By jointly training on multiple tasks, the model can better adapt to various different input types and retrieval tasks. Even if the data for a certain task is scarce, the model can still obtain useful information from other tasks, thereby improving the robustness of the model. Since the model needs to be trained on multiple tasks, it can effectively avoid overfitting to a specific task. The diversity between tasks helps improve the performance of the model in different scenarios. Multi-task training improves the zero-shot learning ability of the model. Even when it has not seen certain specific tasks, the model can still perform reasoning and retrieval through the shared task knowledge.

[0244] Through the above detailed technical solutions and processes, the model can be jointly trained on multiple tasks and optimized according to the characteristics of different tasks, thereby enhancing the retrieval ability and generalization ability of the model.

[0245] Technical objective of instruction tuning: The objective of Instruction Tuning is to enable the model to understand and execute the user's natural language instructions to achieve more flexible and accurate information retrieval. Traditional multi-modal information retrieval systems can usually only process queries in a specific format, such as image or text queries, and lack the ability to understand natural language instructions. Through instruction tuning, this application can closely integrate natural language instructions with the retrieval task, allowing users to describe their retrieval intentions through simple and flexible text instructions, and the model generates accurate retrieval results according to the instructions.

[0246] Instruction tuning mainly involves the following technical details:

[0247] Instruction Parsing and Encoding: Convert the natural language instructions provided by the user into a vector representation that can be understood by the machine through natural language processing (NLP) technology.

[0248] Dynamic Adjustment of Retrieval Tasks: Dynamically adjust the retrieval tasks executed by the model according to the input instructions, so as to select the appropriate retrieval modality (image, text, or a combination of both) according to the instruction content.

[0249] Task Execution: According to the instructions and query input, the model selects the corresponding retrieval task (e.g., text-to-image retrieval, image-to-text retrieval, etc.) and generates retrieval results.

[0250] As Figure 12 shown, the workflow of instruction tuning can be divided into the following steps:

[0251] Instruction Input: The user provides natural language instructions, such as "Find similar images according to the following description" or "Please find the description that matches this picture". These instructions express the user's retrieval intent.

[0252] Instruction Encoding: The model converts the natural language instructions input by the user into a vector representation through a pre-trained language model (such as BERT, GPT, etc.). This vector representation will contain the semantic information of the instructions, enabling the model to understand the user's needs.

[0253] Assume the instruction input by the user is qinst, then the encoded representation of this instruction is:

[0254] where Encoder(·) represents the process of encoding the instruction text into a vector. This encoded vector will be input into the model together with the modal features of the query (such as image or text).

[0255] Retrieval Task Adjustment: According to the encoded instructions, the model dynamically adjusts the type of retrieval task. The information in the instructions usually includes the type of retrieval modality (such as "image" and "text") and the retrieval target (such as "find relevant descriptions" or "find similar images").

[0256] For example, assume the instruction is "Find relevant images according to the following description", the model will match it with the text-to-image retrieval task; if the instruction is "Find the description similar to this image", the model will select the image-to-text retrieval task.

[0257] Task Execution and Retrieval: According to the selection of the task type, the model retrieves relevant information from the corresponding candidate pool according to the input query (text, image, or text-image combination). The model calculates the similarity between the query and the candidate information and selects the most matching result.

[0258] For example, the similarity calculation formula for text-to-image retrieval is:

[0259] Among them, fT(qt) and fI(ci) represent the feature representations of the text query and the image candidate respectively.

[0260] Output retrieval results: Based on the calculated similarity, the model returns the retrieval results most relevant to the query, which may be outputs of images, texts, or combinations of images and texts.

[0261] Loss function for instruction tuning: To enable the model to perform retrieval tasks according to user instructions, instruction tuning needs to be optimized by combining task-specific loss functions. Suppose there is a task set T, and the loss function for each task is Lt. Then the joint loss function for instruction tuning can be expressed as:

[0262] Among them, αt is the weighted coefficient of the task, which can be adjusted according to the difficulty or importance of the task. Lt is the loss function of task t, usually the cross-entropy loss or the contrastive learning loss, depending on the nature of the task.

[0263] For example, if the current task is text-to-image retrieval, the loss function is the contrastive learning loss:

[0264] Among them, C is the candidate pool, and score(qt,ci) is the similarity between the query qt and the candidate ci.

[0265] Enhance the flexibility of the model: Instruction tuning enables the model to understand natural language instructions and perform corresponding retrieval tasks. This flexibility allows users to express complex retrieval requirements through simple language instructions, thus greatly improving the user experience.

[0266] Improve the accuracy of task execution: Through instruction tuning, the model can more accurately identify the user's retrieval intention, avoiding the limitations brought by fixed tasks and input formats in traditional methods. The model can dynamically adjust the type of retrieval task according to the user's instructions and perform the most appropriate retrieval method, thereby improving the relevance and accuracy of the retrieval results.

[0267] Simplify user input: Compared with the traditional complex query input form, users only need to provide simple natural language instructions, which reduces the usage threshold and improves the usability of the system. This enables the multimodal retrieval technology to be more widely applied to ordinary users, rather than being limited to professionals only.

[0268] Improve zero-shot ability: By training the model to understand natural language instructions during the instruction tuning stage, it can infer the retrieval target based on the user's instructions without having explicitly seen certain tasks, which significantly improves the zero-shot learning ability of the model.

[0269] Instruction tuning is an important innovation in this application. By enabling the model to understand and execute users' natural language instructions, it enhances the flexibility and precision of the multimodal information retrieval system. Compared with traditional methods, instruction tuning not only simplifies the user input form but also improves the accuracy of task execution, enabling the system to perform different retrieval tasks according to different instructions, thus meeting the diverse needs of users.

[0270] The solution of this application is evaluated on multiple standard test sets through multi-task training and instruction tuning mechanisms. The results show significant gains in different types of retrieval tasks. The following are several typical test sets and their specific gains:

[0271] Flickr30k (image-text retrieval). On the Flickr30k dataset, the method of this application can better understand the relationship between images and texts through multimodal task training and instruction tuning. In the image-to-text retrieval task, the precision is improved by about 1.8%, and the recall rate is increased by 1.5%. In addition, through instruction tuning, the system can provide more relevant results in multimodal scenarios according to users' natural language instructions.

[0272] COCO (image-text retrieval and image-text combination retrieval). In the tests on the COCO dataset, especially in the image-to-text retrieval and image-text combination retrieval tasks, the system performs excellently. In the image-to-text retrieval, the accuracy is increased by 2%, and in the image-text combination retrieval, the F1 value of the system is increased by 1.2%. The introduction of instruction tuning enables the model to accurately execute users' instructions, thus further improving the retrieval precision.

[0273] Fashion-Gen (image-to-text retrieval). On the Fashion-Gen dataset, especially in the image-to-text retrieval task, the accuracy of the model is increased by 2.1%. Through instruction tuning, the system can better understand the user's query intention when dealing with complex descriptions of fashion images, improving the retrieval effect in specific fields.

[0274] Through the tests on these standard datasets, the solution of this application performs significantly better than traditional methods in multimodal retrieval tasks, and the accuracy, precision, and recall rate are significantly improved.

[0275] It can be understood that in the specific implementation of this application, data related to user information, etc. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0276] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0277] According to another aspect of the embodiments of the present application, there is also provided a media information processing device for implementing the above-mentioned media information processing method. As Figure 13 shown, the device includes:

[0278] A first acquisition module 1302, configured to acquire source media information to be retrieved, where the source media information represents the input information that needs to be retrieved currently;

[0279] A second acquisition module 1304, configured to acquire prompt template information, where the prompt template information is used to indicate the type of retrieval task that needs to be retrieved currently;

[0280] A processing module 1306, configured to input the source media information and the prompt template information into a target retrieval model to determine target media information, where the target retrieval model represents a model obtained by jointly training according to at least two types of retrieval tasks, the loss function used in the training process of the target retrieval model is a joint loss function, the joint loss function is determined by at least two sub-loss functions corresponding one-to-one to at least two types of retrieval tasks, and the target media information represents the media information retrieved from the source media information according to the prompt template information.

[0281] As an optional solution, the above device is further configured to: before inputting the source media information and the prompt template information into the target retrieval model to determine the target media information, acquire a preset set of retrieval task types and sample media information, where the set of retrieval task types includes different cross-modal retrieval task types, and the sample media information includes media information of different modalities; construct an initial retrieval model with shared parameters based on the retrieval task types, where the initial retrieval model includes a first encoding branch and a second encoding branch, and the first encoding branch and the second encoding branch are used to map the sample media information of different modalities into the same embedding space; train the initial retrieval model based on the sample media information to obtain the target retrieval model.

[0282] As an alternative solution, the above-mentioned device is used to train an initial retrieval model based on sample media information in the following manner to obtain a target retrieval model: In each batch, the initial retrieval model is trained based on sample media information in the following manner to obtain a target retrieval model, where the retrieval task type corresponding to each batch is the first retrieval task type: Determine first sample media information associated with the first retrieval task type from the sample media information; Input the first sample media information into the initial retrieval model to obtain first sample predicted media information; Determine a first target sub-loss function corresponding to the first retrieval task type according to the first sample predicted media information; Perform backpropagation on the initial retrieval model based on the first target sub-loss function to update the model parameters of the initial retrieval model; Repeat the above steps until the target retrieval model is generated.

[0283] As an alternative solution, the above-mentioned device is used to determine a first target sub-loss function corresponding to the first retrieval task type according to the first sample predicted media information in the following manner: Obtain a target weight coefficient determined for the first retrieval task type; Determine an initial sub-loss function corresponding to the first retrieval task type according to the first sample predicted media information; Determine the first target sub-loss function based on the initial sub-loss function and the target weight coefficient.

[0284] As an alternative solution, the above-mentioned device is used to train an initial retrieval model based on sample media information in the following manner to obtain a target retrieval model: In each batch, the initial retrieval model is trained based on sample media information in the following manner to obtain a target retrieval model, where the sample media information input in each batch includes media information of different modalities: Determine second sample media information from the sample media information and determine a second retrieval task type corresponding to the second sample media information; Input the second sample media information into the initial retrieval model to obtain second sample predicted media information; Determine a second target sub-loss function corresponding to the second retrieval task type according to the second sample predicted media information; Perform backpropagation on the initial retrieval model based on the second target sub-loss function to update the model parameters of the initial retrieval model; Repeat the above steps until the target retrieval model is generated.

[0285] As an alternative, the above device is used to train an initial retrieval model based on sample media information in the following manner to obtain a target retrieval model: input a first set of sample media information into the initial retrieval model to obtain a first sub-loss function, and input a second set of sample media information into the initial retrieval model to obtain a second sub-loss function, where the first sub-loss function and the second sub-loss function correspond to different retrieval task types; determine a first joint loss function based on the first sub-loss function and the second sub-loss function, and train the initial retrieval model based on the first joint loss function to obtain the target retrieval model, where the first joint loss function represents a function that assigns different weights to different retrieval task types and is determined based on whether the sample media information of different modalities is similar. When the first joint loss function satisfies the first loss condition, the initial retrieval model is determined to be the target retrieval model, and the joint loss function includes the first joint loss function.

[0286] As an alternative, the above device is used to input source media information and prompt template information into the target retrieval model in the following manner to determine target media information: determine a candidate media information set based on the prompt template information, where the candidate media information set includes candidate media information specified by the prompt template information; determine the candidate media information in the candidate media information set that satisfies a preset similarity condition with the source media information as the target media information.

[0287] As an alternative, the above device is used to determine a candidate media information set based on the prompt template information in the following manner: perform feature extraction and splicing on the source media information and the prompt template information to obtain a source feature vector; determine the current retrieval task type based on the source feature vector; determine the candidate media information set according to the current retrieval task type.

[0288] As an alternative, the above device is used to determine the candidate media information in the candidate media information set that satisfies a preset similarity condition with the source media information as the target media information in the following manner: obtain the candidate feature vectors corresponding to the candidate media information in the candidate media information set; determine a target feature vector based on the similarity between the source feature vector and the candidate feature vectors; determine the candidate media information corresponding to the target feature vector as the target media information.

[0289] As an alternative solution, the above-mentioned device is further configured to: obtain a preset set of retrieval task types and sample media information, where the set of retrieval task types includes different cross-modal retrieval task types, and the sample media information includes media information of different modalities; construct an initial retrieval model with shared parameters based on the retrieval task types, where the initial retrieval model includes a first encoding branch, a second encoding branch, and a third encoding branch, the first encoding branch and the second encoding branch are used to map sample media information of different modalities to the same embedding space, and the third encoding branch is used to map prompt template information to the same embedding space; train the initial retrieval model based on the sample media information to obtain a target retrieval model.

[0290] As an alternative solution, the above-mentioned device is configured to train the initial retrieval model based on the sample media information in the following manner to obtain a target retrieval model: input the third set of sample media information into the initial retrieval model to obtain a third sub-loss function, and input the fourth set of sample media information into the initial retrieval model to obtain a fourth sub-loss function, where the third sub-loss function and the fourth sub-loss function correspond to different retrieval task types; determine a second joint loss function based on the third sub-loss function and the fourth sub-loss function, and train the initial retrieval model based on the second joint loss function to obtain a target retrieval model, where the second joint loss function is a function that assigns different weights to different retrieval task types and is determined based on whether the sample media information is similar to the set of sample candidate media information. When the second joint loss function satisfies the second loss condition, the initial retrieval model is determined as the target retrieval model, and the joint loss function includes the second joint loss function.

[0291] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.

[0292] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0293] According to one aspect of the present application, a computer program product is provided, and the computer program product includes a computer program.

[0294] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0295] Figure 14 A block diagram of a computer system for implementing an electronic device according to an embodiment of the present application is schematically shown.

[0296] It should be noted that Figure 14 The computer system 1400 of the shown electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0297] As Figure 14 shown, the computer system 1400 includes a central processing unit 1401 (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1402 (ROM) or a program loaded from a storage section 1408 into a random access memory 1403 (RAM). In the random access memory 1403, various programs and data required for system operation are also stored. The central processing unit 1401, the read-only memory 1402, and the random access memory 1403 are connected to each other via a bus 1404. An input / output interface 1405 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1404.

[0298] The following components are connected to the input / output interface 1405: an input section 1406 including a keyboard, a mouse, etc.; an output section 1407 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1408 including a hard disk, etc.; and a communication section 1409 including a network interface card such as a local area network card, a modem, etc. The communication section 1409 performs communication processing via a network such as the Internet. A drive 1410 is also connected to the input / output interface 1405 as needed. A removable medium 1411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1410 as needed so that a computer program read from it can be installed into the storage section 1408 as needed.

[0299] In particular, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1409, and / or installed from the removable medium 1411. When the computer program is executed by the central processing unit 1401, various functions defined in the system of the present application are executed.

[0300] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1409, and / or installed from the removable medium 1411. When the computer program is executed by the central processing unit 1401, various functions provided by the embodiments of the present application are executed.

[0301] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned method for processing media information is further provided. The electronic device may be Figure 1 the terminal device or server shown. This embodiment takes the electronic device as the terminal device as an example for illustration. As Figure 15 shown, the electronic device includes a memory 1502 and a processor 1504. A computer program is stored in the memory 1502, and the processor 1504 is configured to execute the steps in any one of the above method embodiments through the computer program.

[0302] Optionally, in this embodiment, the above-mentioned electronic device may be at least one network device among multiple network devices in a computer network.

[0303] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the methods in the embodiments of the present application through a computer program.

[0304] Optionally, those of ordinary skill in the art can understand that Figure 15 the structure shown is only schematic, Figure 15 and it does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 15 , or have a different configuration from that shown in Figure 15 .

[0305] Among them, the memory 1502 can be used to store software programs and modules, such as the program instructions / modules corresponding to the media information processing method and apparatus in the embodiments of the present application. The processor 1504 executes various functional applications and data processing by running the software programs and modules stored in the memory 1502, that is, implements the above-mentioned media information processing method. The memory 1502 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1502 may further include a memory remotely disposed relative to the processor 1504, and these remote memories may be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1502 may specifically but not limitedly be used to store information such as text and pictures. As an example, as Figure 15 shown, the above-mentioned memory 1502 may include, but not limited to, the first acquisition module 1302, the second acquisition module 1304, and the processing module 1306 in the above-mentioned media information processing apparatus. In addition, it may also include, but not limited to, other module units in the above-mentioned media information processing apparatus, which will not be elaborated in this example.

[0306] Optionally, the above-mentioned transmission device 1506 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one instance, the transmission device 1506 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one instance, the transmission device 1506 is a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.

[0307] In addition, the above-mentioned electronic device further includes: a display 1508, which is used to display the above-mentioned target media information; and a connection bus 1510, which is used to connect each module component in the above-mentioned electronic device.

[0308] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes can form a peer-to-peer network, and any form of computing device, such as a server, a terminal, and other electronic devices, can become a node in the blockchain system by joining the peer-to-peer network.

[0309] According to one aspect of the present application, a computer-readable storage medium is provided. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the media information processing methods provided in various alternative implementations of the above media information processing aspect.

[0310] Optionally, in this embodiment, the above computer-readable storage medium may be configured to store instructions for executing the methods in various embodiments of the present application.

[0311] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. The program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0312] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0313] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing one or more electronic devices to execute all or part of the steps of the methods described in various embodiments of the present application.

[0314] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0315] In several embodiments provided by the present application, it should be understood that the disclosed application program can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of units or modules may be in an electrical or other form.

[0316] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0317] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0318] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for processing media information, characterized in that Including: Obtain source media information to be retrieved, where the source media information represents the input information that needs to be retrieved currently; Obtain prompt template information, where the prompt template information is used to indicate the type of retrieval task that needs to be retrieved currently; Input the source media information and the prompt template information into a target retrieval model to determine target media information, where the target retrieval model represents a model obtained by jointly training according to at least two types of retrieval tasks, the loss function used in the training process of the target retrieval model is a joint loss function, the joint loss function is determined by at least two sub-loss functions corresponding one by one to the at least two types of retrieval tasks, and the target media information represents the media information obtained by retrieving the source media information according to the prompt template information.

2. The method according to claim 1, wherein Before inputting the source media information and the prompt template information into the target retrieval model to determine the target media information, the method further includes: Obtain a preset set of retrieval task types and sample media information, where the set of retrieval task types includes different cross-modal retrieval task types, and the sample media information includes media information of different modalities; Construct an initial retrieval model with shared parameters based on the retrieval task types, where the initial retrieval model includes a first encoding branch and a second encoding branch, and the first encoding branch and the second encoding branch are used to map the sample media information of different modalities to the same embedding space; Train the initial retrieval model based on the sample media information to obtain the target retrieval model.

3. The method according to claim 2, wherein The training the initial retrieval model based on the sample media information to obtain the target retrieval model includes: Each batch trains the initial retrieval model based on the sample media information in the following manner to obtain the target retrieval model, where the retrieval task type corresponding to each batch is the first retrieval task type: Determine first sample media information associated with the first retrieval task type from the sample media information; Input the first sample media information into the initial retrieval model to obtain first sample predicted media information; Determine a first target sub-loss function corresponding to the first retrieval task type according to the first sample predicted media information; Perform backpropagation on the initial retrieval model based on the first target sub-loss function to update the model parameters of the initial retrieval model; Repeat the above steps until the target retrieval model is generated.

4. The method according to claim 3, wherein The determining the first target sub-loss function corresponding to the first retrieval task type according to the first sample predicted media information includes: Obtain a target weight coefficient determined for the first retrieval task type; Determine an initial sub-loss function corresponding to the first retrieval task type according to the first sample predicted media information; Determine the first target sub-loss function based on the initial sub-loss function and the target weight coefficient.

5. The method according to claim 2, wherein The training the initial retrieval model based on the sample media information to obtain the target retrieval model includes: Each batch trains the initial retrieval model based on the sample media information in the following manner to obtain the target retrieval model, where the sample media information input in each batch includes media information of different modalities: Determine the second sample media information from the sample media information, and determine the second retrieval task type corresponding to the second sample media information; Input the second sample media information into the initial retrieval model to obtain the second sample predicted media information; Determine the second target sub-loss function corresponding to the second retrieval task type according to the second sample predicted media information; Perform backpropagation on the initial retrieval model based on the second target sub-loss function to update the model parameters of the initial retrieval model; Repeat the above steps until the target retrieval model is generated.

6. The method according to claim 2, characterized in that, The training of the initial retrieval model based on the sample media information to obtain the target retrieval model includes: Input the first set of sample media information into the initial retrieval model to obtain the first sub-loss function, and input the second set of sample media information into the initial retrieval model to obtain the second sub-loss function, where the first sub-loss function and the second sub-loss function correspond to different retrieval task types; Determine the first joint loss function based on the first sub-loss function and the second sub-loss function, and train the initial retrieval model based on the first joint loss function to obtain the target retrieval model, where the first joint loss function is a function that assigns different weights to different retrieval task types and is determined based on whether the sample media information of different modalities is similar. When the first joint loss function satisfies the first loss condition, the initial retrieval model is determined as the target retrieval model, and the joint loss function includes the first joint loss function.

7. The method according to claim 1, wherein The input of the source media information and the prompt template information into the target retrieval model to determine the target media information includes: Determine the candidate media information set based on the prompt template information, where the candidate media information set includes the candidate media information specified by the prompt template information; Determine the candidate media information that satisfies the preset similarity condition with the source media information in the candidate media information set as the target media information.

8. The method according to claim 7, characterized in that, The determination of the candidate media information set based on the prompt template information includes: Extract and splice the features of the source media information and the prompt template information to obtain the source feature vector; Determine the current retrieval task type based on the source feature vector; Determine the candidate media information set according to the current retrieval task type.

9. The method according to claim 8, characterized in that, The determination of the candidate media information that satisfies the preset similarity condition with the source media information in the candidate media information set as the target media information includes: Obtain the candidate feature vectors corresponding to the candidate media information in the candidate media information set; Determine the target feature vector based on the similarity between the source feature vector and the candidate feature vectors; Determine the candidate media information corresponding to the target feature vector as the target media information.

10. The method according to claim 1, wherein The method further includes: Obtain a preset set of retrieval task types and sample media information, where the set of retrieval task types includes different cross-modal retrieval task types, and the sample media information includes media information of different modalities; Construct an initial retrieval model with shared parameters based on the retrieval task types, where the initial retrieval model includes a first encoding branch, a second encoding branch, and a third encoding branch. The first encoding branch and the second encoding branch are used to map the sample media information of different modalities to the same embedding space, and the third encoding branch is used to map the prompt template information to the same embedding space; Train the initial retrieval model based on the sample media information to obtain the target retrieval model.

11. The method according to claim 10, characterized in that, The training the initial retrieval model based on the sample media information to obtain the target retrieval model includes: Input the third set of sample media information into the initial retrieval model to obtain a third sub-loss function, and input the fourth set of sample media information into the initial retrieval model to obtain a fourth sub-loss function, where the third sub-loss function and the fourth sub-loss function correspond to different retrieval task types; Determine a second joint loss function based on the third sub-loss function and the fourth sub-loss function, and train the initial retrieval model based on the second joint loss function to obtain the target retrieval model, where the second joint loss function is a function determined by assigning different weights to different retrieval task types and based on whether the sample media information is similar to the sample candidate media information set. When the second joint loss function satisfies the second loss condition, the initial retrieval model is determined to be the target retrieval model, and the joint loss function includes the second joint loss function.

12. A processing device for media information, characterized in that, Including: A first acquisition module for acquiring source media information to be retrieved, where the source media information represents the input information that needs to be retrieved currently; A second acquisition module for acquiring prompt template information, where the prompt template information is used to indicate the retrieval task type that needs to be retrieved currently; A processing module for inputting the source media information and the prompt template information into the target retrieval model to determine target media information, where the target retrieval model represents a model jointly trained according to at least two retrieval task types, and the loss function used in the training process of the target retrieval model is a joint loss function, and the joint loss function is determined by at least two sub-loss functions corresponding one-to-one to the at least two retrieval task types, and the target media information represents the media information retrieved from the source media information according to the prompt template information.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where the program, when run by a processor, executes the method described in any one of claims 1 to 11.

14. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 11.

15. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 11 through the computer program.

Citation Information

Cited By

  • Multi-modal data pairing method and system based on deep learning

    CN120994874A

  • Multi-modal learning method and device under different sensing sources, equipment and medium

    CN121280735A

  • Network security threat identification method and system based on network security knowledge structure perception

    CN121333806A