Content query method and content query model training method
By introducing encoding, querying, reconstruction, and decision-making units into the content query model, the semantic matching problem of the dual-tower model under modal imbalance and noise interference is solved, enabling the generation and recall of high-precision content query results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, the dual-tower model based on dense vectors has insufficient generalization ability when there is an imbalance between text and images, lack of modality, or noise interference, making it difficult to support high-precision semantic matching and resulting in poor accuracy of content query results.
By introducing a content query model that includes a content encoding unit, a query encoding unit, a reconstruction unit, and a decision unit, and using sample reconstruction data and reconstruction weights for training, content vectors and query vectors are generated. Reconstruction weights are then generated through the decision unit, achieving adaptive training and high-quality modality reconstruction, avoiding invalid reconstruction, and improving semantic matching accuracy.
During the training phase, the system corrects representation biases caused by information compression, modality loss, or noise interference, enhances cross-modal semantic core focus, improves recall and ranking quality, and achieves high-precision content query results.
Smart Images

Figure CN121833931A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of artificial intelligence, and in particular to a content query method and a content query model training method. BACKGROUND
[0002] With the explosive growth of Internet content, efficient content query based on semantics has become the core link of search engines, recommendation systems, and intelligent question answering and other applications. Traditional keyword matching methods are difficult to capture the deep semantic association between user queries and content, prompting researchers to widely use double-tower models based on dense vectors for semantic recall.
[0003] Currently, the double-tower model usually only relies on contrast learning for training, resulting in insufficient generalization ability of the generated content vector when facing imbalanced text and images, missing modalities, or noise interference, making it difficult to support high-precision semantic matching, resulting in poor accuracy of content query results. Therefore, there is an urgent need for a content query scheme with high accuracy. SUMMARY
[0004] Therefore, the embodiments of the present specification provide a content query method. One or more embodiments of the present specification also relate to a content query model training method, a content query device, a content query model training device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects in the prior art.
[0005] According to a first aspect of the embodiments of the present specification, a content query method is provided, including: obtaining target query data and content to be queried; inputting the content to be queried into a content encoding unit in a content query model to obtain a content vector, and inputting the target query data into a query encoding unit in the content query model to obtain a query vector, wherein the content query model includes a content encoding unit, a query encoding unit, a reconstruction unit, and a decision unit, the content encoding unit is trained based on sample reconstruction data of sample content, predicted reconstruction data, and reconstruction weights of the reconstruction unit, the predicted reconstruction data is obtained by processing the sample content vector of the sample content by the reconstruction unit, and the decision unit is used to generate the reconstruction weights; and generating a content query result according to the content vector and the query vector.
[0006] According to a second aspect of the embodiments of the present specification, a content query model training method is provided, including: obtaining sample data, wherein the sample data includes sample content, sample query data, and sample reconstruction data; inputting the sample content into a content encoding unit in a content query model to obtain a sample content vector, and inputting the sample query data into a query encoding unit in the content query model to obtain a sample query vector; using a reconstruction unit in the content query model to generate predicted reconstruction data based on the sample content vector, and using a decision unit in the content query model to generate a reconstruction weight of the reconstruction unit; and performing parameter adjustment on the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weight to obtain a trained content query model.
[0007] According to a third aspect of the embodiments of the present specification, a content query device is provided, including: a first obtaining module configured to obtain target query data and to-be-queried content; a first input module configured to input the to-be-queried content into a content encoding unit in a content query model to obtain a content vector, and to input the target query data into a query encoding unit in the content query model to obtain a query vector, wherein the content query model includes the content encoding unit, the query encoding unit, a reconstruction unit, and a decision unit, the content encoding unit is trained based on sample reconstruction data of sample content, predicted reconstruction data, and a reconstruction weight of the reconstruction unit, the predicted reconstruction data is obtained by processing a sample content vector of the sample content by the reconstruction unit, and the decision unit is used to generate the reconstruction weight; and a first generating module configured to generate a content query result according to the content vector and the query vector.
[0008] According to a fourth aspect of the embodiments of the present specification, a content query model training device is provided, including: a second obtaining module configured to obtain sample data, wherein the sample data includes sample content, sample query data, and sample reconstruction data; a second input module configured to input the sample content into a content encoding unit in a content query model to obtain a sample content vector, and to input the sample query data into a query encoding unit in the content query model to obtain a sample query vector; a second generating module configured to use a reconstruction unit in the content query model to generate predicted reconstruction data based on the sample content vector, and to use a decision unit in the content query model to generate a reconstruction weight of the reconstruction unit; and a first adjustment module configured to perform parameter adjustment on the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weight to obtain a trained content query model.
[0009] According to a fifth aspect of the embodiments of the present specification, a computing device is provided, comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method provided in the first aspect or the second aspect.
[0010] According to a sixth aspect of the embodiments of the present specification, a computer readable storage medium is provided, which stores computer programs / instructions, which, when executed by a processor, implement the steps of the method provided in the first aspect or the second aspect.
[0011] According to a seventh aspect of the embodiments of the present specification, a computer program product is provided, comprising computer programs / instructions, which, when executed by a processor, implement the steps of the method provided in the first aspect or the second aspect.
[0012] One embodiment of the present specification provides a content query method, comprising: obtaining target query data and to-be-queried content; inputting the to-be-queried content into a content encoding unit in a content query model to obtain a content vector, and inputting the target query data into a query encoding unit in the content query model to obtain a query vector, wherein the content query model comprises the content encoding unit, the query encoding unit, a reconstruction unit, and a decision unit, the content encoding unit is trained based on sample reconstruction data of sample content, predicted reconstruction data, and reconstruction weights of the reconstruction unit, the predicted reconstruction data is obtained by processing a sample content vector of the sample content by the reconstruction unit, and the decision unit is used for generating the reconstruction weights; and generating a content query result according to the content vector and the query vector. In the training stage of the content query model, on the one hand, the sample content vector is input into the reconstruction unit as a condition to generate the predicted reconstruction data, and the difference between the predicted reconstruction data and the real sample reconstruction data is used as a supervision signal to drive the content encoding unit to learn deep and structured semantic representations that can support high-quality modal reconstruction, effectively correcting the representation deviation caused by information compression, modal loss, or noise interference; on the other hand, the decision unit is introduced to dynamically generate modal-related reconstruction weights, realizing a sample-level adaptive training target: when the graphic and text information of the sample content is complete and the semantics is rich, the content query model automatically enhances the multi-modal joint modeling strength; when a certain modal is missing, low-quality, or insignificant, the reconstruction loss is significantly weakened or even ignored, avoiding forced reconstruction of low-information modal, so that the finally learned sample query vector focuses on the truly transferable cross-modal semantic core, not only accurate in the training distribution, but also showing significantly enhanced robustness when facing modal unbalanced or sparse samples not seen in the training set. Based on this training paradigm, the content query model can efficiently generate high-precision content query results in the content query reasoning stage using a lightweight double-tower encoding structure (content encoding unit and query encoding unit): since the content vector and the query vector have been supervised and adaptively aligned in a unified semantic space, the two can accurately reflect the cross-modal semantic correlation, thereby significantly improving the recall rate and ranking quality. BRIEF DESCRIPTION OF DRAWINGS
[0013] Figure 1 is a flowchart of a content query method provided by one embodiment of the present specification;
[0014] Figure 2 is a processing process schematic diagram of a content encoding unit provided by one embodiment of the present specification;
[0015] Figure 3 is an architecture diagram of a content query system provided by one embodiment of the present specification;
[0016] Figure 4 is a flowchart of a content query model training method provided by one embodiment of the present specification;
[0017] Figure 5 is a processing procedure schematic diagram of a content query model training method provided by an embodiment of the present specification;
[0018] Figure 6 is a flow timing diagram of a content query method provided by an embodiment of the present specification;
[0019] Figure 7 is a structure schematic diagram of a content query device provided by an embodiment of the present specification;
[0020] Figure 8 is a structure schematic diagram of a content query model training device provided by an embodiment of the present specification;
[0021] Figure 9 is a structure block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION
[0022] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced without the specific details, other than in the examples, and it is understood that the scope of the present specification is not limited to the details below.
[0023] The terminology used in one or more embodiments of the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in one or more embodiments of the present specification and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term "and / or" includes any and all combinations of one or more of the associated listed items.
[0024] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal sequence, but are used only to distinguish one piece of information from another. For example, without departing from the scope of one or more embodiments of the present specification, first can be termed second, and similarly, second can be termed first. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining."
[0025] In addition, it should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present specification are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or rejection.
[0026] In one or more embodiments of the present specification, a large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, tens of billions, hundreds of billions, thousands of billions or even tens of billions of model parameters. The large model can also be called a foundation model. Through large-scale unlabeled corpus pre-training, a pre-trained model with hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability. For example, large language model (LLM, Large Language Model), multi-modal pre-training model, etc.
[0027] In actual application, the large model only needs a small amount of sample to fine-tune the pre-trained model and can be applied to different tasks. The large model can be widely used in natural language processing (NLP, Natural Language Processing) and computer vision fields. Specifically, it can be applied to computer vision field tasks such as visual question answering (VQA, Visual Question Answering), image caption (IC, Image Caption), image generation, and natural language processing field tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the large model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0028] First, the terms involved in one or more embodiments of the present specification are explained.
[0029] Dual-Tower Model: A deep learning architecture widely used in recommendation systems, information retrieval, semantic search, etc. The core idea is to encode two objects that need to be matched (such as users and items, queries and content) into vectors through two independent neural networks (towers), and then match or recall in a shared semantic space through similarity calculation.
[0030] InfoNCE (Information Noise Contrastive Estimation): Only requires the positive sample to be more similar than the negative samples in the current batch, and does not require global discriminability. If the model compresses all outputs to a point, as long as the positive sample is slightly "closer", it can deceive the loss function.
[0031] Note: refers to the text, image or other forms of content created by individuals or teams for the purpose of recording information, organizing ideas, reminding, etc. in learning, work or daily life. Notes can be simple lines of text, or complex electronic content containing charts, links, multimedia files and other elements. Notes are widely used in various life and work scenarios, such as knowledge point notes and teaching plan notes in learning and education scenarios, diet record notes, exercise record notes, shopping sharing notes in health and life scenarios, etc.
[0032] Deep self-attention (Transformer): a network structure based on multi-head self-attention mechanism module, mainly used for processing sequence data. The Transformer model includes a stack of repeated encoding units (Encoder) and decoding units (Decoder). This design allows the Transformer to efficiently learn long-term dependencies, making it suitable for a variety of natural language processing tasks, including machine translation, text summarization, question answering systems, etc.
[0033] BERT (Bidirectional Encoder Representations from Transformers) model: an NLP pre-training model. The model learns a large amount of unlabeled text data, capturing deep semantic information in the text, and achieving significant performance improvement in a variety of NLP tasks.
[0034] Sentence embedding model (Sentence Transformer): a deep learning model based on the Transformer architecture, designed to encode entire sentences or text passages into fixed-dimensional vector representations (i.e. sentence vectors). These vectors capture the semantic information of the sentence, making semantically similar sentences closer in the vector space, thereby supporting downstream tasks such as semantic search, text clustering, sentence similarity calculation, etc.
[0035] Diffusion Model: A class of generative models that generate high-quality data (such as images, audio, etc.) by simulating a forward process of gradually adding noise and a backward process of learning denoising.
[0036] Latent Diffusion Model (LDM): An efficient and high-quality image generation diffusion model, which is characterized by conducting the diffusion process in a compressed "latent space" rather than in the pixel space.
[0037] Mixture of Experts (MoE): A sparse activation neural network architecture design paradigm, which divides the model into multiple "expert" sub-networks and uses a learnable "router" to dynamically select the most relevant few experts for computation while keeping the rest inactive, thereby significantly improving model capacity and expressiveness without significantly increasing computational cost.
[0038] Multimodal Large Language Model (MLLM): A large-scale language model that can handle and understand multiple modal inputs (such as text, images, audio, video, etc.). This type of model performs well in visual question answering, image-text generation, cross-modal retrieval, and other tasks.
[0039] Central Processing Unit (CPU): The core general-purpose processor of a computer system, responsible for executing the instructions of the operating system and application programs, and controlling the operation of the entire system. It is designed to efficiently handle various types of serial tasks and complex logic control.
[0040] Graphics Processing Unit (GPU): Originally designed to accelerate graphics rendering, it has now become a widely used coprocessor for parallel computing. Its architecture contains a large number of lightweight computing cores, suitable for simultaneously processing large-scale similar computing tasks.
[0041] In a query system (such as a query engine, a recommendation system, or a question and answer system), in order to efficiently find relevant results from massive content (Documents) to a user query (Query), a "double tower model" architecture is usually used to vectorize the Query and the Documents respectively, and the vector similarity is used for fast retrieval. This process is an indispensable key link in online services (i.e., systems that respond in real time after a user initiates a request). In the double tower model, the Query and the Documents are independently encoded, and there is no cross-attention or deep interaction mechanism between the two towers. The goal of each tower is to compress high-dimensional and complex input (such as a piece of text, a user behavior sequence) into a fixed-dimensional vector (embedding), and through an InfoNCE contrastive loss function, to reduce the distance between semantic positive sample pairs (related Query-Documents) and to increase the distance between negative sample pairs (unrelated Query-Documents).
[0042] Considering that the content usually contains text, images, and other multi-modal data, a MoE architecture can be introduced into the double tower model to improve the processing efficiency and representation quality of the model. The Router in the MoE usually dynamically selects which experts to activate according to the input content, thereby achieving adaptive modeling of different modalities or semantic subspaces. This mechanism naturally supports "on-demand encoding", that is, when a modality information is missing, ambiguous, or redundant, the MoE can allocate less or even zero weight to the relevant experts, achieving weakening or ignoring of the modality. However, once the reconstruction phase (such as abstract reconstruction, image reconstruction, etc.) is entered, most existing frameworks still use a full-modal forced decoding strategy, that is, regardless of whether the encoder considers a modality important or whether the data of a modality is scarce, the decoder must generate the output of all modalities based on the shared representation. This design leads to two key problems. First, waste of computing resources: on the redundant generation of low-information or meaningless modalities, especially when using large decoders, the overhead is significant; second, when the information of a modality in the representation has been suppressed or even nearly lost by the MoE mechanism, the forced decoder "creates something out of nothing" to reconstruct the content, which is easy to introduce hallucinations or false mappings, not only reducing the generation quality, but also possibly polluting the shared representation space and interfering with the performance of downstream tasks.
[0043] Although the traditional multi-modal model can use fixed hyperparameters such as a text hyperparameter (λ_text) and an image hyperparameter (λ_image) to balance the importance of text reconstruction and image reconstruction. However, this approach ignores the large differences that may exist between different sample contents.
[0044] To solve this coding-decoding mismatch problem, in the embodiments of the present specification, the sparsity and adaptability of MoE are extended from the encoding stage to the reconstruction stage, and a "coding-decision-reconstruction" model training and application scheme is proposed. Specifically, target query data and to-be-queried content are obtained; the to-be-queried content is input into a content encoding unit in a content query model to obtain a content vector, and the target query data is input into a query encoding unit in the content query model to obtain a query vector, wherein the content query model includes the content encoding unit, the query encoding unit, a reconstruction unit, and a decision unit, the content encoding unit is trained based on sample reconstruction data of sample content, predicted reconstruction data, and reconstruction weights of the reconstruction unit, the predicted reconstruction data is obtained by processing a sample content vector of the sample content by the reconstruction unit, and the decision unit is used to generate the reconstruction weights; and a content query result is generated according to the content vector and the query vector.
[0045] It is worth noting that the scheme proposed in the embodiments of the present specification can enable the content query model to have the ability to intelligently judge "whether it is worth reconstructing" after encoding is completed, dynamically decide whether to perform and how to perform the subsequent reconstruction task, thereby concentrating the computing resources on the task that is really beneficial to learning high-quality representation, avoiding invalid and harmful reconstruction, and realizing more efficient, robust and semantically consistent multi-modal model learning. At the same time, the functions of the Router or the expert in the MoE architecture can be extended, so that it can not only output the activation weight of the encoding process, but also output the dynamic reconstruction weight for the reconstruction task.
[0046] In the present specification, a content query method is provided, and the present specification also relates to a content query model training method, a content query device, a content query model training device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0047] Referring to Figure 1 , Figure 1 A flowchart of a content query method provided by an embodiment of the present specification is shown, which specifically includes the following steps:
[0048] Step 102: Obtain target query data and to-be-queried content.
[0049] It should be noted that the target query data refers to the original query content input by the user during retrieval, which is used to express information needs. The target query data can be input into the query encoding unit, and its semantics will be encoded into a query vector to participate in subsequent similarity calculation. The target query data is usually text modal data, including but not limited to keywords, natural language questions or intent descriptions. For example, the target query data can be "recommend a light sunscreen suitable for oily skin". The target query data can also be data of other modalities, such as voice, video, image, etc. When the query encoding unit is used to encode the target query data, the target query data can be converted into text and input into the query encoding unit.
[0050] The to-be-queried content refers to the original content waiting to be matched with the target query data. The number of to-be-queried content can be one or more. The to-be-queried content can be input into the content encoding unit and converted into a content vector for subsequent ranking. The to-be-queried content can be single modal, such as structured or unstructured text data, or multi-modal multimedia content, i.e., information containing at least two modalities such as images, text, video, and audio. In addition, the to-be-queried content can also cover other formats of content of any object in the content sharing platform, such as location information, group chat records, product items, virtual resources, etc. The to-be-queried content can come from various application scenarios, such as product detail page text and user comments in an e-commerce platform, recommended notes and interactive comment data under the notes in a content sharing platform, and other user-generated or system-generated content that integrates multiple media forms.
[0051] In actual applications, there are various ways to obtain the target query data and the to-be-queried content, which are selected according to actual conditions, and the embodiments of the present specification do not make any limitation on this. In one possible implementation manner of the present specification, the target query data and the to-be-queried content can be read from the database of the content query system. In another possible implementation manner of the present specification, the target query data and the to-be-queried content can be received from the user through the client.
[0052] Step 104: input the to-be-queried content into the content encoding unit in the content query model to obtain a content vector, and input the target query data into the query encoding unit in the content query model to obtain a query vector, wherein the content query model comprises the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit, the content encoding unit is trained based on sample reconstruction data of sample content, predicted reconstruction data, and reconstruction weights of the reconstruction unit, the predicted reconstruction data is obtained by processing the sample content vector of the sample content by the reconstruction unit, and the decision unit is used to generate the reconstruction weights.
[0053] It should be noted that the content query model refers to an end-to-end neural network architecture, including a content encoding unit, a query encoding unit, a reconstruction unit, and a decision unit. For example, a reconstruction decoder for image reconstruction and a reconstruction decoder for summary reconstruction can be attached to the content encoding tower of the dual-tower model to construct a content query model; or a content query model can be constructed based on a large model. Through the content query model, efficient semantic retrieval can be achieved, and the representation quality can be improved through the reconstruction task. In the training stage, the content query model can jointly optimize the contrast loss (for retrieval) and the reconstruction loss (for image generation and / or summary generation).
[0054] The content encoding unit refers to a neural network module in the content query model responsible for mapping the content to be queried into a dense vector (i.e., a content vector), such as a multi-modal language model decoder (MLLM MoE Decoder) with a MoE architecture. Its key role is not only to perform semantic compression on high-dimensional input, but also to force the sample content vector to encode sufficient cross-modal semantic information through the joint training mechanism of the reconstruction unit and the decision unit, thereby significantly improving the information density, cross-modal alignment capability, and downstream generalization performance of the representation. Taking the MLLM MoE Decoder as an example, the MLLM MoE Decoder can dynamically select experts for processing according to the content of different modalities. The output is a shared multi-modal content vector that integrates text and visual information and preserves the relationship between modalities. The introduction of the MoE architecture enables the content query model to activate only the experts related to the content to be queried, enabling sparse computation and adaptive modeling.
[0055] The query encoding unit refers to a neural network module in the content query model responsible for encoding the target query data into a query vector, such as an LLM Decoder. Through the query encoding unit, the user's intent can be accurately described in the semantic space aligned with the content vector. The query encoding unit usually has the same structure as the content encoding unit or shares some parameters to ensure consistency in the embedding space.
[0056] The reconstruction unit is a generative module in the content query model, which is used to reconstruct the modal content (i.e., sample reconstruction data) such as text or image in the sample content vector output by the content encoding unit, and generate predicted reconstruction data. In the training stage, the reconstruction unit can provide a self-supervised signal by comparing the predicted reconstruction data with the sample reconstruction data to calculate the reconstruction loss, thereby constraining the content encoding unit to learn more complete and structured semantic representations, and forcing the sample content vector to retain key semantics (such as image-text consistency, layout structure, key entities, etc.) sufficient to support modal reconstruction, rather than surface statistical features. In practical applications, the reconstruction unit includes at least one of an image reconstruction unit and a summary reconstruction unit.
[0057] The image reconstruction unit is a generative sub-module, usually composed of a deep neural network (such as CNN, VisionTransformer or LDM, etc.), which functions to automatically generate a corresponding predicted content image as a conditional input of the sample content vector (i.e. the output of the content encoding unit). In the training phase, the image reconstruction unit can provide a cross-modal regularization signal by reconstructing the real image associated with the sample content semantics (i.e. the sample content image), which is used to calculate the image loss and optimize the content encoding unit in reverse, so that the output of the content vector not only contains the text semantics, but also contains the visual features sufficient to support visual reconstruction.
[0058] The abstract reconstruction unit is a generative decoder that generates text conditioned on the sample content vector, which functions to automatically generate a corresponding predicted content abstract as a conditional input of the sample content vector. In the training phase, the abstract reconstruction unit can provide another regularization signal by reconstructing the real abstract of the sample content (i.e. the sample content abstract), which is used to calculate the abstract loss and constrain the content encoding unit to learn more semantically complete representations, effectively alleviating the information bottleneck and representation collapse problem. For example, the abstract reconstruction unit can be a 12-layer Transformer Decoder or a 1.5B LLM Decoder.
[0059] The decision unit is a lightweight neural network that dynamically predicts the reconstruction weights (such as abstract reconstruction weights and image reconstruction weights) of each modality reconstruction task according to the sample content vector. The decision unit can also be called a decision gating network, which can be composed of a linear layer plus a Softmax activation function, or a multi-layer perceptron (MLP, Multi-Layer Perceptron) plus a Softmax activation function. Through the decision unit, sample-level adaptive supervision can be achieved, and the contribution proportion of reconstruction loss (such as image loss and abstract loss) in the total loss can be dynamically adjusted according to the information quality and importance of each modality in the input sample content, thereby avoiding forced reconstruction of low information quantity or missing modalities. For example, when the sample content is a pure text announcement (without valid images), the decision unit can output very low image reconstruction weights, greatly weakening the influence of image loss, thereby preventing the content query model from learning incorrect mappings due to "creating something out of nothing", and improving the purity and robustness of the representation.
[0060] The content vector refers to a low-dimensional dense representation output by the content encoding unit. In an embodiment of the present specification, the content vector is further constrained: it must contain sufficient information to reconstruct the sample reconstruction data. Therefore, the content vector can not only encode the text theme of the content to be queried, but also implicitly integrate the visual information (such as material, style, layout) of the content to be queried. This “reconstructability” requirement significantly improves the discriminability and generalizability of the content vector, enabling it to accurately match relevant content when facing queries containing visual intent.
[0061] The query vector refers to a low-dimensional dense representation output by the query encoding unit, representing the semantic embedding of the target query data. The content vector and the query vector can calculate semantic relevance through cosine similarity or dot product to determine the content query result (such as a recall ranking list).
[0062] In practical applications, the query encoding unit can be trained based only on the sample content vector and the sample query vector, or it can be trained based on the “sample content vector and sample query vector” in combination with “sample reconstruction data, predicted reconstruction data, and reconstruction weights of the reconstruction unit”.
[0063] Step 106: generating a content query result according to the content vector and the query vector.
[0064] It should be noted that the content query result refers to the result returned after calculation and sorting according to the content vector and the query vector. The content query result can be at least one content in multiple content to be queried that meets the target query data, or a part of the content in a content to be queried that meets the target query data. The content query result can be in different forms, such as a content list that meets the target query data, a complete content, a content abstract segment, or a structured response with relevance score, etc. Due to the optimization of the semantic completeness, cross-modal alignment capability, and information density of the content vector through the reconstruction task (such as the abstract reconstruction task, the image reconstruction task) of the reconstruction unit, the content query result is usually more accurate, more discriminative, and more responsive to queries containing visual intent than the traditional dual tower model.
[0065] It is worth noting that although the reconstruction unit and the decision unit are introduced in the content query model training process, they do not participate in the processing in the actual inference (i.e., generating the content query result) stage. The content query model only uses the dual tower structure composed of the lightweight content encoding unit and the query encoding unit to complete the content query, which not only retains the representation advantage brought by multi-modal supervision, but also ensures low latency and high throughput of the query service, achieving the engineering and algorithmic collaborative optimization of enhancing in training and being efficient in inference. Moreover, the content query method proposed in the present specification can be applied to search engines, e-commerce product recommendations, intelligent customer service question and answer matching, personalized news pushing, and other scenarios.
[0066] In actual applications, there are various ways to generate content query results according to the content vector and the query vector, which are selected according to actual conditions, and embodiments of the present specification do not make any limitation on this. In a possible implementation manner of the present specification, the cosine similarity between the query vector and the content vector can be calculated, and the content query results are recalled or sorted according to the cosine similarity score. In another possible implementation manner of the present specification, the dot product of the query vector and the content vector can be used as the relevance score, and the content query results are recalled or sorted according to the relevance score. Compared with the cosine similarity, the inner product retains the vector length information, and can reflect the “confidence” or “importance”.
[0067] By applying the scheme of the embodiments of the present specification, since the content vector and the query vector have been strongly supervised and adaptively aligned in the unified semantic space, the cross-modal semantic correlation can be accurately reflected, and the generation efficiency and accuracy of the content query results can be significantly improved.
[0068] For the training process of the content query model, in an optional embodiment of the present specification, an auxiliary reconstruction task can be introduced for the content encoding unit on the basis of the Query-Documents contrastive learning loss function. The reconstruction task requires the content query model to restore the sample reconstruction data of the sample content based on the sample content vector. That is, the training manner of the above content query model can include the following steps:
[0069] Obtaining sample data, wherein the sample data includes sample content, sample query data and sample reconstruction data;
[0070] Inputting the sample content into the content encoding unit to obtain the sample content vector, and inputting the sample query data into the query encoding unit to obtain the sample query vector;
[0071] Using the reconstruction unit to generate predicted reconstruction data based on the sample content vector, and using the decision unit to generate the reconstruction weight of the reconstruction unit;
[0072] Based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data and the reconstruction weight, the parameters of the content encoding unit, the query encoding unit, the reconstruction unit and the decision unit are adjusted to obtain the trained content query model.
[0073] It should be noted that the sample data refers to supervised data used for training the content query model, and at least includes sample content, sample query data and sample reconstruction data. The number of sample data is usually multiple, and in the same group of sample data, the sample query data and the sample content have a correlation relationship, while the sample query data and the sample content in different groups of sample data may or may not have a correlation relationship.
[0074] Sample content refers to the labeled content sample used in the content query model training stage, usually with real sample reconstruction data. Sample content is input to the content encoding unit to learn semantic representation. Sample content can be single-modal or multi-modal multimedia content. In addition, sample content can also cover other formats of content of any object in the content sharing platform. Sample content can come from various application scenarios, such as product detail page text and user reviews in e-commerce platforms, recommended notes and interactive comment data under the notes in content sharing platforms, and other user-generated or system-generated content that integrates multiple media forms. Sample query data refers to user query text semantically related to sample content, such as keywords in search logs, natural language questions, or queries corresponding to click behavior.
[0075] Sample reconstruction data refers to the true label of the original modal content in the sample content. Sample reconstruction data includes at least one of sample content image and sample content summary. Sample reconstruction data can be used as a supervised target for reconstruction task, compared with predicted reconstruction data to calculate reconstruction loss and provide self-supervised signals for content encoding unit.
[0076] Sample content image refers to a real image semantically associated with sample content, such as product main image, article image, cover image, or information visualization image. Sample content image can provide fine-grained visual information beyond text for content query model, and is a supervision source for cross-modal alignment. Through sample content image, image loss can be calculated to constrain the content vector to contain semantic information that can be mapped to visual space.
[0077] Sample content summary refers to a real summary text paired with sample content, usually written by humans or automatically extracted with high quality, representing the core text semantics of sample content summary, title, theme, entity, etc., and used to measure the "semantic integrity" of sample content vector.
[0078] Sample content vector refers to the vector representation output by the content encoding unit after processing the sample content. On the one hand, sample content vector can be used for contrastive learning (to calculate similarity with sample query vector), and on the other hand, sample content vector can be used as input condition for reconstruction unit. Sample query vector refers to the vector representation output by the query encoding unit after processing sample query data.
[0079] Predicted reconstruction data refers to the reconstructed content generated by the reconstruction unit based on the sample content vector. Predicted reconstruction data includes at least one of predicted content image and predicted content summary. Predicted reconstruction data can reflect the integrity of modal information retained by the sample content vector. For example, if the predicted reconstruction data is close to the sample reconstruction data, it means that the sample content vector encodes effective semantic structure; otherwise, it means that the sample content vector has information loss or bias.
[0080] The predicted content summary refers to a summary text generated by the summary reconstruction unit based on the sample content vector autoregression. The predicted content image refers to an image automatically generated by the image reconstruction unit with the sample content vector as the conditional input. In the embodiments of the present specification, the predicted content image can not be required to be accurately aligned with the sample content image at the pixel level, but is required to be consistent with the sample content image in terms of semantics, structure or key visual information (such as color distribution, object category, texture).
[0081] The reconstruction weight is a dynamic coefficient generated by the decision unit. The reconstruction weight includes at least one of the summary reconstruction weight and the image reconstruction weight. Through the reconstruction weight, sample-level adaptive supervision can be achieved, and the importance of each reconstruction task can be adjusted according to the modal quality of the sample content, avoiding forced reconstruction of low information modalities. For example, when the sample content has no effective image, the image reconstruction weight tends to zero, greatly weakening the influence of the image loss, thereby preventing noise introduction and improving model robustness. Moreover, the reconstruction weight also has interpretability, which can reflect the internal judgment of the content query model on the importance of the modal. For example, if the image reconstruction weight of the sample content is extremely low, it can be traced back to verify whether the image indeed lacks semantic information (such as a pure background image, an icon or an irrelevant illustration), thereby providing intuitive explanation and manual verification basis for the model behavior of the content query model.
[0082] In actual applications, there are various ways to adjust the parameters of the content encoding unit, the query encoding unit, the reconstruction unit and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data and the reconstruction weight to obtain a trained content query model, which are selected according to actual conditions, and the embodiments of the present specification do not make any limitation on this. In a possible implementation manner of the present specification, the contrast loss can be calculated based on the sample content vector and the sample query vector; the reconstruction loss can be calculated based on the reconstruction weight, the sample reconstruction data and the predicted reconstruction data; the total loss can be calculated according to the contrast loss and the reconstruction loss; the parameters of the query encoding unit can be adjusted according to the contrast loss, and the parameters of the content encoding unit, the reconstruction unit and the decision unit can be adjusted according to the total loss, respectively, to obtain a trained content query model. In another possible implementation manner of the present specification, after the contrast loss and the total loss are obtained, the parameters of the content encoding unit, the query encoding unit, the reconstruction unit and the decision unit can be adjusted according to the total loss, respectively, to obtain a trained content query model.
[0083] Exemplarily, the reconstruction unit includes an abstract reconstruction unit and an image reconstruction unit. The decision unit can be located in parallel with the sample content vector, including a simple linear layer and an activation layer after the linear layer. The content query model no longer uses fixed hyperparameters (λ_text and λ_image), but instead generates two dynamic weights, namely the abstract reconstruction weight (λ_text_dynamic) and the image reconstruction weight (λ_image_dynamic), by the decision unit. The range of λ_text_dynamic and λ_image_dynamic is between [0, 1] respectively, which respectively represents the "confidence" or "necessity" of the content query model to perform text reconstruction and image reconstruction on the current sample content vector. In the decision unit, the input of the decision unit (sample content vector and / or auxiliary processing vector) is transformed by the linear layer, which aims to map the original input to a dimension suitable for subsequent operations. The transformed features are then passed through the activation layer, which converts the transformed features into a probability distribution, so that the two values λ_text_dynamic and λ_image_dynamic output respectively represent the probability or confidence of performing text reconstruction and image reconstruction. Since text reconstruction and image reconstruction can be considered important or unimportant at the same time, the sum of λ_text_dynamic and λ_image_dynamic is not necessarily equal to 1.
[0084] It is worth noting that if λ_text_dynamic is close to zero, it means that the content query model considers that the text information of the sample content is not sufficient or suitable for meaningful abstract reconstruction, and the gradient of L_recon_text in this training will be greatly suppressed, or even not calculated at all, realizing the calculation skip; if λ_image_dynamic is close to zero, it means that the content query model considers that the image information of the sample content is not sufficient or suitable for meaningful image reconstruction, and the gradient of L_recon_image in this training will be greatly suppressed, or even not calculated at all, realizing the calculation skip. Therefore, the way of dynamically generating λ_text_dynamic and λ_image_dynamic by the decision unit according to the characteristics of the sample content is more flexible, which can automatically adjust the reconstruction strategy according to the specific situation of each sample content, avoiding the low efficiency or performance decline caused by applying the same reconstruction strategy to all sample contents.
[0085] Further, based on the sample content vector, the sample query vector, the sample content summary, the predicted content summary, the summary reconstruction weight, the sample content image, the predicted content image, and the image reconstruction weight, when adjusting the parameters of the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit, the contrastive loss (L_contrastive) can be calculated based on the sample content vector and the sample query vector; the summary loss (L_recon_text) can be calculated based on the summary reconstruction weight, the sample content summary, and the predicted content summary; the image loss (L_recon_image) can be calculated based on the image reconstruction weight, the sample content image, and the predicted content image; the total loss (L_total) is generated according to the contrastive loss, the image loss, and the summary loss, and the parameters of the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit are adjusted according to the total loss to obtain the trained content query model. Wherein, L_total = L_contrastive + λ_text_dynamic * L_recon_text + λ_image_dynamic * L_recon_image. Through gradient descent, the decision unit will learn a strategy: for text-rich content, output a higher λ_text_dynamic; for image-rich content, output a higher λ_image_dynamic; for content with obvious missing or noisy modality information, output the corresponding reconstruction weight close to zero.
[0086] By applying the scheme of the embodiments of the present specification, on the one hand, the contrastive learning goal between the sample content and the sample query data is used to drive the content query model to accurately align the relevant samples in the unified semantic space; on the other hand, the reconstruction unit is introduced to reconstruct the sample reconstruction data, and the decision unit dynamically generates the reconstruction weight of the modality perception, realizing on-demand supervision (strengthening multi-modal modeling when the text and image information is rich, and automatically weakening the reconstruction loss of a certain modality when the modality is missing or has serious noise), thereby avoiding the representation pollution caused by invalid reconstruction. The finally trained content query model not only learns the content vector and the query vector with high discriminability and high information density, but also exhibits stronger generalization ability and stability when facing heterogeneous, sparse, or multi-noise input in real scenarios.
[0087] In an optional embodiment of the present specification, the processing process of the content encoding unit with the structure of MLLM MoE Decoder is described, that is, the content encoding unit includes a transformation layer, a routing layer, a processing layer, and an output layer, and the processing layer includes a plurality of processing sub-units; the above inputting the sample content into the content encoding unit to obtain the sample content vector can include the following steps:
[0088] The sample content is processed through the transformation layer to obtain a content transformation vector;
[0089] The routing layer processes the content transformation vector to obtain an activation weight corresponding to each processing subunit;
[0090] The target processing subunit processes the content transformation vector to obtain a content processing vector, wherein the target processing subunit is selected from the plurality of processing subunits based on the activation weight;
[0091] The output layer generates a sample content vector based on the content processing vector and the activation weight corresponding to the target processing subunit.
[0092] It should be noted that the transformation layer is an initial processing module of the content encoding unit, which is usually composed of a multi-head attention mechanism (Multi-head Attention) and an add&norm layer, and is used for context-aware semantic fusion of the input sample content to obtain a unified intermediate hidden state (i.e., a content transformation vector). The transformation layer does not involve expert selection, but is a standard Transformer encoding process, which ensures that the content query model can understand the dependency between elements in the sample content and provides high-quality routing input for the routing layer.
[0093] The routing layer (Router) is a learnable lightweight network that receives the content transformation vector and dynamically determines which processing subunits (i.e., experts) should be activated and their corresponding activation weights based on the semantic features of the content transformation vector, thereby achieving sparse and adaptive computation allocation. For example, if the sample content is mainly in the form of charts, the routing layer may activate experts skilled in visual reasoning with high weights; if it is pure text, it may be biased towards language experts. This mechanism enables the model to have an "on-demand calling" capability, improving expression efficiency and capacity. The activation weight is a scalar value calculated by the routing layer for each processing subunit, which represents the "relevance" or "activation degree" of the current input sample content to the expert.
[0094] The processing layer is composed of a plurality of parallel processing subunits, each of which is an independent feedforward neural network (FFN, Feedforward Neural Network) responsible for nonlinear transformation of the input to output a content processing vector. The processing layer is the core of the MoE architecture, which realizes exponential expansion of model capacity through expert division while keeping the computational load of single inference controllable (e.g., only activating Top-2 experts). Under the guidance of the routing layer, only the selected target processing subunit participates in actual computation to process the content transformation vector and generate a more targeted semantic representation. Different processing subunits can implicitly learn different skills (such as "graph-text alignment", "table parsing", "sentiment analysis", etc.), and complex semantic understanding can be achieved through combination.
[0095] The target processing subunit refers to Top-k high-weight experts (e.g., k = 2) selected from all processing subunits according to the activation weights output by the routing layer. This mechanism embodies the sparse activation characteristic of MoE, which not only retains the capacity of large models but also controls the computational cost. The target processing subunit actually performs the calculation task to generate the content processing vector, and the rest of the unselected experts remain dormant and do not participate in the forward propagation.
[0096] The output layer is used to weight and sum the content processing vectors output by the target processing subunit according to their corresponding activation weights to generate the final sample content vector. Weighted aggregation ensures that the content query model can utilize the complementary capabilities of multiple experts while maintaining the continuity and stability of the output.
[0097] By applying the scheme of the embodiments of the present specification, the context-aware content transformation vector is provided through the transformation layer, the routing layer realizes input-driven expert selection, the processing layer enhances the model expression capability through the sparse activation processing subnetwork, and the output layer fuses the expert opinions to generate a unified sample content vector. The entire process not only significantly improves the understanding ability of the content query model for complex and heterogeneous content (such as graphic text layout, tables, and charts), but also effectively controls the computational overhead through the sparsity of MoE; and, the structure naturally supports joint training with subsequent reconstruction units and decision units, so that the content vector can not only maintain the retrieval performance but also encode rich semantics sufficient to support multi-modal reconstruction.
[0098] Referring to Figure 2 , Figure 2 Fig. 1 shows a processing process schematic diagram of a content encoding unit provided by an embodiment of the present specification. The content encoding unit is a Transformer-based MoE decoder, which combines the MoE architecture and the Transformer decoder mechanism to realize adaptive and sparse coding of multi-modal input (such as graphic text). Specifically, the content encoding unit includes a transformation layer, a routing layer, a processing layer, and an output layer. The processing layer includes multiple processing subunits (e.g., processing subunit 1, processing subunit 2, …, processing subunit 7, and processing subunit 8). The input of the content encoding unit is multi-modal sample content, and the output is a sample content vector. Moreover, the output layer can be connected to a decision unit, which can generate summary reconstruction weights and image reconstruction weights based on the sample content vector finally output by the content encoding unit; and / or generate summary reconstruction weights and image reconstruction weights based on the intermediate layer features of the content encoding unit.
[0099] In an optional embodiment of the present specification, the above-mentioned use of the reconstruction unit to generate predicted reconstruction data based on the sample content vector, and the use of the decision unit to generate reconstruction weights of the reconstruction unit, can include the following steps:
[0100] The reconstruction unit comprises an abstract reconstruction unit; the abstract reconstruction unit is used to generate a predicted content abstract based on the sample content vector, and the decision unit is used to generate an abstract reconstruction weight of the abstract reconstruction unit; and / or,
[0101] The reconstruction unit comprises an image reconstruction unit; the image reconstruction unit is used to generate a predicted content image based on the sample content vector, and the decision unit is used to generate an image reconstruction weight of the image reconstruction unit.
[0102] It should be noted that the abstract reconstruction weight is a scalar coefficient dynamically generated by the decision unit, which is used to adjust the contribution intensity of the abstract reconstruction task in the total loss function. The abstract reconstruction weight can adaptively determine whether to apply strong supervision to the predicted content abstract according to the richness of the text information, the semantic integrity or the abstract generability of the sample content; when the sample content lacks effective text content (such as pure pictures or random codes), the abstract reconstruction weight is automatically reduced to avoid forced generation of meaningless abstracts.
[0103] The image reconstruction weight is another scalar coefficient dynamically generated by the decision unit, which is used to control the contribution intensity of the image reconstruction task in the total loss function. The image reconstruction weight can enable the content query model to focus on the truly discriminative visual semantics rather than the surface pixel details, thereby improving the robust encoding ability of the sample content vector for the visual modality and enhancing the adaptability of the system to multi-modal unbalanced data.
[0104] In actual application, if the information of a certain modality in the sample content is scarce (such as image blur or text loss), the corresponding reconstruction weight tends to be zero, and the reconstruction is inhibited; if the information is sufficient, the reconstruction weight is close to 1.
[0105] By applying the scheme of the embodiments of the present specification, the decision unit dynamically evaluates the information value of each modality, and only applies strong reconstruction supervision to the high-confidence modality, effectively avoiding the calculation waste and representation pollution caused by traditional global forced reconstruction. This enables the vector learned by the content encoding unit to not only have high discriminability in the semantic retrieval task, but also accurately retain key information sufficient to support high-quality abstract generation and visual semantic restoration. Ultimately, the content query model can exhibit stronger generalization ability, robustness and interpretability when facing complex content with imbalanced text and image proportions, missing modalities or noise interference in real scenarios.
[0106] In an optional embodiment of the present specification, the above step of generating the reconstruction weight of the reconstruction unit by using the decision unit can comprise the following steps:
[0107] The decision unit is used to generate the reconstruction weight of the reconstruction unit based on the sample content vector; and / or,
[0108] The decision unit generates reconstruction weights of the reconstruction unit based on an auxiliary processing vector, wherein the auxiliary processing vector is an intermediate layer vector of a target processing subunit in the content encoding unit.
[0109] It should be noted that the sample content vector is a dense embedding representation finally output by the content encoding unit, which is used to represent the semantic content of the entire sample content. The sample content vector is located at the end of the encoding process and can reflect the overall semantic integrity, and is suitable as a high-level basis for generating reconstruction weights.
[0110] The auxiliary processing vector refers to the hidden state vector output by the target processing subunit (i.e., the activated MoE expert) at an intermediate layer (such as before or after the FFN) in the content encoding unit. Since the target processing subunit may have enhanced or filtered a specific modality (such as visual or language) during processing, the intermediate output can more early and more sensitively reflect the modality quality, thereby improving the accuracy and timeliness of the reconstruction weight prediction. Whether it is a sample content vector or an auxiliary processing vector, it contains rich text and visual information and reflects the semantic content of the current sample content.
[0111] By applying the scheme of the embodiments of the present specification, the decision unit can not only ensure semantic consistency based on the global sample content vector, but also capture modality-specific details based on the auxiliary processing vector inside the target processing subunit. Through a lightweight decision head (such as a linear layer and an activation layer), the reconstruction weight is predicted, which not only significantly reduces the calculation waste and noise introduction of invalid modalities, but also promotes the content encoding unit to focus on learning the truly generalizable and reconstructable cross-modal core semantics.
[0112] In an optional embodiment of the present specification, the above based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weight, the parameter adjustment of the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit to obtain the trained content query model can include the following steps:
[0113] Based on the sample content vector and the sample query vector, a contrast loss is calculated.
[0114] Based on the reconstruction weight, the sample reconstruction data, and the predicted reconstruction data, a reconstruction loss is calculated.
[0115] According to the contrast loss, the parameter of the query encoding unit is adjusted, and according to the contrast loss and the reconstruction loss, the parameters of the content encoding unit, the reconstruction unit, and the decision unit are adjusted respectively, and the trained content query model is obtained.
[0116] It should be noted that the contrast loss refers to a loss value for measuring the degree of semantic matching between the sample content vector and the sample query vector. The way of calculating the generated loss includes but is not limited to InfoNCE, cross-entropy loss function (CE, Cross-Entropy). In a batch of sample data, the contrast loss can narrow the vector distance between the sample query vector and the relevant sample content, while pushing away the vector distance between the sample query vector and the irrelevant sample content, driving the query encoding unit and the content encoding unit to align in the shared semantic space, and improving the query relevance and matching.
[0117] The reconstruction loss is a supervised signal for measuring the difference between the predicted reconstruction data and the sample reconstruction data, and is usually calculated by cross-entropy, perception loss. The reconstruction loss includes at least one of the image loss and the summary loss. Based on the reconstruction weight, the sample reconstruction data and the predicted reconstruction data, the way of calculating the reconstruction loss includes at least one of "calculating the summary loss based on the summary reconstruction weight, the sample content summary and the predicted content summary" and "calculating the image loss based on the image reconstruction weight, the sample content image and the predicted content image".
[0118] The image loss, which can also be referred to as the image reconstruction loss, refers to a loss value for measuring the difference between the predicted content image and the sample content image. The way of calculating the image loss includes but is not limited to mean square error loss function, perception loss function. The image loss can be used as a regularization signal to constrain the sample content vector to contain the visual semantics of the reconstructable sample content image, preventing information loss.
[0119] The summary loss, which can also be referred to as the summary reconstruction loss, refers to a loss value for measuring the difference between the predicted content summary and the sample content summary. The way of calculating the summary loss includes but is not limited to cross-entropy loss function, negative log-likelihood loss function (NLL, Negative Log-Likelihood). The summary loss can be used as another regularization signal to constrain the sample content vector to contain sufficient semantic information to support accurate summary generation, preventing information loss.
[0120] By applying the scheme of the embodiments of the present specification, the parameters of the query encoding unit are adjusted only by the contrast loss, while the parameters of the content encoding unit are adjusted by both the contrast loss and the reconstruction loss, ensuring that the sample query vector focuses on matching, and the sample content vector is forced to contain information sufficient to reconstruct the real sample reconstruction data.
[0121] In an optional embodiment of the present specification, during the model training stage, in order to pursue maximum representation capability, a sparse large model such as MoE can be used as a content encoding unit; and during the inference stage, due to the requirements of delay, cost and stability, the dynamic routing overhead, expert scheduling complexity and high memory occupation of MoE cannot be tolerated. If the content encoding unit of the MoE architecture is directly deployed, engineering problems such as uneven load, cache invalidation and low batch processing efficiency may be faced. In order to solve this problem, an embodiment of the present specification proposes a scheme for distilling and compressing the content encoding unit. The distilled target content encoding unit has a fixed structure, a determined calculation path, is easy to quantify and accelerate, and is suitable for the pre-computation and batch inference scenarios of a vector retrieval engine. That is, after the above steps of adjusting the parameters of the content encoding unit, the query encoding unit, the reconstruction unit and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data and the reconstruction weight, and obtaining the trained content query model, the following steps can be further included:
[0122] distilling and compressing the content encoding unit to obtain a target content encoding unit;
[0123] inputting the content to be queried into the content encoding unit in the content query model to obtain a content vector, which can include the following steps:
[0124] inputting the content to be queried into the target content encoding unit to obtain a content vector.
[0125] It should be noted that distillation compression is a model compression technology. By letting a "student model" (such as a dense encoder) imitate the output behavior of a "teacher model" (such as an MoE encoder), the model size and inference overhead can be significantly reduced while maintaining high performance. Through distillation compression, the complex content encoding unit (such as MLLM-MoE Decoder) used in the training stage can be converted into a lightweight target content encoding unit, which is convenient for deployment to an online service system.
[0126] The target content encoding unit refers to a lightweight content encoding unit obtained after distillation compression, which is usually a dense neural network (such as a variant of BERT or Sentence Transformer) with a fixed structure. The target content encoding unit inherits the semantic understanding capability of the original complex content encoding unit, but has a smaller parameter size, a determined calculation path, no sparse activation overhead, and is more suitable for high-concurrency deployment on CPU / GPU, saving the overhead of online deployment resources. During the inference stage, the target content encoding unit can replace the content encoding unit to efficiently generate a content vector and support real-time retrieval services.
[0127] In actual applications, there are various ways to distill and compress the content encoding unit to obtain the target content encoding unit, which are selected according to actual conditions, and the embodiments of the present specification do not make any limitation on this. In a possible implementation manner of the present specification, the distance (such as the cosine distance) between the content vectors of the student model (dense encoding unit) and the teacher model (content encoding unit) in the output layer can be minimized through output vector distillation, so that the student model directly imitates the semantic representation space of the teacher. During training, the student model learns to approximate the dense vector output by the teacher, thereby inheriting its retrieval capability. This method is suitable for a double-tower architecture and does not require additional supervision signals. In another possible implementation manner of the present specification, the content vector of the student model is required to be close to the output of the teacher model through contrastive distillation, and the student model is also required to maintain similar positive and negative sample discrimination ability in the contrastive learning task. Specifically, in the InfoNCE contrastive loss, the query-content similarity distribution generated by the teacher model is used as a soft label to guide the student model to learn the same relative ranking relationship. This method can better preserve the discriminative structure of the teacher model in semantic retrieval.
[0128] By applying the scheme of the embodiments of the present specification, the content encoding unit of the high-capacity architecture such as MoE is fully learned for multi-modal semantic and reconstruction capability in the training stage, and the distilled target content encoding unit is deployed in the inference stage, which greatly reduces the calculation resource consumption, memory occupation and response delay without losing the content query accuracy. The final content query system can process complex image-text content and meet the online service demand of high concurrency and low delay.
[0129] In an optional embodiment of the present specification, the above distilling and compressing the content encoding unit to obtain the target content encoding unit can include the following steps:
[0130] Obtaining auxiliary content and auxiliary query data;
[0131] Inputting the auxiliary content into the content encoding unit to obtain an auxiliary content vector, and inputting the auxiliary query data into the query encoding unit to obtain an auxiliary query vector;
[0132] Inputting the auxiliary content into the dense encoding unit to obtain a dense content vector, wherein the dense encoding unit has the same structure as the content encoding unit, and the parameter amount of the dense encoding unit is less than that of the content encoding unit;
[0133] Adjusting the parameters of the dense encoding unit according to the auxiliary content vector, the auxiliary query vector and the dense content vector to obtain the target content encoding unit.
[0134] Note that the auxiliary content is a batch of multi-modal content samples introduced additionally in the distillation stage, used to transfer the knowledge of the teacher model (original content encoding unit). The auxiliary content usually comes from a large-scale corpus (such as web pages, news, product descriptions, etc.), and in an optional embodiment, sample content can also be used as auxiliary content.
[0135] The auxiliary query data is a natural language query text related to the semantic of the auxiliary content or constructed, used to activate the query encoding unit and generate the corresponding auxiliary query vector.
[0136] The content encoding unit is the content encoding unit in the trained content query model (such as MLLM-MoE Decoder), which is the “teacher model” in the distillation compression process. The structure of the content encoding unit may include sparse-activated expert networks, multi-layer routing mechanisms, etc., which have strong representation ability but high inference cost, and are not suitable for direct deployment.
[0137] The query encoding unit is the query encoding unit in the trained content query model, which remains fixed (parameters are not updated) in the distillation compression stage. The query encoding unit can encode the auxiliary query data to generate the auxiliary query vector, which is used to construct the semantic similarity distribution of the teacher model.
[0138] The dense encoding unit is a neural network similar in structure to the content encoding unit but with significantly fewer parameters (such as removing MoE, reducing the number of layers or hidden dimensions of the content encoding unit), which participates in the distillation compression process as a “student model”. The “dense” property of the dense encoding unit means that all parameters are activated in each inference, and the calculation path is determined, which is suitable for efficient batch processing and vector indexing systems. The dense encoding unit learns to generate auxiliary content vectors close to the output of the content encoding unit by imitating the behavior of the content encoding unit, and finally becomes the deployable target content encoding unit.
[0139] The auxiliary content vector is the high-dimensional vector output by the content encoding unit after encoding the auxiliary content, representing the semantic representation of the teacher model, which is the learning goal of the student model. The auxiliary query vector is the vector obtained by encoding the auxiliary query data by the query encoding unit, which is in the same semantic space as the auxiliary content vector. The auxiliary query vector is used to calculate the content-query similarity under the teacher model, construct soft labels (such as similarity distribution), and support contrastive distillation. The dense content vector is the vector output by the dense encoding unit after encoding the same auxiliary content, which is the current prediction result of the student model.
[0140] By applying the scheme of the embodiments of the present specification, the dense coding unit not only learns to fit the auxiliary content vector directly, but also implicitly learns the complex content-query relative relationship in the teacher model through the auxiliary query vector. This dual supervision mechanism enables the dense coding unit to effectively inherit the semantic discrimination ability and multi-modal understanding depth of the content coding unit while greatly compressing the parameter amount. The final target content coding unit has high retrieval accuracy, low reasoning delay and strong deployment compatibility.
[0141] Considering that the model parameter amount of the content query model is relatively large and the operation resources of the client are limited, the content query method proposed in the embodiments of the present specification can be applied to a content query system as shown in Figure 3 , but is not limited thereto. Referring to Figure 3 , Figure 3 An architecture diagram of a content query system provided by one embodiment of the present specification is shown, and the content query system can include a client 302 and a server 304.
[0142] The client 302 is configured to send target query data and to-be-queried content to the server 304.
[0143] The server 304 is configured to input the to-be-queried content into a content coding unit in a content query model to obtain a content vector, and input the target query data into a query coding unit in the content query model to obtain a query vector, wherein the content query model includes the content coding unit, the query coding unit, a reconstruction unit and a decision unit, the content coding unit is trained based on sample reconstruction data of sample content, predicted reconstruction data and reconstruction weights of the reconstruction unit, the predicted reconstruction data is obtained by processing a sample content vector of the sample content by the reconstruction unit, and the decision unit is configured to generate the reconstruction weights; generate a content query result according to the content vector and the query vector; and send the content query result to the client 302.
[0144] The client 302 is further configured to receive the content query result sent by the server 304.
[0145] As shown in Figure 3As shown, the content query model is deployed in the server 304, and the server 304 can be connected to one or more clients 302 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The data transmitted by the client 302 can need to be encoded, transcoded, compressed, and the like before being published to the server 304. The server 304 can establish a communication connection between multiple clients 302, and in a content query scenario, the server 304 is used to provide content query services between multiple clients 302. The multiple clients 302 can respectively act as a sending end or a receiving end, and realize communication through the server 304. A user can interact with the server 304 through the client 302 to receive data sent by other clients 302, or send data to other clients 302, and the like. In a content query scenario, the user can publish a data stream to the server 304 through the client 302, the server 304 generates a content query result according to the data stream, and pushes the content query result to other clients that establish a communication connection.
[0146] The client 302 can be a browser, an application (APP), or a web application such as a HyperText Markup Language 5 (H5) application, or a light application (also known as a small program, a lightweight application), or a cloud application, and the like. The client 302 can be developed based on a software development kit (SDK) provided by the server 304 for a corresponding service, such as a real-time communication (RTC) SDK. The client 302 can be deployed in an electronic device, and needs to depend on a device or an APP in the device for running, and the like. The electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer (PC), and the like. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, mailbox clients, social platform software, and the like. The client 302 can also interact with a user through a user graphical interface to call the content query model, and realize the content query method provided in the embodiments of the present specification.
[0147] The service end 304 can include a server providing various services, for example, a server providing a communication service for a plurality of clients, for example, a server for background training supporting a model used on a client, for example, a server processing data sent by a client, and the like. It should be noted that the service end 304 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of a cloud service, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, and the like. Basic cloud computing services of artificial intelligence technology, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0148] It should be noted that the content query method provided in the embodiments of the present specification is generally executed by the service end, but in other embodiments of the present specification, in the case that the running resources of the client can meet the deployment and running conditions of the content query model, the client can also have similar functions as the service end, so as to execute the content query method provided in the embodiments of the present specification. In other embodiments, the content query method provided in the embodiments of the present specification can also be executed by the client and the service end together.
[0149] Referring to Figure 4 , Figure 4 A flowchart of a content query model training method provided by one embodiment of the present specification is shown, which specifically includes the following steps:
[0150] Step 402: Obtain sample data, wherein the sample data includes sample content, sample query data, and sample reconstruction data.
[0151] Step 404: Input the sample content into the content encoding unit in the content query model to obtain a sample content vector, and input the sample query data into the query encoding unit in the content query model to obtain a sample query vector.
[0152] Step 406: Use the reconstruction unit in the content query model to generate predicted reconstruction data based on the sample content vector, and use the decision unit in the content query model to generate reconstruction weights of the reconstruction unit.
[0153] Step 408: Based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weights, adjust the parameters of the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit to obtain a trained content query model.
[0154] It should be noted that the implementation of steps 402 to 408 can refer to the training method of the content query model described above, and the embodiments of the present specification will not be described again.
[0155] By applying the scheme of the embodiments of the present specification, on the one hand, the comparison learning goal between the sample content and the sample query data is used to drive the content query model to accurately align the relevant samples in the unified semantic space; on the other hand, the reconstruction unit is introduced to reconstruct the sample reconstruction data, and the reconstruction weight of the modal perception is dynamically generated by the decision unit, so as to realize the on-demand supervision (strengthening multi-modal modeling when the graphic information is rich, and automatically weakening the reconstruction loss when a certain modal is missing or the noise is serious), thereby avoiding the representation pollution caused by invalid reconstruction. The content query model finally trained not only learns the content vector and the query vector with high discriminability and high information density, but also exhibits stronger generalization ability and stability when facing heterogeneous, sparse or multi-noise input in real scenes.
[0156] Referring to Figure 5 , Figure 5 A processing process schematic diagram of a content query model training method provided by an embodiment of the present specification is shown, and the content query model training method can be regarded as an adaptive multi-modal reconstruction representation training method based on expert decision. The content query model includes a query encoding unit, a content encoding unit, a decision unit, an abstract reconstruction unit and an image reconstruction unit. The content query model training method includes: obtaining sample data, wherein the sample data includes sample content, sample query data, sample content abstract and sample content image; inputting the sample content into the content encoding unit to obtain a sample content vector, and inputting the sample query data into the query encoding unit to obtain a sample query vector; using the abstract reconstruction unit to generate a predicted content abstract based on the sample content vector, and using the image reconstruction unit to generate a predicted content image based on the sample content vector; using the decision unit to generate an abstract reconstruction weight of the abstract reconstruction unit, and using the decision unit to generate an image reconstruction weight of the image reconstruction unit; calculating a comparison loss based on the sample content vector and the sample query vector; calculating an abstract loss based on the abstract reconstruction weight, the sample content abstract and the predicted content abstract; calculating an image loss based on the image reconstruction weight, the sample content image and the predicted content image; adjusting the parameters of the content query model based on the comparison loss, the abstract loss and the image loss to obtain a trained content query model.
[0157] Referring to Figure 6 , Figure 6 A flow timing diagram of a content query method provided by an embodiment of the present specification is shown, and the server and the client perform data interaction in the content query process.
[0158] The client is configured to send sample data to the server, wherein the sample data comprises sample content, sample query data, and sample reconstruction data.
[0159] The server is configured to input the sample content into a content encoding unit in a content query model to obtain a sample content vector, and input the sample query data into a query encoding unit in the content query model to obtain a sample query vector; generate predicted reconstruction data based on the sample content vector by using a reconstruction unit, and generate reconstruction weights of the reconstruction unit by using a decision unit; and perform parameter adjustment on the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weights, to obtain a trained content query model.
[0160] The client is further configured to send target query data and to-be-queried content to the server.
[0161] The server is further configured to input the to-be-queried content into the content encoding unit in the trained content query model to obtain a content vector, and input the target query data into the query encoding unit in the trained content query model to obtain a query vector; and generate a content query result based on the content vector and the query vector, and send the content query result to the client.
[0162] Corresponding to the content query method embodiments described above, the present specification also provides content query device embodiments, Figure 7 A structural schematic diagram of a content query device is shown in an embodiment of the present specification. As shown in the figure, Figure 7 The device comprises:
[0163] The first obtaining module 702 is configured to obtain target query data and to-be-queried content.
[0164] The first input module 704 is configured to input the to-be-queried content into a content encoding unit in a content query model to obtain a content vector, and input the target query data into a query encoding unit in the content query model to obtain a query vector, wherein the content query model comprises the content encoding unit, the query encoding unit, a reconstruction unit, and a decision unit, the content encoding unit is trained based on sample reconstruction data of sample content, predicted reconstruction data, and reconstruction weights of the reconstruction unit, the predicted reconstruction data is obtained by processing a sample content vector of the sample content by the reconstruction unit, and the decision unit is configured to generate the reconstruction weights.
[0165] The first generation module 706 is configured to generate a content query result based on the content vector and the query vector.
[0166] Optionally, the apparatus further comprises a second adjusting module configured to obtain sample data, wherein the sample data comprises sample content, sample query data and sample reconstruction data; input the sample content into the content encoding unit to obtain a sample content vector, and input the sample query data into the query encoding unit to obtain a sample query vector; generate predicted reconstruction data based on the sample content vector by using the reconstruction unit, and generate reconstruction weights of the reconstruction unit by using the decision unit; and adjust parameters of the content encoding unit, the query encoding unit, the reconstruction unit and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data and the reconstruction weights to obtain a trained content query model.
[0167] Optionally, the content encoding unit comprises a transformation layer, a routing layer, a processing layer and an output layer, the processing layer comprises a plurality of processing subunits; and the second adjusting module is further configured to process the sample content through the transformation layer to obtain a content transformation vector; process the content transformation vector through the routing layer to obtain activation weights corresponding to the plurality of processing subunits respectively; process the content transformation vector through a target processing subunit to obtain a content processing vector, wherein the target processing subunit is selected from the plurality of processing subunits based on the activation weights; and generate the sample content vector based on the content processing vector and the activation weights corresponding to the target processing subunit through the output layer.
[0168] Optionally, the second adjusting module is further configured to: the reconstruction unit comprises an abstract reconstruction unit; generate a predicted content abstract based on the sample content vector by using the abstract reconstruction unit, and generate abstract reconstruction weights of the abstract reconstruction unit by using the decision unit; and / or, the reconstruction unit comprises an image reconstruction unit; generate a predicted content image based on the sample content vector by using the image reconstruction unit, and generate image reconstruction weights of the image reconstruction unit by using the decision unit.
[0169] Optionally, the second adjusting module is further configured to generate the reconstruction weights of the reconstruction unit based on the sample content vector by using the decision unit; and / or, generate the reconstruction weights of the reconstruction unit based on an auxiliary processing vector by using the decision unit, wherein the auxiliary processing vector is an intermediate layer vector of a target processing subunit in the content encoding unit.
[0170] Optionally, the second adjusting module is further configured to calculate a contrast loss based on the sample content vector and the sample query vector; calculate a reconstruction loss based on the reconstruction weights, the sample reconstruction data and the predicted reconstruction data; adjust parameters of the query encoding unit according to the contrast loss, and adjust parameters of the content encoding unit, the reconstruction unit and the decision unit according to the contrast loss and the reconstruction loss respectively to obtain the trained content query model.
[0171] Optionally, the apparatus further comprises a compression module configured to distill compress the content encoding unit to obtain a target content encoding unit; and a first input module 704 further configured to input the content to be queried into the target content encoding unit to obtain a content vector.
[0172] Optionally, the compression module is further configured to obtain auxiliary content and auxiliary query data; input the auxiliary content into the content encoding unit to obtain an auxiliary content vector, and input the auxiliary query data into the query encoding unit to obtain an auxiliary query vector; input the auxiliary content into a dense encoding unit to obtain a dense content vector, wherein the dense encoding unit has the same structure as the content encoding unit, and the parameter amount of the dense encoding unit is less than that of the content encoding unit; and adjust parameters of the dense encoding unit according to the auxiliary content vector, the auxiliary query vector, and the dense content vector to obtain the target content encoding unit.
[0173] By using the light-weight double-tower encoding structure (the content encoding unit and the query encoding unit), the scheme of the embodiment of the present specification can efficiently generate a high-precision content query result: since the content vector and the query vector have been adaptively aligned in the unified semantic space under strong supervision, the two can accurately reflect the cross-modal semantic correlation, thereby significantly improving the recall rate and the sorting quality.
[0174] The above is a schematic scheme of a content query apparatus according to an embodiment of the present specification. It should be noted that the technical scheme of the content query apparatus and the technical scheme of the content query method described above belong to the same concept, and the details of the technical scheme of the content query apparatus that are not described in detail can be referred to the description of the technical scheme of the content query method.
[0175] Corresponding to the content query model training method embodiment described above, the present specification also provides a content query model training apparatus embodiment, Figure 8 Fig. 8 shows a structural schematic diagram of a content query model training apparatus according to an embodiment of the present specification. As shown in the figure, Figure 8 The apparatus comprises:
[0176] A second obtaining module 802 is configured to obtain sample data, wherein the sample data comprises sample content, sample query data, and sample reconstruction data;
[0177] A second input module 804 is configured to input the sample content into a content encoding unit in the content query model to obtain a sample content vector, and input the sample query data into a query encoding unit in the content query model to obtain a sample query vector;
[0178] The second generation module 806 is configured to generate predicted reconstruction data based on the sample content vector by using a reconstruction unit in the content query model, and generate reconstruction weights of the reconstruction unit by using a decision unit in the content query model.
[0179] The first adjustment module 808 is configured to perform parameter adjustment on the content encoding unit, the query encoding unit, the reconstruction unit and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data and the reconstruction weights, to obtain a trained content query model.
[0180] By using the scheme of the embodiments of the present specification, on the one hand, the comparison learning goal between the sample content and the sample query data is used to drive the content query model to accurately align the related samples in the unified semantic space; on the other hand, the reconstruction unit is introduced to reconstruct the sample reconstruction data, and the decision unit is used to dynamically generate the reconstruction weights of the modal perception, so as to realize the on-demand supervision (strengthening the multi-modal modeling when the information of the text and the image is rich, and automatically weakening the reconstruction loss when a certain modal is missing or the noise is serious), thereby avoiding the representation pollution caused by invalid reconstruction. The finally trained content query model not only learns the content vector and the query vector with high discriminability and high information density, but also exhibits stronger generalization ability and stability when facing heterogeneous, sparse or multi-noise input in real scenes.
[0181] The above is a schematic scheme of the content query model training device of the present embodiment. It should be noted that the technical scheme of the content query model training device belongs to the same concept as the technical scheme of the content query model training method described above, and the details of the technical scheme of the content query model training device that are not described in detail can be referred to the description of the technical scheme of the content query model training method.
[0182] Figure 9 A structural block diagram of a computing device is shown. The components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to save data.
[0183] The computing device 900 also includes an access device 940 that enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or combinations of such networks, such as the Internet. The access device 940 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, or the like.
[0184] In one embodiment of the present specification, the above-mentioned components of the computing device 900 and other components not shown in the Figure 9 may be connected to each other, for example, through a bus. It should be understood that Figure 9 The computing device structure diagram shown is only for the purpose of example, and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0185] The computing device 900 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, and the like), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smart watch, smart glasses, and the like), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer. The computing device 900 can also be a mobile or stationary server.
[0186] The processor 920 is configured to execute computer programs / instructions that implement the steps of the content query method or the content query model training method when executed by the processor.
[0187] The above is a schematic scheme of the computing device of the embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the content query method and the content query model training method described above belong to the same concept, and the details of the technical scheme of the computing device that are not described in detail can be seen from the description of the technical scheme of the content query method or the content query model training method.
[0188] An embodiment of the present specification further provides a computer readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the content query method or the content query model training method.
[0189] The above is a schematic scheme of the computer readable storage medium of the embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the content query method and the content query model training method described above belong to the same concept, and the details of the technical scheme of the storage medium that are not described in detail can be seen from the description of the technical scheme of the content query method or the content query model training method.
[0190] An embodiment of the present specification further provides a computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the content query method or the content query model training method.
[0191] The above is a schematic scheme of the computer program product of the embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the content query method and the content query model training method described above belong to the same concept, and the details of the technical scheme of the computer program product that are not described in detail can be seen from the description of the technical scheme of the content query method or the content query model training method.
[0192] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in an order other than that described in the embodiments and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing can be advantageous or possible.
[0193] The computer readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution package, etc. It should be noted that the computer readable medium can include appropriate contents according to the requirements of patent practice, for example, according to the patent practice in some regions, the computer readable medium does not include an electrical carrier signal and a telecommunications signal.
[0194] It should be noted that, for the foregoing method embodiments, in order to facilitate description, each is described as a combination of a series of acts, but those skilled in the art should know that the present specification embodiments are not limited to the order of the acts described, because according to the present specification embodiments, certain steps can be performed in other orders or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the acts and modules involved are not necessarily essential to the present specification embodiments.
[0195] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0196] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details and limit the present invention to the specific embodiments described. Obviously, according to the content of the present specification embodiments, many modifications and changes can be made. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present specification embodiments, so that those skilled in the art can well understand and use the present specification. The present specification is limited only by the claims and their full scope and equivalents.
Claims
1. A content query method, characterized in that, include: Retrieve the target query data and the content to be queried; The content to be queried is input into the content encoding unit in the content query model to obtain a content vector, and the target query data is input into the query encoding unit in the content query model to obtain a query vector. The content query model includes the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit. The content encoding unit is trained based on the sample reconstruction data of the sample content, the predicted reconstruction data, and the reconstruction weights of the reconstruction unit. The predicted reconstruction data is obtained by the reconstruction unit processing the sample content vector of the sample content. The decision unit is used to generate the reconstruction weights. Based on the content vector and the query vector, generate content query results.
2. The method according to claim 1, characterized in that, The training methods for the content query model include: Obtain sample data, wherein the sample data includes the sample content, sample query data, and sample reconstruction data; The sample content is input into the content encoding unit to obtain the sample content vector, and the sample query data is input into the query encoding unit to obtain the sample query vector; Using the reconstruction unit, predictive reconstruction data is generated based on the sample content vector, and using the decision unit, reconstruction weights of the reconstruction unit are generated. Based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weight, the parameters of the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit are adjusted to obtain the trained content query model.
3. The method according to claim 2, characterized in that, The content encoding unit includes a transformation layer, a routing layer, a processing layer, and an output layer, and the processing layer includes multiple processing sub-units; The step of inputting the sample content into the content encoding unit to obtain the sample content vector includes: The sample content is processed through the transformation layer to obtain a content transformation vector; The routing layer processes the content transformation vector to obtain the activation weights corresponding to the plurality of processing sub-units. The target processing subunit processes the content transformation vector to obtain a content processing vector, wherein the target processing subunit selects the content processing vector from the plurality of processing subunits based on the activation weights. The output layer generates the sample content vector based on the content processing vector and the activation weights corresponding to the target processing subunit.
4. The method according to claim 2, characterized in that, The process of generating predicted reconstruction data based on the sample content vector using the reconstruction unit, and generating reconstruction weights for the reconstruction unit using the decision unit, includes: The reconstruction unit includes a summary reconstruction unit; using the summary reconstruction unit, a predicted content summary is generated based on the sample content vector, and using the decision unit, the summary reconstruction weights of the summary reconstruction unit are generated; and / or, The reconstruction unit includes an image reconstruction unit; using the image reconstruction unit, a predicted content image is generated based on the sample content vector, and using the decision unit, the image reconstruction weights of the image reconstruction unit are generated.
5. The method according to claim 2, characterized in that, The step of generating the reconstruction weights for the reconstruction unit using the decision unit includes: Using the decision unit, based on the sample content vector, the reconstruction weights of the reconstruction unit are generated; and / or, Using the decision unit, the reconstruction weights of the reconstruction unit are generated based on the auxiliary processing vector, wherein the auxiliary processing vector is the intermediate layer vector of the target processing subunit in the content encoding unit.
6. The method according to claim 2, characterized in that, The step of adjusting the parameters of the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weights to obtain a trained content query model includes: Calculate the contrast loss based on the sample content vector and the sample query vector; Based on the reconstruction weights, the sample reconstruction data, and the predicted reconstruction data, the reconstruction loss is calculated. Based on the contrast loss, the parameters of the query encoding unit are adjusted, and based on the contrast loss and the reconstruction loss, the parameters of the content encoding unit, the reconstruction unit, and the decision unit are adjusted respectively to obtain the trained content query model.
7. The method according to any one of claims 2 to 6, characterized in that, After adjusting the parameters of the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weights to obtain the trained content query model, the method further includes: The content encoding unit is distilled and compressed to obtain the target content encoding unit; The step of inputting the content to be queried into the content encoding unit in the content query model to obtain the content vector includes: The content to be queried is input into the target content encoding unit to obtain the content vector.
8. The method according to claim 7, characterized in that, The step of distilling and compressing the content encoding unit to obtain the target content encoding unit includes: Obtain auxiliary content and auxiliary query data; The auxiliary content is input into the content encoding unit to obtain an auxiliary content vector, and the auxiliary query data is input into the query encoding unit to obtain an auxiliary query vector; The auxiliary content is input into the dense coding unit to obtain a dense content vector, wherein the dense coding unit has the same structure as the content coding unit, and the number of parameters of the dense coding unit is less than the number of parameters of the content coding unit. Based on the auxiliary content vector, the auxiliary query vector, and the dense content vector, the parameters of the dense encoding unit are adjusted to obtain the target content encoding unit.
9. A method for training a content query model, characterized in that, include: Obtain sample data, wherein the sample data includes sample content, sample query data, and sample reconstruction data; The sample content is input into the content encoding unit in the content query model to obtain the sample content vector, and the sample query data is input into the query encoding unit in the content query model to obtain the sample query vector; Using the reconstruction unit in the content query model, predictive reconstruction data is generated based on the sample content vector, and using the decision unit in the content query model, the reconstruction weight of the reconstruction unit is generated. Based on the sample content vector, the sample query vector, the sample reconstruction data, the predicted reconstruction data, and the reconstruction weight, the parameters of the content encoding unit, the query encoding unit, the reconstruction unit, and the decision unit are adjusted to obtain the trained content query model.
10. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that, It stores a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.