Method for training a recall model, recall method, related device and computer program

The method enhances multimedia recall accuracy by pre-training and fine-tuning a recall model with text pairs and feedback data, addressing low accuracy issues in existing systems to provide precise multimedia prediction.

JP2025531799AActive Publication Date: 2025-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025514203
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-05-11
Filing Date
2024-03-25
Publication Date
2025-09-25
Estimated Expiration
2044-03-25

AI Technical Summary

Technical Problem

Existing multimedia recall systems suffer from low accuracy, leading to multiple interactions between the user and the multimedia server due to irrelevant recall results, increasing search time and user frustration.

Method used

A method involving pre-training and fine-tuning a recall model using text pairs that associate multimedia resource identifiers with description information, followed by fine-tuning with feedback data to enhance the model's ability to predict relevant multimedia.

Benefits of technology

Improves multimedia recall accuracy by learning the association between resource identifiers and description information, enabling precise prediction of related multimedia, thus reducing unnecessary interactions and enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025531799000001_ABST
    Figure 2025531799000001_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology and discloses a method for training a recall model, a recall method, and a related device. The training method includes the steps of: obtaining a plurality of first text pairs including a first question text and a first answer text; pre-training a recall model according to the first question text and the first answer text among the plurality of first text pairs; obtaining a plurality of second text pairs including a second question text questioning a resource identifier of related multimedia corresponding to the multimedia and the second answer text being a resource identifier of related multimedia for the question of the second question text; and fine-tuning the pre-trained recall model according to the second question text and the second answer text among the plurality of second text pairs. The trained recall model can accurately recall the related multimedia corresponding to the multimedia, thereby improving the recall accuracy of the multimedia.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to Chinese Patent Application No. 2023105250958, entitled "Method for training a recall model, a recall method and related devices," filed with the State Intellectual Property Office of the People's Republic of China on May 11, 2023, the entire contents of which are incorporated herein by reference.

[0002] This application relates to the field of artificial intelligence, and more particularly to training recall models and recall for multimedia. [Background technology]

[0003] With the development of multimedia technology, the number of multimedia such as audio and video is rapidly increasing. Accurately recalling multimedia from a large amount of multimedia can effectively reduce the time it takes for users to search for multimedia. If the recalled multimedia is not highly relevant to the multimedia the user actually wants, i.e., if the multimedia recall accuracy is not high, the user will need to search multiple times, resulting in multiple interactions between the terminal accessed by the user and the multimedia server. Therefore, how to improve the accuracy of multimedia recall has become a technical issue that related technologies must urgently address. Summary of the Invention [Means for solving the problem]

[0004] In view of the above problems, embodiments of the present application provide a method for training a recall model, a recall method and related devices to improve the recall accuracy of multimedia.

[0005] According to one aspect of an embodiment of the present application, there is provided a method for training a recall model, the method including: acquiring a plurality of first text pairs, the first text pairs including a first question text generated according to description information of multimedia and targeting a resource identifier of the multimedia as a question, and a first answer text that is a resource identifier for a question of the first question text; pre-training a recall model according to the first question text and the first answer text among the plurality of first text pairs; acquiring a plurality of second text pairs, the second text pair including a second question text targeting a resource identifier of related multimedia as a question, and the second answer text that is a resource identifier for related multimedia for a question of the second question text, wherein the related multimedia for the second text pair corresponds to the multimedia for the first text pair; and fine-tuning the pre-trained recall model according to the second question text and the second answer text among the plurality of second text pairs.

[0006] According to one aspect of an embodiment of the present application, there is provided a recall method, including the steps of: obtaining a resource identifier of a target multimedia; generating a target question text according to the resource identifier of the target multimedia, with related resource identifiers as the query target, wherein the related resource identifiers refer to the resource identifiers of related multimedia corresponding to the target multimedia; generating a target answer text corresponding to the target question text according to the target question text using a recall model obtained by training according to the above-mentioned method for training a recall model, wherein the target answer text includes resource identifiers of related multimedia corresponding to the target multimedia; and determining a recall result of the target multimedia according to the resource identifiers in the target answer text.

[0007] According to one aspect of the embodiment of the present application, there is provided a first obtaining module configured to perform a step of obtaining a plurality of first text pairs, the first text pairs including a first question text generated according to description information of multimedia and targeting a resource identifier of the multimedia as a question, and a first answer text that is a resource identifier for a question of the first question text; a pre-training module configured to perform a step of pre-training a recall model according to the first question text and the first answer text of the plurality of first text pairs; and a step of obtaining a plurality of second text pairs, the second text pairs including a first question text and a first answer text that is a resource identifier for a question of the first question text. is provided an apparatus for training a recall model, the apparatus including: a second acquisition module configured to perform a step of: including a second question text questioning a resource identifier of related multimedia; and a second answer text which is a resource identifier of related multimedia for a question of the second question text, wherein the related multimedia for the second text pair and the multimedia for the first text pair correspond to each other; and a fine-tuning module configured to perform a step of fine-tuning a pre-trained recall model according to the second question text and the second answer text of a plurality of second text pairs.

[0008] According to one aspect of an embodiment of the present application, there is provided a recall device, including: a third acquisition module configured to perform a step of acquiring a resource identifier of a target multimedia; a target question text generation module configured to perform a step of generating a target question text, in accordance with the resource identifier of the target multimedia, with related resource identifiers as the query target, wherein the related resource identifiers refer to resource identifiers of related multimedia corresponding to the target multimedia; a target answer text determination module configured to perform a step of generating a target answer text corresponding to the target question text according to the target question text using a recall model obtained by training according to the above-mentioned method for training a recall model, wherein the target answer text includes resource identifiers of related multimedia corresponding to the target multimedia; and a recall result determination module configured to perform a step of determining a recall result of the target multimedia according to the resource identifiers in the target answer text.

[0009] According to one aspect of an embodiment of the present application, there is provided an electronic device including a processor and a memory in which a computer program is stored, the electronic device realizing the above-described method for training a recall model or the above-described recall method when the computer program is executed by the processor.

[0010] According to one aspect of an embodiment of the present application, there is provided a computer-readable storage medium having a computer program stored thereon, the computer program being configured to, when executed by a processor, implement the above-described method for training a recall model or the above-described recall method.

[0011] According to one aspect of an embodiment of the present application, there is provided a computer program product including a computer program that, when executed by a processor, causes the method for training a recall model as described above or the recall method as described above to be implemented.

[0012] In this application, a recall model is first pre-trained using a plurality of first text pairs. Among the first text pairs, a first question text is generated according to multimedia description information and targets a multimedia resource identifier, while a first answer text is a multimedia resource identifier. Therefore, pre-training allows the recall model to learn the association between the multimedia resource identifier and the multimedia description information, thereby determining a feature characterization of the multimedia resource identifier through the characteristics of the multimedia description information. After pre-training is complete, the pre-trained recall model is fine-tuned using a second question text targeting a related multimedia resource identifier corresponding to the multimedia, and a second answer text including the related multimedia resource identifier corresponding to the multimedia. In this way, the recall model can use the association relationship between the multimedia resource identifiers and the multimedia description information learned in the pre-training stage to learn the characteristic commonalities between the reference multimedia and the corresponding related multimedia based on the multimedia resource identifiers and the corresponding related multimedia resource identifiers in the fine-tuning stage, so that in the subsequent application process, the recall model can accurately predict the corresponding related multimedia resource identifiers according to the reference multimedia resource identifiers, and can also accurately recall the corresponding related multimedia, thereby effectively ensuring the correlation between the recalled related multimedia and the reference multimedia.In addition, the method of the present application can convert the multimedia recall task into a text generation task. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a schematic diagram illustrating an application scene of the present application shown in one embodiment of the present application; [Figure 2]1 is a flowchart of a method for training a recall model presented in one embodiment of the present application. [Figure 3] FIG. 1 is a schematic diagram illustrating the training of a recall model as presented in one embodiment of the present application. [Figure 4] FIG. 1 is a schematic diagram of a recall model according to an embodiment of the present application. [Figure 5] FIG. 1 is a schematic diagram illustrating encoding and decoding processes according to the BART model. [Figure 6] FIG. 1 is a schematic diagram illustrating a transformer model. [Figure 7] 2 is a flowchart of an embodiment of step 220 shown in an embodiment of the present application. [Figure 8] 2 is a flowchart of an embodiment of step 240 shown in an embodiment of the present application. [Figure 9] 1 is a flowchart of a recall method presented in one embodiment of the present application. [Figure 10] 4 is a flowchart of a recall method presented in another embodiment of the present application. [Figure 11] FIG. 1 is a block diagram of a recall model training device as presented in one embodiment of the present application. [Figure 12] FIG. 1 is a block diagram of a recall device shown in one embodiment of the present application. [Figure 13] FIG. 1 shows a schematic configuration diagram of a computer system suitable for realizing an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0014]

[0023] Exemplary embodiments will now be described in more detail with reference to the drawings. However, exemplary embodiments may be embodied in various forms and should not be construed as being limited to the examples set forth herein. Rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of exemplary embodiments to those skilled in the art.

[0015] It should be noted that "multiple" as referred to in this specification means two or more. "And / or" describes an association relationship between related objects and indicates that three relationships may exist. For example, A and / or B can indicate three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the related objects before and after it are in an "OR" relationship.

[0016] Before specifically describing the method of the present application, the terms used in the present application will be explained as follows.

[0017] A sequence-to-sequence (Seq2Seq) model is a neural network model that maps sequences to sequences. Sequence-to-sequence models are primarily used to improve machine translation techniques by mapping sentences (phrase sequences) in one language to corresponding sentences in another language.

[0018] Text generation refers to generating understandable text from non-verbal expressions. Depending on the classification of non-verbal expressions, text generation can be done in three ways: "text to text," "data to text," and "image to text."

[0019] Pre-training refers to using as much training data as possible to extract as many common features as possible from it, thereby reducing the learning burden of the model for a particular task.

[0020] The principle of fine tuning is to use the known model structure and known model parameters to modify the parameters of the output layer as the output layer of the current task, and fine-tune the parameters of several network layers above the lowest layer (i.e., the output layer), thereby effectively utilizing the strong generalization ability of deep neural networks and avoiding the need to design complex models and reduce training time.

[0021] The key to prompt learning is to convert the problem to be solved into the same format as the pre-trained task through a template. For example, for the text "I missed the bus today," we can construct a template "I missed the bus today. I felt so [MASK][MASK]" and use a Masked Language Model (MLM) to predict the sentiment word to identify the sentiment polarity. We can also construct a prefix "English: I missed the bus today. Chinese: [MASK][MASK]" and then use a generative model to obtain the corresponding Chinese translation.

[0022] The Transformer model is a deep learning model that uses self-attention, which allows different weights to be assigned to different parts of the input data depending on their importance. This model is primarily used in the fields of natural language processing (NLP) and computer vision (CV).

[0023] The attention mechanism is a solution to a problem proposed by imitating human attention. Simply put, it is about quickly sorting out high-value information from a large amount of information. It is mainly used to solve the problem of how difficult it is to obtain a reasonable vector representation when the input sequence of a long short-term memory (LSTM) model / recursive neural network (RNN) model is long. The main method is to retain the intermediate results of the LSTM, train a new model, and associate the intermediate results with the output to achieve the goal of filtering information.

[0024] The Bidirectional Encoder Representations from Transformers (BERT) model is a pre-trained language representation model that emphasizes generating deep bidirectional language representations by employing a novel masking language model, rather than pre-training using conventional unidirectional language models or by concatenating two unidirectional language models at a shallow level.

[0025] Generative Pre-trained Transformer (GPT) is an autoregressive language model that uses deep learning to generate natural language that is understandable to humans.

[0026] A Knowledge Graph is essentially a knowledge database with a directed graph structure, called a semantic network. A Knowledge Graph is a data structure consisting of entities, relationships, and attributes.

[0027] Overfitting refers to the difference between training error and test error being too large, i.e., the model complexity is higher than the actual problem, and the model can make accurate predictions on the training set but not on the test set.

[0028] Artificial intelligence (AI) refers to theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to imitate and extend human intelligence, thereby sensing the environment, acquiring knowledge, and utilizing that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science, aiming to understand the essence of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. In other words, AI is a technology that studies the design principles and implementation methods of various intelligent machines, endowing machines with sensing, reasoning, and decision-making capabilities. AI software technology mainly encompasses several areas, including computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. In the method of this application, natural language processing technology is used to convert a multimedia recall task into a text generation task, thereby improving the accuracy of multimedia recall.

[0029] 1 is a schematic diagram illustrating an application scenario of the present application shown in one embodiment of the present application. As shown in FIG. 1, the application scenario includes a server 120. The server 120 may be a physical server or a cloud server, but is not specifically limited thereto.

[0030] The server 120 can train a recall model according to the method for training a recall model of the present application, which includes step S1 of pre-training the recall model with a plurality of first text pairs and step S2 of fine-tuning the pre-trained recall model with a plurality of second text pairs. Based on the fine-tuned recall model, step S3 can be performed to invoke the recall model to perform a recall process for each multimedia in the multimedia library, i.e., the recall model can determine recall results corresponding to the multimedia according to the resource identifiers corresponding to the multimedia, and associate the resource identifiers of the multimedia with the recall results corresponding to the multimedia and store them in a recall dataset. The recall results corresponding to the multimedia indicate related multimedia related to the multimedia. In a specific embodiment, step S3 can be performed offline.

[0031] Based on the recall data set, the server 120 can provide a search and recall service for the terminal. In this case, the application scenario further includes the terminal 110, which may be a smartphone, tablet, laptop, desktop computer, in-vehicle terminal, smart TV, etc., but is not specifically limited here.

[0032] The terminal 110 is communicatively connected to the server 120 via a wired or wireless network, and the server 120 performs search recall according to the following procedure: In step S41, a multimedia search request is received. In step S42, matching with a second multimedia is performed. Specifically, the multimedia search request sent from the terminal 110 includes a search keyword, and the second multimedia matching the search keyword is determined by matching the search keyword with the description information of each multimedia in the multimedia library. In step S43, a recall result corresponding to the second multimedia is determined. Specifically, the recall result corresponding to the second multimedia is obtained from the recall dataset. In step S44, the second multimedia and the recall result corresponding to the second multimedia are transmitted.

[0033] In some embodiments, training the recall model, invoking the recall model to determine the recall result of each multimedia, and providing the search recall service for the terminal may be performed by the same electronic device, e.g., all performed by server 120, or may be performed by separate electronic devices, but are not specifically limited herein.

[0034] The detailed implementation of the technical methods of the embodiments of the present application will be described in detail below.

[0035] 2 is a flowchart of a method for training a recall model according to an embodiment of the present application. The method may be executed by an electronic device with processing capabilities, such as a server, but is not specifically limited thereto. As shown in FIG. 2, the method includes at least steps 210 to 240, which will be described in detail as follows:

[0036] Step 210: Obtain a plurality of first text pairs, each of which includes a first question text generated according to the description information of the multimedia and targeting a resource identifier of the multimedia as a question, and a first answer text that is a resource identifier for the question of the first question text.

[0037] The first text pair refers to a text pair for pre-training a recall model. The question text in the first text pair that represents a question is called the first question text, and the answer text in the first text pair that represents an answer to the question is called the first answer text.

[0038] The multimedia may be video (long video, medium video, etc.), audio, audio video, etc., such as drama, documentary, movie, variety, animation, cartoon, etc. The description information of the multimedia is used to indicate basic attributes of the multimedia. The basic attributes of the multimedia may be the multimedia name (drama name, documentary name, movie name, animation name, etc.), running time, language, screenwriter, director, profile, starring, genre, etc.

[0039] In some embodiments, the multimedia description information is stored in the form of a knowledge graph, that is, the multimedia description information can be represented as a directed graph. Taking the movie "Chen Xianling" as an example, the schema of the multimedia description information can be shown as follows:

[0040] { "channel": "movie", "alias": [ "Chen Xing-ling's Spinoff" ], "year": "2019", "area": ​​[ "inland" ], "language": [ "Standard language" ], "summary": "Near Mount Qishan is Fufeng City, once known as the 'City That Never Goes Dark'. The two finally solved the mystery, captured the mastermind, and brought peace to the people." "produce": [ "Guizhou XX Transmission Co., Ltd.", "Guangzhou YY Media Co., Ltd." ], "series_name": "Chen Xiling", "kgid": "kg_41753519", "serial_version": "0", "kgid_name": "Chen X Ling's Living Soul", "english_title": "The Living Dead", "publish_time_in_source": "2020-06-27 00:00:00", "season_num": "", "entity_type": "movie", "actors": [ { "id": "151110", "name": "YuXX", "type": "leading" }, { "id": "8253725", "name": "Zheng XX", "type": "leading" }, { "id": "1541022", "name": "King XX", "type": "leading" }.

[0041] As can be seen from the above, in the multimedia-compatible schema of the movie "Chen Xingling," different types of description information are represented by description fields (one description field is used to represent one basic attribute) and description field values. The description fields are, for example, the description field "year" representing the running time, the description field "summary" representing the profile, the description field "actors" representing the actors, etc. In other embodiments, the multimedia description information may include more or fewer description fields than those described above, but this is not specifically limited here.

[0042] A resource identifier is a text string used to uniquely identify a piece of multimedia, and may be a Content Identification (CID) for the multimedia. The Content Identifier for the multimedia is obtained by cryptographically processing (e.g., hashing) the content of the multimedia. The Content Identifier is equivalent to a "content fingerprint" for the multimedia.

[0043] The first question text is directed to a resource identifier of a multimedia resource, i.e., the first question text is posed to obtain a resource identifier of a multimedia resource. For example, the first question text may be "What is the CID of XX?", and "XX" in the first question text may be at least one qualifier for limiting the multimedia resource, for example, "XX" may be a basic attribute such as the name of the multimedia resource, starring roles, running time, screenwriter, director, genre, etc.

[0044] In some embodiments, the first question text and the first answer text in each first text pair may be determined according to the following steps A1 to A3.

[0045] Step A1: Obtain description information of multimedia and resource identifiers corresponding to the multimedia.

[0046] Step A2: Generate a first question text that targets the resource identifier of the multimedia as a question according to the value of at least one description field of the description information.

[0047] As described above, one description field in the description information of the multimedia is used to represent one basic attribute, and the value of one description field in the description information represents the attribute content of the multimedia in the corresponding attribute. In step A2, the multimedia is limited by the value of at least one description field in the description information, and a first question text is generated based on the limited multimedia, with the resource identifier of the multimedia as the query target. For example, the first question text may be "What is the CID corresponding to the movie "Chen X-ling's Extra Edition?" In this first question text, "Chen X-ling's Extra Edition" is the value of the description field representing the movie name.

[0048] In some embodiments, step A2 includes the following steps B1 to B3, which are described in detail below.

[0049] Step B1: Obtain a first question template that queries the resource identifier, where the first question template indicates at least one description field.

[0050] In this application, a question template for generating a first question text is referred to as a first question template, and this first question template is preset. The number of first question templates is not limited, and multiple first question templates may be used to ensure the richness of the generated first question text.

[0051] The first question template indicates a common question content in the first question text, such as "What is the corresponding CID?", and further indicates at least one description field for limiting the multimedia. For example, the description field for limiting the multimedia indicated in the first question template may be a description field indicating the name of the multimedia, a description field indicating the starring role of the multimedia, a description field indicating the director of the multimedia, etc. The common question content in different first question templates may be the same, but the indicated description fields may be different, so that multiple first question texts can be generated using different first question templates based on the description information of the same multimedia.

[0052] Step B2: Obtain the value of each description field indicated in the first question template from the description information of the multimedia.

[0053] Based on the first question template, the value of each description field indicated by the first question template can be obtained from the description information of each multimedia in the multimedia library. For example, if the description fields indicated by the first question template include a description field indicating a starring role, the value of the description field indicating a starring role is obtained from the description information of the multimedia, and this value represents the starring role of the multimedia.

[0054] Step B3: Combine the obtained description field value with the first question template to obtain a first question text.

[0055] In step B3, the value of the acquired description field is embedded in the position of the description field in the first question template to obtain a corresponding first question text.

[0056] For example, the first question template may be "What is the CID corresponding to [description field II] starring [description field I]?", where description field I refers to the description field representing the star, and description field II refers to the description field representing the multimedia name. Based on this first question template, a first question text such as "What is the CID corresponding to YYYY starring Liu XX?" can be generated.

[0057] For example, the first question template may be, "What is the CID corresponding to [description field II] directed by [description field III]?" Here, description field III refers to the description field representing the director. Based on this first question template, a first question text such as "What is the CID corresponding to ZZZ directed by Zhang XX?" can be generated.

[0058] It should be noted that the first question templates listed above are merely examples and are not considered to limit the scope of use of the present application. In specific embodiments, more first question templates may be set to enrich the format of the first question text.

[0059] In this way, the first question text can be efficiently generated based on multimodal description information using the values ​​of the description fields indicated by the first question template. In addition, since there is more than one first question template, the richness of the generated first question text can be further ensured, thereby improving the quality of pre-training for the recall model.

[0060] Step A3: The multimedia resource identifier is set as a first answer text corresponding to the first question text.

[0061] For example, for a first question text "What is the CID corresponding to Director Zhang XX's ZZZ?", the first answer text corresponding to this first question text is a resource identifier corresponding to the multimedia "Director Zhang XX's ZZZ", for example, klv6811ljzbhs8k.

[0062] For example, for the first question text "What is the CID for YYYY starring Liu XX?", the corresponding first answer text is mzc0020035l5vcf, where "mzc0020035l5vcf" is the resource identifier for the multimedia YYYY starring Liu XX.

[0063] In some embodiments, when the first question text has different numbers of question targets, the first question template may be divided into a preferred scene template and an intentional scene template. Here, the preferred scene template targets one question target, i.e., targets one multimedia resource identifier, while the intentional scene template targets at least two multimedia resource identifiers. It may also be understood that the preferred scene template defines one multimedia resource in its description field, while the intentional scene template defines at least two multimedia resource identifiers in its description field.

[0064] For example, in the above example, the first question template corresponding to the two first question texts, "What is the CID corresponding to YYYY starring Liu XX?" and "What is the CID corresponding to ZZZ directed by Zhang XX?", can be regarded as the preferred scene template.

[0065] The intentional scene template may be, for example, “What is the CID corresponding to the [description field IV] movie series?”, where description field IV is a description field representing the series name of the series to which the multimedia belongs. Based on this intentional scene template, a first question text such as “What is the CID corresponding to the Chen Xiling movie series?” can be generated, and the first answer text corresponding to this first question text is mzc00200lpxf8uq.

[0066] For each multimedia in the multimedia library, the description information of the multimedia can be used to generate a plurality of first text pairs for each multimedia according to the above procedure.

[0067] In this way, since the description fields can describe the description information in different dimensions, the value of at least one description field of the description information can accurately represent the multimedia, thereby further increasing the relevance between the constructed first question text and the corresponding multimedia, and improving the generation quality of the first question text.

[0068] Step 220: Pre-train a recall model according to a first question text and a first answer text of the plurality of first text pairs.

[0069] The recall model is a sequence-to-sequence model constructed through a neural network, i.e., the recall model can map an input sequence to an output sequence. In this application, during the pre-training process, the input sequence of the recall model is the first question text, and the output sequence is the answer text indicating the multimedia resource identifiers.

[0070] During the pre-training process, a first question text is input to a recall model, and the recall model performs a semantic coding process on the first question text. Then, a decoding process is performed according to the semantic coding result to output a predicted multimedia resource identifier. Next, a first loss is calculated according to the predicted output multimedia resource identifier and the resource identifier in the first answer text corresponding to the first question text, and the weighting parameters of the recall model are reverse-adjusted according to the first loss.

[0071] In some embodiments, a termination condition for pre-training can be set in advance. The termination condition for pre-training may be, for example, that the number of iterations of pre-training reaches a first threshold, or that the loss function in the pre-training stage has converged, but is not limited thereto. If it is determined that the termination condition for pre-training has been reached during the pre-training process, the pre-training is stopped.

[0072] Since the first question text is generated according to the multimedia description information, the first answer text is a multimedia resource identifier, and the first question text and the first answer text can be used to pre-train a recall model, allowing the recall model to learn the feature representation of the multimedia resource identifier, i.e., to learn the association relationship between the multimedia resource identifier and the multimedia description information, so that the recall model can perceive and memorize the multimedia corresponding description information of different resource identifiers, and thus the multimedia resource identifier can be described by the multimedia description information.

[0073] Step 230: Obtain a plurality of second text pairs, each of which includes a second question text that queries a resource identifier of related multimedia corresponding to the multimedia, and a second answer text that is a resource identifier of related multimedia for the question of the second question text.

[0074] The second text pair is a text pair for fine-tuning the recall model, and the question text representing the question in the second text pair is called the second question text, and the answer text representing the answer in the second text pair is called the second answer text.

[0075] Related multimedia corresponding to one multimedia A refers to multimedia that is related to that multimedia A in one or more dimensions, such as multimedia that is highly similar to that multimedia A (e.g., high content similarity, the same type, etc.), or multimedia that is related to the content of that multimedia A, or multimedia that is of high interest to users along with multimedia A, or multimedia that is likely to attract users' attention along with multimedia A.

[0076] The second question text queries about resource identifiers of related multimedia corresponding to the multimedia, i.e., the second question text is posed to obtain resource identifiers of related multimedia corresponding to the multimedia. In some embodiments, the second question text indicates reference multimedia. For example, if the second question text queries about resource identifiers of related multimedia corresponding to multimedia A, the reference multimedia is multimedia A. In some embodiments, the reference multimedia can be indicated by the resource identifier of the reference multimedia in the second question text. For example, the second question text may be, "Which CID do users who search for CID mzc002000mqs1cp also tend to click on?" The resource identifier "mzc002000mqs1cp" in this second question text is used to narrow down the reference multimedia, and the second question text is posed to obtain resource identifiers of related multimedia corresponding to the multimedia with the resource identifier "mzc002000mqs1cp."

[0077] In some embodiments, the second question text and the second answer text in the second text pair can be determined according to the following steps C1 to C3, which are described in more detail below.

[0078] Step C1: Obtain multimedia feedback data indicating at least two multimedia whose feedback operations have been triggered within a set time.

[0079] The feedback operation may be an interactive operation on the multimedia when the multimedia is displayed, such as a click operation, a like operation, a favorite operation, a forward operation, etc., but is not specifically limited thereto. In some embodiments, after displaying a thumbnail of the multimedia on the user interface, a client operation log indicating the user's feedback operation on the multimedia thumbnail is collected, and then at least two multimedia items that triggered a feedback operation within a set time period can be determined according to the operation log within a certain period. However, if it is necessary to collect the client operation log on the multimedia thumbnail, user permission or consent must be obtained for any item, and the collection, handling, and processing of the operation log must comply with the relevant laws, regulations, and standards of the relevant country and region.

[0080] In some embodiments, in the scene of a multimedia searched by a user, based on the matching searched multimedia (for ease of distinction, the searched multimedia is referred to as a third multimedia), at least one fourth multimedia highly similar to the third multimedia is further matched from the multimedia library, and the third multimedia and at least one fourth multimedia are pushed to the search originator, so that the third multimedia and at least one fourth multimedia can be displayed on the search result display page. Then, from the operation log collected from the client, other multimedia clicked by the user within a set time after clicking the third multimedia can be determined. In the search scene, the third multimedia matching the search term is what the user pays the most attention to and is interested in, but when the user clicks on the fourth multimedia displayed on the search result display page, it can be seen that the user is also paying attention to the fourth multimedia that triggered the click operation. In this case, the multimedia feedback data is a log of multiple click operations on the search result display page, and the multimedia feedback data reflects at least two multimedia that triggered click operations within the set time.

[0081] Step C2: Generate a second query text according to a resource identifier corresponding to a first multimedia among the at least two multimedia, the second query text targeting the resource identifier of the related multimedia corresponding to the first multimedia, where the related multimedia corresponding to the first multimedia includes at least one multimedia among the at least two multimedia excluding the first multimedia.

[0082] In other words, in step C2, one of the at least two multimedia for which a feedback operation is triggered within the set time indicated by the multimedia feedback data is regarded as a first multimedia, and the other multimedia among the at least two multimedia, excluding the first multimedia, are regarded as related multimedia corresponding to the first multimedia.

[0083] In the second question text, a resource identifier corresponding to the first multimedia is used to characterize or limit the first multimedia, and based on this, a query is made to a resource identifier of a related multimedia corresponding to the first multimedia. For example, the second question text may be "What is the CID of the multimedia that would also interest a user who is interested in CID XX?", where "XX" in the second question text refers to the CID of the first multimedia.

[0084] In a search scenario, that is, when the multimedia feedback data is a log of multiple click operations on a search result page, the third multimedia searched for matching can be the first multimedia, and correspondingly, the fourth multimedia clicked indicated by the click operation log can be the related multimedia of the third multimedia.

[0085] In some embodiments, step C1 includes steps D1 and D2, which are described in more detail below.

[0086] Step D1: Obtain a second question template that targets the resource identifier of the related multimedia as a question target.

[0087] The second question template refers to a question template set for the second question text. This second question template may be set in advance, and there may be one or more. Similarly, this second question template indicates a common question content in the second question text. That is, when multiple second question texts are generated according to the same second question template, the common question content among the multiple second question texts is the same. The second question template also indicates the location of a resource identifier in the second question text of the reference multimedia. When the second question text queries the resource identifier of related multimedia corresponding to multimedia A, the reference is the resource identifier of multimedia A.

[0088] The second question template may be "What multimedia CIDs would also interest users who are interested in CID XX?" Alternatively, for example, the second question template may be "What CIDs do users who search for CID:XX tend to click on?" Alternatively, for example, the second question template may be "What CIDs do users who have focused on CID:XX also pay attention to?" The position of "XX" in the second question template shown above is the position of the referenced multimedia resource identifier. Of course, the above is merely an example of the second question template, and is not considered to limit the scope of use of the present application.

[0089] Step D2: Combine a resource identifier corresponding to a first multimedia of the at least two multimedia with a second question template to obtain a second question text.

[0090] As described above, the second question template indicates the location of the resource identifier in the second question text of the referenced multimedia, so the second question text can be obtained by embedding the resource identifier corresponding to the first multimedia as a reference in the corresponding location in the second question template.

[0091] In this way, the second question template can be used to quickly and conveniently generate second question texts, and since there is more than one second question template, the richness of the generated second question texts can be further ensured, and the quality of pre-training for the recall model can be improved.

[0092] Step C3: The resource identifier of the related multimedia corresponding to the first multimedia is set as the second answer text corresponding to the second question text.

[0093] A first multimedia can be associated and determined according to the multimedia feedback data, and at least one other multimedia, excluding the first multimedia, among the at least two multimedia for which a feedback operation indicated by the multimedia feedback data has been triggered is set as the associated multimedia of the first multimedia, and a resource identifier of the associated multimedia corresponding to the first multimedia can be associated and determined based on the stored resource identifier of each multimedia, so that a second answer text corresponding to the second question text can be associated and determined.

[0094] According to the above steps C1 to C3, a second question-and-answer pair can be determined, which includes the second question text "What CID do users who searched for cid:mzc002000mqs1 cp also tend to click on?" and the second answer text "mzc0020028aguo0." "mzc002000mqs1cp" in the second question text is the resource identifier of the reference multimedia (i.e., the resource identifier of the first multimedia), and "mzc0020028aguo0" in the second answer text is the resource identifier of the related multimedia corresponding to the first multimedia.

[0095] The second text pair serves to enable the recall model to learn the association knowledge between different multimedia during the training process, and at least two of the multimedia feedback data can appropriately indicate the association between these multimedia based on the feedback operation within a set time, for example, the association between the first multimedia and other multimedia. Therefore, the second text pair determined in this manner can effectively improve the training quality of the recall model.

[0096] In some embodiments, the second text pairs are determined using multimedia feedback data over multiple time periods (e.g., multiple time periods over the past 30 days) to ensure a sufficient number of second text pairs, which provides sufficient training samples for the fine-tuning stage.

[0097] Step 240: Fine-tune the pre-trained recall model according to the second question text and the second answer text of the plurality of second text pairs.

[0098] In the fine-tuning process, the input sequence of the pre-trained recall model is the second question text, and the output sequence is an answer text indicating resource identifiers of related multimedia corresponding to the multimedia. In the fine-tuning process, the second question text is input to the pre-trained recall model, and the pre-trained recall model performs a semantic coding process on the second question text, then performs a decoding process according to the semantic coding result, and outputs predicted resource identifiers of the multimedia. Next, a second loss is calculated according to the predicted output resource identifiers of the related multimedia and the resource identifiers in the second answer text corresponding to the second question text, and the weighting parameters of the recall model are further adjusted inversely according to the second loss.

[0099] After pre-training, the recall model learns the association relationship between the multimedia resource identifier and the multimedia description information, and uses the features corresponding to the multimedia description information to construct a feature representation of the multimedia resource identifier. Then, a second question text targeting the resource identifier of the associated multimedia corresponding to the multimedia and a second answer text including the resource identifier of the associated multimedia corresponding to the multimedia are used to fine-tune the pre-trained recall model. In this way, the recall model can use the association relationship between the multimedia resource identifier and the multimedia description information learned in the pre-training stage to learn the characteristic commonalities between the reference multimedia and the associated multimedia corresponding to the multimedia based on the resource identifier of the multimedia and the resource identifier of the associated multimedia corresponding to the multimedia in the fine-tuning stage. Therefore, in the subsequent application process, the recall model can accurately recall the associated multimedia corresponding to the multimedia according to the resource identifier of the reference multimedia, i.e., accurately predict the resource identifier of the associated multimedia corresponding to the multimedia.

[0100] In some embodiments, after pre-training the recall model, in order to shorten the training time and improve the training efficiency, the weighting parameters of some network layers in the recall model can be reverse-tuned according to the second loss during the fine-tuning process. Specifically, since the parameters of the network layers closest to the output of the recall model are directly related to the recall task of the recall model, the weighting parameters of the lowest network layer (i.e., the output layer) and the network layers above the lowest network layer can be reverse-tuned according to the second loss during the fine-tuning phase. In this way, it takes less time to adjust the weighting parameters of some network layers in the recall model than to adjust the weighting parameters of all network layers of the recall model, thereby reducing the time required for fine-tuning. In this way, reverse-tuning the weighting parameters of the lowest network layer (i.e., the output layer) and the network layers above the lowest network layer in the recall model not only utilizes the powerful generalization ability of deep neural networks, but also avoids the design of complex models and long training times.

[0101] In some embodiments, a termination condition for fine tuning can be set in advance. The termination condition for fine tuning may be, for example, that the number of iterations of fine tuning reaches a second threshold, or that the loss function in the fine tuning stage has converged, but is not limited thereto. If it is determined during the fine tuning process that the termination condition for fine tuning has been reached, the fine tuning is stopped.

[0102] Figure 3 is a schematic diagram illustrating the training of a recall model according to an embodiment of the present application. Figure 3 illustrates two first text pairs and two second text pairs. In one first text pair, the left text is the first question text, and the right text is the first answer text. The two texts in the same dashed box (or in the same column) belong to the same text pair (i.e., either the first text pair or the second text pair). The above description can be used to refer to the pre-training of the recall model using the first text pair and the fine-tuning of the recall model after pre-training using the second text, and therefore, further description will not be provided here.

[0103] In this application, a recall model is first pre-trained using multiple first text pairs. The first question text of the first text pair is generated according to multimedia description information, and the question is directed to a multimedia resource identifier. The first answer text is also a multimedia resource identifier. Therefore, pre-training the recall model using a priori knowledge of multimedia is equivalent to having the recall model perceive and memorize the multimedia description information. Pre-training allows the recall model to learn the association between the multimedia resource identifier and the multimedia description information, thereby determining a feature representation of the multimedia resource identifier through the characteristics of the multimedia description information. After pre-training is completed, fine-tuning is performed on the pre-trained recall model using a second question text that queries the resource identifier of the associated multimedia corresponding to the multimedia, and a second answer text that includes the resource identifier of the associated multimedia corresponding to the multimedia. In this way, the recall model can use the association between the multimedia resource identifier and the multimedia description information learned in the pre-training stage to learn characteristic commonalities between the reference multimedia and the associated multimedia corresponding to the multimedia in the fine-tuning stage based on the multimedia resource identifier and the associated multimedia resource identifier. In the subsequent application process, the recall model can accurately recall the relevant multimedia corresponding to the multimedia according to the resource identifier of the reference multimedia, that is, can accurately predict the resource identifier of the relevant multimedia corresponding to the multimedia.In addition, the method of this application can convert the multimedia recall task into a text generation task, so that there is no need to process the video frames or audio frames in the multimedia, which can simplify the multimedia recall task and improve recall efficiency.

[0104] In some embodiments, the multimedia data table can be adjusted to facilitate the construction of the first and second text pairs. Because the original multimedia data table includes multimedia values ​​for each description field but does not include resource identifiers for the multimedia, resource identifiers for the multimedia can be added to the multimedia data table before pre-training, and the information in the multimedia data table can be periodically fully updated according to the occurrence of multimedia in the multimedia library, e.g., when multimedia is added or removed from the multimedia library. The first and second text pairs can then be constructed based on the data in the multimedia data table. After pre-training the recall model using the first text pairs, the recall model can learn feature representations for each resource identifier in the multimedia data table, where the feature representations for the resource identifiers are embodied by features of the descriptive information for the multimedia corresponding to the resource identifier. However, if the number of multimedia in the multimedia library is large, the number of resource identifiers in the multimedia data table will also be large. As such, the training and online estimation speeds may be slow, and cold booting may be difficult. In practice, however, it has been found that maintaining the number of resource identifiers at the millions level can meet the training time length and online estimation speed required when learning feature representations of resource identifiers.

[0105] FIG. 4 is a schematic diagram of a recall model according to an embodiment of the present application. As shown in FIG. 4, the recall model includes an encoder network 410 and a decoder network 420. Here, the encoder network 410 performs semantic coding on an input sequence and outputs a semantic coding sequence. The decoder network decodes the semantic coding sequence output from the encoder network to obtain an output sequence. Specifically, in the pre-training stage, the input sequence is a first question text, and the output sequence is a predicted multimedia resource identifier. In the fine-tuning stage, the input sequence is a second question text, and the output sequence is a predicted multimedia resource identifier.

[0106] In some embodiments, the recall model may be a BART (Bidirectional and Auto-Regressive Transformers) model. The BART model incorporates the bidirectional encoding characteristics of the BERT model and the left-to-right decoding characteristics of the GPT model and is built on a standard sequence-to-sequence transformer model. Therefore, the BART model is more suitable for text generation scenarios than the BERT model. Compared to the GPT model, the BART model also provides more bidirectional context information. Figure 5 is a schematic diagram illustrating the encoding and decoding process of the BART model. As shown in Figure 5, after an input sequence is input to an encoder network 410, the encoder network 410 bidirectionally encodes the input sequence and outputs a semantic coding sequence. Then, the decoder network 420 performs autoregressive decoding (i.e., unidirectional decoding from left to right) to obtain an output sequence. In the BART model, the input sequence of the encoder network does not need to be aligned with the output sequence of the decoder network; the input sequence of the encoder network can be preprocessed. Here, preprocessing means, for example, replacing characters at certain positions in the input sequence with mask symbols. For example, in the input sequence of Figure 5, the characters after "A" and the characters after "B" are both replaced with mask symbols.

[0107] The BART model employs an attention mechanism and a Transformer model structure. In the application scenario of this application, considering the millions of data volumes and resource consumption in the multimedia library, the encoder network in the recall model includes an encoder in a three-layer Transformer model, and the decoder network in the recall model includes a decoder in a three-layer Transformer model. Figure 6 is a schematic diagram of a Transformer model. As shown in Figure 6, the encoder in the Transformer model includes a multi-head attention layer, a first summation and normalization layer, a feedforward neural network layer, and a second summation and normalization layer. A residual connection is established between the multi-head attention layer and the first summation and normalization layer, and a residual connection is established between the input of the feedforward neural network layer and the second summation and normalization layer. The decoder in the Transformer model includes a masked multi-head attention layer, a third summation and normalization layer, a multi-head attention layer, a fourth summation and normalization layer, a feedforward neural network layer, and a fifth summation and normalization layer, where a residual connection is established between the input of the masked multi-head attention layer and the third summation and normalization layer, a residual connection is established between the input of the multi-head attention layer and the fourth summation and normalization layer, and a residual connection is established between the input of the feedforward neural network layer and the fifth summation and normalization layer.

[0108] In some embodiments, as shown in FIG. 7, step 220 includes the following steps 710 through 740, which are described in more detail below.

[0109] Step 710: Perform a semantic coding process on the first question text using the encoder network to obtain a first semantic coding sequence corresponding to the first question text.

[0110] Specifically, the encoder network performs semantic coding processing on the first question text based on an attention mechanism (e.g., a multi-head attention mechanism), thereby making full use of the context information in the first question text and ensuring the accuracy of the obtained first semantic coding sequence.

[0111] Step 720: Decode the first semantic coding sequence using a decoder network to obtain a predicted answer text corresponding to the first question text.

[0112] The predicted answer text decoded from the decoder network includes a multimedia resource identifier for the predicted question of the first question text.

[0113] Step 730: Calculate a first loss according to the predicted answer text corresponding to the first question text and the corresponding first answer text.

[0114] The loss function of the recall model in the pre-training stage can be preset. For ease of distinction, the loss function set for the recall model in the pre-training stage is referred to as the first loss function. The first loss function may be a cross-entropy loss function, an absolute value loss function, a mean-variance loss function, or the like, but is not specifically limited here. Then, the predicted answer text corresponding to the first question text and the corresponding first answer text can be substituted into the first loss function to calculate the first loss. This first loss reflects the difference between the predicted answer text corresponding to the first question text and the first answer text corresponding to the first question text.

[0115] In one specific embodiment, the first loss function may be a cross-entropy loss function, which is shown in Equation 1 below.

[0116]

number

[0117]

number

[0118] Step 740: Inversely adjust the weighting parameters of the encoder network and the decoder network according to the first loss.

[0119] In a particular embodiment, a gradient descent method may be used to adjust weighting parameters of the encoder network and the decoder network according to the first loss so as to minimize the first loss function.

[0120] For each first sample pair, pre-training of the recall model is repeated according to the procedure shown in steps 710 to 740 until the end condition of pre-training is reached.

[0121] In some embodiments, step 240 includes steps 810 through 840, as shown in FIG. 8, which are described in more detail below.

[0122] Step 810: Perform a semantic coding process on the second question text using the pre-trained encoder network to obtain a second semantic coding sequence corresponding to the second question text.

[0123] Step 820: Decode the second semantic coding sequence using a pre-trained decoder network to obtain a predicted answer text corresponding to the second question text.

[0124] The predicted answer text corresponding to the second question text includes a resource identifier of the relevant multimedia for the predicted second question text.

[0125] Step 830: Calculate a second loss according to the predicted answer text corresponding to the second question text and the corresponding second answer text.

[0126] Similarly, the second loss function of the recall model in the fine-tuning stage can be preset. The second loss function can be set according to actual needs, but is not specifically limited here. Then, the predicted answer text corresponding to the second question text and the corresponding second answer text are substituted into the second loss function to calculate a second loss. This second loss reflects the difference between the predicted answer text corresponding to the second question text and the corresponding second answer text.

[0127] In some embodiments, to mitigate the overfitting problem, the second loss function may be a label-smoothed cross-entropy loss function. The label-smoothed cross-entropy loss function is the same as Equation 1, except that q i is determined according to the following equation 3.

[0128]

number

number

[0129] Label smoothing is a normalization technique used to mitigate overfitting. Label smoothing cross-entropy loss functions smooth the probability distribution of true labels, preventing the model from overpredicting a single class during training and reducing the risk of overfitting. Specifically, label smoothing cross-entropy loss functions can be thought of as converting the probability distribution of true labels from a one-hot vector into a smoothed probability distribution. This smoothed probability distribution allows the recall model to focus on the distribution of data during training, rather than focusing too much on a particular category. This improves model robustness and reduces the risk of overfitting. Label smoothing cross-entropy loss functions also provide a certain normalization effect. Label smoothing cross-entropy loss functions smooth the probability distribution of true labels during training, thereby reducing model complexity and further reducing the risk of overfitting.

[0130] Step 840: Back-adjust the weighting parameters of some network layers in the pre-trained recall model according to the second loss.

[0131] In some embodiments, the weighting parameters of some network layers in the recall model can be adjusted using a gradient descent method according to the second loss to minimize the second loss function. In some embodiments, in step 840, only the weighting parameters of the decoder network in the recall model can be adjusted, or the weighting parameters of some network layers can be set before adjusting the output layer and the output layer in the decoder network. This can reduce the amount of parameter adjustment and shorten the training time of the recall model.

[0132] For each second sample pair, fine-tuning of the recall model is repeated according to the procedure shown in steps 810 to 840 until a fine-tuning termination condition is reached. After the fine-tuning is completed, the recall model can be applied to an online application to accurately recall related multimedia corresponding to multimedia based on the resource identifier of the multimedia.

[0133] 9 is a flowchart of a recall method according to an embodiment of the present application. The recall method is executed by an electronic device such as a server, and includes the following steps 910 to 940, as shown in FIG. 9, which will be described in detail below.

[0134] Step 910: Obtain the resource identifier of the target multimedia.

[0135] The target multimedia refers to the multimedia used to determine the recall results. In some embodiments, each multimedia in the multimedia library can be used as the target multimedia to determine the recall results for each multimedia in the multimedia library according to the method of the present application.

[0136] Step 920: Generate a target question text according to the resource identifier of the target multimedia, taking the related resource identifier as the question target, where the related resource identifier refers to the resource identifier of the related multimedia corresponding to the target multimedia.

[0137] In some embodiments, a target question text can be generated according to the second question template and the resource identifier of the target multimedia. Specifically, the resource identifier of the target multimedia can be embedded in the second question template at the position representing the resource identifier of the reference multimedia, thereby obtaining the target question text. The format of the second question template may be as described above, and will not be further described here.

[0138] Step 930: Using the recall model obtained by training the recall model training method according to any one of the above embodiments, generate a target answer text corresponding to the target question text according to the target question text, where the target answer text includes a resource identifier of related multimedia corresponding to the target multimedia.

[0139] In step 930, the target question text is input to an encoder network in the recall model, and the encoder network performs semantic coding on the target question text to obtain a semantic coding sequence corresponding to the target question text. Then, the decoder network in the recall model decodes the semantic coding sequence corresponding to the target question text to obtain a target answer text.

[0140] Step 940: Determine the recall result of the target multimedia according to the resource identifier in the target answer text.

[0141] According to the correspondence relationship between the multimedia and the resource identifier, the multimedia corresponding to the resource identifier in the target answer text can be determined. The multimedia corresponding to the resource identifier in the target answer text is the related multimedia corresponding to the target multimedia. In step 940, the related multimedia corresponding to the target multimedia determined based on the resource identifier in the target answer text can be used as the recall result of the target multimedia.

[0142] In this application, the recall task for multimedia is converted into a text generation task, that is, for the target multimedia to be recalled, a target question text is generated that queries the related resource identifiers according to the resource identifiers of the target multimedia, and then a trained recall model is invoked to generate a target answer text for the target question text according to the target question text. The target question text queries the related resource identifiers, and the target answer text generated by the trained recall model includes resource identifiers corresponding to the related multimedia corresponding to the target multimedia, so that the related multimedia corresponding to the target multimedia can be determined according to the resource identifiers corresponding to the related multimedia corresponding to the target multimedia, and the recall of the related multimedia corresponding to the target multimedia can be realized.

[0143] Because the recall model was obtained through pre-training using multiple first text pairs and fine-tuning using multiple second text pairs, the recall model can accurately learn the association between resource identifiers and multimedia description information, allowing the features of the multimedia description information to be used as feature representations of the resource identifiers. Thus, when the recall model is used for multimedia recall, by learning the association between resource identifiers and multimedia description information, it is possible to recall multimedia that is highly related to the feature representations of the resource identifiers (i.e., to the features of the multimedia description information). This ensures the association between the recalled multimedia and the reference target multimedia, improving the accuracy of multimedia recall.

[0144] In some embodiments, after step 940, the method further includes associating and storing the resource identifiers of the target multimedia with the recall results of the target multimedia in a recall dataset.

[0145] In some embodiments, as shown in FIG. 10, the method includes the following steps:

[0146] Step 1010: Obtain a multimedia search request including the search keyword.

[0147] Search keywords may be terms used to limit the multimedia to be searched, such as starring, multimedia name, director, script, genre, and the like.

[0148] Step 1020: Match the multimedia according to the search keyword to determine a second multimedia that matches the search keyword.

[0149] In some embodiments, multimedia matching can be performed based on a maintained multimedia data table, which includes at least the value of the multimedia in each description field. That is, the multimedia data table includes description information for each multimedia. Then, the search keyword can be matched with the description information of the multimedia in the multimedia data table to determine multimedia that match the search keyword, i.e., the second multimedia. However, the number of second multimedia determined to match the search keyword in the multimedia search request may be one or more.

[0150] Step 1030: Obtain a recall result corresponding to the second multimedia from the recall dataset.

[0151] The recall data set stores recall results corresponding to multiple multimedia, so after determining a first media, the recall result corresponding to the first multimedia can be obtained from the recall data set according to the resource identifier of the first multimedia.

[0152] In some embodiments, steps 910 to 940 and the procedure for storing the recall results of the multimedia in the recall dataset can be performed offline. In this way, in the process of providing a search and recall service online, it is not necessary to call the recall model to determine the recall results of the second multimedia only when the second multimedia is determined. Instead, by calling the recall model in advance offline to determine and store the recall results of each multimedia, the recall results of the multimedia can be read directly from the recall dataset in the process of providing a search and recall service online, thereby improving the online service efficiency of the search and recall service and shortening the response time.

[0153] In some other embodiments, if the server's computing power is sufficient to meet the response time length requirement, when matching and determining the second multimedia, the second multimedia can be used as the target multimedia, and then the recall result of the second multimedia can be determined according to the procedures of steps 920 to 940.

[0154] Step 1040: Send the second multimedia and the recall results corresponding to the second multimedia to the originator of the multimedia search request.

[0155] Based on the embodiment corresponding to Figure 10, by submitting a multimedia search request once, not only can second multimedia matched based on search keywords be returned to the sender of the multimedia search request, but also related multimedia corresponding to the second multimedia can be returned. This can avoid the need to submit a multimedia search request again when a user needs to search for multimedia similar to or of the same type as the second multimedia, thereby reducing the number of interactions between the terminal and the server and improving the user experience.

[0156] In some embodiments, in a multimedia search scenario, the user is notified in advance whether it is necessary to add related multimedia corresponding to the second multimedia to be searched to the search results, and if the user's permission or consent is obtained, when searching for multimedia according to the embodiment corresponding to Figure 10, the second multimedia and the recall results corresponding to the second multimedia can be sent to the sender of the multimedia search request. On the other hand, if the user does not agree to add related multimedia corresponding to the second multimedia to be searched to the search results, there is no need to send the recall results of the second multimedia to the sender of the multimedia search request, and only the second multimedia determined to be matched is sent.

[0157] The following introduces an apparatus embodiment of the present application, which can implement the method in the above embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the above method embodiment of the present application.

[0158] 11 is a block diagram of a recall model training apparatus according to an embodiment of the present application, which may be disposed in an electronic device to implement the recall model training method according to the present application. As shown in FIG. 11 , the recall model training apparatus includes: a first acquisition module 1110 configured to perform a step of acquiring a plurality of first text pairs, the first text pairs being generated according to multimedia description information, and including a first question text questioning a resource identifier of the multimedia and a first answer text being a resource identifier for the question of the first question text; a pre-training module 1120 configured to perform a step of pre-training a recall model according to the first question text and the first answer text of the plurality of first text pairs; a second acquisition module 1130 configured to perform a step of acquiring a plurality of second text pairs, the second question text questioning a resource identifier of related multimedia and the second answer text being a resource identifier for the question of the second question text, wherein the related multimedia for the second text pair corresponds to the multimedia for the first text pair; and a fine-tuning module 1140 configured to perform a step of fine-tuning the pre-trained recall model according to the second question text and the second answer text of the plurality of second text pairs.

[0159] In some embodiments, when the recall model includes an encoder network and a decoder network, the pre-training module 1120 includes: a first semantic coding unit configured to perform a semantic coding process on the first question text using the encoder network to obtain a first semantic coding sequence corresponding to the first question text; a first decoding unit configured to perform a decoding process on the first semantic coding sequence using the decoder network to obtain a predicted answer text corresponding to the first question text; a first loss determination unit configured to calculate a first loss according to the predicted answer text corresponding to the first question text and the corresponding first answer text; and a first adjustment unit configured to perform an inverse adjustment of weighting parameters of the encoder network and the decoder network according to the first loss.

[0160] In some embodiments, the fine-tuning module 1140 includes: a second semantic coding unit configured to perform a semantic coding process on the second question text using a pre-trained encoder network to obtain a second semantic coding sequence corresponding to the second question text; a second decoding unit configured to perform a decoding process on the second semantic coding sequence using a pre-trained decoder network to obtain a predicted answer text corresponding to the second question text; a second loss determination unit configured to calculate a second loss according to the predicted answer text corresponding to the second question text and the corresponding second answer text; and a second adjustment unit configured to perform an inverse adjustment of weighting parameters of some network layers in the pre-trained recall model according to the second loss.

[0161] In some embodiments, the recall model training apparatus further includes: a fourth acquisition module configured to perform a step of acquiring description information of the multimedia and a resource identifier corresponding to the multimedia; a first question text identification module configured to perform a step of generating a first question text that targets the multimedia resource identifier as a question according to a value of at least one description field of the description information; and a first answer text determination module configured to perform a step of determining the multimedia resource identifier as a first answer text corresponding to the first question text.

[0162] In some embodiments, the first question text determination module includes: a first question template acquisition unit configured to perform a step of acquiring a first question template, the first question template having the resource identifier as a query target and indicating at least one description field; a first acquisition unit configured to perform a step of acquiring values ​​of each description field indicated by the first question template from the multimedia description information; and a first combination unit configured to perform a step of combining the acquired values ​​of the description fields with the first question template to acquire a first question text.

[0163] In some embodiments, the recall model training device further includes: a second acquisition unit configured to perform the step of acquiring multimedia feedback data indicating at least two multimedia for which a feedback operation was triggered within a set time; a second question text determination unit configured to perform the step of generating a second question text according to a resource identifier corresponding to a first multimedia of the at least two multimedia, the second question text setting a resource identifier of a related multimedia corresponding to the first multimedia as a question target, wherein the related multimedia corresponding to the first multimedia includes at least one multimedia of the at least two multimedia excluding the first multimedia; and a second answer text determination unit configured to perform the step of setting the resource identifier of the related multimedia corresponding to the first multimedia as a second answer text corresponding to the second question text.

[0164] In some embodiments, the second question text determination unit includes: a second question template acquisition unit configured to perform a step of acquiring a second question template that targets a resource identifier of the related multimedia as a question target; and a second combination unit configured to perform a step of combining a resource identifier corresponding to a first multimedia of the at least two multimedia with the second question template to obtain a second question text.

[0165] 12 is a block diagram of a recall device according to an embodiment of the present application. This recall device may be disposed in an electronic device to implement the recall method of the present application. As shown in FIG. 12, the recall device includes: a third acquisition module 1210 configured to acquire a resource identifier of a target multimedia; a target question text generation module 1220 configured to generate a target question text according to the resource identifier of the target multimedia, the target question text targeting a related resource identifier as a query, where the related resource identifier refers to the resource identifier of the related multimedia corresponding to the target multimedia; a target answer text determination module 1230 configured to generate a target answer text corresponding to the target question text according to the target question text using a recall model obtained by training according to any one of the methods for training a recall model described above, the target answer text including the resource identifier of the related multimedia corresponding to the target multimedia; and a recall result determination module 1240 configured to determine a recall result of the target multimedia according to the resource identifier in the target answer text.

[0166] In some embodiments, the recall device further includes an association storage module configured to perform the step of associating and storing the resource identifier of the target multimedia and the recall result of the target multimedia in the recall dataset.

[0167] In some embodiments, the recall device further includes: a fifth acquisition module configured to perform a step of acquiring a multimedia search request including a search keyword; a matching module configured to perform a step of matching multimedia according to the search keyword and determining a second multimedia that matches the search keyword; a recall result acquisition module configured to perform a step of acquiring recall results corresponding to the second multimedia from the recall dataset; and a transmission module configured to perform a step of transmitting the second multimedia and the recall results corresponding to the second multimedia to the originator of the multimedia search request.

[0168] 13 is a schematic diagram of a computer system suitable for realizing an electronic device according to an embodiment of the present application. Note that the computer system 1300 of the electronic device shown in FIG. 13 is merely an example and does not limit the functionality and scope of use of the embodiment of the present application. This electronic device may be used to execute the method for training a recall model according to the present application, or may be used to execute the recall method according to the present application.

[0169] 13, a computer system 1300 includes a central processing unit (CPU) 1301 that can perform various appropriate operations and processes, such as executing the methods in the above-described embodiments, in accordance with a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage unit 1308 into a random access memory (RAM) 1303. The RAM 1303 also stores various programs and data necessary for the operation of the system. The CPU 1301, ROM 1302, and RAM 1303 are connected to one another via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0170] The I / O interface 1305 includes an input unit 1306 including a keyboard, a mouse, etc., an output unit 1307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc., a storage unit 1308 including a hard disk, etc., and a communication unit 1309 including a network interface card such as a LAN (Local Area Network) card and a modem. The communication unit 1309 performs communication processing via a network such as the Internet. A driver 1310 is also connected to the I / O interface 1305 as needed. Removable media 1311 such as a magnetic disk, optical disk, magneto-optical disk, and semiconductor memory are installed in the driver 1310 as needed, thereby facilitating the installation of computer programs read from the removable media 1311 in the storage unit 1308 as needed.

[0171] In particular, according to an embodiment of the present application, the procedures described above with reference to the flowcharts may be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the methods illustrated in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network via the communication unit 1309 and / or installed from removable media 1311. When the computer program is executed by the central processing unit (CPU) 1301, various functions defined in the system of the present application are performed.

[0172] It should be noted that the computer-readable medium described in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used in conjunction with or to instruct the execution of a system, apparatus, or device. On the other hand, in this application, a computer-readable signal medium may include a propagating data signal in baseband or as part of a carrier wave that carries computer-readable program code. Such propagating data signals may be in various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transmit a program that can be used in conjunction with, or that can instruct the execution of, a system, apparatus, or device. The program code contained in the computer-readable medium may be transmitted over any suitable medium, including, but not limited to, wirelessly, wired, etc., or any suitable combination of the above.

[0173] The flowcharts and block diagrams in the accompanying drawings illustrate possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, including one or more executable instructions for implementing a given logical function. It should be noted that, in some alternative embodiments, the functions noted in the blocks may occur in a different order from the order noted in the accompanying drawings. For example, two blocks shown in succession may actually be executed essentially in parallel or in the reverse order, as determined by the related functionality. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented in a dedicated hardware-based system that performs a given function or operation, or in a combination of dedicated hardware and computer instructions.

[0174] The units described in the embodiments of the present application may be implemented in the form of software or hardware, and the units described herein may be located in a processor, in which the names of these units do not necessarily limit the units themselves.

[0175] In another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments or may exist independently of the electronic device, carrying a computer program, the computer-readable storage medium storing instructions for which, when executed by a processor, realizes the method for training a recall model or the method for recall according to any one of the above embodiments.

[0176] According to one aspect of the present application, there is provided a computer program product including a computer program stored in a computer-readable storage medium, which causes a processor of a computing device to read the computer program from the computer-readable storage medium and execute the computer program, thereby causing the computing device to perform the method for training a recall model or the method for recall according to any one of the above embodiments.

[0177] It should be noted that although the above detailed description refers to several modules or units of a device for performing operations, such division is not mandatory. In fact, according to an embodiment of the present application, the features and functions of two or more of the modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied in multiple modules or units.

[0178] From the description of the above embodiments, it is easy for those skilled in the art to understand that the exemplary embodiments described herein can be realized by software or by combining necessary hardware with software. Therefore, the technical methods according to the embodiments of the present application can be embodied in the form of a software product. The software product can be stored on a non-volatile storage medium (which can be a CD-ROM, a USB disk, a mobile hard disk, etc.) or a network, which includes a plurality of instructions for causing a computing device (which can be a personal computer, a server, a touch terminal, a network device, etc.) to execute the method according to the embodiments of the present application.

[0179] Other embodiments of the present application will be readily apparent to those skilled in the art after considering this specification and practicing the embodiments disclosed herein. This application is intended to cover any modifications, uses, or adaptations of the present application in accordance with the general principles of the present application, including common knowledge or customary technical means known in the art but not disclosed herein.

[0180] It should be understood that the present application is not limited to the exact construction described above and illustrated in the accompanying drawings, and that various modifications and variations can be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims. [Explanation of symbols]

[0181] 110 Terminal 120 servers 410 Encoder Network 420 Decoder Network 1110 First Acquisition Module 1120 Pre-training Module 1130 Second Acquisition Module 1140 Fine Tuning Module 1210 Third Acquisition Module 1220 Target Question Text Generation Module 1230 Target Answer Text Determination Module 1240 Recall Outcome Determination Module 1300 Computer Systems 1301 CPU 1302 ROM 1303 RAM 1304 Bus 1305 I / O Interface 1306 Input section 1307 Output section 1308 Storage section 1309 Communications Department 1310 driver 1311 Removable Media

Claims

1. 1. A method for training a recall model executed by an electronic device, comprising: obtaining a plurality of first text pairs, each of the first text pairs including a first question text generated according to description information of the multimedia and targeting a resource identifier of the multimedia as a question, and a first answer text that is a resource identifier for a question of the first question text; pre-training a recall model according to a first question text and a first answer text of a plurality of first text pairs; obtaining a plurality of second text pairs, each of the second text pairs including a second question text that queries a resource identifier of related multimedia and a second answer text that is a resource identifier of related multimedia for a question of the second question text, wherein the related multimedia related to the second text pair corresponds to the multimedia related to the first text pair; fine-tuning the pre-trained recall model according to a second question text and a second answer text of the plurality of second text pairs; How to train a recall model, including:

2. If the recall model includes an encoder network and a decoder network, The step of pre-training a recall model according to a first question text and a first answer text of a plurality of first text pairs includes: performing a semantic coding process on the first question text using the encoder network to obtain a first semantic coding sequence corresponding to the first question text; decoding the first semantic coding sequence with the decoder network to obtain a predicted answer text corresponding to the first question text; calculating a first loss according to a predicted answer text corresponding to the first question text and a corresponding first answer text; inversely adjusting weighting parameters of the encoder network and the decoder network according to the first loss; 2. The method of training a recall model of claim 1, comprising:

3. The step of fine-tuning the pre-trained recall model according to a second question text and a second answer text of a plurality of second text pairs includes: performing a semantic coding process on the second question text using a pre-trained encoder network to obtain a second semantic coding sequence corresponding to the second question text; decoding the second semantic coding sequence with a pre-trained decoder network to obtain a predicted answer text corresponding to the second question text; calculating a second loss according to a predicted answer text corresponding to the second question text and the corresponding second answer text; Inversely adjusting weighting parameters of some network layers in the pre-trained recall model according to the second loss; 3. The method of training a recall model of claim 2, comprising:

4. Before the step of obtaining a plurality of first text pairs, further obtaining description information of multimedia and resource identifiers corresponding to the multimedia; generating the first question text targeting the multimedia resource identifier according to a value of at least one description field of the description information; determining the multimedia resource identifier as a first answer text corresponding to the first question text; 4. A method for training a recall model according to claim 1, comprising:

5. generating the first question text targeting the multimedia resource identifier in accordance with a value of at least one description field of the description information, obtaining a first question template, the first question template targeting a resource identifier and indicating at least one description field; obtaining values ​​of each description field indicated by the first question template from multimedia description information; combining the obtained descriptive field values ​​with the first question template to obtain the first question text; 5. The method of training a recall model of claim 4, comprising:

6. Before the step of obtaining a plurality of second text pairs, further obtaining multimedia feedback data indicating at least two multimedia for which a feedback operation has been triggered within a set time; generating a second question text according to a resource identifier corresponding to a first multimedia among the at least two multimedia, the second question text targeting a resource identifier of a related multimedia corresponding to the first multimedia as a question target, wherein the related multimedia corresponding to the first multimedia includes at least one multimedia among the at least two multimedia, excluding the first multimedia; a step of setting a resource identifier of related multimedia corresponding to the first multimedia as a second answer text corresponding to the second question text; 4. A method for training a recall model according to claim 1, comprising:

7. The step of generating a second question text according to a resource identifier corresponding to a first multimedia of the at least two multimedia, the second question text being directed to a resource identifier of a related multimedia corresponding to the first multimedia as a query target, comprises: obtaining a second question template that targets the resource identifiers of the related multimedia; combining a resource identifier corresponding to a first multimedia of the at least two multimedia with the second question template to obtain the second question text; 7. The method of training a recall model of claim 6, comprising:

8. 1. A recall method performed by an electronic device, comprising: obtaining a resource identifier of the target multimedia; generating a target question text according to the resource identifier of the target multimedia, the target question text being directed to a related resource identifier, wherein the related resource identifier refers to a resource identifier of a related multimedia corresponding to the target multimedia; generating a target answer text corresponding to the target question text according to the target question text using a recall model obtained by training according to the method of any one of claims 1 to 7, wherein the target answer text includes resource identifiers of related multimedia corresponding to the target multimedia; determining a recall result of the target multimedia according to a resource identifier in the target answer text; How to train a recall model, including:

9. After the step of determining a recall result of the target multimedia according to a resource identifier in the target answer text, further: a step of associating the resource identifier of the target multimedia with the recall result of the target multimedia and storing the association result in a recall data set; 9. The method of training a recall model of claim 8, comprising:

10. obtaining a multimedia search request including a search keyword; Matching multimedia according to the search keyword to determine a second multimedia that matches the search keyword; obtaining a recall result corresponding to the second multimedia from the recall dataset; sending the second multimedia and a recall result corresponding to the second multimedia to an originator of the multimedia search request; 10. The method of training a recall model of claim 9, comprising:

11. a first acquiring module configured to execute a step of acquiring a plurality of first text pairs, the first text pairs including a first question text generated according to multimedia description information and targeting a resource identifier of the multimedia as a question, and a first answer text that is a resource identifier for a question of the first question text; a pre-training module configured to pre-train a recall model according to a first question text and a first answer text of the plurality of first text pairs; a second acquisition module configured to perform the steps of: acquiring a plurality of second text pairs, the second text pairs including a second question text questioning a resource identifier of related multimedia and a second answer text being a resource identifier of related multimedia to a question of the second question text, wherein the related multimedia for the second text pair corresponds to the multimedia for the first text pair; a fine-tuning module configured to perform fine-tuning on the pre-trained recall model according to a second question text and a second answer text of the plurality of second text pairs; A training device for recall models, including:

12. a third acquisition module configured to perform the step of acquiring a resource identifier of the target multimedia; a target question text generation module configured to perform a step of generating a target question text according to a resource identifier of the target multimedia, the target question text being directed to a related resource identifier, wherein the related resource identifier refers to a resource identifier of a related multimedia corresponding to the target multimedia; a target answer text determination module configured to perform the steps of: generating a target answer text corresponding to the target question text according to the target question text by a recall model trained according to the method of any one of claims 1 to 7, wherein the target answer text includes resource identifiers of related multimedia corresponding to the target multimedia; a recall result determination module configured to perform the step of determining a recall result of the target multimedia according to a resource identifier in the target answer text; , including a recall device.

13. An electronic device including a processor and a memory in which a computer program is stored, An electronic device that, when the computer program is executed by the processor, causes the device to implement the method according to any one of claims 1 to 7 or the method according to any one of claims 8 to 10.

14. A computer-readable storage medium on which a computer program is stored, A computer-readable storage medium that, when the computer program is executed by a processor, causes the method according to any one of claims 1 to 7 to be implemented, or causes the method according to any one of claims 8 to 10 to be implemented.

15. A computer program product comprising a computer program, A computer program product, which, when the computer program is executed by a processor, causes the method according to any one of claims 1 to 7 to be implemented, or causes the method according to any one of claims 8 to 10 to be implemented.

Citation Information

Patent Citations

  • Method for searching for item based on burying similarity, computer device, and computer program

    JP2022173084A

  • Recommending a Content Item

    JP2022539802A

  • Custom Compilation Videos

    US20210011939A1