Model training method and apparatus, and information retrieval method and apparatus
By selecting and training the encoder in the search model, fixing some parameters, and reducing the number of model training times, the problems of high computing resource consumption and management costs in information retrieval in different modalities are solved, and efficient information retrieval is achieved.
Patent Information
- Application Number
- PCT/CN2024/127718
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-24
AI Technical Summary
When the prior art realizes information retrieval between different modes, it is necessary to build multiple machine learning models, resulting in high computing resource consumption and increased management costs.
A model training method is adopted to obtain the encoder in the search model, encode the training sample and match the results, fix some encoder parameters, and only train other encoders to reduce the number of models and the number of training times.
It improves model training efficiency, reduces model number and management costs, and improves the convenience of information retrieval.
Smart Images

Figure CN2024127718_24072025_PF_FP_ABST
Abstract
Description
A method and device for model training and information retrieval Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method and device for model training and information retrieval. Background Art
[0002] With the rapid development and widespread attention of online businesses, the forms (i.e., modalities) in which users can access media content are becoming increasingly diverse. For example, users can browse text, listen to music, view pictures, and watch videos through the Internet. With the integration of artificial intelligence technology into the Internet, it can not only bring users a richer business experience, but also more effectively protect the security of users' information and personal privacy data.
[0003] Currently, users can use data of one modality to retrieve matching data of another modality. For example, if a user enters a descriptive text, the server or cloud platform can use the descriptive text to search for images that match the text content of the descriptive text and return them to the user. In this example, the descriptive text is data of one modality, and the returned image is data of another modality.
[0004] However, in practical applications, if pairwise retrieval between n modalities is required, a corresponding machine learning model must be built for each two modalities and trained to enable cross-modality retrieval. However, since this approach requires building a model for each two modalities, it consumes a large amount of computing resources and significantly increases model management costs.
[0005] Therefore, how to avoid the consumption of more computing resources and reduce the management cost of the model is an urgent problem to be solved.
[0006] Summary of the Invention
[0007] This specification provides a method and apparatus for model training and information retrieval to avoid the consumption of computing resources in information retrieval and reduce the management cost of the information retrieval model.
[0008] This manual adopts the following technical solutions.
[0009] This specification provides a model training method, including: obtaining a retrieval model, wherein the retrieval model includes multiple encoders, different encoders are used to encode data of different modalities; selecting a first encoder and a second encoder from the retrieval model, and obtaining a first training sample in the modality corresponding to the first encoder, and obtaining a second training sample in the modality corresponding to the second encoder; inputting the first training sample and the second training sample into the retrieval model, so as to encode the first training sample by the first encoder to obtain a first encoding result, and to encode the second training sample by the second encoder to obtain a second encoding result, and determining a matching result between the first training sample and the second training sample based on the first encoding result and the second encoding result; training the first encoder and the second encoder based on the matching result; fixing the network parameters of the trained first encoder, and for each other encoder in the retrieval model except the first encoder and the second encoder, inputting the obtained third training sample in the modality corresponding to the other encoder into the other encoder, so as to train the other encoder based on the encoding result output by the other encoder and the encoding result output by the fixed trained first encoder.
[0010] Optionally, the data of each modality includes at least one of text data, image data, video data, audio data and business sequence data.
[0011] Optionally, the first encoder and the second encoder are trained according to the matching result, including: determining labeling information between the first training sample and the second training sample, the labeling information being used to represent the actual matching relationship between the first training sample and the second training sample; if it is determined according to the labeling information that the first training sample and the second training sample match, the first encoder and the second encoder are trained with minimizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the labeling information and the matching result as optimization goals; if it is determined according to the labeling information that the first training sample and the second training sample do not match, the first encoder and the second encoder are trained with maximizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the labeling information and the matching result as optimization goals.
[0012] Optionally, obtaining a first training sample in the modality corresponding to the first encoder and obtaining a second training sample in the modality corresponding to the second encoder include: obtaining the first training sample in the modality corresponding to the first encoder, and inputting the first training sample into a preset generation model, so that the generation model outputs data in the modality corresponding to the second encoder based on the first training sample as the generated data corresponding to the first training sample; and obtaining the second training sample in the modality corresponding to the second encoder based on the generated data corresponding to the first training sample.
[0013] This specification provides a method for information retrieval, comprising: in response to an information retrieval request, determining the modality corresponding to retrieval data and data retrieved based on the retrieval data as a target modality, the modality corresponding to the retrieval data being different from the target modality; determining each candidate data corresponding to the retrieval data under the target modality; inputting the retrieval data and each candidate data into a pre-trained retrieval model to determine an encoding result corresponding to the retrieval data through an encoder corresponding to the modality of the retrieval data in the retrieval model, and determining an encoding result corresponding to each candidate data through an encoder corresponding to the target modality in the retrieval model, the retrieval model being trained by a model training method; and determining an information retrieval result for the retrieval data based on the encoding result corresponding to the retrieval data and the encoding results corresponding to each candidate data.
[0014] This specification provides a model training device, including: an acquisition module for acquiring a retrieval model, wherein the retrieval model includes a plurality of encoders, and different encoders are used to encode data of different modalities; a selection module for selecting a first encoder and a second encoder from the retrieval model, and obtaining a first training sample under the modality corresponding to the first encoder, and obtaining a second training sample under the modality corresponding to the second encoder; an encoding module for inputting the first training sample and the second training sample into the retrieval model, so as to encode the first training sample through the first encoder to obtain a first encoding result, and to encode the second training sample through the second encoder to obtain a first encoding result. The invention relates to a method for fixing the network parameters of the first encoder after training and inputting the obtained third training sample under the corresponding mode of each other encoder except the first encoder and the second encoder into the other encoder for training the other encoder according to the encoding result output by the other encoder and the encoding result output by the fixed trained first encoder.
[0015] Optionally, the data of each modality includes at least one of text data, image data, video data, audio data and business sequence data.
[0016] Optionally, the first training module is specifically used to determine the labeling information between the first training sample and the second training sample, where the labeling information is used to represent the actual matching relationship between the first training sample and the second training sample; if it is determined that the first training sample and the second training sample match according to the labeling information, the first encoder and the second encoder are trained with the optimization goals of minimizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the labeling information and the matching result; if it is determined that the first training sample and the second training sample do not match according to the labeling information, the first encoder and the second encoder are trained with the optimization goals of maximizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the labeling information and the matching result.
[0017] Optionally, the selection module is specifically used to obtain a first training sample under the modality corresponding to the first encoder, and input the first training sample into a preset generation model, so that the generation model outputs data under the modality corresponding to the second encoder based on the first training sample as the generation data corresponding to the first training sample; and obtains a second training sample under the modality corresponding to the second encoder based on the generation data corresponding to the first training sample.
[0018] This specification provides an information retrieval device, including: a response module, used to determine the modality corresponding to the retrieval data and the data retrieved based on the retrieval data in response to an information retrieval request, as a target modality, the modality corresponding to the retrieval data being different from the target modality; a determination module, used to determine each candidate data corresponding to the retrieval data under the target modality; an input module, used to input the retrieval data and each candidate data into a pre-trained retrieval model, so as to determine the encoding result corresponding to the retrieval data through an encoder corresponding to the modality of the retrieval data in the retrieval model, and to determine the encoding result corresponding to each candidate data through an encoder corresponding to the target modality in the retrieval model, wherein the retrieval model is trained by a model training method; a retrieval module, used to determine the information retrieval result for the retrieval data based on the encoding result corresponding to the retrieval data and the encoding result corresponding to each candidate data.
[0019] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned model training and information retrieval methods.
[0020] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned model training and information retrieval methods are implemented.
[0021] At least one of the above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects: in the model training and information retrieval methods provided in this specification, a retrieval model can be obtained, and the retrieval model contains several encoders, and different encoders are used to encode data of different modalities. Then, a first encoder and a second encoder are selected from the retrieval model, and a first training sample under the modality corresponding to the first encoder is obtained, and a second training sample under the modality corresponding to the second encoder is obtained. The first training sample and the second training sample are input into the retrieval model, so that the first training sample is encoded by the first encoder to obtain a first encoding result, and the second training sample is encoded by the second encoder to obtain a second encoding result, and the matching result between the first training sample and the second training sample is determined by the first encoding result and the second encoding result. Then, based on the matching result, the first encoder and the second encoder can be trained, the network parameters of the trained first encoder are fixed, and for each other encoder in the retrieval model except the first encoder and the second encoder, the third training sample obtained in the corresponding mode of the other encoder is input into the other encoder to train the other encoder based on the encoding result output by the other encoder and the encoding result output by the fixed trained first encoder.
[0022] From the above content, it can be seen that the model training and information retrieval method provided in this specification can greatly improve the efficiency of model training compared to the conventional method: each two modalities need to build a model, train this model, and when implementing the retrieval of these two modalities, it is necessary to use this model for retrieval. Since the encoder of one modality can be used as a bridge to connect the encoders of other modalities, the other modalities do not need to be trained with each other, reducing the number of models and the number of training times. The more modalities there are, the more this method can avoid more model training processes, the number of stored models, and improve the convenience of using models for information retrieval in practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0024] FIG1 is a flow chart of a model training method in this specification;
[0025] FIG2 is a schematic diagram of a feature code provided in this specification;
[0026] FIG3 is a flow chart of an information retrieval method in this specification;
[0027] FIG4 is a schematic diagram of a model training device provided in this specification;
[0028] FIG5 is a schematic diagram of an information retrieval device provided in this specification;
[0029] FIG6 is a schematic diagram of an electronic device provided in this specification corresponding to FIG1 or FIG3 . DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0031] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0032] FIG1 is a flow chart of a model training method in this specification, which specifically includes the following steps S100 to S108.
[0033] S100: Acquire a retrieval model, wherein the retrieval model includes a plurality of encoders, and different encoders are used to encode data of different modalities.
[0034] In actual applications, the service platform needs to train a retrieval model for information retrieval. The retrieval model is used to provide retrieval services for users and can be used by the service platform for information recommendation services. The service platform can also use the retrieval model in other information retrieval-related services.
[0035] Based on this, a retrieval model to be trained can be obtained, in which there are several encoders, and different encoders are used to encode data of different modalities.
[0036] The modalities in this specification may include text modalities, image modalities, video modalities, audio modalities, and service sequence modalities. Accordingly, the data of each modality includes text data, image data, video data, audio data, and service sequence data. Service sequence data can refer to sensor sequence data representing user motion characteristics collected by the user's device (such as a mobile phone, wristband, etc.), or it can refer to the user's operation records in the app provided by the service platform through the device (such as a mobile phone, tablet, etc.).
[0037] S102: Select a first encoder and a second encoder from the retrieval model, and obtain a first training sample in a modality corresponding to the first encoder, and obtain a second training sample in a modality corresponding to the second encoder.
[0038] S104: Input the first training sample and the second training sample into the retrieval model, encode the first training sample through the first encoder to obtain a first encoding result, and encode the second training sample through the second encoder to obtain a second encoding result, and determine the matching result between the first training sample and the second training sample through the first encoding result and the second encoding result.
[0039] S106: Training the first encoder and the second encoder according to the matching result.
[0040] The training process of the retrieval model can be divided into two parts. In the first part, the two encoders in the retrieval model need to be trained. In the second part, the parameters of one of the two encoders trained in the retrieval model in the first part need to be fixed to train the remaining encoders. Therefore, the structure of the retrieval model can be shown in Figure 2.
[0041] FIG2 is a schematic diagram of the structure of a retrieval model provided in this specification.
[0042] Based on this, during the first part of the training process, the service platform can select a first encoder and a second encoder from the retrieval model to be trained, and obtain a first training sample under the modality corresponding to the first encoder, and obtain a second training sample under the modality corresponding to the second encoder. The first training sample mentioned here refers to the data under the modality corresponding to the first encoder, and similarly, the second training sample refers to the data under the modality corresponding to the second encoder.
[0043] Then, the first training sample and the second training sample can be input into the above-mentioned retrieval model, so that the first training sample is encoded by the first encoder to obtain a first encoding result, and the second training sample is encoded by the second encoder to obtain a second encoding result. The first encoding result and the second encoding result are used to determine the matching result between the first training sample and the second training sample. The first encoder and the second encoder are then trained based on the matching result.
[0044] The matching results mentioned above can be used to indicate whether the first training sample predicted by the retrieval model matches the second training sample, or the degree of matching between the first training sample predicted by the retrieval model and the second training sample. To illustrate the concept of matching mentioned in this specification, for example, for an image and a piece of text, if the text is related to the image, it can be considered that the text matches the image.
[0045] During training, the labeling information between the first training sample and the second training sample can be determined, and the labeling information can be used to represent the actual matching relationship between the first training sample and the second training sample. Then, the first encoder and the second encoder can be trained with minimizing the difference between the labeling information and the matching result as the optimization goal.
[0046] In the training process of the first encoder and the second encoder, there may be positive samples (i.e., matching first training samples and second training samples) and negative samples (i.e., mismatching first training samples and second training samples). When the first encoder and the second encoder are trained using the positive samples and the negative samples:
[0047] If the first training sample and the second training sample are determined to match based on the above-mentioned annotation information, the first encoder and the second encoder can be trained with the optimization goals of minimizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the annotation information and the matching result.
[0048] If it is determined based on the labeling information that the first training sample and the second training sample do not match, the first encoder and the second encoder are trained with the optimization goals of maximizing the difference between the first encoding result and the second encoding result and minimizing the difference between the labeling information and the matching result.
[0049] That is, during training, the distance between the encoding results of two training samples in the positive sample can be shortened, and the distance between the encoding results of two training samples in the negative sample can be increased.
[0050] S108: Fix the network parameters of the trained first encoder, and for each other encoder in the retrieval model except the first encoder and the second encoder, input the obtained third training sample under the corresponding modality of the other encoder into the other encoder, so as to train the other encoder according to the encoding result output by the other encoder and the encoding result output by the fixed trained first encoder.
[0051] After completing the first part of the training process in the above steps, the network parameters of the trained first encoder can be fixed, and for each other encoder in the retrieval model other than the first encoder and the second encoder, the obtained third training samples in the corresponding modality of the other encoder are input into the other encoder to train the other encoder based on the encoding results output by the other encoder and the encoding results output by the fixed trained first encoder. In this training process, the training samples in the corresponding modality of the first encoder used can be different from those used in the first part.
[0052] That is to say, the above process uses the first encoder as an intermediate bridge between the second encoder and other encoders. The first part of the training process fully trains the first encoder and the second encoder in the retrieval model so that the feature extraction of the retrieval model for the modality corresponding to the first encoder and the modality corresponding to the second encoder can be aligned.
[0053] In the second part, the parameters of the first encoder are fixed and the other encoders are trained. Since the other encoders are trained while the first encoder remains unchanged, not only the features of the mode corresponding to the first encoder and the corresponding modes of other encoders can be aligned, but also the features of the mode corresponding to the second encoder and the corresponding modes of other encoders can be aligned.
[0054] Therefore, if retrieval between different modalities is required, retrieval between the modalities corresponding to the two encoders can be achieved through any two encoders.
[0055] When the parameters of the trained first encoder are fixed, the training of other encoders is similar to the training of the first encoder and the second encoder.
[0056] Among them, the sample pairs corresponding to the first encoder and other encoders contain corresponding data under the modality corresponding to the first encoder and data under the modalities corresponding to other encoders. The sample pairs may contain positive samples and negative samples. For positive samples, the difference between the encoding results output by the trained first encoder and the encoding results output by other encoders can be minimized. For negative samples, the difference between the encoding results output by the trained first encoder and the encoding results output by other encoders can be maximized. At the same time, for each sample pair, the difference between the annotation information and the matching results determined by the encoding results output by the trained first encoder and the encoding results output by other encoders can be minimized. Among them, the annotation information mentioned here is the actual matching relationship between the two data in the sample pair.
[0057] It should be noted that the first encoder in the retrieval model can be an encoder corresponding to the text modality, the second encoder can be an encoder corresponding to the image modality, and the other encoders can be encoders corresponding to the audio modality, video modality, and business sequence modality, respectively. Of course, the first encoder, the second encoder, and the other encoders can be replaced with encoders corresponding to any modality. For example, the first encoder can also be an encoder corresponding to the image modality. The second encoder can be an encoder corresponding to the video modality, and the other encoders can be encoders corresponding to the remaining modalities. The specific modalities corresponding to the encoders can be set according to actual needs.
[0058] In this specification, the complete training samples may include training samples for training the first encoder and the second encoder, as well as training samples for training the first encoder and other encoders. For both types of training samples, they are in the form of sample pairs. In this method, sample pairs can be automatically generated by generating a model.
[0059] Specifically, a first training sample under the modality corresponding to the first encoder can be obtained, and the first training sample can be input into a preset generation model so that the generation model outputs data under the modality corresponding to the second encoder based on the first training sample as the generation data corresponding to the first training sample, and based on the generation data corresponding to the first training sample, a second training sample under the modality corresponding to the second encoder is obtained. The generation model mentioned here can be an existing model for generating text, images, audio or video.
[0060] That is, a generative model can generate data in one modality based on data in another modality. For example, a generative model can generate matching images based on input text. The text and the matching images can constitute positive samples, while the text and other images generated by the generative model (images that do not match the text) can constitute negative samples. Of course, a generative model can also be specified to generate data in another modality that matches data in one modality, and data in another modality that does not match data in the first modality.
[0061] This method yields a unified neural network model (retrieval model) encompassing text / image / audio / video / sequence encoders. For each modality, features need only be extracted once using the corresponding encoder. During feature retrieval, cross-modal feature comparisons can be performed based on the features extracted by this model.
[0062] Because text features are used as anchors during training, features from text, images, videos, audio, and sequences are aligned in feature space. In addition to being applicable to feature retrieval from text, images, audio, video, and sequences, this model is also applicable to pairwise feature retrieval between various modalities, such as audio-image and image-sequence.
[0063] The above is an explanation of this method from the perspective of model training. The following is an explanation of this method from the perspective of actual model use.
[0064] FIG3 is a flowchart of an information retrieval method in this specification, which specifically includes the following steps S300 to S306 .
[0065] S300: In response to an information retrieval request, determine retrieval data and a modality corresponding to data retrieved based on the retrieval data as a target modality, where the modality corresponding to the retrieval data is different from the target modality.
[0066] S302: Determine candidate data corresponding to the search data in the target mode.
[0067] S304: Input the retrieval data and the candidate data into a pre-trained retrieval model to determine the encoding result corresponding to the retrieval data through the encoder corresponding to the modality of the retrieval data in the retrieval model, and determine the encoding result corresponding to the candidate data through the encoder corresponding to the target modality in the retrieval model. The retrieval model is trained by the above-mentioned model training method.
[0068] S306: Determine an information retrieval result for the retrieval data according to the encoding result corresponding to the retrieval data and the encoding results corresponding to each candidate data.
[0069] In this specification, a retrieval model is obtained after training. If there is a retrieval requirement between any two modalities, it can be completed through the encoders of these two modalities in the retrieval model.
[0070] Specifically, in response to the information retrieval request, the service platform may determine the retrieval data and the modality corresponding to the data to be retrieved based on the retrieval data as the target modality. The modality corresponding to the retrieval data is different from the target modality.
[0071] Then, candidate data corresponding to the search data in the target modality can be determined. The candidate data mentioned here can be all the data in the target modality, or part of the data in the target modality selected in a certain way.
[0072] The search data and each candidate data are input into a pre-trained search model, and an encoding result corresponding to the search data is determined by an encoder corresponding to the modality of the search data in the search model. The encoding result corresponding to each candidate data is determined by an encoder corresponding to the target modality in the search model. Based on the encoding result corresponding to the search data and the encoding results corresponding to each candidate data, an information retrieval result for the search data is determined. The search model is trained using the above-mentioned model training method.
[0073] Among them, the correlation between the encoding result corresponding to the retrieval data and the encoding results corresponding to each candidate data (such as determined by a correlation coefficient or determined by a neural network module in a retrieval model, etc.) can be used to determine which candidate data or candidate data matches the retrieval data, thereby obtaining an information retrieval result for the retrieval data.
[0074] In actual applications, there are many scenarios for pairwise retrieval between text and image, text and audio, text and video, video and image, and video and audio. Here we explain the retrieval scenarios between business sequence data and other modal data. For example, for sensor sequence data collected by mobile phones, bracelets, etc. that can represent the user's motion status, relevant videos, audios, and texts can be retrieved based on this sequence data and recommended to the user (for example, if the user is exercising, relevant post-exercise relaxation videos, music suitable for exercise, and other data can be retrieved and recommended to the user).
[0075] From the above content, it can be seen that the business execution and information retrieval method provided in this specification can greatly improve the efficiency of model training compared to the conventional method: each two modalities need to build a model, train this model, and when implementing the retrieval of these two modalities, it is necessary to use this model for retrieval. Since the encoder of one modality can be used as a bridge to connect the encoders of other modalities, the other modalities do not need to be trained with each other, reducing the number of models and the number of training times. The more modalities there are, the more this method can avoid more model training processes, the number of stored models, and improve the convenience of using models for information retrieval in practice.
[0076] It should be noted that, for the sake of ease of description, the execution entity of this method is described as a service platform in the above content. The execution entity of this method can be a server, a computer, etc., which is not limited here.
[0077] The above is a method for model training and information retrieval provided in one or more embodiments of this specification. Based on the same idea, this specification also provides a device for model training and information retrieval, as shown in Figures 4 and 5.
[0078] FIG4 is a schematic diagram of a model training device provided in this specification, which specifically includes:
[0079] An acquisition module 401 is used to acquire a retrieval model, wherein the retrieval model includes a plurality of encoders, and different encoders are used to encode data of different modalities;
[0080] A selection module 402 is configured to select a first encoder and a second encoder from the retrieval model, and obtain a first training sample in a modality corresponding to the first encoder, and obtain a second training sample in a modality corresponding to the second encoder;
[0081] an encoding module 403 configured to input the first training sample and the second training sample into the retrieval model, encode the first training sample using the first encoder to obtain a first encoding result, and encode the second training sample using the second encoder to obtain a second encoding result, and determine a matching result between the first training sample and the second training sample based on the first encoding result and the second encoding result;
[0082] A first training module 404, configured to train the first encoder and the second encoder according to the matching result;
[0083] The second training module 405 is used to fix the network parameters of the trained first encoder, and for each other encoder in the retrieval model except the first encoder and the second encoder, input the third training sample obtained in the corresponding mode of the other encoder into the other encoder, so as to train the other encoder according to the encoding result output by the other encoder and the encoding result output by the fixed trained first encoder.
[0084] Optionally, the data of each modality includes at least one of text data, image data, video data, audio data and business sequence data.
[0085] Optionally, the first training module 404 is specifically used to determine the labeling information between the first training sample and the second training sample, where the labeling information is used to represent the actual matching relationship between the first training sample and the second training sample; if it is determined that the first training sample and the second training sample match according to the labeling information, the first encoder and the second encoder are trained with the optimization goals of minimizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the labeling information and the matching result; if it is determined that the first training sample and the second training sample do not match according to the labeling information, the first encoder and the second encoder are trained with the optimization goals of maximizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the labeling information and the matching result.
[0086] Optionally, the selection module 402 is specifically used to obtain a first training sample under the modality corresponding to the first encoder, and input the first training sample into a preset generation model, so that the generation model outputs data under the modality corresponding to the second encoder based on the first training sample as the generation data corresponding to the first training sample; and obtains a second training sample under the modality corresponding to the second encoder based on the generation data corresponding to the first training sample.
[0087] FIG5 is a schematic diagram of an information retrieval device provided in this specification, which specifically includes:
[0088] A response module 501 is configured to, in response to an information retrieval request, determine retrieval data and a modality corresponding to data retrieved based on the retrieval data as a target modality, wherein the modality corresponding to the retrieval data is different from the target modality;
[0089] A determination module 502 is configured to determine candidate data corresponding to the search data in the target modality;
[0090] An input module 503 is configured to input the search data and each candidate data into a pre-trained search model, so as to determine an encoding result corresponding to the search data using an encoder corresponding to the modality of the search data in the search model, and to determine an encoding result corresponding to each candidate data using an encoder corresponding to the target modality in the search model, wherein the search model is trained using a model training method;
[0091] The retrieval module 504 is configured to determine an information retrieval result for the retrieval data according to the encoding result corresponding to the retrieval data and the encoding results corresponding to each candidate data.
[0092] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above-mentioned model training and information retrieval methods.
[0093] This specification also provides a schematic structural diagram of the electronic device shown in Figure 6. As shown in Figure 6, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned model training and information retrieval methods. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0094] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0095] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0096] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0097] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0098] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0099] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0100] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0102] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0103] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0104] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0105] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0106] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0107] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing nodes connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage nodes.
[0108] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0109] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for model training, comprising: Obtaining a retrieval model, wherein the retrieval model includes a plurality of encoders, and different encoders are used to encode data of different modalities; Selecting a first encoder and a second encoder from the retrieval model, and obtaining a first training sample in the modality corresponding to the first encoder, and obtaining a second training sample in the modality corresponding to the second encoder; Inputting the first training sample and the second training sample into the retrieval model, so as to encode the first training sample through the first encoder to obtain a first encoding result, and encode the second training sample through the second encoder to obtain a second encoding result, and determining a matching result between the first training sample and the second training sample based on the first encoding result and the second encoding result; Training the first encoder and the second encoder according to the matching result; Fixing the network parameters in the trained first encoder, and for each other encoder in the retrieval model except the first encoder and the second encoder, inputting the obtained third training sample in the modality corresponding to the other encoder into the other encoder, so as to train the other encoder based on the encoding result output by the other encoder and the encoding result output by the fixed and trained first encoder.
2. The method according to claim 1, wherein the data of each modality includes: At least one of text data, image data, video data, audio data, and service sequence data.
3. The method according to claim 1, wherein training the first encoder and the second encoder according to the matching result comprises: Determining annotation information between the first training sample and the second training sample, where the annotation information is used to represent the actual matching relationship between the first training sample and the second training sample; If it is determined according to the annotation information that the first training sample and the second training sample match, then training the first encoder and the second encoder with the optimization objective of minimizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the annotation information and the matching result; If it is determined according to the annotation information that the first training sample and the second training sample do not match, then training the first encoder and the second encoder with the optimization objective of maximizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the annotation information and the matching result.
4. The method according to claim 1, wherein obtaining the first training sample in the modality corresponding to the first encoder and obtaining the second training sample in the modality corresponding to the second encoder comprises: Obtaining the first training sample in the modality corresponding to the first encoder, and inputting the first training sample into a preset generation model, so that the generation model outputs data in the modality corresponding to the second encoder Based on the first training sample, as the generated data corresponding to the first training sample; Obtain the second training sample in the modality corresponding to the second encoder according to the generated data corresponding to the first training sample.
5. A method for information retrieval, comprising: In response to an information retrieval request, determine retrieval data and the modality corresponding to the data retrieved based on the retrieval data as the target modality, where the modality corresponding to the retrieval data is different from the target modality; Determine each candidate data corresponding to the retrieval data in the target modality; Input the retrieval data and each candidate data into a pre-trained retrieval model, so as to determine the encoding result corresponding to the retrieval data through the encoder corresponding to the modality of the retrieval data in the retrieval model, and determine the encoding results corresponding to each candidate data through the encoder corresponding to the target modality in the retrieval model, where the retrieval model is trained by the method according to any one of claims 1 to 4; Determine the information retrieval result for the retrieval data according to the encoding result corresponding to the retrieval data and the encoding results corresponding to each candidate data.
6. A device for model training, comprising: An acquisition module, configured to acquire a retrieval model, where the retrieval model includes a plurality of encoders, and different encoders are used to encode data in different modalities; A selection module, configured to select a first encoder and a second encoder from the retrieval model, acquire a first training sample in the modality corresponding to the first encoder, and acquire a second training sample in the modality corresponding to the second encoder; An encoding module, configured to input the first training sample and the second training sample into the retrieval model, so as to encode the first training sample through the first encoder to obtain a first encoding result, and encode the second training sample through the second encoder to obtain a second encoding result, and determine the matching result between the first training sample and the second training sample through the first encoding result and the second encoding result; A first training module, configured to train the first encoder and the second encoder according to the matching result; A second training module, configured to fix the network parameters in the trained first encoder, and for each other encoder in the retrieval model except the first encoder and the second encoder, input the obtained third training sample in the modality corresponding to the other encoder into the other encoder, so as to train the other encoder according to the encoding result output by the other encoder and the encoding result output by the fixed trained first encoder.
7. The device according to claim 6, wherein the data of each modality includes: At least one of text data, image data, video data, audio data, and service sequence data.
8. The apparatus according to claim 6, wherein the first training module is specifically configured to determine the annotation information between the first training sample and the second training sample, where the annotation information is used to represent the actual matching relationship between the first training sample and the second training sample; if it is determined according to the annotation information that the first training sample and the second training sample match, then minimizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the annotation information and the matching result are used as the optimization objectives to train the first encoder and the second encoder; if it is determined according to the annotation information that the first training sample and the second training sample do not match, then maximizing the difference between the first encoding result and the second encoding result, and minimizing the difference between the annotation information and the matching result are used as the optimization objectives to train the first encoder and the second encoder.
9. The apparatus according to claim 6, wherein the selection module is specifically configured to obtain the first training sample in the modality corresponding to the first encoder, and input the first training sample into a preset generation model, so that the generation model outputs data in the modality corresponding to the second encoder based on the first training sample as the generated data corresponding to the first training sample; and obtain the second training sample in the modality corresponding to the second encoder according to the generated data corresponding to the first training sample.
10. An information retrieval apparatus, comprising: a response module, configured to respond to an information retrieval request, and determine the retrieval data and the modality corresponding to the data retrieved based on the retrieval data as the target modality, where the modality corresponding to the retrieval data is different from the target modality; a determination module, configured to determine each candidate data corresponding to the retrieval data in the target modality; an input module, configured to input the retrieval data and each candidate data into a pre-trained retrieval model, so as to determine the encoding result corresponding to the retrieval data through the encoder corresponding to the modality of the retrieval data in the retrieval model, and determine the encoding results corresponding to each candidate data through the encoder corresponding to the target modality in the retrieval model, where the retrieval model is trained by the method according to any one of claims 1 to 4; a retrieval module, configured to determine an information retrieval result for the retrieval data according to the encoding result corresponding to the retrieval data and the encoding results corresponding to each candidate data.
11. A computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1 to 5 above.
12. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor, when executing the program, implements the method according to any one of claims 1 to 5 above.
Citation Information
Patent Citations
Similar picture retrieval method and device based on multi-modal pre-training and electronic equipment
CN114461839A
Content generation method and device, electronic equipment and computer readable storage medium
CN117194702A
Method for improving efficiency of adapting large language model to multi-modal task
CN117194989A
Multi-modal large language model training method and system based on multi-modal encoder
CN117218498A
Model training and information retrieval method and device
CN118153627A