Multimodal large language model training method, correlation calculation method, and label generation method
Through the two-stage training method, the multimodal large language model can learn image features related to search terms and text description information at the same time, solving the problem that multimodal feature vectors are not accurate enough in the prior art, and realizing more accurate feature vector generation.
Patent Information
- Application Number
- PCT/CN2024/129489
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-11-01
- Publication Date
- 2025-06-19
AI Technical Summary
The existing multimodal large language model cannot learn image features related to search terms and text description information at the same time in search scenarios, resulting in the generated multimodal feature vectors being inaccurate enough.
Using a two-stage training method, first, the multimodal large language model is trained based on the image information and text description information of the sample product, so that it can learn to enhance the image features related to the text description information; then, the second training is performed based on the image information and text description information of the search terms and sample product, so that the model can learn to enhance the image features related to the search terms and text description information.
The generated second multimodal feature vector is more accurate and can better meet the application needs in search scenarios.
Smart Images

Figure CN2024129489_19062025_PF_FP_ABST
Abstract
Description
Multimodal large language model training method, correlation calculation and label generation method
[0001] This disclosure claims priority to a Chinese patent application filed with the Patent Office of China on December 14, 2023, with application number 202311724117.X and application name “Multimodal large language model training method, correlation calculation and label generation method,” the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0002] The present disclosure relates to the field of artificial intelligence technology, and in particular to a multimodal large language model training method, correlation calculation method, and label generation method. Background Art
[0003] The Multimodal Large Language Model (MLLM) is a natural language processing model based on deep learning that can process various types of data such as text, images, and audio to complete natural language tasks such as search, intelligent question answering, and translation.
[0004] At present, the multimodal large language model in the related technology can only generate a multimodal feature vector based on text information and image information. In the search scenario, text information usually includes search terms and text description information of the product, and the multimodal large language model trained using the related technology can only learn the image features related to the search terms and generate a multimodal feature vector with enhanced image features related to the search terms, or learn the image features related to the text description information and generate a multimodal feature vector with enhanced image features related to the text description information. It is unable to simultaneously learn the image features related to the search terms and text description information in the search scenario, resulting in the generated multimodal feature vector being inaccurate. Therefore, there is an urgent need to train a new multimodal large language model to generate a more accurate multimodal feature vector to meet the application needs in the search scenario.
[0005] Summary of the Invention
[0006] The disclosed embodiments provide a multimodal large language model training method, correlation calculation, and label generation method. The multimodal large language model trained by this method can simultaneously learn image features related to search terms and text description information. The generated second multimodal feature vector is more accurate and meets the application requirements of search scenarios. The technical solution is as follows:
[0007] In a first aspect, a multimodal large language model training method is provided, the method comprising:
[0008] Obtaining a sample search term and a sample product, wherein the sample product has image information and text description information;
[0009] Invoking a pre-trained multimodal large language model to process the image information and text description information of the sample product to obtain a sample text feature vector, a sample image feature vector, and a sample first multimodal feature vector of the sample product, wherein image features related to the text description information in the sample first multimodal feature vector are enhanced;
[0010] Training the pre-trained multimodal large language model based on the sample text feature vector, the sample image feature vector, and the sample first multimodal feature vector of the sample product to obtain a multimodal large language model;
[0011] Invoking the large multimodal language model to process the sample search term and the image information and text description information of the sample product to obtain a sample search term feature vector, a first multimodal feature vector of the sample product, and a sample second multimodal feature vector, wherein image features related to the text description information and the search term in the sample second multimodal feature vector are enhanced;
[0012] Based on the sample search term feature vector and the first multimodal feature vector and the sample second multimodal feature vector of the sample product, the multimodal large language model is trained to obtain a trained multimodal large language model. The trained multimodal large language model is used to generate a second multimodal feature vector of the product based on the search term and the image information and text description information of the product.
[0013] In a second aspect, a correlation calculation method is provided, wherein the method applies the trained multimodal large language model described in the first aspect, and the method comprises:
[0014] Obtaining a search term and candidate products found based on the search term, wherein the candidate products have image information and text description information;
[0015] Calculating a semantic relevance score between the search term and the candidate product based on the search term and the text description information;
[0016] Invoking the trained multimodal large language model to process the search term, the image information, and the text description information to obtain a second multimodal feature vector and a search term feature vector for the candidate product;
[0017] Calculating a graphic-text relevance score between the search term and the candidate product based on the search term feature vector and the second multimodal feature vector;
[0018] Based on the semantic relevance score and the image-text relevance score, a total relevance score between the search term and the candidate product is calculated.
[0019] In another embodiment of the present disclosure, the calculating, based on the search term feature vector and the second multimodal feature vector, a graphic-text relevance score between the search term and the candidate product includes:
[0020] The cosine similarity between the search term feature vector and the second multimodal feature vector is calculated to obtain the image-text relevance score.
[0021] In another embodiment of the present disclosure, calculating the total relevance score between the search term and the candidate product based on the semantic relevance score and the image-text relevance score includes:
[0022] The semantic relevance score and the image-text relevance score are weightedly calculated to obtain the total relevance score.
[0023] In a third aspect, a tag generation method is provided, wherein the method applies the trained multimodal large language model described in the first aspect, and the method comprises:
[0024] Calling the trained multimodal large language model to process the search term, image information, and text description information corresponding to the candidate product to obtain a second multimodal feature vector;
[0025] Obtaining a first instruction template corresponding to a preset level category to which the candidate product belongs, the first instruction template being used to describe attributes of the candidate product to be output;
[0026] Based on the second multimodal feature vector and the first instruction template, a first label for the candidate product under the preset level category is generated.
[0027] In a fourth aspect, a multimodal large language model training device is provided, the device comprising:
[0028] A first acquisition module is used to acquire sample search terms and sample products, wherein the sample products have image information and text description information;
[0029] a first processing module, configured to call a pre-trained multimodal large language model to process the image information and text description information of the sample product to obtain a sample text feature vector, a sample image feature vector, and a sample first multimodal feature vector of the sample product, wherein image features related to the text description information in the sample first multimodal feature vector are enhanced;
[0030] A first training module is configured to train the pre-trained multimodal large language model based on the sample text feature vector, the sample image feature vector, and the sample first multimodal feature vector of the sample product to obtain a multimodal large language model;
[0031] a second processing module, configured to call the multimodal large language model to process the sample search term and the image information and text description information of the sample product to obtain a sample search term feature vector, a first multimodal feature vector of the sample product, and a sample second multimodal feature vector, wherein image features related to the text description information and the search term in the sample second multimodal feature vector are enhanced;
[0032] The second training module is used to train the multimodal large language model based on the sample search term feature vector and the first multimodal feature vector and the sample second multimodal feature vector of the sample product to obtain a trained multimodal large language model. The trained multimodal large language model is used to generate a second multimodal feature vector of the product based on the search term and the image information and text description information of the product.
[0033] In a fifth aspect, a correlation calculation device is provided, wherein the device applies the trained multimodal large language model described in the first aspect, and the device comprises:
[0034] An acquisition module, configured to acquire a search term and candidate products found based on the search term, wherein the candidate products have image information and text description information;
[0035] A first calculation module, configured to calculate a semantic relevance score between the search term and the candidate product based on the search term and the text description information;
[0036] a processing module, configured to call the trained multimodal large language model to process the search term, the image information, and the text description information to obtain a second multimodal feature vector and a search term feature vector for the candidate product;
[0037] A second calculation module, configured to calculate a graphic-text relevance score between the search term and the candidate product based on the search term feature vector and the second multimodal feature vector;
[0038] The third calculation module is used to calculate the total relevance score between the search term and the candidate product based on the semantic relevance score and the image-text relevance score.
[0039] In a sixth aspect, a label generation device is provided, wherein the device applies the trained multimodal large language model described in the first aspect, and the device comprises:
[0040] a processing module, configured to call the trained multimodal large language model to process the search terms, image information, and text description information corresponding to the candidate products to obtain a second multimodal feature vector;
[0041] a first acquisition module, configured to acquire a first instruction template corresponding to a preset level category to which the candidate product belongs, wherein the first instruction template is used to describe attributes to be output of the candidate product;
[0042] The first generating module is configured to generate a first label for the candidate product under the preset level category based on the second multimodal feature vector and the first instruction template.
[0043] In the seventh aspect, an electronic device is provided, comprising a processor and a memory; the memory stores at least one program code; the at least one program code is used to be called and executed by the processor to implement the multimodal large language model training method described in the first aspect, or the correlation calculation method described in the second aspect, or the label generation method described in the third aspect.
[0044] In an eighth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor, it can implement the multimodal large language model training method described in the first aspect, or the correlation calculation method described in the second aspect, or the label generation method described in the third aspect.
[0045] In the ninth aspect, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, it can implement the multimodal large language model training method described in the first aspect, or the correlation calculation method described in the second aspect, or the label generation method described in the third aspect.
[0046] The technical solutions provided by the embodiments of the present disclosure have the following beneficial effects:
[0047] A two-stage training method is adopted to train the pre-trained multimodal large language model, so that the trained multimodal large language model can simultaneously learn how to enhance the image features related to the search terms and text description information. In the first training stage, the pre-trained multimodal large language model is trained based on the sample image feature vector, sample text feature vector and sample first multimodal feature vector generated by the image information and text description information of the sample products. During the training process, the pre-trained multimodal large language model learns how to enhance the image features related to the text description information. After the training in the first training stage, a multimodal large language model is obtained. In the second training stage, based on the multimodal large language model trained in the first training stage, the multimodal large language model is trained using the image information and text description information of the sample search terms and sample products. During the training process, the multimodal large language model learns how to enhance the image features related to the search terms and text description information. After training in the second training stage, a trained multimodal large language model is obtained. The image features related to the text description information and the search terms in the second multimodal feature vector generated by the trained multimodal large language model are enhanced. Compared with only learning image features related to text description information or image features related to search terms, the multimodal large language model trained by the present invention learns more comprehensive knowledge in the search scenario, and the generated second modal feature vector is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0049] FIG1 is a schematic diagram of the structure of a pre-trained multimodal large language model provided by an embodiment of the present disclosure;
[0050] FIG2 is a schematic diagram of an application process of a multimodal large language model in a search scenario provided by an embodiment of the present disclosure;
[0051] FIG3 is a schematic diagram of a process for performing correlation calculation based on a multimodal large language model provided by an embodiment of the present disclosure;
[0052] FIG4 is a schematic diagram of a process for generating product labels based on a multimodal large language model according to an embodiment of the present disclosure;
[0053] FIG5 is a schematic diagram of an implementation environment involved in the multimodal large language model training method, correlation calculation method, and label generation method provided in an embodiment of the present disclosure;
[0054] FIG6 is a flowchart of a multimodal large language model training method provided by an embodiment of the present disclosure;
[0055] FIG7 is a flowchart of a correlation calculation method provided by an embodiment of the present disclosure;
[0056] FIG8 is a flowchart of a label generation method provided by an embodiment of the present disclosure;
[0057] FIG9 is a schematic diagram of the structure of a multimodal large language model training device provided by an embodiment of the present disclosure;
[0058] FIG10 is a schematic structural diagram of a correlation calculation device provided by an embodiment of the present disclosure;
[0059] FIG11 is a schematic structural diagram of a label generation device provided by an embodiment of the present disclosure;
[0060] FIG12 shows a structural block diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0061] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.
[0062] It should be understood that the terms "each," "plurality," and "any" used in the embodiments of the present disclosure include two or more, each refers to each of the corresponding plurality, and any refers to any one of the corresponding plurality. For example, if a plurality of words includes 10 words, each refers to each of the 10 words, and any refers to any one of the 10 words.
[0063] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse. For example, the search terms and user preference information involved in this disclosure are all obtained with full authorization.
[0064] Before implementing the embodiments of the present disclosure, the terms involved in the embodiments of the present disclosure are first explained.
[0065] Large language models (LLMs) are deep learning models trained using large amounts of text data. They can generate natural language text or understand the meaning of text. Large language models can handle a variety of natural language tasks, such as text classification, intelligent question-answering, and conversation, and are a key path to artificial intelligence.
[0066] Blip2 (Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models) is a multimodal Transformer model that addresses the computationally expensive end-to-end training of Vision-Language Pre-training (VLP) models.
[0067] CLIP (Contrastive Language-Image Pre-Training) is a cross-modal pre-training model for contrast-based image-text learning.
[0068] Instruction learning is used to provide prompts or instructions to large language models, thereby inspiring them to obtain better output results.
[0069] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0070] With the research and advancement of artificial intelligence technology, it has been studied and applied in a variety of fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robotics, smart healthcare, and smart customer service. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solutions provided in the embodiments of this disclosure involve artificial intelligence technologies such as natural language processing, which will be specifically explained through the subsequent embodiments. Natural language processing is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, the language people use in daily life, and is closely related to the study of linguistics. Natural language processing technologies generally include text processing, semantic understanding, machine translation, robotic question answering, knowledge graphs, and other technologies.
[0071] In the search scenario, the products searched based on the search terms usually have image information and text description information (for example, title, product details page information, category, style, model, etc. to which the product belongs). To facilitate subsequent applications, a multimodal large language model will be called to fuse the product's image information and related text information (text description information or search terms) to obtain a multimodal feature vector in which the image features related to the text information are enhanced. The multimodal large language model in the related art only performs one stage of training, either learning how to enhance the image features related to the search terms, or learning how to enhance the image features related to the text description information. In fact, the image of the product is not only related to the search terms, but also to the text description information of the product itself, while the multimodal large language model in the related art can only generate a multimodal fusion feature vector in which the image features related to one type of text information (search terms or text description information) are enhanced. Therefore, the multimodal large language model trained by the related art is not accurate enough to meet the application requirements in the search scenario.
[0072] To this end, the present disclosure provides a method for training a large multimodal language model. The method adopts a two-stage training approach. In the first training stage, a pre-trained large multimodal language model is trained based on the image information and text description information of sample products, so that the large multimodal language model can learn how to enhance the image features related to the text description information, thereby generating a first multimodal feature vector with higher accuracy. In the second training stage, the multimodal feature vector trained in the first training stage is trained based on the image information and text description information of sample search terms and sample products, so that the trained large multimodal language model can learn how to enhance the image features related to the text description information and search terms. Specifically, it learns how to enhance the features related to the search terms in the first multimodal feature vector, thereby generating a second multimodal feature vector with higher accuracy.
[0073] FIG1 shows a schematic structural diagram of a multimodal large language model involved in an embodiment of the present disclosure. Referring to FIG1 , the multimodal large language model includes an image encoder, a text encoder, a Q-Former structure, and the like.
[0074] The image encoder encodes the image information to produce an image feature vector. This image information includes the main product image, its size, contrast, and color. The text encoder encodes the text description (or search term) to produce a text feature vector (or search term feature vector). The image feature vector encoded by the image encoder and the text feature vector (or search term feature vector) encoded by the text encoder have the same dimensionality, for example, 768 dimensions, to facilitate the Q-Former architecture to fuse the image feature vector and text feature vector (or search term feature vector).
[0075] The input of the Q-Former structure is two eigenvectors, and the output is a multimodal eigenvector after the fusion of the two eigenvectors, and the dimensions of the input eigenvector and the output eigenvector are the same. The Q-Former structure includes at least one self-attention module, at least one cross-attention module and at least one forward feedback module. The self-attention module is used to calculate the correlation between each element and other elements in the input eigenvector. The cross-attention module is used to calculate the attention between the two input eigenvectors to calculate the correlation between the two eigenvectors. The forward feedback module is used to enhance the nonlinear fitting ability of the Q-Former structure. It can be understood that the Q-Former structure shown in Figure 1 (including two cross-attention modules, one self-attention module and two forward feedback modules) is only an exemplary structure. According to actual computing requirements, the number of cross-attention modules, self-attention modules and forward feedback modules can be adjusted, and the positions of the cross-attention modules and self-attention modules can be adjusted.
[0076] The input and output of the multimodal large language model shown in Figure 1 are different at different stages. The following describes the different stages in detail.
[0077] In the first training stage, the corresponding multimodal large language model is a pre-trained multimodal large language model. For any sample product, the image information of the sample product is input into the image encoder of the pre-trained multimodal large language model, encoded by the image encoder, and a sample image feature vector is output. The text description information of the sample product is input into the text encoder of the pre-trained multimodal large language model, encoded by the text encoder, and a sample text feature vector is output. The sample image feature vector and the sample text feature vector are input into the Q-Former structure of the pre-trained multimodal large language model, processed in sequence by the cross-attention module, the self-attention module, the cross-attention module and the forward feedback module, and the sample first multimodal feature vector is output. Then, based on the sample first multimodal feature vectors, sample image feature vectors and sample text feature vectors of multiple sample products, the pre-trained multimodal large language model is trained.
[0078] In the second training phase, the corresponding multimodal large language model is the multimodal large language model trained in the first training phase. For any sample product, the image information of the sample product is input into the image encoder of the multimodal large language model trained in the first training phase, encoded by the image encoder, and an image feature vector is output. The text description information of the sample image is input into the text encoder of the multimodal large language model trained in the first training phase, encoded by the text encoder, and a text feature vector is output. The image feature vector and the text feature vector are input into the Q-Former structure of the multimodal large language model trained in the first training phase, processed in sequence by the cross-attention module, the self-attention module, the cross-attention module, and the forward feedback module, and the first multimodal feature vector is output. Next, the sample search term is input into the text encoder of the multimodal large language model trained in the first training phase, encoded by the text encoder, and the sample search term feature vector is output. Next, the sample search term feature vector and the first multimodal feature vector are input into the Q-Former structure of the multimodal large language model trained in the first training phase. They are processed sequentially through the cross-attention module, the self-attention module, the cross-attention module, and the forward feedback module to output the sample second multimodal feature vector. The multimodal large language model trained in the first training phase is then trained based on the sample second modal feature vectors of multiple sample products to obtain a trained multimodal large language model.
[0079] In a search scenario, for candidate products based on a search term, a relevance score can be calculated between the search term and the candidate products. This allows the product candidates to be ranked in descending order of relevance scores, thereby improving user click-through rates and conversion rates. Currently, related technologies calculate the relevance score between a search term and a candidate product by obtaining a search term feature vector corresponding to the search term and a text feature vector corresponding to the candidate product's text description. The semantic relevance score between the search term feature vector and the text feature vector is then calculated as the relevance score between the search term and the candidate product. This method only considers the semantic relevance score between the search term and the candidate product's text description, resulting in a relatively single dimension of relevance calculation and inaccurate results. In particular, when a candidate product's title and image do not match, or the title is cluttered or inconsistent with the candidate product's content, or is too short, the calculated semantic relevance score itself is unreliable and cannot represent the relevance between the search term and the candidate product.
[0080] In order to improve the accuracy of the correlation calculation results, the embodiment of the present disclosure considers the image-text correlation between the image information of the candidate product and the search term on the basis of the original correlation calculation, and adds a correlation calculation dimension. Specifically, based on the trained multimodal large language model, a second multimodal feature vector and a search term feature vector of the candidate product are generated, and then based on the second multimodal feature vector and the search term feature vector, the image-text correlation score between the candidate product and the search term is calculated, and then the semantic correlation score and the image-text correlation score are weighted to obtain the total correlation score between the candidate product and the search term. The calculation method of the embodiment of the present disclosure characterizes the correlation between the candidate product and the search term from multiple angles, and the correlation calculation result is more reliable. Furthermore, the canonical correlation between the search term and the user preference information is also considered, and then the total correlation score between the candidate product and the search term is obtained by weighted calculation of the semantic correlation score, the image-text correlation score and the canonical correlation score. Compared with only calculating the semantic correlation score, or only calculating the semantic correlation score and the image-text correlation score, the calculation result is more accurate.
[0081] In the search scenario, for candidate products searched based on search terms, appropriate labels can also be generated for the searched candidate products. The generated labels can be words that represent attributes, such as style, design, model, etc., or phrases containing specific scenarios, such as European and American style clothes, Japanese and Korean style clothes, summer clothes, winter clothes, etc., so as to better display products according to user portraits and search results to improve user click-through rates and conversion rates. When generating labels for candidate products, the related technology usually trains a discriminant model based on a deep neural network, and then processes the relevant information of the candidate product (for example, main picture, title, model, etc.) by calling the discriminant model to obtain the probability of each candidate label, and then uses the label with the highest probability as the label of the candidate product. The related technology relies heavily on the manual pre-definition of labels, and the number and content of the manually pre-defined labels are fixed and cannot be dynamically adjusted according to the needs of users in the search scenario. It can be seen that the label generation method of the related technology is not flexible enough and has certain limitations.
[0082] In order to be able to flexibly generate labels for candidate products based on user needs, the embodiment of the present disclosure pre-sets different instruction templates (prompts) for different levels of categories based on actual application scenarios, and generates a second multimodal feature vector of the candidate product based on the trained multimodal large language model. Then, based on the instruction template corresponding to a certain level of category of the candidate product, the large language model is guided. The large language model performs instruction learning based on the instruction template, and combines the title, image and other information of the candidate product indicated by the second multimodal feature vector to generate appropriate labels for the candidate product. The method provided by the embodiment of the present disclosure does not require manual pre-definition of labels for the product, and the label generation method is more flexible.
[0083] The multimodal large language model trained by the embodiment of the present disclosure can be used to calculate the correlation between the candidate products searched based on the search terms and the search terms, and can also be used to generate labels for the candidate products. Figure 2 shows a schematic diagram of the application process of the trained multimodal large language model. Referring to Figure 2, for any candidate product, the text description information of the candidate product is input into the text encoder of the trained multimodal large language model, and the text feature vector of the candidate product is output, and the image information of the candidate product is input into the image encoder of the trained multimodal large language model, and the image feature vector of the candidate product is output. The text feature vector and the image feature vector of the candidate product are then fused to obtain a first multimodal feature vector. Next, the search term is input into the text encoder of the trained multimodal large language model, and the search term feature vector is output. Then, the first multimodal feature vector and the search term feature vector are fused to obtain a second multimodal feature vector. Based on the obtained second multimodal feature vector, correlation calculation can be performed, and labels can also be generated for candidate products.
[0084] It should be noted that the above description uses the example of fusing the text feature vector and the image feature vector and then inputting the search term into the text encoder. The text description information and the search term can also be input into the text encoder in sequence, and after the text feature vector and the search term feature vector are output, the text feature vector and the image feature vector are fused. The embodiment of the present disclosure does not limit the timing of inputting the search term into the text encoder.
[0085] FIG3 shows a schematic diagram of a calculation process for correlation calculation based on a multimodal large language model provided by an embodiment of the present disclosure. Referring to FIG3 , the correlation calculation process includes a pre-process and an application process. The pre-process is the process of training a pre-trained multimodal large language model to obtain a trained multimodal large language model. Specifically, the image information and text information of sample search terms and sample products are used to train the pre-trained large language model to obtain a multimodal large language model. The application process is the process of calculating the image-text correlation score between the candidate product and the search term based on the search term feature vector and the second multimodal feature vector generated by the trained multimodal large language model, and then calculating the total correlation score of the candidate product. Specifically, when multiple candidate products are searched based on the search term, for any candidate product, the trained multimodal large language model can be called to process the search term and the image information and text information of the candidate product to obtain the search term feature vector and the second multimodal feature vector. Then, based on the search term feature vector and the second multimodal feature vector, the image-text correlation score between the candidate product and the search term is calculated. At the same time, a semantic relevance score between the candidate product and the search term can be calculated based on the search term and the candidate product's textual description. A canonical relevance score between the candidate product and the search term can also be calculated based on the search term and user preference information. Finally, a weighted calculation is performed on the semantic relevance score, canonical relevance score, and image-text relevance score to obtain the total relevance score between the candidate product and the search term.
[0086] FIG4 shows a schematic diagram of a generation process for generating labels for candidate products based on a multimodal large language model according to an embodiment of the present disclosure. Referring to FIG4 , the label generation process includes a pre-process and an application process. The pre-process is the process of training a pre-trained multimodal large language model to obtain a trained multimodal large language model. Specifically, the pre-trained large language model is trained using sample search terms and image information and text information of sample products to obtain a multimodal large language model. The application process is the process of generating labels for candidate products based on a second multimodal feature vector generated by an instruction template and a trained multimodal large language model. Specifically, when multiple candidate products are found based on a search term, for any candidate product, the trained multimodal large language model can be called to process the search term and the image information and text information of the candidate product to obtain a second multimodal feature vector. Then, the instruction template corresponding to the candidate product is obtained, and then the large language model is called to process the second multimodal feature vector and the instruction template to generate a corresponding label for the candidate product.
[0087] Please refer to Figure 5, which shows the implementation environment involved in the multimodal large language model training method, correlation calculation method and label generation method provided by the embodiment of the present disclosure. Referring to Figure 5, the implementation environment includes: terminal 501, server 502 and server 503.
[0088] Among them, terminal 501 can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. Server 502 and server 503 can be independent physical servers, or a server cluster or distributed system composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The above-mentioned terminal 501, server 502 and server 503 can be directly or indirectly connected through wired or wireless communication, and the embodiments of the present disclosure do not make specific limitations on this.
[0089] For the multimodal large language model training process, terminal 501 is used to obtain the sample search terms input by the user and send a search request to server 502, so that server 502 returns multiple sample products searched, each sample product has image information and text information. Server 502 has strong computing power and can search for multiple sample products based on sample search terms. Server 503 also has strong computing power and can train a multimodal large language model based on sample search terms, image information and text information of multiple sample products. The specific training process is not described in detail here. After the multimodal large language model training is completed, it can be directly deployed on server 503 and can be transplanted to other servers.
[0090] For the relevance calculation process, the terminal 501 is used to obtain the search term input by the user and send a search request to the server 502, so that the server 502 returns multiple candidate products for the search, each candidate product having image information and text information. The server 502 has strong computing power and can search for multiple candidate products based on the search term. If multiple large language models including a trained multimodal large language model are deployed on the server 502, the server 502 can calculate the total relevance score between the search term and each candidate product based on the search term, the text description information and image information of each candidate product, and the user preference information, and then return the multiple candidate products found in the search to the terminal 501 in descending order of the total relevance score. If multiple large language models other than the trained multimodal large language model are deployed on server 502, and a trained multimodal large language model is deployed on server 503, server 502 may call the multimodal large language model deployed on server 503 to calculate the image relevance score between the search term and each candidate product, and call the deployed large language model to calculate the semantic relevance score between the search term and each candidate product, and calculate the canonical relevance score between the search term and each candidate product based on user preference information, search terms, etc., and then calculate the total relevance score between the search term and each candidate product based on the semantic relevance score, image-text relevance score, and image-text relevance score between the search term and each candidate product.
[0091] For the label generation process, the terminal 501 is used to obtain the search term input by the user and send a search request to the server 502 so that the server 502 returns multiple candidate products found in the search, each candidate product having image information and text information. The server 502 has strong computing power and can search for multiple candidate products based on the search term. If multiple large language models including a trained multimodal large language model are deployed on the server 502, and instruction templates corresponding to different levels of categories are stored, the server 502 can call the multimodal large language model to process the search term, the text description information and image information of each candidate product, obtain the second multimodal feature vector corresponding to each candidate product, and obtain the instruction template corresponding to a certain level of category for each candidate product. Then, the large language model is called to process the second multimodal feature vector and instruction template corresponding to each candidate product to obtain a label for each candidate product. If instruction templates corresponding to categories of different levels are stored on server 502, but a trained multimodal large language model is not deployed, and a trained multimodal large language model is deployed on server 503, server 502 can call the multimodal large language model deployed on server 503 to process the search terms, the text description information and the image information of each candidate product to obtain the second multimodal feature vector corresponding to each candidate product, and obtain the instruction template corresponding to a certain level category of each candidate product. Then, the large language model is called to process the second multimodal feature vector and the instruction template corresponding to each candidate product to obtain the label of each candidate product.
[0092] Based on the implementation environment shown in FIG5 , the embodiment of the present disclosure provides a multimodal large language model training method. Taking the server 503 executing the embodiment of the present disclosure as an example, referring to FIG6 , the method flow provided by the embodiment of the present disclosure includes:
[0093] 601. Obtain sample search terms and sample products.
[0094] Among them, the sample search terms and sample products are training samples for training the multimodal large language model. The sample search terms can be input by the user, and the sample products can be obtained by searching based on the sample search terms. When searching for sample products based on the sample search terms, the search can be conducted in combination with at least one of the original text of the sample search terms, the semantic vector of the sample search terms, the user preference information, and the rewritten words of the sample search terms (i.e., words with the same or similar semantics as the sample search terms). Since the multimodal large language model to be trained in the embodiment of the present disclosure needs to fuse the image information and text description information of the products, the products searched based on each search term can be screened manually or by other machine methods, retaining products with both image information and text description information, or deleting products with image information but no text description information, or with text description information but no image information, and finally obtaining multiple sample products.
[0095] To facilitate the management of training samples, the sample products corresponding to each sample search term can be regarded as a training batch, and the training batches corresponding to multiple sample search terms can be combined into a training sample set. Each training batch includes multiple sample products corresponding to the sample search term, and these sample products include positive sample products and negative sample products. In different training stages, the meanings of the positive sample products and negative sample products included in each training batch are different. In the first training stage, the positive sample products are sample products whose image information and text information come from the same product; the negative sample products are sample products whose image information and text information come from different products. The negative sample products can be obtained by negative sampling of multiple sample products searched by multiple sample search terms, that is, the image information of the sample product searched by one sample search term is combined with the text description information of the sample product searched based on another sample search term. In the second training stage, the positive sample products are the sample products that match the sample search terms, that is, the sample products searched based on the sample search terms corresponding to the training batch; the negative sample products are the sample products that do not match the sample search terms, that is, the sample products that are not searched based on the sample search terms corresponding to the training batch. The negative sample products can be obtained by negative sampling of multiple sample products searched by multiple sample search terms, that is, combining the sample products searched by one sample search term with another sample search term.
[0096] Furthermore, in order to better complete the training task (i.e., objective function) set when training the multimodal large language model in the embodiment of the present disclosure, for both the first training stage and the second training stage, one positive sample product and multiple negative sample products can be combined into a training batch.
[0097] Furthermore, to facilitate subsequent calculations, each sample product can undergo image normalization and text normalization for its corresponding image information. Image normalization includes scaling and resizing. Text normalization includes removing stop words (e.g., the, an, that), normalizing word forms (e.g., unifying singular and plural forms), and more.
[0098] 602. Call a pre-trained multimodal large language model to process the image information and text description information of the sample product to obtain a sample text feature vector, a sample image feature vector, and a sample first multimodal feature vector of the sample product.
[0099] Among them, the pre-trained multimodal large language model can be Blip2, CLIP, etc.
[0100] Specifically, a pre-trained multimodal large language model is called to process the image information and text description information of the sample product to obtain a sample text feature vector, a sample image feature vector, and a sample first multimodal feature vector of the sample product, including:
[0101] 6021. Call the pre-trained multimodal large language model to extract features from the text description information of the sample product to obtain a sample text feature vector of the sample product.
[0102] In the present disclosure, the pre-trained multimodal large language model includes a text encoder, which is used to encode text description information. When the text description information of each sample product is input into the text encoder, the text encoder extracts and encodes the text features to obtain a sample text feature vector for each sample product.
[0103] 6022. Perform feature extraction on the image information of the sample product to obtain a sample image feature vector of the sample product.
[0104] The pre-trained multimodal large language model disclosed herein includes an image encoder for encoding image information. When the image information of each sample product is input into the image encoder, the image encoder extracts and encodes the image features to obtain a sample image feature vector for each sample product.
[0105] 6023. Fuse the sample text feature vector and the sample image feature vector of the sample product to obtain a sample first multimodal feature vector of the sample product.
[0106] The pre-trained multimodal large language model disclosed herein includes a Q-Former structure, which is used to fuse input feature vectors. For each sample product, when the sample text feature vector and sample image feature vector of the sample product are input into the Q-Former structure, the sample text feature vector and the sample image feature vector are processed by the cross-attention module, self-attention module, forward feedback module, etc. in the Q-Former structure to obtain the sample first multimodal feature vector of the sample product. The sample first multimodal feature vector is a feature vector obtained by fusing the sample text feature vector and the sample image feature vector. The image features related to the text description information in the sample first multimodal feature vector are enhanced. When the sample first multimodal feature vector is used to train the pre-trained multimodal large language model, the trained multimodal large language model can learn how to enhance the image features related to the text description information in the image feature vector.
[0107] 603. Based on the sample text feature vector, the sample image feature vector, and the sample first multimodal feature vector of the sample product, the pre-trained multimodal large language model is trained to obtain the multimodal large language model.
[0108] The sample products are labeled with labels that represent the correlation between the image information and the text description information, including relevant labels and irrelevant labels. If the sample product is a positive sample product, the label labeled with the sample product is the relevant label; if the sample product is a negative sample product, the label labeled with the sample product is the irrelevant label.
[0109] Specifically, based on the sample text feature vector, sample image feature vector, and sample first multimodal feature vector of the sample product, the pre-trained multimodal large language model is trained to obtain the multimodal large language model, including:
[0110] 6031. Based on the sample first multimodal feature vector of the sample product, obtain a label representing the correlation between the image information and text description information of the sample product.
[0111] The embodiment of the present disclosure adds a fully connected layer on the basis of the pre-trained multimodal large language model. The fully connected layer corresponds to a binary classification training task. By inputting the first multimodal feature vector of each sample product into the fully connected layer, a label representing the correlation between the image information and text description information of each sample product can be output.
[0112] 6032. Based on the first total objective loss function, the sample text feature vector and the sample image feature vector of the sample product, and the generated label and the annotation label corresponding to the sample product, the pre-trained multimodal large language model is trained to obtain a multimodal large language model.
[0113] To better adapt to downstream tasks, the disclosed embodiment adds two training tasks to each of the two training stages, namely, image-text matching task (ITM) and image-text contrastive learning task (ITC). The tasks aim to establish matching and correspondence between images and texts, align the representations of images and texts, maximize their mutual information, and thus obtain a multimodal feature vector.
[0114] The image-text matching task is a binary classification task that requires the model to predict whether an image-text pair matches or does not match. For the first training phase, the goal is to ensure that the label representing the correlation between the image information and the text description for each sample product is consistent with the generated label. For the second training phase, the goal is to ensure that the label representing the correlation between the sample search term and the first multimodal feature vector for each sample product is consistent with the generated label.
[0115] The goal of the image-text comparison task is to maximize the similarity of positive samples and minimize the similarity of negative samples. For the first training phase, the goal of the image-text matching task is to minimize the distance between the relevant sample image feature vectors and the sample text feature vectors, and minimize the distance between the irrelevant sample image feature vectors and the sample text feature vectors. For the second training phase, the goal of the image-text matching task is to minimize the distance between the relevant sample search term feature vectors and the first multimodal feature vectors, and minimize the distance between the irrelevant sample search term feature vectors and the first multimodal feature vectors.
[0116] Different total target loss functions are set in the two training stages for the above two training tasks. Among them, the first training stage corresponds to the first total target loss function, which includes the first image-text matching loss function and the first image-text contrast loss function. The first image-text matching loss function needs to complete the image-text matching loss task of the first training stage, and the first image-text contrast loss function needs to complete the image-text contrast learning task of the first training stage; the second training stage corresponds to the second total target loss function, which includes the second image-text matching loss function and the second image-text contrast loss function. The second image-text matching loss function needs to complete the image-text matching loss task of the second training stage, and the second image-text contrast loss function needs to complete the image-text contrast learning task of the second training stage.
[0117] In the first training phase, the pre-trained multimodal large language model is trained based on the first overall objective loss function, the sample text feature vectors and sample image feature vectors of the sample products, and the generated labels and annotated labels corresponding to the sample products to obtain the multimodal large language model, including the following steps:
[0118] 60321. Input the generated label and the annotated label corresponding to the sample product into the first image-text matching loss function to obtain the first image-text matching loss function value.
[0119] 60322. Input the sample text feature vector and the sample image feature vector into the first image-text contrast loss function to obtain the first image-text contrast loss function value.
[0120] 60323. Determine a first total target loss function value based on the first image-text matching loss function value and the first image-text comparison loss function value.
[0121] Obtain the corresponding weight values of the first image-text matching loss function and the first image-text contrast loss function in the first total target loss function, perform weighted calculation on the first image-text matching loss function value and the first image-text contrast loss function value, and obtain the first total target loss function value.
[0122] 60324. Based on the first total objective loss function value, adjust the model parameters of the pre-trained multimodal large language model to obtain the multimodal large language model.
[0123] The optimization goal of the first training phase is to minimize the value of the first total objective loss function. If the value of the first total objective loss function is greater than a first preset threshold, the model parameters of the pre-trained multimodal large language model need to be adjusted. Then, based on the pre-trained multimodal large language model after the model parameters are adjusted, the image information and text information of each sample product are processed to obtain the sample image feature vector, sample text feature vector, and sample first multimodal feature vector of each sample product. Then, based on the sample image feature vector, sample text feature vector, and sample first multimodal feature vector of each sample product, the pre-trained multimodal large language model after the model parameters are adjusted is trained until the function value of the first total objective loss function is less than the first preset threshold, or the number of training times reaches the number threshold. The pre-trained multimodal large language model obtained at the end of the training is used as the multimodal large language model trained in the first training phase (for the convenience of subsequent description, referred to as the multimodal large language model), and then it is retrained in the second training phase.
[0124] 604. Call a large multimodal language model to process the sample search terms and the image information and text description information of the sample products to obtain a sample search term feature vector and a first multimodal feature vector and a sample second multimodal feature vector for each sample product.
[0125] The image features related to the text description and search terms in the second multimodal feature vector of the sample are enhanced. In the second training phase, the multimodal large language model trained in the first phase is trained in batches. Each training batch includes multiple sample products, including positive and negative samples.
[0126] Specifically, a multimodal large language model is called to process the image information and text description information of the sample search term and the sample product to obtain the sample search term feature vector and the first multimodal feature vector and the sample second multimodal feature vector of the sample product, including:
[0127] 6041. Call the multimodal large language model to extract features from the text description information of the sample product to obtain the text feature vector of the sample product.
[0128] The multimodal large language model trained in the first training phase includes a text encoder, which encodes text description information. When the text description information of each sample product is input into the text encoder, the text encoder extracts and encodes the text features to obtain a text feature vector for each sample product.
[0129] 6042. Perform feature extraction on the image information of the sample product to obtain an image feature vector of the sample product.
[0130] The multimodal large language model trained in the first training phase includes an image encoder, which is used to encode image information. When the image information of each sample product is input into the image encoder, the image encoder extracts and encodes the image features, generating an image feature vector for each sample product.
[0131] 6043. Fuse the text feature vector and the image feature vector of the sample product to obtain a first multimodal feature vector of the sample product.
[0132] The large multimodal language model trained in the first training phase includes a Q-Former structure, which is used to fuse input feature vectors. For each sample product, the text feature vector and image feature vector are input into the Q-Former structure. These vectors are then processed through the Q-Former's cross-attention module, self-attention module, and forward feedback module to generate the first multimodal feature vector for the sample product.
[0133] 6044. Perform feature extraction on the sample search term to obtain a feature vector of the sample search term.
[0134] The sample search term is input into the text encoder of the multimodal large language model trained in the first training phase. The text encoder extracts the search term features and encodes them to obtain the sample search term feature vector.
[0135] 6045. Fuse the sample search term feature vector with the first multimodal feature vector of the sample product to obtain a sample second multimodal feature vector of the sample product.
[0136] The sample search term feature vector and the first multimodal feature vector of each sample product are respectively input into the Q-Former structure. The sample search term feature vector and the first multimodal feature vector of each sample product are processed by the cross-attention module, self-attention module, forward feedback module, etc. in the Q-Former structure to obtain the sample second multimodal feature vector of each sample product.
[0137] 605. Based on the sample search term feature vector and the first multimodal feature vector and the second multimodal feature vector of the sample product, a multimodal large language model is trained to obtain a trained multimodal large language model.
[0138] Each sample product is annotated with a label representing its relevance to the sample search term. Specifically, based on the sample search term feature vector and the first multimodal feature vector and the second multimodal feature vector of each sample product, a large multimodal language model is trained to obtain a trained large multimodal language model, including the following steps:
[0139] 6051. Based on the sample second multimodal feature vector of each sample product, obtain a label representing the correlation between each sample product and the sample search term.
[0140] The embodiment of the present disclosure adds a fully connected layer on the basis of the multimodal large language model. The fully connected layer corresponds to a binary classification training task. By inputting the second multimodal feature vector of each sample product into the fully connected layer, a label representing the correlation between each sample product and the sample search term can be output.
[0141] 6052. Based on the second total objective loss function, the first multimodal feature vector and the sample search term feature vector of each sample product, and the generated label and the annotated label corresponding to each sample product, the multimodal large language model is trained to obtain a trained multimodal large language model.
[0142] The second overall objective loss function includes a second image-text matching loss function and a second image-text comparison loss function. In the second training phase, the multimodal large language model is trained based on the second overall objective loss function, the first multimodal feature vector and sample search term feature vector of each sample product, and the generated label and annotated label corresponding to each sample product to obtain a trained multimodal large language model, including the following steps:
[0143] 60521. Input the generated label and the annotated label corresponding to each sample product into the second image-text matching loss function to obtain the second image-text matching loss function value.
[0144] 60522. Input the first multimodal feature vector and the sample search term feature vector of each sample product into the second image-text contrast loss function to obtain the second image-text contrast loss function value.
[0145] In the embodiment of the present disclosure, the second image-text contrast loss function can be expressed as:
[0146] Among them, L q is the second image-text contrast loss function, q is the search term feature vector, k+ is the first multimodal feature vector of the positive sample items in the training batch, that is, the first multimodal feature vector related to the search term feature vector, ki is the first multimodal feature vector of the negative sample items in the training batch, that is, the first multimodal feature vector unrelated to the search term feature vector, and t is the temperature coefficient (custom hyperparameter).
[0147] 60523. Determine a second total target loss function value based on the second image-text matching loss function value and the second image-text comparison loss function value.
[0148] According to the weight values corresponding to the second image-text matching loss function and the second image-text contrast loss function in the second total objective loss function, the second image-text matching loss function value and the second image-text contrast loss function value are weightedly calculated to obtain the second total objective loss function value.
[0149] 60524. Based on the second total objective loss function value, adjust the model parameters of the multimodal large language model to obtain a trained multimodal large language model.
[0150] The optimization goal of the second training stage is to minimize the second total objective loss function value. If the second total objective loss function value is greater than the second preset threshold, it is necessary to adjust the model parameters of the multimodal large language model, and then based on the multimodal large language model after the model parameters are adjusted, the sample search terms and the image information and text information of each sample product are processed to obtain the sample search term feature vector, the first multimodal feature vector of each sample product, and the sample second multimodal feature vector. Then, based on the sample search term feature vector, the first multimodal feature vector of each sample product, and the sample second multimodal feature vector, the parameter-adjusted multimodal large language model is trained until the function value of the second total objective loss function is less than the second preset threshold, or the number of training times reaches the number threshold. The multimodal large language model obtained at the end of the training is used as the multimodal large language model trained in the second training stage, that is, the trained multimodal large language model described in the embodiment of the present disclosure. The trained multimodal large language model is used to generate the second multimodal feature vector of the product based on the search terms and the image information and text description information of the product.
[0151] After the two training stages mentioned above, the trained multimodal large language model has learned how to extract image features related to search terms and text description information from the product image information. The second multimodal feature vector generated by calling the trained multimodal large language model is more accurate and can reflect the characteristics of the product after the image and text are integrated from multiple angles.
[0152] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.
[0153] Based on the implementation environment shown in FIG5 , an embodiment of the present disclosure provides a correlation calculation method. This method applies the trained multimodal large language model shown in FIG6 . Taking server 502 executing the embodiment of the present disclosure as an example, referring to FIG7 , the method flow provided by the embodiment of the present disclosure includes:
[0154] 701. Obtain a search term and candidate products found based on the search term.
[0155] In an e-commerce search scenario, when a user wants to purchase a product, they can enter a user-defined search term related to the product on a client (e.g., an application, mini-program, or webpage) providing the product purchase service. In response to the user's search operation on the client, the terminal generates a search request containing the search term and sends it to the server. Upon receiving the search request, the server searches based on the search term in the request, obtaining multiple candidate products. Each candidate product has image information and text description information, allowing for descriptions of the candidate product from different perspectives. When describing these candidate products based on the image information and text description information, it is found that these candidate products have varying relevance to the search term, with some having greater relevance to the search term and others having less. If the products are simply displayed to the user in the order in which they were searched, the less relevant candidate products may be displayed at the front, while the more relevant candidate products may be displayed at the back. Users will then browse these candidate products in descending order of their display position. Because the top-ranked candidate products have little relevance to the search term and fail to meet the user's search needs, users are unlikely to continue browsing these candidate products, resulting in low click-through and conversion rates for these candidate products. To improve recognition accuracy in e-commerce search, the correlation between the search term and multiple candidate products can be calculated. This allows the products to be displayed to users in descending order of relevance, improving the user experience while increasing user click-through and conversion rates for the candidate products.
[0156] 702. Calculate a semantic relevance score between the search term and the candidate product based on the search term and the text description information.
[0157] Among them, the semantic relevance score is used to reflect the degree of relevance between the search term and the candidate product from a semantic perspective. The larger the semantic relevance score, the higher the degree of relevance between the search term and the candidate product from a semantic perspective; the smaller the semantic relevance score, the lower the degree of relevance between the search term and the candidate product from a semantic perspective. For any candidate product, when calculating the semantic relevance score between the search term and the candidate product, the text description information of the search term and the candidate product can be input into the text encoder of the large language model respectively. After encoding by the text encoder, the search term feature vector and the text feature vector are output, and then the similarity between the search term feature vector and the text feature vector is calculated. The similarity calculation result is then used as the semantic relevance score between the search term and the candidate product. The above-mentioned large language model can be any large language model with a text encoder in the related art (the reason why the multimodal large language model trained in the embodiment of the present disclosure is not used here is that any model has a certain calculation accuracy. If one model is used to calculate different relevance scores, the calculation result will be limited by the accuracy of the model. If different models are used to calculate different relevance scores, the accuracy of the relevance score can be improved by adjusting the weight values corresponding to different models). When calculating the similarity between the search term feature vector and the text feature vector, the cosine similarity formula may be used for calculation.
[0158] 703. Call the trained multimodal large language model to process the search term, image information, and text description information to obtain a second multimodal feature vector and a search term feature vector of the candidate product.
[0159] In the embodiment of the present disclosure, the dimensions of the second multimodal feature vector and the search term feature vector are relatively high. In order to reduce the amount of calculation in the subsequent image-text correlation calculation, the second multimodal feature vector and the search term feature vector can be subjected to dimensionality reduction processing. In one possible implementation, a fully connected layer can be added on the basis of the trained multimodal large language model, and the fully connected layer can be used to perform a linear transformation on the second multimodal feature vector and the search term feature vector, so as to directly output a second multimodal feature vector and a search term feature vector with a lower dimension. For example, the second multimodal feature vector and the search term feature vector can be reduced from 768 dimensions to 256 dimensions. In another possible implementation, the structure of the trained multimodal large language model can be improved, and a fully connected layer can be added to the above-mentioned trained multimodal large language model. When the trained multimodal large language model is called to process the search term, image information and text description information, the second multimodal feature vector and the search term feature vector with a lower dimension can be directly obtained. Of course, other methods may also be used to perform dimensionality reduction processing on the second multimodal feature vector and the search term feature vector, which will not be described one by one in the embodiment of the present disclosure.
[0160] 704. Calculate the image-text relevance score between the search term and the candidate product based on the search term feature vector and the second multimodal feature vector.
[0161] Among them, the image-text relevance score is used to reflect the degree of relevance between the search term and the candidate product in terms of image and text, that is, the degree of relevance between the semantics of the search term and the image of the candidate product. The larger the image-text relevance score, the higher the degree of relevance between the search term and the candidate product in terms of image and text; the smaller the image-text relevance score, the lower the degree of relevance between the search term and the candidate product in terms of image and text. When calculating the image-text relevance score between the search term and the candidate product based on the search term feature vector and the second multimodal feature vector, the cosine similarity between the search term feature vector and the second multimodal feature vector can be calculated, and the obtained cosine similarity value can be used as the image-text relevance score. Among them, the calculation formula of cosine similarity is:
[0162] Among them, Similarity is the cosine similarity value, A is the search term feature vector, B is the second multimodal feature vector, A i is the element on the i-th dimension in the search word feature vector, B i is the element on the i-th dimension in the second multimodal feature vector, and n is the number of dimensions of the search term feature vector (or the second multimodal feature vector).
[0163] 705. Based on the semantic relevance score and the image-text relevance score, calculate the total relevance score between the search term and the candidate product.
[0164] In the disclosed embodiment, different weight values can be pre-configured for the semantic relevance score and the image-text relevance score, respectively. Then, based on the weight values corresponding to the semantic relevance score and the image-text relevance score, the semantic relevance score and the image-text relevance score are weighted and calculated to obtain the total relevance score between the search term and the candidate product. The weight values configured for the semantic relevance score and the image-text relevance score are not fixed and can be dynamically adjusted based on the application results in the actual search scenario. When configuring different weight values for the semantic relevance score and the image-text relevance score, the following two methods are included but not limited to:
[0165] The first method calculates the proportion of semantic relevance scores and image-text relevance scores across categories. Offline fitting is performed based on this proportion to obtain the weights corresponding to the semantic relevance scores and image-text relevance scores for each category. In practice, this method can be used to obtain the weights corresponding to the semantic relevance scores and image-text relevance scores for each category based on the category to which the candidate product belongs.
[0166] Second, a model can be trained to determine the weights corresponding to the semantic relevance score and the image-text relevance score. This model can be used to determine the weights corresponding to the semantic relevance score and the image-text relevance score, respectively, when applied.
[0167] Furthermore, considering that different users have different preferences, user preferences will also affect the user's clicks and conversions on candidate products. In view of this, the canonical correlation score between the search term and the candidate product can be calculated based on the search term and user preference information. The canonical correlation score is used to reflect the degree of relevance between the search term and the candidate product from the perspective of user preference. If it is determined based on the text description information of the candidate product that the candidate product meets the user's preference, it can be determined that the canonical correlation score between the search term and the candidate product is high; if it is determined based on the text description information of the candidate product that the candidate product does not meet the user's preference, it can be determined that the canonical correlation score between the search term and the candidate product is low. Based on the obtained semantic correlation score, image-text correlation score and canonical correlation score, the total correlation score between the search term and the candidate product can be obtained by weighted calculation of the semantic correlation score, image-text correlation score and canonical correlation score.
[0168] Compared to existing relevance calculation methods in some e-commerce scenarios, the disclosed embodiment, based on a multimodal large language model, simultaneously considers product image information, text description information, and more, resulting in a more robust relevance score and more accurate calculation results. Because the relevance score no longer relies on the semantic relevance score between the search term and the text description information, even if there are some image-text mismatches, or product titles are cluttered or too short, better calculation results can be obtained by calculating the image-text relevance score and the regular relevance score, and adjusting the weights of the different relevance scores.
[0169] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.
[0170] Based on the implementation environment shown in FIG2 , an embodiment of the present disclosure provides a tag generation method. This method applies the multimodal large language model trained in FIG6 . Taking the server 502 in FIG1 as an example, referring to FIG8 , the method flow provided by the embodiment of the present disclosure includes:
[0171] 801. Call the trained multimodal large language model to process the search terms, image information, and text description information corresponding to the candidate product to obtain a second multimodal feature vector.
[0172] In search scenarios, candidate products are typically tagged so they can be displayed based on user profiles and search results, thereby increasing click-through and conversion rates. Before generating labels for these candidate products, images can be normalized, such as by scaling and resizing them. Text descriptions can also be normalized, such as by removing stop words and unifying capitalization, to serve as input for subsequent models.
[0173] For any candidate product, a trained large multimodal language model is invoked to process the search term, image information, and text description information corresponding to the candidate product to obtain a second multimodal feature vector. Specifically, the title, category, product details page information, model number, etc. included in the text description information can be concatenated, then the text encoder of the trained large multimodal language model is invoked to encode the text to obtain a text feature vector. The image encoder of the trained large multimodal language model is then invoked to encode the image information to obtain an image feature vector. The text feature vector and image feature vector are then input into the Q-Former of the trained large multimodal language model to output a first multimodal feature vector. Next, the search term is input into the text encoder of the trained large multimodal language model to encode the text, outputting a search term feature vector. The search term feature vector and the first multimodal feature vector are then input into the Q-Former of the trained large multimodal language model to output a second multimodal feature vector corresponding to the candidate product. The image features related to the search term and text description information in this second multimodal feature vector are enhanced, reflecting the various attributes of the candidate product.
[0174] Furthermore, in the embodiment of the present disclosure, the dimension of the second multimodal feature vector is relatively high. In order to reduce the amount of calculation in subsequent calculations, the second multimodal feature vector can be subjected to dimensionality reduction processing. In one possible implementation, a fully connected layer can be added on the basis of the trained multimodal large language model, and the second multimodal feature vector can be linearly transformed by using the fully connected layer, so as to directly output a second multimodal feature vector with a lower dimension, for example, reducing the second multimodal feature vector from 768 dimensions to 256 dimensions. In another possible implementation, the structure of the trained multimodal large language model can be improved, and a fully connected layer can be added to the above-mentioned trained multimodal large language model. When the trained multimodal large language model is called to process the search terms, image information and text description information, a second multimodal feature vector with a lower dimension can be directly obtained. Of course, other methods can also be used to reduce the dimensionality of the second multimodal feature vector, which will not be described one by one in the embodiment of the present disclosure.
[0175] It should be noted that the above takes the example of generating labels for candidate products based on the search terms input by the user. If the candidate products for which labels need to be generated are not searched based on the search terms input by the user, but are identified by the merchant, then the search terms corresponding to the candidate products can be the search terms given by the server that are most likely to correspond to the candidate products.
[0176] 802. Obtain a first instruction template corresponding to the preset level category to which the candidate product belongs.
[0177] Considering that products in different categories have different labels—for example, clothing products might be labeled by style, while sports products might be labeled by material, weight, brand, etc.—the disclosed embodiments can classify products into different levels of categories for different application scenarios and set different instruction templates for each level. Each level of category has at least one instruction template. The instruction templates are designed natural language instructions that describe the attributes of the candidate products to be output, guiding the model to generate the corresponding attributes.
[0178] For example, for clothing products, the configured instruction template can be:
[0179] 1. The material of this product is (please output the material of the product). This instruction template is used to guide the model to generate the material of the clothing;
[0180] 2. Please give the style of the corresponding clothes according to the information (Please generate the corresponding product style according to the provided information). This instruction template is used to guide the model to generate clothing styles.
[0181] For example, in the 3C electronics field, the configured instruction template can be:
[0182] Please output the performance parameters of this product. This instruction template is used to guide the model to generate the performance parameters of the product.
[0183] Furthermore, to facilitate subsequent applications, the correspondence between categories at different levels and instruction templates may be stored.
[0184] In the disclosed embodiment, for any candidate product, the preset level category to which the candidate product belongs can be obtained from the text description information of the candidate product. Then, based on the preset level category to which the candidate product belongs, the first instruction template corresponding to the preset level category to which the candidate product belongs is obtained from the correspondence between different level categories and instruction templates. The preset level category can be set as needed and can be a first-level category, a second-level category, etc.
[0185] 803. Generate a first label for the candidate product under a preset level category based on the second multimodal feature vector and the first instruction template.
[0186] In the embodiment of the present disclosure, the second multimodal feature vector can reflect the different attributes of the candidate product, and the first instruction template is used to indicate the attributes of the candidate product to be output. When the second multimodal feature vector and the first instruction template are input into the large language model, the large language model processes the second multimodal feature vector through a multi-layer transformer according to the instructions of the first instruction template, and outputs the word with the highest probability in the context indicated by the first instruction template as the first label generated for the candidate product. The number of the first labels is the same as the number of the first instruction templates. If there is one first instruction template, one first label is generated. If there are at least two first instruction templates, at least two labels are generated. In order to distinguish the first labels generated based on each first instruction template, after outputting the word with the highest probability in the context indicated by the first instruction template based on each first instruction template, an end symbol can be added to the end of the word to indicate the end of one output. The end symbol can be <eof>By using the method provided by the embodiment of the present disclosure, at least one first label can be generated for each candidate product, and by combining these first labels, a final label for each candidate product can be obtained.
[0187] Considering that large language models have certain accuracy issues, in order to improve the reliability of the first label generated for the candidate product, the first label corresponding to the candidate product can be manually evaluated. If the evaluation result of the first label is reliable, the first label corresponding to the candidate product is retained; if the evaluation result of the first label is unreliable, the user can re-set the corresponding instruction template for the preset category to which the candidate product belongs, that is, the second instruction template. The server obtains the second instruction template entered by the user for the preset level category and updates the first instruction template corresponding to the preset level category in the stored correspondence to the second instruction template. Optionally, duplicate labels can be filtered out during the evaluation process to avoid interference from duplicate labels to users.
[0188] Furthermore, after the first instruction template corresponding to the preset level category in the corresponding relationship is updated to the second instruction template, a second label of the candidate product under the preset level category will be generated based on the second multimodal feature vector and the second instruction template of the candidate product. Specifically, the large language model can be called again to process the second multimodal feature vector and the second instruction template of the candidate product to obtain the second label of the candidate product under the preset level category. Of course, the generated second label also needs to be manually evaluated. If the evaluation result is unreliable, the user needs to set the corresponding instruction template for the preset category to which the candidate product belongs again until the evaluation result of the generated label is reliable.
[0189] The disclosed embodiments can construct different instruction templates based on the needs of actual scenarios. Using instruction / instruction learning, the downstream large language model is guided to generate targeted labels for products based on instruction templates, combined with the product's text description information and image information, thereby greatly saving the cost of manual labeling. Compared with the traditional method of matching existing template libraries based on neural networks, this method generates more diverse labeling effects.
[0190] In a search scenario, by using the labels generated for commodities according to the embodiments of the present disclosure and combining them with user portraits, products with corresponding or similar labels can be displayed to users, thereby improving the user's click-through rate and conversion rate.
[0191] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.
[0192] Please refer to FIG9 , which shows a schematic diagram of the structure of a multimodal large language model training device provided by an embodiment of the present disclosure. The device can be implemented by software, hardware, or a combination of both, and becomes all or part of an electronic device. The device includes:
[0193] The first acquisition module 901 is used to acquire sample search terms and sample products, where the sample products have image information and text description information;
[0194] A first processing module 902 is configured to call a pre-trained multimodal large language model to process the image information and text description information of the sample product to obtain a sample text feature vector, a sample image feature vector, and a sample first multimodal feature vector of the sample product, wherein the image features related to the text description information in the sample first multimodal feature vector are enhanced;
[0195] A first training module 903 is configured to train a pre-trained multimodal large language model based on the sample text feature vector, the sample image feature vector, and the sample first multimodal feature vector of the sample product to obtain a multimodal large language model;
[0196] A second processing module 904 is configured to call a large multimodal language model to process the sample search term and the image information and text description information of the sample product to obtain a sample search term feature vector, a first multimodal feature vector of the sample product, and a sample second multimodal feature vector. Image features related to the text description information and the search term in the sample second multimodal feature vector are enhanced.
[0197] The second training module 905 is used to train the multimodal large language model based on the sample search term feature vector and the first multimodal feature vector and the sample second multimodal feature vector of the sample product to obtain a trained multimodal large language model. The trained multimodal large language model is used to generate the second multimodal feature vector of the product based on the search term and the image information and text description information of the product.
[0198] In another embodiment of the present disclosure, the first processing module 902 is used to call a pre-trained multimodal large language model to perform feature extraction on the text description information of the sample product to obtain a sample text feature vector of the sample product; perform feature extraction on the image information of the sample product to obtain a sample image feature vector of the sample product; and fuse the sample text feature vector and the sample image feature vector of the sample product to obtain a sample first multimodal feature vector of the sample product.
[0199] In another embodiment of the present disclosure, the sample product is annotated with a label representing the correlation between the image information and the text description information. The first training module 903 is used to obtain the label representing the correlation between the image information and the text description information of the sample product based on the sample first multimodal feature vector of the sample product; based on the first total objective loss function, the sample text feature vector and the sample image feature vector of the sample product, and the generated label and the annotated label corresponding to the sample product, the pre-trained multimodal large language model is trained to obtain the multimodal large language model.
[0200] In another embodiment of the present disclosure, the first total target loss function includes a first image-text matching loss function and a first image-text contrast loss function. The first training module 903 is used to input the generated label and the annotated label corresponding to the sample product into the first image-text matching loss function to obtain the first image-text matching loss function value; input the sample text feature vector and the sample image feature vector into the first image-text contrast loss function to obtain the first image-text contrast loss function value; determine the first total target loss function value based on the first image-text matching loss function value and the first image-text contrast loss function value; and adjust the model parameters of the pre-trained multimodal large language model based on the first total target loss function value to obtain the multimodal large language model.
[0201] In another embodiment of the present disclosure, the second processing module 904 is used to call the multimodal large language model, perform feature extraction on the text description information of the sample product to obtain the text feature vector of the sample product; perform feature extraction on the image information of the sample product to obtain the image feature vector of the sample product; fuse the text feature vector and the image feature vector of the sample product to obtain the first multimodal feature vector of the sample product; perform feature extraction on the sample search term to obtain the sample search term feature vector; and fuse the sample search term feature vector with the first multimodal feature vector of the sample product to obtain the sample second multimodal feature vector of the sample product.
[0202] In another embodiment of the present disclosure, the sample product is annotated with a label representing the correlation between the sample product and the sample search term. The second training module 905 is used to obtain the label representing the correlation between the sample product and the sample search term based on the sample second multimodal feature vector of the sample product; based on the second total objective loss function, the first multimodal feature vector and the sample search term feature vector of the sample product, and the generated label and the annotated label corresponding to the sample product, the multimodal large language model is trained to obtain a trained multimodal large language model.
[0203] In another embodiment of the present disclosure, the second total objective loss function includes a second image-text matching loss function and a second image-text contrast loss function. The second training module 905 is used to input the generated label and the annotated label corresponding to the sample product into the second image-text matching loss function to obtain the second image-text matching loss function value; input the first multimodal feature vector and the sample search term feature vector of the sample product into the second image-text contrast loss function to obtain the second image-text contrast loss function value; determine the second total objective loss function value based on the second image-text matching loss function value and the second image-text contrast loss function value; and adjust the model parameters of the multimodal large language model based on the second total objective loss function value to obtain a trained multimodal large language model.
[0204] Please refer to FIG10 , which shows a schematic diagram of the structure of a correlation calculation device provided by an embodiment of the present disclosure. The device uses the trained multimodal large language model in FIG6 . The device can be implemented by software, hardware, or a combination of both and can become all or part of an electronic device. The device includes:
[0205] Acquisition module 1001, for acquiring search terms and candidate products found based on the search terms, wherein the candidate products have image information and text description information;
[0206] A first calculation module 1002 is configured to calculate a semantic relevance score between the search term and the candidate product based on the search term and the text description information;
[0207] Processing module 1003 is used to call the trained multimodal large language model to process the search term, image information, and text description information to obtain a second multimodal feature vector and a search term feature vector for the candidate product;
[0208] A second calculation module 1004 is configured to calculate a graphic-text relevance score between the search term and the candidate product based on the search term feature vector and the second multimodal feature vector;
[0209] The third calculation module 1005 is used to calculate the total relevance score between the search term and the candidate product based on the semantic relevance score and the image-text relevance score.
[0210] In another embodiment of the present disclosure, the second calculation module 1004 is configured to calculate the cosine similarity between the search term feature vector and the second multimodal feature vector to obtain a picture-text relevance score.
[0211] In another embodiment of the present disclosure, the third calculation module 1005 is configured to perform weighted calculation on the semantic relevance score and the image-text relevance score to obtain a total relevance score.
[0212] In another embodiment of the present disclosure, the apparatus further comprises:
[0213] A fourth calculation module is used to calculate the canonical correlation score between the search term and the product based on the text description information and the user preference information;
[0214] The fifth calculation module is used to perform weighted calculation on the semantic relevance score, the image-text relevance score and the regular relevance score to obtain the total relevance score between the search term and the candidate product.
[0215] Please refer to FIG11 , which shows a schematic diagram of the structure of a label generation device provided by an embodiment of the present disclosure. The device uses the trained multimodal large language model in FIG6 . The device can be implemented through software, hardware, or a combination of both, and becomes all or part of an electronic device. The device includes:
[0216] Processing module 1101 is used to call the trained multimodal large language model to process the search terms, image information, and text description information corresponding to the candidate product to obtain a second multimodal feature vector;
[0217] A first acquisition module 1102 is configured to acquire a first instruction template corresponding to a preset level category to which the candidate product belongs, the first instruction template being used to describe attributes of the candidate product to be output;
[0218] The first generating module 1103 is configured to generate a first label for the candidate product under a preset level category based on the second multimodal feature vector and the first instruction template.
[0219] In another embodiment of the present disclosure, the first acquisition module 1102 is configured to acquire the first instruction template from the correspondence between different level categories and instruction templates based on the preset level category to which the candidate product belongs.
[0220] In another embodiment of the present disclosure, the apparatus further comprises:
[0221] A second acquisition module is configured to acquire a second instruction template input by a user for a preset level category if the evaluation result of the first tag is unreliable;
[0222] The updating module is used to update the first instruction template corresponding to the preset level category in the corresponding relationship to the second instruction template.
[0223] In another embodiment of the present disclosure, the apparatus further comprises:
[0224] The second generating module is configured to generate a second label for the candidate product under a preset level category based on the second multimodal feature vector and the second instruction template.
[0225] FIG12 shows a block diagram of an electronic device 1200 according to an exemplary embodiment of the present disclosure. Generally, the electronic device 1200 includes a processor 1201 and a memory 1202 .
[0226] The processor 1201 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1201 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1201 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1201 may also include an artificial intelligence processor, which is used to process computing operations related to machine learning.
[0227] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory computer-readable storage media. For example, the non-transitory computer-readable storage medium may be a CD-ROM (Compact Disc Read-Only Memory), ROM, RAM (Random Access Memory), magnetic tape, floppy disk, and optical data storage device. The computer-readable storage medium stores at least one computer program, which, when executed, can implement the above-mentioned multimodal large language model generation method, correlation calculation method, or label generation method.
[0228] Of course, the electronic device described above may also include other components, such as input / output interfaces and communication components. The input / output interface provides an interface between the processor and a peripheral interface module, which may be an output device, an input device, etc. The communication component is configured to facilitate wired or wireless communication between the electronic device and other devices.
[0229] Those skilled in the art will understand that the structure shown in FIG12 does not constitute a limitation on the electronic device 1200 , and may include more or fewer components than shown, or combine certain components, or adopt a different component arrangement.
[0230] An embodiment of the present disclosure provides a computer-readable storage medium, which stores at least one computer program. When the at least one computer program is executed by a processor, it can implement the above-mentioned multimodal large language model training method, or correlation calculation method, or label generation method.
[0231] An embodiment of the present disclosure provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it can implement the above-mentioned multimodal large language model training method, or correlation calculation method, or label generation method.
[0232] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0233] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.< / eof>
Claims
1. A multimodal large language model training method, wherein: The method comprises: Acquire sample search terms and sample products, wherein the sample products have image information and text description information; Calling a pre-trained multimodal large language model to process the image information and text description information of the sample product to obtain a sample text feature vector, a sample image feature vector, and a sample first multimodal feature vector of the sample product, wherein the image features related to the text description information in the sample first multimodal feature vector are enhanced; Based on the sample text feature vector, the sample image feature vector and the sample first multimodal feature vector of the sample product, the pre-trained multimodal large language model is trained to obtain a multimodal large language model; Calling the multimodal large language model to process the sample search term and the image information and text description information of the sample product to obtain a sample search term feature vector and a first multimodal feature vector and a sample second multimodal feature vector of the sample product, wherein image features related to the text description information and the search term in the sample second multimodal feature vector are enhanced; Based on the sample search term feature vector and the first multimodal feature vector and the sample second multimodal feature vector of the sample product, the multimodal large language model is trained to obtain a trained multimodal large language model. The trained multimodal large language model is used to generate a second multimodal feature vector of the product based on the search term and the image information and text description information of the product.
2. The method according to claim 1, wherein: The calling of the pre-trained multimodal large language model to process the image information and text description information of the sample product to obtain a sample text feature vector, a sample image feature vector and a sample first multimodal feature vector of the sample product includes: Calling the pre-trained multimodal large language model to perform feature extraction on the text description information of the sample product to obtain a sample text feature vector of the sample product; Extracting features from the image information of the sample product to obtain a sample image feature vector of the sample product; The sample text feature vector and the sample image feature vector of the sample commodity are fused to obtain a sample first multimodal feature vector of the sample commodity.
3. The method according to claim 1, wherein: The sample product is annotated with a label representing the correlation between the image information and the text description information, and the pre-trained multimodal large language model is trained based on the sample text feature vector, the sample image feature vector and the sample first multimodal feature vector of the sample product to obtain the multimodal large language model, including: Based on the sample first multimodal feature vector of the sample product, obtaining a label representing the correlation between the image information and the text description information of the sample product; Based on the first total objective loss function, the sample text feature vector and the sample image feature vector of the sample commodity The pre-trained multimodal large language model is trained based on the generated labels and the labeled labels corresponding to the sample products to obtain the multimodal large language model.
4. The method according to claim 3, wherein: The first total objective loss function includes a first image-text matching loss function and a first image-text contrast loss function. The pre-trained multimodal large language model is trained based on the first total objective loss function, the sample text feature vector and the sample image feature vector of the sample product, and the generated label and the annotated label corresponding to the sample product to obtain the multimodal large language model, including: Inputting the generated label and the annotated label corresponding to the sample product into the first image-text matching loss function to obtain a first image-text matching loss function value; Inputting the sample text feature vector and the sample image feature vector into the first image-text contrast loss function to obtain a first image-text contrast loss function value; Determine a first total target loss function value based on the first image-text matching loss function value and the first image-text comparison loss function value; Based on the first total objective loss function value, the model parameters of the pre-trained multimodal large language model are adjusted to obtain the multimodal large language model.
5. The method according to claim 1, wherein: The calling of the multimodal large language model to process the sample search term and the image information and text description information of the sample product to obtain the sample search term feature vector and the first multimodal feature vector and the sample second multimodal feature vector of the sample product includes: Calling the multimodal large language model to perform feature extraction on the text description information of the sample product to obtain a text feature vector of the sample product; Extracting features from the image information of the sample product to obtain an image feature vector of the sample product; Fusing the text feature vector and the image feature vector of the sample product to obtain a first multimodal feature vector of the sample product; Performing feature extraction on the sample search term to obtain a feature vector of the sample search term; The sample search term feature vector is fused with the first multimodal feature vector of the sample product to obtain a sample second multimodal feature vector of the sample product.
6. The method according to claim 1, wherein: The sample product is annotated with a label representing the correlation with the sample search term, and the multimodal large language model is trained based on the sample search term feature vector and the first multimodal feature vector and the sample second multimodal feature vector of the sample product to obtain a trained multimodal large language model, including: Based on the sample second multimodal feature vector of the sample product, obtaining a label representing the correlation between the sample product and the sample search term; Based on the second total objective loss function, the first multimodal feature vector and the sample search term feature vector of the sample product, and the generated labels and the annotated labels corresponding to the sample product, the multimodal large language model is trained to obtain a trained multimodal large language model.
7. The method according to claim 6, wherein: The second total objective loss function includes a second image-text matching loss function and a second image-text contrast loss function. The multimodal large language model is trained based on the second total objective loss function, the first multimodal feature vector of the sample product and the sample search term feature vector, and the generated label and the annotated label corresponding to the sample product to obtain the trained multimodal large language model, including: Inputting the generated label and the annotated label corresponding to the sample product into the second image-text matching loss function to obtain a second image-text matching loss function value; Inputting the first multimodal feature vector of the sample product and the sample search term feature vector into the second image-text contrast loss function to obtain a second image-text contrast loss function value; Determining a second total target loss function value based on the second image-text matching loss function value and the second image-text comparison loss function value; Based on the second total objective loss function value, the model parameters of the multimodal large language model are adjusted to obtain a trained multimodal large language model.
8. A method for calculating correlation, wherein: The method applies the trained multimodal large language model according to any one of claims 1 to 7, and the method comprises: Acquire a search term and candidate commodities searched based on the search term, wherein the candidate commodities have image information and text description information; Calculating a semantic relevance score between the search term and the candidate product based on the search term and the text description information; Calling the trained multimodal large language model to process the search term, the image information, and the text description information to obtain a second multimodal feature vector and a search term feature vector of the candidate product; Calculating a picture-text relevance score between the search term and the candidate product based on the search term feature vector and the second multimodal feature vector; Based on the semantic relevance score and the image-text relevance score, a total relevance score between the search term and the candidate product is calculated.
9. The method according to claim 8, wherein: The method further comprises: Calculating a canonical correlation score between the search term and the candidate product based on the text description information and the user preference information; The semantic relevance score, the image-text relevance score and the regular relevance score are weightedly calculated to obtain the total relevance score.
10. A label generation method, wherein: The method applies the trained multimodal large language model according to any one of claims 1 to 7, and the method comprises: Calling the trained multimodal large language model to process the search terms, image information, and text description information corresponding to the candidate products to obtain a second multimodal feature vector; Acquire a first instruction template corresponding to a preset level category to which the candidate product belongs, the first instruction template being used to describe the attributes of the candidate product to be output; Based on the second multimodal feature vector and the first instruction template, a first label for the candidate product under the preset level category is generated.
11. The method according to claim 10, wherein: The step of obtaining the first instruction template corresponding to the preset level category to which the candidate product belongs includes: Based on the preset level category to which the candidate commodity belongs, the first instruction template is obtained from the correspondence between different level categories and instruction templates.
12. The method according to claim 11, wherein: After generating the first label of the candidate product under the preset level category based on the second multimodal feature vector and the first instruction template, the method further includes: If the evaluation result of the first tag is unreliable, obtaining a second instruction template input by the user for the preset level category; The first instruction template corresponding to the preset level category in the corresponding relationship is updated to the second instruction template.
13. The method according to claim 12, wherein: After the first instruction template corresponding to the preset level category in the corresponding relationship is updated to the second instruction template, the method further includes: Based on the second multimodal feature vector and the second instruction template, a second label for the candidate product under the preset level category is generated.
14. An electronic device, wherein: It comprises a processor and a memory; the memory stores at least one program code; the at least one program code is used to be called and executed by the processor to implement the multimodal large language model training method as described in any one of claims 1 to 7, or the correlation calculation method as described in claim 8 or 9, or the label generation method as described in any one of claims 10 to 13.
15. A computer-readable storage medium, wherein: The computer-readable storage medium stores at least one computer program, which, when executed by the processor, can implement the multimodal large language model training method as described in any one of claims 1 to 7, or the correlation calculation method as described in claim 8 or 9, or the label generation method as described in any one of claims 10 to 13.
16. A computer program product, wherein: The computer program product includes a computer program, which, when executed by a processor, can implement the multimodal large language model training method as described in any one of claims 1 to 7, or the correlation calculation method as described in claim 8 or 9, or the label generation method as described in any one of claims 10 to 13.
Citation Information
Patent Citations
Multi-modal large language model training method, correlation calculation method and label generation method
CN118113901A
Resource searching method and device, computer equipment and storage medium
CN113377976A
Combined commodity retrieval method and system based on multi-modal pre-training model
CN114840705A
Object detection method, commodity detection method and similarity prediction model training method
CN115272722A
Commodity understanding method, device and equipment
CN116310437A
Cited By
Multi-modal panoramic image blind quality evaluation method and system based on AI generation description
CN120356071A
Intelligent data query method and system based on big data
CN120429479A
Zero sample learning-based dangerous event detection method and system in driving scene
CN120431550A
Image processing method and device, image display method and device, equipment and storage medium
CN120765355A
Multi-modal data semantic retrieval method and device, equipment and storage medium
CN120950705A