Multi-modal search large model optimization method and device, cluster and storage medium

By expanding query seeds to construct search terms, and using graphic and text pairs to optimize multimodal search models, the problem of uneven effects in different scenarios is solved, and higher universality and generalization are achieved.

CN120030043APending Publication Date: 2025-05-23HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311862579.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2023-12-29
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The effects of existing search models in different application scenarios are uneven, and iterative optimization relies on manual written instruction data, limiting their universality and generalization.

Method used

By obtaining query seeds, including general high-frequency content and scene content, it is expanded to build search terms, using these words to search in the data source for image search, obtain graphic pairs, and optimize the multimodal search model based on these graphic pairs.

Benefits of technology

The universality and generalization of multimodal search models are improved, and the optimized model performs more consistently and has better results in different scenarios without relying on manual instruction data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030043A_ABST
    Figure CN120030043A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal search large model optimization method and device, a cluster and a storage medium, and belongs to the technical field of AI. The method comprises the steps that a query seed is obtained, the query seed comprises general high-frequency content and / or scene content, the scene content comprises high-frequency content and / or key content of a scene applied by a multi-modal search large model, a first retrieval word corresponding to the query seed is constructed on the basis of a first expansion word obtained by expanding the query seed, and the first retrieval word corresponds to the query seed; in the first data source, image retrieval is conducted on the first retrieval word, image-text pairs corresponding to the first retrieval word are obtained, each image-text pair comprises an image and description information of the image, the large multi-modal search model is optimized based on the image-text pairs corresponding to the first retrieval word, and the optimized large multi-modal search model is obtained. By adopting the method provided by the invention, the universality and generalization of multi-modal search large model optimization are improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese patent application No. 202311560273.7 filed on November 20, 2023, entitled “Method and device for self-optimization of large model effects”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a method, device, cluster and storage medium for multimodal search large model optimization. Background Art

[0003] The search big model is a model trained based on a large amount of data. It has many advantages such as generality, scalability and emerging capabilities. It is an important step in achieving general artificial intelligence. With the rapid development of the search big model, it can cope with various natural language processing tasks, including artificial intelligence dialogue, context understanding, machine translation, question answering and text generation.

[0004] When training a large search model, it is trained based on a large amount of data so that the model has "general knowledge" capabilities. Since the distribution of training data for different large search models is different, such as web page data, conversation data, books, news data, and scientific data, the direct effects in different scenarios will be uneven. Therefore, when the large search model is applied to different application scenarios, the data of the scenario is used to fine-tune and iterate the large search model so that the large search model can be better applied to the corresponding scenario.

[0005] In the process of using the search large model, in order to make the search large model have better reasoning performance, the search large model is usually iteratively optimized. During the iterative optimization, people write a small amount of instruction data, use the small amount of instruction data to generate new instruction data, and then generate training samples based on the new instruction data, and use the training samples to update the search large model. Since the generation of training samples relies on the small amount of instruction data written by people, the ability of the search large model depends on the small amount of instruction data, which is limited by the quality and diversity of the instruction data, hindering the universality and generalization of the iterative optimization of the search large model. Summary of the invention

[0006] The present application provides a method, device, cluster and storage medium for multimodal search large model optimization, which can improve the versatility and generalization of multimodal search large model optimization.

[0007] In a first aspect, the present application provides a method for optimizing a multimodal search large model, the method comprising: obtaining a query seed, the query seed comprising general high-frequency content and / or scene content, the scene content comprising high-frequency content and / or key content of the scene in which the multimodal search large model is applied; constructing a first search term corresponding to the query seed based on a first extended term obtained by expanding the query seed; performing an image search on the first search term in a first data source to obtain an image-text pair corresponding to the first search term, each image-text pair comprising an image and description information of the image; optimizing the multimodal search large model based on the image-text pair corresponding to the first search term to obtain an optimized multimodal search large model.

[0008] In the scheme shown in this application, a query seed is constructed, the query seed is used to expand the search terms, the search terms are used to search, and multiple image-text pairs are obtained. Based on the image-text pairs obtained by the search, the multimodal search model is optimized. In this way, the image-text pairs used for optimization are of better quality and have diversity. When optimizing, there is no need to rely on the instruction data written by people. The ability of the optimized multimodal search model will not rely on the instruction data, so that the optimization of the multimodal search model has universality and generalization, and can also improve the optimization effect.

[0009] In an optional manner, the query seed acquisition includes: acquiring a candidate seed; constructing a second search term corresponding to the candidate seed based on a second extended term obtained by expanding the candidate seed; performing a text search on the second search term in a second data source to obtain a consulting text corresponding to the second search term; constructing a test set corresponding to the candidate seed based on the consulting text; using an evaluation model to infer test samples in the test set to obtain inference results of the test samples; determining a performance indicator corresponding to the candidate seed based on the inference results of the test samples and a label of the test samples; if the performance indicator is lower than a first threshold, adding the candidate seed to the query seed.

[0010] In the solution shown in the present application, when constructing the query seed, the weak points of the multimodal search large model are automatically identified, so that when optimizing the multimodal search large model, the data generalization of the weak points is enhanced, and the multimodal search large model is optimized in a targeted manner.

[0011] In an optional manner, the method further includes: obtaining a test sample corresponding to the candidate seed in a test set of the evaluation model, and adding the test sample to the test set corresponding to the candidate seed.

[0012] In an optional manner, the multimodal search big model is a big model for searching text by image or searching image by image, the general high-frequency content includes a general high-frequency heat map, and the scene content includes a high-frequency heat map and / or key pictures of the scene in which the multimodal search big model is applied; based on the first extended term obtained by expanding the query seed, a first search term corresponding to the query seed is constructed, including: converting the query seed into a query term; expanding the query term to obtain a first extended term corresponding to the query term; and combining the query term and the first extended term to form a first search term.

[0013] In the solution shown in the present application, in a large model of searching for text with images or searching for images with images, the query seed is a picture, and the query seed is converted into a query term to expand the query seed to obtain the first search term used for the search.

[0014] In an optional manner, the multimodal search big model is a big model for searching images based on text, the general high-frequency content includes general high-frequency hot words, and the scene content includes high-frequency hot words and / or key words of the scene in which the multimodal search big model is applied; based on the first extended word obtained by expanding the query seed, the first search word corresponding to the query seed is constructed, including: expanding the query seed to obtain the first extended word corresponding to the query seed; combining the query seed and the first extended word to form a first search word.

[0015] In the solution shown in the present application, in the large model of searching images with text, the query seed is a word, and the query seed is expanded to obtain the first search word used for the search.

[0016] In an optional manner, based on the image-text pairs corresponding to the first search term, the multimodal search big model is optimized to obtain the optimized multimodal search big model, including: adding the image-text pairs corresponding to the first search term to the incremental training set; if the number of image-text pairs in the incremental training set is less than the second threshold, expanding the first extended term to obtain the third extended term, updating the third extended term to the next round of first search term, until the number of image-text pairs in the incremental training set is equal to the second threshold, and using the incremental training set to optimize the multimodal search big model to obtain the optimized multimodal search big model.

[0017] In the solution shown in the present application, in order to obtain enough image-text pairs, the extended words are expanded again to obtain more image-text pairs. This will enable the optimized multi-model search large model to have stronger reasoning capabilities when optimizing the large model.

[0018] In an optional manner, based on the image-text pairs corresponding to the first search term, the multimodal search big model is optimized to obtain the optimized multimodal search big model, including: filtering the image-text pairs corresponding to the first search term to obtain filtered image-text pairs; based on the filtered image-text pairs, the multimodal search big model is optimized to obtain the optimized multimodal search big model.

[0019] In the solution shown in the present application, after the image-text pairs are searched and obtained, the image-text pairs are filtered so that the quality of the image-text pairs used for optimization is better, thereby making the reasoning ability of the optimized multi-model search large model stronger.

[0020] In an optional manner, the image-text pairs corresponding to the first search term are filtered, including one or more of the following: among the image-text pairs corresponding to the first search term, based on the relevance of the first search term and the image, the image-text pairs corresponding to the first search term are filtered; based on the relevance of the image and the description information in the image-text pairs corresponding to the first search term, the image-text pairs corresponding to the first search term are filtered; the image-text pairs corresponding to the first search term are deduplicated; or, among the image-text pairs corresponding to the first search term, the image-text pairs containing images having a resolution lower than a third threshold are deleted.

[0021] In an optional manner, based on the image-text pairs corresponding to the first search term, the multimodal search big model is optimized to obtain the optimized multimodal search big model, including: determining the images whose description information quality in the image-text pairs corresponding to the first search term does not meet the requirements; generating description information for the determined images to obtain image-text pairs whose quality meets the requirements; based on the image-text pairs whose quality meets the requirements, the multimodal search big model is optimized to obtain the optimized multimodal search big model.

[0022] In the scheme shown in the present application, after searching for image-text pairs, the description information is regenerated for pictures whose description information quality of the image-text pairs does not meet the requirements, rather than directly deleting the image-text pairs. This not only makes the quality of the optimized image-text pairs better, but also makes the number of image-text pairs larger, thereby making the reasoning ability of the optimized multi-model search large model stronger.

[0023] In an optional manner, the multimodal search big model is optimized based on the image-text pair corresponding to the first search term to obtain the optimized multimodal search big model, including: sampling the first image-text pair from the original training set of the multimodal search big model; optimizing the multimodal search big model based on the first image-text pair and the image-text pair corresponding to the first search term to obtain the optimized multimodal search big model.

[0024] In the solution shown in this application, when optimizing the multimodal search large model, the image-text pairs in the original training set are also used, so that the optimized multimodal search large model will not forget the reasoning ability that has been learned before, that is, avoid the forgetting effect.

[0025] In an optional manner, the query seed also includes bad cases and / or low-frequency content in the original training set of the multimodal search large model. In this way, optimization can also be performed for bad cases and low-frequency content.

[0026] In a second aspect, the present application provides a device for multimodal search large model optimization, which has the function of implementing the above-mentioned first aspect and any optional method of the first aspect. The device includes at least one module, and at least one module is used to implement the multimodal search large model optimization method provided by the above-mentioned first aspect and any optional method of the first aspect.

[0027] In some embodiments, the modules in the device for multimodal search of large model optimization are implemented by software, and the modules in the device for multimodal search of large model optimization are program modules. In other embodiments, the modules in the device for multimodal search of large model optimization are implemented by hardware or firmware.

[0028] In the third aspect, the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory, the processor of the at least one computing device being used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method for multimodal search large model optimization provided in the first aspect and any optional manner of the first aspect.

[0029] In a fourth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method for multimodal search large model optimization provided by the first aspect and any optional method of the first aspect.

[0030] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on at least one computing device in a computing device cluster, enables the at least one computing device to execute the method for multimodal search large model optimization provided in the first aspect and any optional method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a schematic diagram of a system architecture provided by an exemplary embodiment of the present application;

[0032] Figure 2It is a flowchart of a method for multi-modal search large model optimization provided by an exemplary embodiment of the present application;

[0033] Figure 3 It is a flowchart of a method for multi-modal search large model optimization provided by another exemplary embodiment of the present application;

[0034] Figure 4 It is a schematic diagram of the effect of an optimized multi-modal search large model provided by an exemplary embodiment of the present application;

[0035] Figure 5 It is a flowchart of a method for obtaining a query seed provided by an exemplary embodiment of the present application;

[0036] Figure 6 It is a flowchart of a method for multi-modal search large model optimization provided by another exemplary embodiment of the present application;

[0037] Figure 7 It is a structural schematic diagram of a device for multi-modal search large model optimization provided by an exemplary embodiment of the present application;

[0038] Figure 8 is a schematic diagram of the structure of a computing device provided by an exemplary embodiment of the present application;

[0039] Fig. 9 is a schematic diagram of the structure of a computing device cluster provided by an exemplary embodiment of the present application;

[0040] Fig.10 It is a schematic diagram of computing device connections provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0041] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0042] With the rapid development of AI big models, multimodal search big models have come into being. Multimodal search includes using text to search for images, or using images to search for text. It can cope with various natural language processing tasks, including artificial intelligence dialogue, context understanding, machine translation, question and answer, or text generation.

[0043] In order to make the multimodal search large model have better performance, it is necessary to optimize the multimodal search large model. In the embodiment of the present application, the query seed and the extended results of the query seed are used to retrieve and optimize the image-text pairs used by the multimodal search large model. The image-text pairs are used to optimize the multimodal search large model without manual participation and do not rely on instruction data written by people. In addition, the query seed and the extended results are used to retrieve the image-text pairs, so that the image-text pairs used for optimization are of better quality and have diversity. Therefore, the scheme of the embodiment of the present application is adopted to make the optimization of the multimodal search large model universal and generalizable.

[0044] The application scenarios of the embodiments of the present application are described below.

[0045] The embodiments of the present application can be applied to the application scenario of searching images by text, and optimize the large model of searching images by text. Among them, searching images by text refers to searching images by text. For example, the user inputs text, and the multimodal search large model calculates the relevance between the text and each image, and outputs images related to the text based on the relevance.

[0046] The embodiments of the present application can also be applied to the application scenario of searching for text by image, and optimize the large model of searching for text by image. Among them, searching for text by image means searching for text by image. For example, when a user inputs an image, the multimodal search large model calculates the relevance between the image and each text, and outputs the text related to the image based on the relevance.

[0047] The embodiments of the present application can also be applied to the application scenario of searching images by images, and optimize the large model of searching images by images. Among them, searching images by images refers to searching for images using images. For example, a user inputs an image, and the large multimodal search model converts the image into text, calculates the correlation between the text and each image, and based on the correlation, outputs images related to the text, thereby obtaining images related to the input image. Alternatively, the embodiments of the present application also provide a preprocessing module, and the user inputs an image, and the preprocessing module converts the image into text, and the large multimodal search model calculates the correlation between the text and each image, and based on the correlation, outputs images related to the text, thereby obtaining images related to the input image.

[0048] The system architecture of the embodiment of the present application is described below.

[0049] Figure 1 Provides a system architecture of an embodiment of the present application. Figure 1, the system architecture includes a public cloud 101 and a terminal device 102. The public cloud 101 is connected to the terminal device 102 via a wired or wireless network. The public cloud 101 is an entity that uses basic resources to provide cloud services to users in a cloud computing model. The public cloud 101 can also be considered as a cloud environment. The public cloud 101 includes a cloud data center, which includes a large number of basic resources owned by a cloud service provider. The large number of basic resources include computing resources, storage resources, and network resources. The computing resources included in the cloud data center can be a computing device cluster, and the computing device cluster includes at least one computing device. The computing device can be a server, etc. The terminal device 102 is a device used by the user, such as a computer or a tablet computer.

[0050] When a user uses a cloud service, the user can input an optimization instruction through an application programming interface (API) or an interactive interface (the interactive interface can be a graphical user interface (GUI)), and the terminal device 102 sends the optimization instruction to the public cloud 101. The public cloud 101 optimizes the multimodal search large model. Alternatively, when a user uses a cloud service, the user inputs optimization configuration information through an API or an interactive interface, and the optimization configuration information includes a configured optimization time point (such as an optimization cycle, etc.), and the terminal device 102 sends the optimization configuration information to the public cloud 101. The public cloud 101 optimizes the multimodal search large model based on the optimization time point configured by the optimization configuration information.

[0051] Optionally, the user may also input the query seed and other information mentioned later for use by the public cloud 101 when optimizing the multimodal search model. The user may also input the relevant files of the multimodal search model for use by the public cloud 101 when optimizing the multimodal search model, and the public cloud 101 returns the optimized multimodal search model to the terminal device 102.

[0052] In another system architecture, the process of multimodal search large model optimization can be implemented on the terminal device.

[0053] The following describes the execution subject of the embodiment of the present application.

[0054] The execution subject of the method for multimodal search large model optimization is a device for multimodal search large model optimization. Optionally, the device is a hardware device, and the hardware device is a computing device. Optionally, the device is a software device, such as a set of software programs running on the computing device. The following describes the method flow with the execution subject being a computing device.

[0055] The following describes the method flow of the embodiment of the present application. Figure 2A method flow for multi-modal search large model optimization is provided, see steps 201 to 204 .

[0056] Step 201, query seed construction: obtain the query seed.

[0057] Among them, the query seed includes general high-frequency content and / or scene content, and the scene content includes high-frequency content and / or key content of the scene applied by the multimodal search large model. Among them, the general high-frequency content includes the high-frequency content searched within the first time period closest to the current time on the Internet, and the high-frequency content of the scene includes the high-frequency content searched within the second time period closest to the current time. The key content includes the content that people focus on in the scene, or it can also be understood as the content that people are most likely to search for in the scene. The first time period and the second time period can be set according to actual needs. For example, the first time period and the second time period are the same as the optimization period.

[0058] In this embodiment, the user inputs the optimization instruction through the terminal device, and the computing device receives the optimization instruction sent by the terminal device to obtain the query seed.

[0059] Alternatively, the computing device stores an optimization cycle, and obtains the query seed whenever the optimization cycle is reached.

[0060] Optionally, when acquiring the query seed, the computing device may acquire high-frequency content within a first time period closest to the current time from the Internet, obtain general high-frequency content, and determine the scenario to which the multimodal search large model is to be applied, and determine the high-frequency content and / or key content within a second time period closest to the current time in the scenario. Among them, the computing device obtains the identifier of the scenario to which the multimodal search large model is applied from the optimization instruction, or the computing device receives the identifier sent by the terminal device, or obtains the stored identifier. Alternatively, when acquiring the query seed, the computing device receives the query seed input by the user through the terminal device. For example, the optimization instruction carries the query seed.

[0061] Step 202, search term construction: based on the first expanded term obtained by expanding the query seed, construct a first search term corresponding to the query seed.

[0062] In this embodiment, the computing device expands the query seed to obtain an expanded term, which is called the first expanded term. If the query seed is a term, the first expanded term is merged with the query seed to form the first search term corresponding to the query seed. For example, the query seed is "okra", and the first expanded term obtained after expansion is "okra flower", "how to grow okra" and "cold okra", etc. For another example, the query seed is "apple", and the first expanded term obtained after expansion is "red apple", "green apple" and "red apple in A", etc. If the query seed is a picture, the first expanded term is merged with the term corresponding to the query seed to form the first search term corresponding to the query seed.

[0063] Step 203, data retrieval: In the first data source, perform image retrieval on the first search term to obtain image-text pairs corresponding to the first search term, each image-text pair including an image and description information of the image.

[0064] In this embodiment, the first data source includes one or more of an Internet search engine, data in a customer's vertical domain, data in a customer's private domain, or an original test set of a multimodal search large model, wherein when using the customer's private domain data, access rights have been obtained from the customer, and the original test set is a test set used by the multimodal search large model during training. The computing device performs image retrieval on each first search term in the first data source to obtain a picture-text pair corresponding to each first search term. For example, the first search terms include "okra flower", "how to grow okra", and "cold okra", etc., and through the Internet search for "okra flower", "how to grow okra", and "cold okra", etc., obtain the picture-text pair corresponding to "okra flower", the picture-text pair corresponding to "how to grow okra", and the picture-text pair corresponding to "cold okra". Each picture-text pair includes a picture and description information of the picture. The description information of the picture is used to describe the scene and content of the picture, and the description information of the picture is the data already in the first data source.

[0065] Step 204, incremental training: based on the image-text pair corresponding to the first search term, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

[0066] In this embodiment, after obtaining the image-text pair corresponding to the first search term, the image-text pair is added to the incremental training set. When the number of image-text pairs in the incremental training set is less than the second threshold, the first extended term is expanded to obtain an extended term, which is called the third extended term. The third extended term is updated to the first search term in step 203, and steps 203 to 204 are continued until the number of image-text pairs in the incremental training set reaches the second threshold. The second threshold is set according to actual needs. For example, the second threshold is 10,000.

[0067] The computing device uses the image-text pairs in the incremental training set to optimize the multimodal search large model to obtain the optimized multimodal search large model. Alternatively, the computing device obtains the original training set of the multimodal search large model, and the original training set is the training set used to train the multimodal search large model, and the original training set includes image-text pairs. In the original training set, the first image-text pair is sampled, and the number of the first image-text pairs is set according to actual needs. For example, there are 500 million image-text pairs in the original training set, and 100,000 image-text pairs are sampled from the 500 million image-text pairs. Then the first image-text pair and the image-text pairs in the incremental training set are mixed to obtain a mixed training set, and the mixed training set is used to optimize the multimodal search large model to obtain the optimized multimodal search large model.

[0068] Alternatively, after obtaining the first search term, the image-text pair corresponding to the first search term is mixed with the first image-text pair to obtain a mixed training set, and the mixed training set is used to optimize the multimodal search large model to obtain an optimized multimodal search large model.

[0069] In step 204, when optimizing the multimodal search large model, the forgetting effect of the incremental training of the model can be avoided by using the first image-text pair.

[0070] In addition, the computing device can also introduce an adapter based on the structure of the multimodal search large model to specifically carry out the optimization process using the incremental training set.

[0071] In an optional manner, when sampling the first image-text pair from the original training set, the first image-text pair may be sampled in a random and uniformly distributed manner.

[0072] The following uses the multimodal search model as an example to illustrate the solution. Figure 3 Steps 301 to 310 in the process shown.

[0073] Step 301, query seed construction: obtain the query seed.

[0074] In this embodiment, in the scenario of searching for images with text, general high-frequency content includes general high-frequency hot words, and general high-frequency hot words include words with a search frequency higher than a certain threshold among the hot words searched in the first time period closest to the current time on the Internet. General high-frequency hot words are also called general high-frequency hot words. The scene content includes high-frequency hot words and / or key words in the scenario applied by the multimodal search large model. High-frequency hot words refer to words with a search frequency higher than a certain threshold among the hot words searched in the second time period closest to the current time in the application scenario. The threshold is set according to actual needs. Key words refer to the content that people focus on in the application scenario.

[0075] The process of obtaining the query seed in step 301 refers to the description of step 201 and will not be repeated here.

[0076] In an optional manner, the query seed also includes bad cases and / or low-frequency words in the original training set, where bad cases refer to words that are incorrectly searched by the multimodal search large model, and low-frequency words refer to words in the original training set whose frequency of appearance is lower than the frequency threshold, and the frequency threshold is set according to actual needs. The computing device can obtain bad cases fed back by various channels, including but not limited to customers and crowd testing. The computing device can count the frequency of appearance of words in the original training set, obtain words whose frequency of appearance is lower than the frequency threshold, and determine them as low-frequency words in the original training set.

[0077] Step 302, search term construction: expand the query seed to obtain a first expanded term corresponding to the query seed, and combine the query seed and the first expanded term to form a first search term.

[0078] In this embodiment, after obtaining the query seed, the query seed is expanded to obtain the first expanded term corresponding to the query seed. The first expanded term has a qualitative improvement in diversity and quantity relative to the query seed. The first expanded term is then merged with the query seed to form the first search term. For example, the query seed is "radish", which is expanded to "white radish", "carrot", "radish flower" and "radish meatballs", etc., and the first search terms are "radish", "white radish", "carrot", "radish flower" and "radish meatballs", etc.

[0079] Optionally, the extension method includes but is not limited to one or more of a search engine query extension function, a synonym list, a near-synonymous word list, a local expert definition document, or a knowledge graph.

[0080] Optionally, when a query seed is expanded using multiple expansion methods, the same words may be obtained through expansion. When forming the first search term, only one of the same words needs to be retained.

[0081] Step 303, data retrieval: In the first data source, perform image retrieval on the first search term to obtain image-text pairs corresponding to the first search term, each image-text pair including an image and description information of the image.

[0082] The description of step 303 refers to the description of step 203 and will not be repeated here.

[0083] Step 304, data filtering: filtering the image-text pairs corresponding to the first search term to obtain filtered image-text pairs.

[0084] In this embodiment, after the image-text pairs corresponding to the first search term are obtained, the image-text pairs are filtered using one or more of the following methods.

[0085] Method 1, based on the correlation between the first search term and the picture in the corresponding picture-text pair, the picture-text pair corresponding to the first search term is filtered. For example, the first search term and the picture in the corresponding picture-text pair are input into the correlation model to obtain the correlation between the first search term and the picture in the corresponding picture-text pair, and the picture-text pair to which the picture has a correlation lower than the first correlation threshold is deleted. For another example, the picture in the picture-text pair corresponding to the first search term is converted into text, the correlation between the first search term and the text is calculated, and the picture-text pair corresponding to the text with a correlation lower than the first correlation threshold is deleted. In this way, the picture-text pairs can be filtered from the perspective of the correlation between the first search term and the picture. It should be noted here that although in step 303, the picture-text pair determined to be related to the first search term, a search algorithm is used, and a filtering algorithm is used here to measure the correlation and filter the picture-text pairs.

[0086] Method 2: Based on the correlation between the image and the description information in the image-text pair corresponding to the first search term, the image-text pairs corresponding to the first search term are filtered. For example, for each image-text pair corresponding to the first search term, the image-text pair includes an image and description information of the image, the image is input into the description information generation model to obtain the first description information of the image, the correlation between the first description information and the description information of the image is calculated, the correlation corresponding to each image-text pair is obtained, and the image-text pairs whose correlation is lower than the second correlation threshold are deleted. For another example, for each image-text pair corresponding to the first search term, the description information of the image and the image is input into the correlation model to obtain the correlation between the description information of the image and the image, the correlation corresponding to each image-text pair is obtained, and the image-text pairs whose correlation is lower than the second correlation threshold are deleted. In this way, the image-text pairs can be filtered from the perspective of image-text correlation.

[0087] Method three, deduplicate the image-text pairs corresponding to the first search term. For example, for every two image-text pairs, use the perceptual hash algorithm to calculate the Hamming distance of the images, determine the image-text pairs to which the images whose Hamming distance is less than or equal to the first value belong as duplicate image-text pairs, determine the image-text pairs to which the images whose Hamming distance is greater than the first value and less than the second value belong as similar image-text pairs, retain only one copy of the duplicate image-text pairs, and retain only one copy of the similar image-text pairs. Here, the perceptual hash algorithm is used to calculate the similarity between images, and other algorithms, such as cosine distance, etc., can also be used, which is not limited in the embodiments of the present application.

[0088] Method 4: Delete the image-text pairs with lower resolutions from the image-text pairs corresponding to the first search term. For example, determine the resolution of the images in each image-text pair, and delete the image-text pairs with resolutions lower than a third threshold.

[0089] In this way, filtering the image-text pairs corresponding to the first search term can improve the quality of the image-text pairs, thereby achieving good optimization effects when subsequently optimizing the multimodal search large model.

[0090] Step 305, description information correction: determine the pictures whose description information quality does not meet the requirements in the picture-text pair corresponding to the first search term, generate description information for the determined pictures, and obtain the picture-text pair whose quality meets the requirements.

[0091] The quality of the description information does not meet the requirements, which means that the description information does not correspond to the content of the picture.

[0092] In this embodiment, there may be pictures whose description information quality does not meet the requirements in the picture-text pairs corresponding to the first search term. The computing device obtains the picture-text pairs filtered out in the above method 2. If the resolution of the pictures in these picture-text pairs is still relatively high, the pictures in the picture-text pairs are determined as pictures whose description information quality does not meet the requirements. The pictures are input into the description information generation model to obtain the description information of the pictures, and the newly generated description information is used to replace the original description information to obtain the picture-text pairs whose quality meets the requirements.

[0093] In this way, for the image-text pairs with poor description information, the description information is regenerated instead of being deleted, so more image-text pairs can be obtained.

[0094] In addition, in step 305, the description information is generated using a generation model of the description information. In order to make the generated description information more accurate, an interactive interface can be provided to the user. The computing device feeds back the image and the generated description information to the terminal device used by the user. The terminal device displays the image and the generated description information on the interactive interface. The user can proofread the description information to provide accurate description information to the computing device.

[0095] Step 306, constructing an incremental training set: adding the image-text pairs processed in steps 304 and 305 to the incremental training set.

[0096] Step 307: determine whether the number of image-text pairs in the incremental training set reaches a second threshold.

[0097] The second threshold is an empirical value, which can be used to optimize the multimodal search large model once.

[0098] Step 308: If the second threshold is reached, execute step 310.

[0099] Step 309, expanding the incremental training set: if the second threshold is not reached, the first extended term is expanded to obtain a third extended term, and the third extended term is updated as the first search term for the next round, until the number of image-text pairs in the incremental training set reaches the second threshold, and step 310 is executed.

[0100] In this embodiment, if the number of image-text pairs in the incremental training set does not reach the second threshold, the first extended term is expanded to obtain an extended term, which is called the third extended term. The third extended term is updated to the first search term in the next round, and the process goes to step 303 to step 307 until it is determined that the number of image-text pairs in the incremental training set reaches the second threshold.

[0101] Step 310, incremental training: using the incremental training set, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

[0102] In this embodiment, the computing device uses the image-text pairs in the incremental training set to optimize the multimodal search large model to obtain the optimized multimodal search large model. Alternatively, the computing device obtains the original training set of the multimodal search large model. In the original training set, the first image-text pairs are sampled in a random uniform distribution form, and the number of the first image-text pairs is set according to actual needs. Then, the first image-text pairs and the image-text pairs in the incremental training set are mixed to obtain a mixed training set, and the mixed training set is used to optimize the multimodal search large model to obtain the optimized multimodal search large model.

[0103] The optimization process is as follows: inputting the vector of the image and the vector of the description information in the image-text pair into the multimodal search large model, so that the distance between the vector of the image and the vector of the description information in the same image-text pair becomes smaller and smaller.

[0104] The present application also provides Figure 3 The effect description of the process shown in the multimodal gallery is Figure 3 The process shown can improve the search effect of the multimodal search large model and perform targeted repairs on bad cases. The index comparison of the optimized multimodal search large model is good, same, bad (GSB) = 70:42:4, where good means 70 better than before optimization, average means 42 the same as before optimization, and bad means 4 worse than before optimization. Figure 4The comparison chart of the multimodal search model before and after optimization shows that for most query terms, the search results of the optimized multimodal search model are significantly better than those of the multimodal search model before optimization. Only for a few query terms, the search results of the optimized multimodal search model are worse than those of the multimodal search model before optimization. For example, the multimodal search model could search for various pictures of sunflowers before optimization, but after optimization, it could search for various pictures of okra in addition to various pictures of sunflowers.

[0105] exist Figure 3 In the process shown, in order to optimize the weak categories when optimizing the multimodal search model, the processing process of step 301 is: obtain candidate seeds, build a test set corresponding to the candidate seeds based on the expansion results of the candidate seeds, determine that the performance index corresponding to the candidate seeds is lower than the first threshold based on the test set and the evaluation model, and add the candidate seeds to the query seeds. For the specific process, see Figure 5 Steps 501 to 505 in the process shown.

[0106] Step 501, seed construction: obtaining candidate seeds.

[0107] In this embodiment, the process of obtaining candidate seeds is the same as the process of obtaining query seeds in the above text, except that the candidate seeds may not include bad cases. For example, the candidate seeds include high-frequency hot words and / or key words in the application scenario of the large model to be optimized.

[0108] Step 502, expansion construction: based on the second expanded term obtained by expanding the candidate seed, construct a second search term corresponding to the candidate seed.

[0109] In this embodiment, after obtaining the candidate seed, the candidate seed is expanded to obtain the second expanded term corresponding to the candidate seed. The second expanded term has a qualitative improvement in diversity and quantity relative to the candidate seed. The second expanded term is then merged with the candidate seed to form a second search term. For example, the candidate seed is "customs clearance", which can be expanded to "customs declaration company" and "customs clearance of personal belongings" through an Internet search engine or a local expert definition document, and can also be expanded through a synonym table to obtain synonyms such as "customs clearance" and "customs declaration".

[0110] Optionally, the extension method includes but is not limited to one or more of a search engine query extension function, a synonym list, a near-synonymous word list, a local expert definition document, or a knowledge graph.

[0111] It should be noted that the method of expanding the candidate seeds is the same as the method of expanding the query seeds, which will not be described in detail here.

[0112] Step 503, data retrieval: in the second data source, perform a text search on the second search term to obtain a consultation text corresponding to the second search term.

[0113] In this embodiment, the second data source is the same as or different from the first data source mentioned above. The second data source includes one or more of an Internet search engine, data from a customer's vertical domain, data from a customer's private domain, or the original test set of a multimodal search large model. The computing device performs an image search on each second search term in the second data source to obtain a consulting text corresponding to each second search term. For example, by searching for "customs declaration company" and "customs clearance" through an Internet search engine, the corresponding consulting text is obtained, and the consulting text includes "the name of the customs declaration company", "the meaning of customs clearance" and "the process of customs clearance".

[0114] Step 504, obtaining a test set: constructing a test set corresponding to the candidate seed based on the consultation text.

[0115] In this embodiment, the consultation texts with low quality are filtered out, and the consultation texts with low relevance to the second search term are filtered out, and then the filtered consultation texts are deduplicated to improve the diversity of the test set obtained later. Among them, data filtering is divided into quality filtering and relevance filtering.

[0116] Quality filtering includes but is not limited to deleting ambiguous consulting texts, deleting incoherent consulting texts, etc.

[0117] The relevance filtering process is: calculating the relevance between the second search term and the corresponding consulting text, deleting consulting texts with relevance below a certain threshold, or deleting a certain number of consulting texts with low relevance rankings. For example, the second search term and the corresponding consulting text are input into the relevance model to obtain the relevance between the second search term and the corresponding consulting text.

[0118] The deduplication process is: use semantics and keywords to determine similar or identical consulting texts in the consulting texts, retain one copy of the identical consulting texts, and retain one copy of the similar consulting texts.

[0119] Then, the consultation text after the above processing is converted into the format required by the evaluation model to obtain the test set corresponding to the candidate seed. For example, the evaluation model is a large language model (LLM), and the self-instruction method of the large language model is used to convert the consultation text into the format required by the large language model. The test set can also be called a supervised fine-tuning (SFT) test set. The format required by the large language model is represented by 3 key-value pairs. The key of the first key-value pair represents the task type, and the value represents the specific task type, such as a classification task. The second key-value pair is the input. The key of the second key-value pair represents the question, and the value represents the second search term. The third key-value pair is the output. The key of the third key-value pair represents the answer, and the value represents the consultation text.

[0120] In an optional approach, the evaluation model and the multimodal search large model can be two models in a model collection.

[0121] In an optional manner, the computing device may also obtain a test set of the evaluation model, search for relevant data of the candidate seed in the test set, and add the relevant data of the candidate seed to the test set corresponding to the candidate seed.

[0122] In an optional method, after obtaining the test set corresponding to the candidate seed, in order to make the test set corresponding to the candidate seed more accurate, an interactive interface can be provided to the user, and the computing device feeds back the test set to the terminal device used by the user. The terminal device displays the test set on the interactive interface, and the user can proofread the test set to provide an accurate test set to the computing device.

[0123] Step 505, classification evaluation: use the evaluation model to infer the test samples in the test set to obtain the inference result of the test sample, and determine the performance indicator corresponding to the candidate seed based on the inference result of the test sample and the label of the test sample. If the performance indicator is lower than the first threshold, the candidate seed is added to the query seed.

[0124] In this embodiment, each test sample in the test set is input into the evaluation model to obtain the inference result corresponding to each test sample. Then the inference result corresponding to each test sample is compared with the label to obtain the performance index of the test set corresponding to the candidate seed. For example, the value in the input of the test set is input into the evaluation model to obtain the inference result, and the inference result is compared with the value in the output to obtain the performance index of the test set corresponding to the candidate seed. If the performance index is lower than the first threshold, the candidate seed is added to the query seed as a weak category, and the multimodal search large model is optimized. If the performance index is greater than or equal to the first threshold, it means that the candidate seed is not a weak point and does not need to be used to optimize the multimodal search large model.

[0125] In an optional manner, the performance indicator includes effect indicators such as bilingual evaluation under study (BLEU) and / or recall-oriented under study forgisting evaluation (Rouge). When the performance indicator includes the BLEU indicator and the Rouge indicator, the performance indicator is equal to the average of the BLEU indicator and the Rouge indicator.

[0126] exist Figure 5 In the process shown, based on the construction of the retrieved test set, weak points are automatically identified. When optimizing the multimodal search large model, weak points can be targeted for identification, and data generalization of weak points can be enhanced to reduce optimization costs and improve optimization efficiency.

[0127] It should be noted that since the big language model is the basis of the multimodal search big model, the weaknesses of the big language model recognition are generally also the weaknesses of the multimodal search big model. Alternatively, since images and texts are convertible into each other, the weaknesses of the big language model recognition are generally also the weaknesses of the multimodal search big model. Figure 5 The process shown is equivalent to the "leakage detection" process when optimizing a large multi-modal search model. Figure 3 The process shown is equivalent to the "gap filling" process during multimodal search large model optimization.

[0128] In addition, if the candidate seed is added to the query seed, in order to save search resources, the second expanded term can be directly obtained as the first expanded term.

[0129] The following describes the solution using a multimodal search model that uses images to search for text or images to search for images. Figure 6 Steps 601 to 610 in the process shown.

[0130] Step 601, query seed construction: obtain the query seed.

[0131] In this embodiment, in the scenario of searching for images with text, general high-frequency content includes general high-frequency heat maps, and general high-frequency heat maps include images whose search frequency is higher than a certain threshold in the heat maps searched within the first time period closest to the current time on the Internet. The scene content includes high-frequency heat maps and / or key images of the scenario applied by the multimodal search large model. High-frequency heat maps refer to images whose search frequency is higher than a certain threshold in the heat maps searched within the second time period closest to the current time in the application scenario. Key images include images that people focus on in the application scenario, which can also be understood as images that are most likely to be searched.

[0132] The process of obtaining the query seed in step 601 refers to the description of step 201 and will not be repeated here.

[0133] In an optional manner, the query seed also includes bad cases and / or low-frequency images in the original training set, where bad cases refer to images that are incorrectly searched by the multimodal search large model, and low-frequency images refer to images in the original training set that appear less frequently than a frequency threshold. The computing device can obtain bad cases from various channels, including but not limited to customers and crowd testing. The computing device can count the frequency of occurrence of images in the original training set, obtain images that appear less frequently than a frequency threshold, and determine them as low-frequency images in the original training set.

[0134] Step 602, search term construction: convert the query seed into a query term, expand the query term to obtain a first expanded term corresponding to the query term, and combine the query term and the first expanded term to form a first search term.

[0135] In this embodiment, after obtaining the query seed, the vector of the query seed is converted into a corresponding text vector, the word corresponding to the text vector is the query word, and the query word is expanded to obtain the first expanded word corresponding to the query seed. The first expanded word has a qualitative improvement in diversity and quantity compared to the query seed.

[0136] Step 603, data retrieval: In the first data source, perform image retrieval on the first search term to obtain image-text pairs corresponding to the first search term, each image-text pair including an image and description information of the image.

[0137] Step 604, data filtering: filtering the image-text pairs corresponding to the first search term to obtain filtered image-text pairs.

[0138] Step 605, description information correction: determine the pictures whose description information quality does not meet the requirements in the picture-text pair corresponding to the first search term, generate description information for the determined pictures, and obtain the picture-text pair whose quality meets the requirements.

[0139] Step 606, constructing an incremental training set: adding the image-text pairs processed in steps 604 and 605 to the incremental training set.

[0140] Step 607: determine whether the number of image-text pairs in the incremental training set reaches a second threshold.

[0141] Step 608: If the second threshold is reached, execute step 610.

[0142] Step 609, expand the incremental training set: if the second threshold is not reached, expand the first extended term to obtain a third extended term, update the third extended term to the first search term of the next round, execute steps 603 to 606, until the number of image-text pairs in the incremental training set reaches the second threshold, and execute step 610.

[0143] Step 610, incremental training: using the incremental training set, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

[0144] In an optional manner, the query seed in step 601 includes a weak category, and the process of step 601 is similar to Figure 5 Similar to the process shown in Figure 5 The difference between the processes shown is that the candidate seeds include the general high-frequency content including the general high-frequency heat map, the scene content includes the high-frequency heat map and / or key pictures of the scenario in which the multimodal search large model is applied, and the candidate seeds are first converted into candidate words, and then expansion and subsequent processing are performed.

[0145] In the embodiment of the present application, for any given query seed, an incremental training set is constructed based on the retrieval incremental training set construction method to optimize the multimodal search large model. In particular, when the query seed includes weak points, the weak points can be targeted for identification, and the optimization direction is flexible and controllable, which can reduce the optimization cost and improve the optimization efficiency.

[0146] In addition, the query seeds include common high-frequency content and / or scene content, which introduces new external data and ensures bad case repair and model generalization to a certain extent.

[0147] In addition, when optimizing multimodal search large models, there is no need for human participation in the entire process, which shortens the cycle and reduces costs.

[0148] The present application also provides a device for multi-modal search and large model optimization, such as Figure 7 As shown, the device comprises:

[0149] The query seed construction module 710 is used to obtain a query seed, wherein the query seed includes general high-frequency content and / or scene content, wherein the scene content includes high-frequency content and / or key content of the scene to which the multimodal search large model is applied, and can be specifically used to implement the query seed construction function of step 201 and execute the implicit steps included in step 201;

[0150] A search term construction module 720, which is used to construct a first search term corresponding to the query seed based on a first expanded term obtained by expanding the query seed, and can be specifically used to implement the search term construction function of step 202 and execute the implicit steps included in step 202;

[0151] A search module 730 is used to perform image search for the first search term in the first data source to obtain image-text pairs corresponding to the first search term, each image-text pair including an image and description information of the image, which can be used to implement the search function of step 203 and execute the implicit steps included in step 203;

[0152] The incremental training module 740 is used to optimize the multimodal search model based on the image-text pair corresponding to the first search term to obtain an optimized multimodal search model, which can be specifically used to implement the incremental training function of step 204 and execute the implicit steps included in step 204.

[0153] In an optional manner, the query seed construction module 710 is used to:

[0154] Get candidate seeds;

[0155] Based on the second expanded term obtained by expanding the candidate seed, construct a second search term corresponding to the candidate seed;

[0156] In the second data source, a text search is performed on the second search term to obtain a consultation text corresponding to the second search term;

[0157] Based on the consultation text, construct a test set corresponding to the candidate seed;

[0158] Using the evaluation model to infer the test samples in the test set to obtain the inference results of the test samples;

[0159] Determining a performance indicator corresponding to the candidate seed based on the inference result of the test sample and the label of the test sample;

[0160] If the performance indicator is lower than a first threshold, the candidate seed is added to the query seed.

[0161] In an optional manner, the query seed construction module 710 is further used to obtain the test samples corresponding to the candidate seeds in the test set of the evaluation model, and add them to the test set corresponding to the candidate seeds.

[0162] In an optional manner, the multimodal search big model is a big model for searching text by image or searching images by image, the general high-frequency content includes a general high-frequency heat map, and the scene content includes a high-frequency heat map and / or key pictures of the scene to which the multimodal search big model is applied;

[0163] The search term construction module 720 is used to:

[0164] Converting the query seed into a query term;

[0165] Expanding the query term to obtain the first expanded term corresponding to the query term;

[0166] The query term and the first expanded term are combined to form the first search term.

[0167] In an optional manner, the multimodal search big model is a big model for searching images by text, the general high-frequency content includes general high-frequency hot words, and the scene content includes high-frequency hot words and / or key words of the scene to which the multimodal search big model is applied;

[0168] The search term construction module 720 is used to:

[0169] Expanding the query seed to obtain the first expanded term corresponding to the query seed;

[0170] The query seed and the first expanded term are combined to form the first search term.

[0171] In an optional manner, the incremental training module 740 is used to:

[0172] Adding the image-text pair corresponding to the first search term to the incremental training set;

[0173] If the number of image-text pairs in the incremental training set is less than a second threshold, the first extended term is expanded to obtain a third extended term, and the third extended term is updated to the first search term in the next round until the number of image-text pairs in the incremental training set is equal to the second threshold. The incremental training set is used to optimize the multimodal search model to obtain an optimized multimodal search model.

[0174] In an optional manner, the incremental training module 740 is used to:

[0175] Filtering the image-text pairs corresponding to the first search term to obtain filtered image-text pairs;

[0176] Based on the filtered image-text pairs, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

[0177] In an optional manner, the incremental training module 740 is configured to perform one or more of the following:

[0178] Among the image-text pairs corresponding to the first search term, filtering the image-text pairs corresponding to the first search term based on the relevance between the first search term and the image;

[0179] filtering the image-text pairs corresponding to the first search term based on the correlation between the image and the description information in the image-text pairs corresponding to the first search term;

[0180] Deduplication processing is performed on the image-text pairs corresponding to the first search term; or among the image-text pairs corresponding to the first search term, image-text pairs to which images having an image resolution lower than a third threshold are deleted.

[0181] In an optional manner, the incremental training module 740 is used to:

[0182] Determine the pictures whose description information quality does not meet the requirements in the picture-text pair corresponding to the first search term;

[0183] Generate description information for the determined pictures to obtain picture-text pairs that meet the quality requirements;

[0184] Based on the image-text pairs that meet the quality requirements, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

[0185] In an optional manner, the incremental training module 740 is used to:

[0186] Sampling a first image-text pair from an original training set of the multimodal search large model;

[0187] Based on the first image-text pair and the image-text pair corresponding to the first search term, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

[0188] In an optional manner, the query seed also includes bad cases and / or low-frequency content in the original training set of the multimodal search large model.

[0189] Among them, the query seed construction module 710, the search term construction module 720, the search module 730 and the incremental training module 740 can all be implemented by software, or can be implemented by hardware. Exemplarily, the implementation of the query seed construction module 710 is introduced below by taking the query seed construction module 710 as an example. Similarly, the implementation of the search term construction module 720, the search module 730 and the incremental training module 740 can refer to the implementation of the query seed construction module 710.

[0190] As an example of a software functional unit, the query seed construction module 710 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the query seed construction module 710 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region can include multiple AZs.

[0191] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0192] As an example of a hardware functional unit, the query seed construction module 710 may include at least one computing device, such as a server, etc. Alternatively, the query seed construction module 710 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0193] The multiple computing devices included in the query seed construction module 710 can be distributed in the same region or in different regions. The multiple computing devices included in the query seed construction module 710 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the query seed construction module 710 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0194] It should be noted that, in other embodiments, the query seed construction module 710 can be used to execute any step in the method for multimodal search large model optimization, the search term construction module 720 can be used to execute any step in the method for multimodal search large model optimization, the search module 730 can be used to execute any step in the method for multimodal search large model optimization, and the incremental training module 740 can be used to execute any step in the method for multimodal search large model optimization. The steps that the query seed construction module 710, the search term construction module 720, the search module 730 and the incremental training module 740 are responsible for implementing can be specified as needed. The query seed construction module 710, the search term construction module 720, the search module 730 and the incremental training module 740 respectively implement different steps in the method for multimodal search large model optimization to realize all the functions of the device for multimodal search large model optimization.

[0195] The present application embodiment also provides a computing device 100. Figure 8As shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 100.

[0196] The bus 102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The bus 102 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 102 may include a path for transmitting information between various components of the computing device 100 (eg, the memory 106, the processor 104, and the communication interface 108).

[0197] The processor 104 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0198] The memory 106 may include a volatile memory, such as a random access memory (RAM). The memory 106 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0199] The memory 106 stores executable program codes, and the processor 104 executes the executable program codes to respectively implement the functions of the query seed construction module 710, the search term construction module 720, the search module 730, and the incremental training module 740, thereby implementing the multimodal search large model optimization method. That is, the memory 106 stores computer instructions for executing the multimodal search large model optimization method.

[0200] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0201] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0202] like Fig. 9 As shown, the computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the multi-modal search large model optimization method.

[0203] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also respectively store partial instructions for executing the method for multimodal search large model optimization. In other words, the combination of one or more computing devices 100 may jointly execute instructions for executing the method for multimodal search large model optimization.

[0204] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the apparatus for multimodal search large model optimization. That is, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more modules among the query seed construction module 710, the search term construction module 720, the search module 730 and the incremental training module 740.

[0205] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.10 A possible implementation is shown. Fig.10 As shown, two computing devices are connected via a network, and the two computing devices include a first computing device 100A and a second computing device 100B. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the first computing device 100A stores instructions for executing the functions of the query seed construction module 710. At the same time, the memory 106 in the second computing device 100B stores instructions for executing the functions of the search term construction module 720, the search module 730, and the incremental training module 740.

[0206] Fig.10The connection method between the computing device clusters shown can be considered to be that the method for optimizing the multimodal search large model provided in the present application may require interaction with the user, so it is considered to hand over the functions implemented by the query seed construction module 710 to the first computing device 100A for execution.

[0207] It should be understood that Fig.10 The functions of the first computing device 100A shown in FIG. 1 may also be completed by multiple computing devices 100. Similarly, the functions of the second computing device 100B may also be completed by multiple computing devices 100.

[0208] The embodiment of the present application also provides a computer program product containing instructions, which are computer instructions. The computer program product can be a software or program product containing instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device performs a method for multimodal search large model optimization.

[0209] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to perform a method for multimodal search large model optimization.

[0210] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-modal search method for large model optimization, It is characterized in that The method comprises: Acquire a query seed, wherein the query seed includes general high-frequency content and / or scene content, wherein the scene content includes high-frequency content and / or key content of a scene to which the multimodal search large model is applied; constructing a first search term corresponding to the query seed based on a first expanded term obtained by expanding the query seed; In the first data source, performing an image search for the first search term to obtain image-text pairs corresponding to the first search term, each image-text pair including an image and description information of the image; Based on the image-text pair corresponding to the first search term, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

2. The method according to claim 1, It is characterized in that The step of obtaining the query seed comprises: Get candidate seeds; Based on the second expanded term obtained by expanding the candidate seed, construct a second search term corresponding to the candidate seed; In the second data source, a text search is performed on the second search term to obtain a consultation text corresponding to the second search term; Based on the consultation text, construct a test set corresponding to the candidate seed; Using the evaluation model to infer the test samples in the test set to obtain the inference results of the test samples; Determining a performance indicator corresponding to the candidate seed based on the inference result of the test sample and the label of the test sample; If the performance indicator is lower than a first threshold, the candidate seed is added to the query seed.

3. The method according to claim 2, It is characterized in that The method further comprises: The test samples corresponding to the candidate seeds are obtained from the test set of the evaluation model, and added to the test set corresponding to the candidate seeds.

4. The method according to any one of claims 1 to 3, It is characterized in that The multimodal search big model is a big model for searching text by image or searching images by image, the general high-frequency content includes a general high-frequency heat map, and the scene content includes a high-frequency heat map and / or key pictures of the scene to which the multimodal search big model is applied; The step of constructing a first search term corresponding to the query seed based on the first expanded term obtained by expanding the query seed includes: Converting the query seed into a query term; Expanding the query term to obtain the first expanded term corresponding to the query term; The query term and the first expanded term are combined to form the first search term.

5. The method according to any one of claims 1 to 3, It is characterized in that The multimodal search big model is a big model for searching images with text, the general high-frequency content includes general high-frequency hot words, and the scene content includes high-frequency hot words and / or key words of the scene to which the multimodal search big model is applied; The step of constructing a first search term corresponding to the query seed based on the first expanded term obtained by expanding the query seed includes: Expanding the query seed to obtain the first expanded term corresponding to the query seed; The query seed and the first expanded term are combined to form the first search term.

6. The method according to any one of claims 1 to 5, It is characterized in that The step of optimizing the multimodal search model based on the image-text pair corresponding to the first search term to obtain an optimized multimodal search model includes: Adding the image-text pair corresponding to the first search term to the incremental training set; If the number of image-text pairs in the incremental training set is less than a second threshold, the first extended term is expanded to obtain a third extended term, and the third extended term is updated to the first search term in the next round until the number of image-text pairs in the incremental training set is equal to the second threshold. The incremental training set is used to optimize the multimodal search model to obtain an optimized multimodal search model.

7. The method according to any one of claims 1 to 6, It is characterized in that The step of optimizing the multimodal search model based on the image-text pair corresponding to the first search term to obtain an optimized multimodal search model includes: Filtering the image-text pairs corresponding to the first search term to obtain filtered image-text pairs; Based on the filtered image-text pairs, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

8. The method according to claim 7, It is characterized in that The filtering of the image-text pairs corresponding to the first search term includes one or more of the following: Among the image-text pairs corresponding to the first search term, filtering the image-text pairs corresponding to the first search term based on the relevance between the first search term and the image; filtering the image-text pairs corresponding to the first search term based on the correlation between the image and the description information in the image-text pairs corresponding to the first search term; De-duplication processing is performed on the image-text pairs corresponding to the first search term; or, Among the image-text pairs corresponding to the first search term, image-text pairs to which images having image resolutions lower than a third threshold belong are deleted.

9. The method according to any one of claims 1 to 8, It is characterized in that The step of optimizing the multimodal search model based on the image-text pair corresponding to the first search term to obtain an optimized multimodal search model includes: Determine the pictures whose description information quality does not meet the requirements in the picture-text pair corresponding to the first search term; Generate description information for the determined pictures to obtain picture-text pairs that meet the quality requirements; Based on the image-text pairs that meet the quality requirements, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

10. The method according to any one of claims 1 to 9, It is characterized in that The step of optimizing the multimodal search model based on the image-text pair corresponding to the first search term to obtain an optimized multimodal search model includes: Sampling a first image-text pair from an original training set of the multimodal search large model; Based on the first image-text pair and the image-text pair corresponding to the first search term, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

11. The method according to claim 1, It is characterized in that The query seeds also include bad cases and / or low-frequency content in the original training set of the multimodal search large model.

12. A device for multi-modal search and large model optimization, It is characterized in that The device comprises: A query seed construction module, used to obtain a query seed, wherein the query seed includes general high-frequency content and / or scene content, and the scene content includes high-frequency content and / or key content of the scene to which the multimodal search large model is applied; A search term construction module, configured to construct a first search term corresponding to the query seed based on a first expanded term obtained by expanding the query seed; A search module, configured to perform image search for the first search term in a first data source to obtain image-text pairs corresponding to the first search term, each image-text pair including an image and description information of the image; The incremental training module is used to optimize the multimodal search model based on the image-text pair corresponding to the first search term to obtain an optimized multimodal search model.

13. The device according to claim 12, It is characterized in that The query seed construction module is used to: Get candidate seeds; Based on the second expanded term obtained by expanding the candidate seed, construct a second search term corresponding to the candidate seed; In the second data source, a text search is performed on the second search term to obtain a consultation text corresponding to the second search term; Based on the consultation text, construct a test set corresponding to the candidate seed; Using the evaluation model to infer the test samples in the test set to obtain the inference results of the test samples; Determining a performance indicator corresponding to the candidate seed based on the inference result of the test sample and the label of the test sample; If the performance indicator is lower than a first threshold, the candidate seed is added to the query seed.

14. The device according to claim 13, It is characterized in that The query seed construction module is further used to obtain the test samples corresponding to the candidate seeds in the test set of the evaluation model, and add them to the test set corresponding to the candidate seeds.

15. The device according to any one of claims 12 to 14, It is characterized in that The multimodal search big model is a big model for searching text by image or searching image by image, the general high-frequency content includes a general high-frequency heat map, and the scene content includes a high-frequency heat map and / or key pictures of the scene to which the multimodal search big model is applied; The search term construction module is used to: Converting the query seed into a query term; Expanding the query term to obtain the first expanded term corresponding to the query term; The query term and the first expanded term are combined to form the first search term.

16. The device according to any one of claims 12 to 14, It is characterized in that The multimodal search big model is a big model for searching images with text, the general high-frequency content includes general high-frequency hot words, and the scene content includes high-frequency hot words and / or key words of the scene to which the multimodal search big model is applied; The search term construction module is used to: Expanding the query seed to obtain the first expanded term corresponding to the query seed; The query seed and the first expanded term are combined to form the first search term.

17. The device according to any one of claims 12 to 16, It is characterized in that The incremental training module is used to: Adding the image-text pair corresponding to the first search term to the incremental training set; If the number of image-text pairs in the incremental training set is less than a second threshold, the first extended term is expanded to obtain a third extended term, and the third extended term is updated to the first search term in the next round until the number of image-text pairs in the incremental training set is equal to the second threshold. The incremental training set is used to optimize the multimodal search model to obtain an optimized multimodal search model.

18. The device according to any one of claims 12 to 17, It is characterized in that The incremental training module is used to: Filtering the image-text pairs corresponding to the first search term to obtain filtered image-text pairs; Based on the filtered image-text pairs, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

19. The device according to any one of claims 12 to 18, It is characterized in that The incremental training module is used to: Determine the pictures whose description information quality does not meet the requirements in the picture-text pair corresponding to the first search term; Generate description information for the determined pictures to obtain picture-text pairs that meet the quality requirements; Based on the image-text pairs that meet the quality requirements, the multimodal search large model is optimized to obtain an optimized multimodal search large model.

20. A computing device cluster, It is characterized in that comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 11.

21. A computer-readable storage medium, It is characterized in that The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Retrieval enhancement system and method based on multi-modal interaction agent

    CN120950554A