Method, device and equipment for screening cue words and readable medium

By providing sample images and prompt word sets to multimodal models, generating prediction tags and evaluating their quality, the problem of high dependence and long development cycles in the prior art is solved, and an efficient and higher-quality model training and inference process is achieved.

CN120296151APending Publication Date: 2025-07-11JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410045136.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art requires professional participation in the process of machine learning model training and inference, which is costly and long development cycle, and the low quality of prompt words leads to poor use of the model and it is difficult to achieve the functions expected by users.

Method used

A method and device are provided to generate prediction tags by providing a multimodal model with sample images and a set of prompt words, and evaluate the quality of the prompt word set based on the comparison of the predicted tags and reference tags, thereby selecting high-quality prompt words during model training and inference.

Benefits of technology

It improves the efficiency and quality of the model training and inference process, reduces the dependence of professionals, simplifies the model construction process, and improves the model processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296151A_ABST
    Figure CN120296151A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cue word screening method and device, equipment and a medium. The method comprises the following steps: providing a group of sample images and a group of cue words for a trained multi-modal model, wherein the group of cue words is used for indicating an analysis strategy of the trained multi-modal model for analyzing the group of sample images; obtaining a group of prediction labels generated by the multi-modal model based on the group of cue words, wherein the group of prediction labels indicate a prediction analysis result of the group of sample images; determining first evaluation information for the set of cues based on a comparison of the set of prediction tags and the set of reference tags for the set of sample images; and at least based on the first evaluation information, determining the use of a group of cue words in the training process of another multi-modal model or in the reasoning process of another trained multi-modal model. Therefore, prompt words with higher quality can be selected to be input into other multi-modal models with higher quality and efficiency for model reasoning or training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to methods, apparatuses, devices, and computer-readable storage media for prompt filtering. Background Art

[0002] With the development of computer technology, machine learning technology has gradually matured. In the implementation process of machine learning technology, neural networks, deep learning models, etc. are often relied on to provide functional support. Deep learning has achieved many results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech, recommendation and personalization technology, and other related fields.

[0003] Deep learning enables machines to imitate human activities such as audiovisual and thinking, solves many complex pattern recognition problems, and has made great progress in the applications of various fields. Therefore, how to more conveniently implement model training and obtain a model that meets the usage requirements is worthy of attention and an urgent need. Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for prompt filtering is provided. The method includes: providing a set of sample images and a set of prompt words to a first multimodal model, where the set of prompt words is used to indicate an analysis strategy for the first multimodal model to analyze a set of sample images; obtaining a set of predicted labels generated by the multimodal model based on the set of prompt words, where the set of predicted labels indicates a predicted analysis result for the set of sample images; determining first evaluation information for the set of prompt words based on a comparison between the set of predicted labels and a set of reference labels for the set of sample images; and determining the use of the set of prompt words during the training process of a second multimodal model or during the inference process of a trained third multimodal model at least based on the first evaluation information.

[0005] In a second aspect of the present disclosure, a device for prompt filtering is provided. The device includes: a first providing module configured to provide a set of sample images and a set of prompt words to a first multimodal model, where the set of prompt words is used to indicate an analysis strategy for the first multimodal model to analyze a set of sample images; a result obtaining module configured to obtain a set of predicted labels generated by the multimodal model based on the set of prompt words, where the set of predicted labels indicates a predicted analysis result for the set of sample images; an evaluation determining module configured to determine first evaluation information for the set of prompt words based on a comparison between the set of predicted labels and a set of reference labels for the set of sample images; and a use determining module configured to determine the use of the set of prompt words during the training process of a second multimodal model or during the inference process of a trained third multimodal model at least based on the first evaluation information.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method of the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium, which can be executed by a processor to execute the method according to the first aspect of the present disclosure.

[0008] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the following, with reference to the accompanying drawings and referring to the following detailed description, the above and other features, advantages and aspects of the various implementations of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A schematic diagram showing the architecture of a first multimodal model according to some embodiments of the present disclosure;

[0012] Figure 3 A flowchart showing a process for prompt screening according to some embodiments of the present disclosure;

[0013] Figure 4 A block diagram showing a device for prompt screening according to some embodiments of the present disclosure; and

[0014] Figure 5 A block diagram showing an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0017] It should be noted that in the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0018] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to relevant laws and regulations.

[0019] For example, when receiving a user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0020] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] As used herein, the term "model" can learn the corresponding association relationship between input and output from training data, so that after training, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably in this article.

[0022] "Neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs. It generally includes an input layer, an output layer, and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications usually include many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer. The input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.

[0023] Generally, machine learning can roughly be divided into three stages, namely, the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, the test inputs are applied to the trained model to test whether the model can provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs based on the parameter values obtained from training and determine the corresponding outputs.

[0024] Example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0025] For deep learning technologies, it is often necessary to determine the sample data for training the model according to actual needs, and use these sample data to train the model to obtain a model that can provide the functions expected by users. Therefore, the acquisition efficiency and quality of the sample data are particularly important. In addition, with the development of deep learning technologies, in order to improve the processing quality of the model, it is also allowed to guide the processing direction of the model by adding auxiliary information and guiding information, so as to reduce the processing difficulty of the model for the input data and improve the processing quality. For example, for a large language model (LM), the analysis direction of the LM can be guided by inputting prompt words, so that the LM can more clearly understand the user's intention and output results that better meet the user's needs.

[0026] In related solutions, it is usually possible to select a solution that can be implemented by an algorithm after communication between algorithm experts and business personnel about business requirements. Then, a large amount of data is collected, and common classification algorithms such as Convolutional Neural Network (CNN), Resnet, and detection algorithms such as the YOLO series of networks are used. Then, the corresponding models and networks need to be trained. However, in this process, it is often necessary for professional personnel such as algorithm experts to intervene and operate to achieve the goal. This not only incurs high personnel costs but also has problems such as a large data volume requirement and a long development cycle. In addition, in scenarios where the model can be guided based on prompts in order to expect the model to provide functions, such as in the usage scenario of LM, the quality of the LM usage may also fall short of expectations due to, for example, low-quality prompts, and it is often difficult to achieve the functions expected by users.

[0027] The present disclosure provides a method for prompt screening. In the solution of the present disclosure, a set of sample images and a set of prompt words are provided to a first multimodal model. The set of prompt words is used to indicate the analysis strategy for the first multimodal model to analyze a set of sample images; a set of predicted labels generated by the multimodal model based on the set of prompt words is obtained, and the set of predicted labels indicates the predicted analysis results for the set of sample images; first evaluation information for the set of prompt words is determined based on the comparison between the set of predicted labels and a set of reference labels for the set of sample images; and the use of the set of prompt words during the training process of a second multimodal model or during the inference process of a trained third multimodal model is determined at least based on the first evaluation information.

[0028] Thus, by evaluating the constructed prompts, prompts with higher quality can be selected to input other multimodal models with higher quality and efficiency for model inference or training.

[0029] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In Figure 1 In the environment 100, a first multimodal model 115 can be maintained in the electronic device 110. In an embodiment of the present disclosure, a set of sample images 120 and prompts in a set of prompt word sets 130-1, 130-2,..., 130-N can be provided to the first multimodal model 115 for processing, and evaluation information 140 for each set of prompt word sets is generated based on the combination of the processing results. For ease of description, the set of prompt word sets 130-1, 130-2,..., 130-N can be collectively referred to or individually referred to as the set of prompt word sets 130. The number of the set of prompt word sets 130 can be one or more.

[0030] In some embodiments, a set of sample images 120 may include at least one sample image. In the case where a set of sample images 120 includes multiple sample images, the multiple sample images may be images for the same training theme. For example, the theme may be to review whether an image includes specific content, such as whether the image includes pornographic, violent, bloody, etc. content. Taking the theme of reviewing whether an image includes bloody content, or in other words, whether the image is a bloody type of image, the sample images in a set of sample images 120 may include "bloody" content (for example, an image including human bleeding content can be regarded as including "bloody" content and belongs to a bloody image) or non-"bloody" content that is similar to "bloody" content. Correspondingly, the sample images in a set of sample images 120 can be correspondingly indicated by the labels of the content included in the sample images. For example, the label can indicate that the content included in the sample image is "bloody" or "non-bloody". In some scenarios, such a label may also be referred to as a reference label.

[0031] The set of prompt words 130 can be used as the input of the first multimodal model 115 at the same time to indicate the processing strategy of the first multimodal model 115 for a set of sample images 120. For example, it indicates whether the first multimodal model 115 identifies whether the sample images in a set of sample images 120 include target content. Thus, the set of prompt words 130 can be used to control the processing direction of the first multimodal model 115, so that the first multimodal model 115 can process a set of sample images 120 in a pre-determined manner and strategy. Further, the electronic device 110 can determine the evaluation information 140 corresponding to the set of prompt words 130 based on the processing result of the first multimodal model 115 for the set of prompt words 130. Correspondingly, the electronic device 110 can select whether to use the set of prompt words 130 for the training process of the second multimodal model 116 or use the set of prompt words 130 in the inference process of the trained third multimodal model 117 based on the evaluation information 140.

[0032] In Figure 1 , the electronic device 110 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device may refer to any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server includes but is not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and so on.

[0033] It should be understood that Figure 1The components and arrangements in the illustrated environment 100 are merely examples, and computing systems suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. Implementations of this disclosure are not limited in this regard.

[0034] Some example embodiments of this disclosure will be described below with continued reference to the accompanying drawings. For ease of understanding, it may be described in conjunction with the above Figure 1 For example, the process for prompt filtering can be implemented by the electronic device 110.

[0035] The electronic device 110 provides a set of sample images 120 and a set of prompt words 130 to the first multimodal model 115. In an embodiment of this disclosure, the set of prompt words 130 is used to indicate the analysis strategy for the first multimodal model 115 to analyze a set of sample images 120. Thus, the first multimodal model 115 can be utilized to process a set of sample images 120 according to the set of prompt words 130, or rather, to analyze a set of sample images 120 according to the analysis strategy guided by the set of prompt words 130, and obtain a corresponding set of predicted labels. The set of predicted labels indicates the predicted analysis results for a set of sample images 120, or rather, processes a set of sample images 120 according to the analysis strategy indicated by the set of prompt words 130 to generate corresponding predicted analysis results. In some embodiments, the prompt words in the set of prompt words 130 can be used alone or in combination with other prompt words, thereby enriching the prompt words and analysis strategies.

[0036] For ease of understanding, a set of image samples 120 will be used as an example for the theme of reviewing whether an image is "bloody" hereinafter. Correspondingly, the prompt words in the set of prompt words 130 can be used to guide the first multimodal model 115 to analyze and process the sample images from multiple strategies such as whether the content included in the sample images is bloody content, whether there are objects that can be determined as "bloody content" in the sample images, and whether the sample images are bloody images, etc., to generate a set of predicted labels corresponding to the set of prompt words 130. For example, when prompting whether there are objects that can be determined as "bloody content" in the sample images, the predicted label can be the predicted result of "included" or "not included".

[0037] In some embodiments, the first multimodal model 115 can be a zero-shot learning model or a few-shot learning model. A zero-shot learning model is a model constructed based on the zero-shot learning approach. Zero-shot learning can solve the classification problem of image objects without available training data for classes or tasks and only providing descriptions of classes or tasks. After mapping visual data to semantic representation data, a zero-shot learning model can identify new objects based on the semantic descriptions of the labeled data. Thus, using a zero-shot learning model can not only reduce the initial construction cost of the model but also filter out prompts with relatively abundant information, or in other words, prompts that are more easily understood by the model. The first multimodal model 115 can also be a few-shot learning model. A few-shot learning model is a model obtained through few-shot learning (FSL), which is a machine learning method that trains a dataset with limited information. The few-shot learning approach can build an accurate machine learning model with less training data. Since the dimension of the input data is a factor determining resource costs (such as time cost, computing cost, etc.), the data analysis / machine learning cost can be reduced by using few-shot learning.

[0038] In some embodiments, the prompt set 130 can include multiple groups of prompts or prompts in different forms. Different groups of prompts can indicate different analysis strategies, or different groups of prompts indicating the same analysis strategy can also have different natural language expressions. Thus, by evaluating these prompts separately, a prompt set that is more conducive to achieving the training and usage purposes can be determined.

[0039] In some embodiments, the prompt set provided to the first multimodal model 115 can be determined based on at least one of the following four scenarios as examples.

[0040] In the first solution, the prompt set 130 includes a first set of prompts. Each prompt in the first set of prompts instructs the first multi-modal model 115 to determine whether the corresponding sample image includes a target object. For example, in the scenario of reviewing bloody images, the target object can be "blood", or for example, more specific "blood" flowing out due to injuries, etc. For example, the electronic device 110 can input a sample image and a prompt from the first set of prompts. For example, the sample image can be a "scratched skin" picture, and the prompt can be "Is there anyone injured and bleeding in the picture". Accordingly, after processing based on the sample image and the prompt, the first multi-modal model 115 can obtain a prediction label indicating whether the target object (i.e., "blood" flowing out due to human injury) is included in the sample image. Further, the electronic device 110 continues to input a sample image (the newly input sample image can be the same as or different from the previous one) and a prompt different from the previous one, so as to use the first multi-modal model 115 to generate a new prediction label under the guidance of the newly input prompt. Finally, after obtaining the prediction labels corresponding to each prompt in the first set of prompts, a set of prediction labels is formed. Thus, subsequent independent and targeted evaluation of each prompt can be carried out.

[0041] In the second solution, the prompt set 130 includes a second set of prompts and a third set of prompts. Each prompt in the second set of prompts instructs the first multi-modal model 115 to generate description information of the corresponding sample image (for example, the content of the image expressed in text form). The electronic device 110 can first use each prompt in the second set of prompts to guide the first multi-modal model 115 to generate description information of the corresponding sample image in the set of sample images 120. Accordingly, the prediction content generated by the first multi-modal model 115 can be referred to as an intermediate prediction label.

[0042] For example, the prompt words in the second set of prompt words can be "describe the content in the image". For example, in the above example, the sample image can also be "scratched skin", and the first multimodal model 115 can generate an intermediate prediction label based on, for example, "describe the content in the image". The intermediate prediction label is, for example, "There is a person in the picture wearing a gray T-shirt, with the left arm injured and bleeding, and the right hand covering the left arm with a gauze...". Further, the electronic device 110 provides the intermediate prediction label and the prompt words in the third set of prompt words to the first multimodal model 115. Each prompt word in the third set of prompt words instructs the first multimodal model 115 to determine whether the predicted description information includes a target object such as "blood". For example, the prompt word in the third set of prompt words can be "Is there anyone injured and bleeding in the picture?". Accordingly, after processing based on the sample image and the prompt words, the first multimodal model 115 can obtain a prediction label indicating whether the target object (i.e., "blood" flowing out due to a person's injury) is included in the sample image. Similarly, the electronic device 110 can process the combination of the prompt words in each second set of prompt words and the prompt words in the third set of prompt words in a similar manner to obtain a set of prediction labels. Each prediction label in the set of prediction labels indicates whether the corresponding sample image includes the target object. Thus, it is possible to evaluate the prompt words obtained in a combined form, such as the prompt words guided through multiple rounds of conversations.

[0043] In the third solution, the prompt word set 130 includes a fourth set of prompt words and a fifth set of prompt words. Each prompt word in the fourth set of prompt words instructs the first multimodal model 115 to determine whether the corresponding sample image includes a target object (e.g., "blood"), and each prompt word in the fifth set of prompt words instructs the first multimodal model to provide a judgment criterion for the target object. In some embodiments, the judgment criterion can be preset based on the content of the target object. For example, when the target object is "blood", the judgment criterion can be, for example, "If a red patch appears on the human skin, it is considered blood; if a red patch appears on the clothes, it is considered suspected blood; if it appears in other positions, it is considered non-blood".

[0044] Accordingly, the electronic device 110 may guide the multimodal model 115 to process a set of sample images 120 based on the prompt words in the fourth set of prompt words and the prompt words in the fifth set of prompt words. For example, the prompt words in the fourth set of prompt words may be "Is there anyone injured and bleeding in the picture?", and the prompt words in the fifth set of prompt words may be "If a red patch appears on the human skin, it is considered blood". Similarly, the electronic device 110 may process the combination of the prompt words in each fourth set of prompt words and the prompt words in the fifth set of prompt words in a similar manner to obtain a set of predicted labels. Each predicted label in the set of predicted labels indicates whether the corresponding sample image includes the target object under the judgment criterion. Thus, the judgment criterion can be introduced to refine and optimize the prompt words, so as to obtain more forms of prompt words through combination.

[0045] In the fourth solution, the prompt word set 130 includes a set of analysis examples and a sixth set of prompt words. Each analysis example in the set of analysis examples includes an example image, example prompt words for the example image, and an example label for the example image. Each prompt word in the sixth set of prompt words instructs the first multimodal model 115 to refer to the corresponding analysis example to determine whether the corresponding sample image includes the target object (e.g., "blood"). For example, the example image in the analysis example may be an image of "scratched arm skin", the corresponding example prompt words may be "What is the red area on the person's arm in the picture?", and the example label may be "blood". Accordingly, the electronic device 110 may use the prompt words in the sixth set of prompt words to guide the first multimodal model 115 to refer to the corresponding analysis example. For example, the electronic device 110 may provide the first multimodal model 115 with a sample image of "scratched thigh skin", and based on the prompt word in the sixth set of prompt words "Refer to 'The image is an image of scratched arm skin. Based on the prompt word What is the red area on the person's arm in the picture, the result is blood', to analyze what is the red area on the person's thigh in the picture?". Similarly, the electronic device 110 may process the combination of multiple analysis examples and the prompt words in the sixth set of prompt words in a similar manner to obtain a set of predicted labels. Each predicted label in the set of predicted labels indicates whether the corresponding sample image includes the target object. Thus, more auxiliary information can be provided for the first multimodal model 115 through example information to help it better understand the prompt words.

[0046] For ease of understanding, reference may also be made to Figure 2 。 Figure 2 FIG. 10 shows a schematic diagram of the architecture 200 of a first multimodal model according to some embodiments of the present disclosure.

[0047] In block 213, after obtaining a set of sample images and a first set of prompt words 211, the first multimodal model 115 may process them via a convolutional neural network or a transformer model.

[0048] In some embodiments, the first multimodal model 115 may also be configured with a preprocessing module and / or a postprocessing module. For example, the preprocessing module and / or the postprocessing module may uniformly scale the images to a fixed size and add positional encoding information to the text content. Exemplarily, in architecture 200, there may also be included box 212 and / or box 214 to implement preprocessing and / or postprocessing of the images and text content.

[0049] Further, at box 250, the first multimodal model 115 may obtain a set of predicted labels 215 obtained by processing a set of sample images 120 guided by a first set of prompt words based on the processing results generated by a convolutional neural network or a transformer model.

[0050] Similarly, the first multimodal model 115 may, in a manner similar to obtaining a set of predicted labels 215, obtain a corresponding set of predicted labels 225 based on a set of obtained sample images, a second set of prompt words, and a third set of prompt words 221. Obtain a corresponding set of predicted labels 235 based on a set of sample images, a fourth set of prompt words, and a fifth set of prompt words 231. Obtain a corresponding set of predicted labels 245 based on a set of sample images, a set of analysis examples, and a sixth set of prompt words 241. In some embodiments, the first multimodal model 115 may also optimize the processing process based on a set of sample images, a second set of prompt words, and a third set of prompt words 221, the processing process based on a set of sample images, a fourth set of prompt words, and a fifth set of prompt words 231, and the processing process based on a set of sample images, a set of analysis examples, and a sixth set of prompt words 241 by adding a preprocessing module and / or a postprocessing module.

[0051] In an embodiment of the present disclosure, the electronic device 110 may determine first evaluation information for a set of prompt words based on a comparison between a set of prediction tags and a set of reference tags for a set of sample images (e.g., whether "blood" is included in the corresponding sample image). In some embodiments, the reference tags in a set of reference tags may be divided into two categories. For example, in the scenario of whether blood is included in a sample image, the first type of reference tag may indicate that blood is included in the sample image. Correspondingly, the second type of reference tag may indicate that blood is not included in the sample image. Thus, the types of sample images can be distinguished by different types of reference tags. In some embodiments, the first evaluation information includes recall information, and the recall information is determined based on the comparison result between a set of prediction tags and the first type of tags in a set of reference tags. For example, taking "the first type of tag" to indicate that the sample image "includes blood" as an example. For a sample image whose prediction tag indicates "includes blood", if its corresponding reference tag is also "includes blood", it is considered that the prediction tag indicates correctly. Correspondingly, the electronic device 110 may determine the recall information based on the ratio of the number of correctly indicated prediction tags to the total data of a set of prediction tags. Thus, the recall information can be used to characterize the accuracy, or rather, the recall rate, of guiding the first multimodal model 115 based on the set of prompt words 130.

[0052] In some embodiments, the first evaluation information includes precision information, and the precision information is determined based on the comparison result between a set of prediction tags and the second type of tags in a set of reference tags. For example, taking "the second type of tag" to indicate that the sample image "does not include blood" as an example. For a sample image whose prediction tag indicates "does not include blood", if its corresponding reference tag is also "does not include blood", it is considered that the prediction tag indicates correctly. Correspondingly, the electronic device 110 may determine the precision information based on the ratio of the number of correctly indicated prediction tags to the total data of a set of prediction tags. Thus, the precision information can be used to characterize the precision of guiding the first multimodal model 115 based on the set of prompt words 130.

[0053] In some embodiments, the electronic device 110 may continuously generate recall information and precision information as the first multimodal model 115 processes a set of sample images 120 and a set of prompt words 130. For example, for the used guiding words, taking "the first type of tag" to indicate that the sample image "includes blood" and "the second type of tag" to indicate that the sample image "does not include blood" as an example. For the recall information, if the first multimodal model 115 discriminates correctly, the TP count is incremented by 1, and if the model judges incorrectly, the FN count is incremented by 1; for the precision information, if the first multimodal model 115 discriminates correctly, the TN count is incremented by 1, and if the model judges incorrectly, the FP count is incremented by 1. At this time, the recall information can be calculated based on the following formula (1):

[0054] Recall information = TP / (TP + FN) (1)

[0055] Calculate the precision information based on the following formula (2):

[0056] Precision information = TN / (TN + FP) (2)

[0057] In an embodiment of the present disclosure, the electronic device 110 may determine the use of the set of prompt words during the training process of the second multimodal model 116 or during the inference process of the trained third multimodal model 117 at least based on the evaluation information 140 (e.g., the first evaluation information for the set of prompt words 130). In some embodiments, the second multimodal model 116 may be a model similar to the first modal model 115, or a model with a larger model scale and higher complexity than the first modal model 140. For example, the second multimodal model 116 may be a pre-configured general model. Accordingly, the first multimodal model with a weaker model can quickly complete the evaluation of the usage quality of the set of prompt words 130 to screen out prompt words that are more suitable for training the second multimodal model 116. For example, the electronic device 110 may determine the set of prompt words 130 whose recall information and / or precision information values exceed a preset threshold as the available set of prompt words, and use the available set of prompt words to implement the training of the second multimodal model 116. For example, the second multimodal model 116 is trained using the available set of prompt words and a set of sample images 120 so that the second multimodal model 116 has the ability to review bloody images. In some embodiments, the evaluated set of prompt words 130 may also be used again for the first multimodal model 115. For example, when the first multimodal model 115 is subsequently used to process other input images, a set of prompt words with higher quality can be selected according to the priority order. Thus, the processing quality of the first multimodal model 115 can be improved by using "prompt words that can be more guided by the first multimodal model 115". For example, in a scenario where the first multimodal model 115 is also used for image review, the electronic device 110 may determine the desired prompt words to use based on the priority order.

[0058] Similarly, the electronic device 110 can also screen out the set of prompt words during the inference process of the trained third multimodal model 117 in a similar manner (e.g., determining the set of prompt words 130 whose recall information and / or precision information values exceed a preset threshold as the available set of prompt words). For example, the electronic device 110 may use the available set of prompt words to instruct the LM so that the LM can complete the bloody image review task more accurately and with higher quality.

[0059] It should be understood that the set of prompt words provided to the first multi-modal model 115 (e.g., the set of prompt words 130) can be defined according to any one of the four schemes in the above examples.

[0060] In some embodiments, the electronic device 110 may further continue to provide a set of sample images 120 and another set of prompt words 130 to the first multi-modal model 115 to determine the second evaluation information for another set of prompt words. The other set of prompt words 130 is used to indicate the analysis strategy for the first multi-modal model to analyze a set of sample images, and the other set of prompt words 130 is different in natural language expression from the previously used set of prompt words 130. For example, different natural languages can be used in the other set of prompt words to express, for example, the above-mentioned first set of prompt words. For example, in the previously used set of prompt words 130, the example of the prompt word in the first set of prompt words is "Is there anyone injured and bleeding in the picture", then in the other set of prompt words, the prompt word for achieving the same purpose can be expressed as, for example, "Is there a person with bleeding and injury in the picture", etc. Accordingly, the electronic device 110 can use the first multi-modal model 115 to determine the second evaluation information for another set of prompt words based on the way of processing the set of prompt words 130.

[0061] Accordingly, the other set of prompt words can be different from the above-mentioned set of prompt words (e.g., the previously provided set of prompt words can be the set of prompt words 130-1, and the other set of prompt words can be the set of prompt words 130-2), and can be defined, for example, according to any one of the four exemplary schemes described above.

[0062] Furthermore, the electronic device 110 can determine the usage priority of the set of prompt words 130 and the other set of prompt words during the training process of the second multi-modal model 116 or during the inference process of the third multi-modal model 117 based on the comparison of the first evaluation information and the second evaluation information. For example, the electronic device 110 can determine the usage priority of the set of prompt words 130 and the other set of prompt words based on the numerical size of the recall information. For example, the party corresponding to the larger numerical value of the recall information can have a higher usage priority, or rather, be used more preferentially. Another example is that the party corresponding to the larger numerical value of the precision information can have a higher usage priority, or rather, be used more preferentially. Thus, the usage order of multiple sets of prompt words can be determined based on the evaluation information, which is convenient for users to select and use according to their needs later. Of course, according to actual application needs, the first multi-modal model can be selected to evaluate more sets of prompt words 130, and how to use these sets of prompt words can be determined based on the determined evaluation information.

[0063] In some embodiments, if the electronic device 110 determines that the set of prompt words 130 is used during the training of the second multimodal model 116 or during the inference process of the third multimodal model 117. The electronic device 110 may determine a set of input images for the second multimodal model 116 or the third multimodal model 117 based on a corresponding set of sample images 120 of the set of prompt words 130, and the set of input images is of the same type as the set of sample images 120. Specifically, when the electronic device 110 determines that the set of prompt words 130 is used, it may provide the type of the set of sample images for generating the evaluation information 140 of the set of prompt words 130. Accordingly, the user can select a set of input images according to this type for the training of the second multimodal model 116 or the use of the third multimodal model 117. Thus, input expansion can be performed based on the type of a set of sample images, so that the user can select a suitable set of prompt words (e.g., the set of prompt words 130) according to the type of the input images to be processed as expected. Thus, the usage scenarios and scope of the available set of prompt words can be enhanced to improve their usage value.

[0064] In some embodiments, the electronic device 110 may also record the sample images corresponding to the prompt words and the predicted labels corresponding to the prompt words. Thus, the user can retrieve the required prompt words according to the needs based on the recorded results. For example, provide sample images to find available prompt words that can be used to guide the recognition of the sample images (e.g., by means of image search). For example, for a set of prompt words determined for the purpose of reviewing whether an image is a bloody image, it can be used to guide, for example, the third multimodal model 117 to evaluate whether other types of newly input images are bloody images.

[0065] Thus, the overall process method of selecting prompt words using a multimodal model in this solution can facilitate different users to quickly and low-costly build their own business models (e.g., by guiding the multimodal model with high-quality prompt words to achieve rapid construction). Without the need for users to learn professional knowledge and train models, they only need to prepare several pictures and prompt texts to intuitively understand the training and usage effects of the models.

[0066] Subsequently, according to an embodiment of the present disclosure, a set of sample images and a set of prompting words are provided to a trained multi-modal model, where the set of prompting words is used to indicate an analysis strategy for the trained multi-modal model to analyze the set of sample images; a set of predicted labels generated by the multi-modal model based on the set of prompting words is obtained, and the set of predicted labels indicates a predicted analysis result for the set of sample images; first evaluation information for the set of prompting words is determined based on a comparison between the set of predicted labels and a set of reference labels for the set of sample images; and the use of the set of prompting words during the training process of another multi-modal model or during the inference process of a trained another multi-modal model is determined at least based on the first evaluation information. Thus, by evaluating the constructed prompting words, higher-quality prompting words can be selected to be input into other multi-modal models with higher quality and efficiency for model inference or training.

[0067] Figure 3 FIG. 4 shows a flowchart of a process 300 for prompting word screening according to some embodiments of the present disclosure. The process 300 can be implemented at the electronic device 110.

[0068] At block 310, the electronic device 110 provides a set of sample images and a set of prompting words to the first multi-modal model.

[0069] In an embodiment of the present disclosure, the set of prompting words is used to indicate an analysis strategy for the first multi-modal model to analyze a set of sample images.

[0070] At block 320, the electronic device 110 obtains a set of predicted labels generated by the multi-modal model based on the set of prompting words.

[0071] In an embodiment of the present disclosure, the set of predicted labels indicates a predicted analysis result for a set of sample images.

[0072] At block 330, the electronic device 110 determines first evaluation information for the set of prompting words based on a comparison between the set of predicted labels and a set of reference labels for the set of sample images.

[0073] At block 340, the electronic device 110 determines the use of the set of prompting words during the training process of the second multi-modal model or during the inference process of the trained third multi-modal model at least based on the first evaluation information.

[0074] In some embodiments, the first multi-modal model includes a zero-shot learning model or a few-shot learning model.

[0075] In some embodiments, the set of prompting words includes at least one of the following: a first set of prompting words, where each prompting word in the first set of prompting words instructs the first multimodal model to determine whether the corresponding sample image includes a target object; a second set of prompting words, where each prompting word in the second set of prompting words instructs the first multimodal model to generate descriptive information for the corresponding sample image; a third set of prompting words, where each prompting word in the third set of prompting words instructs the first multimodal model to determine whether the predicted descriptive information includes a target object; a fourth set of prompting words, where each prompting word in the fourth set of prompting words instructs the first multimodal model to determine whether the corresponding sample image includes a target object; a fifth set of prompting words, where each prompting word in the fifth set of prompting words instructs the first multimodal model to provide a judgment criterion for the target object; a sixth set of prompting words, where each prompting word in the sixth set of prompting words instructs the first multimodal model to refer to the corresponding analysis example and determine whether the corresponding sample image includes a target object.

[0076] In some embodiments, if the set of prompting words includes the first set of prompting words, each prediction label in a set of prediction labels indicates whether the corresponding sample image includes a target object.

[0077] In some embodiments, if the set of prompting words includes the second set of prompting words and the third set of prompting words, each prediction label in a set of prediction labels indicates whether the corresponding sample image includes a target object.

[0078] In some embodiments, providing a set of sample images and a set of prompting words to the first multimodal model includes: providing a set of sample images and the second set of prompting words to the first multimodal model; obtaining a set of intermediate prediction labels generated by the first multimodal model based on the second set of prompting words, where each intermediate prediction label indicates the predicted descriptive information of the corresponding sample image; and providing a set of intermediate prediction labels and the third set of prompting words to the first multimodal model.

[0079] In some embodiments, if the set of prompting words includes the fourth set of prompting words and the fifth set of prompting words, each prediction label in a set of prediction labels indicates whether the corresponding sample image includes a target object under the judgment criterion.

[0080] In some embodiments, if the set of prompting words includes a set of analysis examples and the sixth set of prompting words, and each analysis example in the set of analysis examples includes an example image, example prompting words for the example image, and an example label for the example image, each prediction label in a set of prediction labels indicates whether the corresponding sample image includes a target object.

[0081] In some embodiments, process 300 further includes: providing a set of sample images and another set of prompt words to a first multimodal model to determine second evaluation information for the another set of prompt words, where the another set of prompt words is used to indicate an analysis strategy for the first multimodal model to analyze a set of sample images, and the another set of prompt words is different from the set of prompt words in natural language expression; and determining a usage priority of the set of prompt words and the another set of prompt words during the training process of the second multimodal model or during the inference process of a third multimodal model based on a comparison between the first evaluation information and the second evaluation information.

[0082] In some embodiments, the first evaluation information includes at least one of recall information and precision information, where the recall information is determined based on a comparison result between a set of predicted labels and first-type labels in a set of reference labels, and the precision information is determined based on a comparison result between a set of predicted labels and second-type labels in a set of reference labels.

[0083] In some embodiments, process 300 further includes: if it is determined that the set of prompt words is used during the training process of the second multimodal model or during the inference process of the third multimodal model, determining a set of input images for the second multimodal model or the third multimodal model based on the set of sample images corresponding to the set of prompt words, where the set of input images is of the same type as the set of sample images.

[0084] Figure 4 FIG. shows a block diagram of an apparatus 400 for prompt word screening according to some embodiments of the present disclosure. The apparatus 400 may be implemented as or included in an electronic device 110.

[0085] The apparatus 400 includes a first providing module 410 configured to provide a set of sample images and a set of prompt words to a first multimodal model, where the set of prompt words is used to indicate an analysis strategy for the first multimodal model to analyze a set of sample images; a result obtaining module 420 configured to obtain a set of predicted labels generated by the multimodal model based on the set of prompt words, where the set of predicted labels indicates a predicted analysis result for a set of sample images; an evaluation determining module 430 configured to determine first evaluation information for the set of prompt words based on a comparison between a set of predicted labels and a set of reference labels for a set of sample images; and a usage determining module 440 configured to determine the usage of the set of prompt words during the training process of the second multimodal model or during the inference process of a trained third multimodal model at least based on the first evaluation information.

[0086] In some embodiments, the first multimodal model includes a zero-shot learning model or a few-shot learning model.

[0087] In some embodiments, the set of prompting words includes at least one of the following: a first set of prompting words, where each prompting word in the first set of prompting words instructs the first multimodal model to determine whether the corresponding sample image includes a target object; a second set of prompting words, where each prompting word in the second set of prompting words instructs the first multimodal model to generate descriptive information for the corresponding sample image; a third set of prompting words, where each prompting word in the third set of prompting words instructs the first multimodal model to determine whether the predicted descriptive information includes a target object; a fourth set of prompting words, where each prompting word in the fourth set of prompting words instructs the first multimodal model to determine whether the corresponding sample image includes a target object; a fifth set of prompting words, where each prompting word in the fifth set of prompting words instructs the first multimodal model to provide a judgment criterion for the target object; a sixth set of prompting words, where each prompting word in the sixth set of prompting words instructs the first multimodal model to refer to the corresponding analysis example to determine whether the corresponding sample image includes a target object.

[0088] In some embodiments, if the set of prompting words includes the first set of prompting words, each prediction label in a set of prediction labels indicates whether the corresponding sample image includes a target object.

[0089] In some embodiments, if the set of prompting words includes the second set of prompting words and the third set of prompting words, each prediction label in a set of prediction labels indicates whether the corresponding sample image includes a target object.

[0090] In some embodiments, the first providing module 410 is further configured to: provide a set of sample images and a second set of prompting words to the first multimodal model; obtain a set of intermediate prediction labels generated by the first multimodal model based on the second set of prompting words, where each intermediate prediction label indicates the predicted descriptive information of the corresponding sample image; and provide a set of intermediate prediction labels and a third set of prompting words to the first multimodal model.

[0091] In some embodiments, if the set of prompting words includes the fourth set of prompting words and the fifth set of prompting words, each prediction label in a set of prediction labels indicates whether the corresponding sample image includes a target object under the judgment criterion.

[0092] In some embodiments, if the set of prompting words includes a set of analysis examples and the sixth set of prompting words, and each analysis example in the set of analysis examples includes an example image, example prompting words for the example image, and an example label for the example image, each prediction label in a set of prediction labels indicates whether the corresponding sample image includes a target object.

[0093] In some embodiments, the apparatus 400 further includes: a second providing module configured to provide a set of sample images and another set of prompt words to the first multimodal model to determine second evaluation information for the another set of prompt words, where the another set of prompt words is used to indicate an analysis strategy for the first multimodal model to analyze a set of sample images, and the another set of prompt words is different from the set of prompt words in natural language expression; and based on a comparison between the first evaluation information and the second evaluation information, determine the usage priority of the set of prompt words and the another set of prompt words during the training process of the second multimodal model or during the inference process of the third multimodal model.

[0094] In some embodiments, the first evaluation information includes at least one of recall information and precision information, where the recall information is determined based on a comparison result between a set of predicted labels and first type labels in a set of reference labels, and the precision information is determined based on a comparison result between a set of predicted labels and second type labels in a set of reference labels.

[0095] In some embodiments, the apparatus 400 further includes: an input determining module configured to, if it is determined that the set of prompt words is used during the training process of the second multimodal model or during the inference process of the third multimodal model, determine a set of input images for the second multimodal model or the third multimodal model based on the set of sample images corresponding to the set of prompt words, and the set of input images is of the same type as the set of sample images.

[0096] The modules included in the apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units in the apparatus 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0097] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 5 The electronic device 500 shown is merely exemplary and should not impose any limitation on the functions and scope of the embodiments described herein.

[0098] As Figure 5As shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be an actual or virtual processor and be capable of performing various processes according to the programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.

[0099] The electronic device 500 generally includes multiple computer storage media. Such media may be any available media accessible to the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 may be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 may be a removable or non-removable medium and may include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.

[0100] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 , a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules that are configured to perform the various methods or actions of the various embodiments of the present disclosure.

[0101] The communication unit 540 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 500 may be implemented by a single computing cluster or multiple computer machines that are capable of communicating through a communication connection. Thus, the electronic device 500 may operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0102] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed through the communication unit 540. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0103] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having one or more computer instructions stored thereon, wherein the one or more computer instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.

[0104] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0105] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing apparatus, a device is produced that implements the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing apparatus, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0106] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0107] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0108] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the field to understand the implementations disclosed herein.

Claims

1. A method for prompt screening, comprising: Providing a set of sample images and a set of prompts to a first multi-modal model, where the set of prompts is used to indicate the analysis strategy for the first multi-modal model to analyze the set of sample images; Obtaining a set of predicted labels generated by the multi-modal model based on the set of prompts, where the set of predicted labels indicates the predicted analysis results for the set of sample images; Determining first evaluation information for the set of prompts based on a comparison between the set of predicted labels and a set of reference labels for the set of sample images; And Determining the use of the set of prompts during the training process of a second multi-modal model or during the inference process of a trained third multi-modal model at least based on the first evaluation information.

2. The method according to claim 1, wherein the first multi-modal model includes a zero-shot learning model or a few-shot learning model.

3. The method according to claim 1, wherein the set of prompts includes at least one of the following: A first set of prompts, where each prompt in the first set of prompts indicates that the first multi-modal model determines whether the corresponding sample image includes a target object; A second set of prompts, where each prompt in the second set of prompts indicates that the first multi-modal model generates descriptive information for the corresponding sample image; A third set of prompts, where each prompt in the third set of prompts indicates that the first multi-modal model determines whether the predicted descriptive information includes a target object; A fourth set of prompts, where each prompt in the fourth set of prompts indicates that the first multi-modal model determines whether the corresponding sample image includes a target object; A fifth set of prompts, where each prompt in the fifth set of prompts indicates that the first multi-modal model provides a judgment criterion for the target object; A sixth set of prompts, where each prompt in the sixth set of prompts indicates that the first multi-modal model refers to the corresponding analysis example to determine whether the corresponding sample image includes a target object.

4. The method according to claim 3, wherein if the set of prompts includes the first set of prompts, each predicted label in the set of predicted labels indicates whether the corresponding sample image includes the target object.

5. The method according to claim 3, wherein if the set of prompts includes the second set of prompts and the third set of prompts, each predicted label in the set of predicted labels indicates whether the corresponding sample image includes the target object.

6. The method according to claim 5, wherein providing a set of sample images and a set of prompts to a first multi-modal model includes: Providing the set of sample images and the second set of prompts to the first multi-modal model; Obtaining a set of intermediate predicted labels generated by the first multi-modal model based on the second set of prompts, where each intermediate predicted label indicates the predicted descriptive information for the corresponding sample image; And Providing the set of intermediate predicted labels and the third set of prompts to the first multi-modal model.

7. The method according to claim 3, wherein if the set of prompting words includes a fourth set of prompting words and a fifth set of prompting words, each prediction label in the set of prediction labels indicates whether the target object is included in the corresponding sample image under the judgment criterion.

8. The method according to claim 3, wherein if the set of prompting words includes a set of analysis examples and a sixth set of prompting words, and each analysis example in the set of analysis examples includes an example image, example prompting words for the example image, and an example label of the example image, each prediction label in the set of prediction labels indicates whether the target object is included in the corresponding sample image.

9. The method according to claim 1, further comprising: Providing the set of sample images and another set of prompting words to the first multi-modal model to determine second evaluation information for the another set of prompting words, where the another set of prompting words is used to indicate an analysis strategy for the first multi-modal model to analyze the set of sample images, and the another set of prompting words is different from the set of prompting words in natural language expression; And Based on a comparison of the first evaluation information and the second evaluation information, determining a usage priority of the set of prompting words and the another set of prompting words during the training process of the second multi-modal model or during the inference process of the third multi-modal model.

10. The method according to claim 1, wherein the first evaluation information includes at least one of recall information and precision information, wherein the recall information is determined based on a comparison result between the set of prediction labels and the first type of labels in the set of reference labels, and the precision information is determined based on a comparison result between the set of prediction labels and the second type of labels in the set of reference labels.

11. The method according to claim 1, further comprising: If it is determined that the set of prompting words is used during the training process of the second multi-modal model or during the inference process of the third multi-modal model, determining a set of input images for the second multi-modal model or the third multi-modal model based on the set of sample images corresponding to the set of prompting words, where the set of input images is of the same type as the set of sample images.

12. An apparatus for prompting word screening, comprising: A first providing module configured to provide a set of sample images and a set of prompting words to a first multi-modal model, where the set of prompting words is used to indicate an analysis strategy for the first multi-modal model to analyze the set of sample images; A result obtaining module configured to obtain a set of prediction labels generated by the multi-modal model based on the set of prompting words, where the set of prediction labels indicates a prediction analysis result for the set of sample images; An evaluation determining module configured to determine first evaluation information for the set of prompting words based on a comparison between the set of prediction labels and a set of reference labels for the set of sample images; And A usage determining module configured to determine the usage of the set of prompting words during the training process of the second multi-modal model or during the inference process of the trained third multi-modal model at least based on the first evaluation information.

13. An electronic device, comprising: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Image processing method and device based on large model, medium, electronics and product

    CN121392544A