Visual big language model training method and device and image analysis method and device

By determining the key entities in the image and generating corresponding negative sample images during the training process of the visual large language model, the problem that positive and negative sample pairs in the prior art cannot provide a clear causal basis, and the accuracy of the model in causal reasoning tasks is improved.

CN119992250APending Publication Date: 2025-05-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510058785.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During the training process of existing visual large language models, positive and negative samples cannot provide clear reasoning causal basis, which leads to the model's easy to regard non-critical entities as reasoning causal basis when performing causal reasoning tasks, which reduces the accuracy of image analysis results.

Method used

By identifying key entities that have a causal relationship with the image subject from the positive sample image and replacing these key entities with the target entity to generate negative sample images, training the visual large language model ensures that the model can clearly learn the causal relationship of the key entities.

Benefits of technology

This method can improve the accuracy of visual large language models in causal reasoning tasks, ensuring that the model does not regard non-critical entities as reasoning causal basis, thereby improving the accuracy of image analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992250A_ABST
    Figure CN119992250A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a visual large language model and an image analysis method and device, and relates to the technical field of computers, in particular to the application fields of machine learning, large language models, visual large language models, risk control and security, automatic driving, smart home, medical image analysis, visual content analysis, generative artificial intelligence and the like. According to the specific implementation scheme, the method comprises the steps of determining a key entity having a causal relationship with an image theme of a positive sample image from the positive sample image; replacing the key entity in the positive sample image by using the target entity to obtain a negative sample image; wherein the target entity does not have a causal relationship with the image theme of the positive sample image; and training the visual large language model by using the positive sample image and the negative sample image to obtain a trained visual large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to application fields such as machine learning, large language models, visual large language models, risk control and security, autonomous driving, smart homes, medical image analysis, visual content analysis, and generative artificial intelligence, and specifically to a training method, image analysis method, and device for a visual large language model. Background Art

[0002] The Visual Big Language Model is a deep learning model that deeply integrates computer vision and natural language processing technologies. It can not only understand and analyze image content, but also cleverly generate relevant natural language descriptions or instructions, thus building a seamless communication bridge between vision and language. At present, the Visual Big Language Model is often entrusted with important tasks due to its powerful functional characteristics to perform complex causal reasoning tasks. Summary of the invention

[0003] The present invention provides a training method for a large visual language model, an image analysis method and a device.

[0004] According to a first aspect of the present disclosure, a method for training a visual large language model is provided, comprising:

[0005] Determining from the positive image a key entity that has a causal relationship with the image subject of the positive image;

[0006] Using the target entity, the key entity in the positive sample image is replaced to obtain a negative sample image; wherein there is no causal relationship between the target entity and the image subject of the positive sample image;

[0007] The visual large language model is trained using the positive sample images and the negative sample images to obtain a trained visual large language model.

[0008] According to a second aspect of the present disclosure, there is provided an image analysis method, comprising:

[0009] Get the target image;

[0010] The target image is input into the trained visual big language model to obtain an image analysis result for the target image using the trained visual big language model.

[0011] According to a third aspect of the present disclosure, a training device for a visual large language model is provided, comprising:

[0012] A key entity determination unit, used to determine a key entity having a causal relationship with an image subject of the positive sample image from the positive sample image;

[0013] An entity replacement unit, used to replace a key entity in a positive sample image with a target entity to obtain a negative sample image; wherein there is no causal relationship between the target entity and an image subject of the positive sample image;

[0014] The model training unit is used to train the visual large language model using the positive sample images and the negative sample images to obtain a trained visual large language model.

[0015] According to a fourth aspect of the present disclosure, there is provided an image analysis device, comprising:

[0016] An image acquisition unit, used for acquiring a target image;

[0017] The image analysis unit is used to input the target image into the trained visual large language model to obtain an image analysis result for the target image using the trained visual large language model.

[0018] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0019] at least one processor;

[0020] a memory communicatively coupled to the at least one processor;

[0021] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided in the first aspect of the present disclosure.

[0022] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method provided according to the first aspect of the present disclosure.

[0023] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method provided according to the first aspect of the present disclosure when executed by a processor.

[0024] The use of the present disclosure can improve the accuracy of image analysis results.

[0025] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0027] Figure 1A flowchart of a method for training a large visual language model provided in an embodiment of the present disclosure;

[0028] Figure 2 An example diagram of the training process of a visual large language model provided in an embodiment of the present disclosure;

[0029] Figure 3 A schematic diagram of a flow chart of an image analysis method provided in an embodiment of the present disclosure;

[0030] Figure 4 A schematic diagram of an application scenario of a method for training a large visual language model provided in an embodiment of the present disclosure;

[0031] Figure 5 A schematic diagram of an application scenario of an image analysis method provided by an embodiment of the present disclosure;

[0032] Figure 6 A schematic structural block diagram of a training device for a visual large language model provided in an embodiment of the present disclosure;

[0033] Figure 7 A schematic structural block diagram of an image analysis device provided in an embodiment of the present disclosure;

[0034] Figure 8 A schematic structural block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0036] As described in the background technology, at present, Visual Large Language Model (VLLM) is often entrusted with important tasks to perform complex causal reasoning tasks due to its powerful functional characteristics. However, the inventors have found that at present, when training VLLM, positive sample images and negative sample images are usually obtained from the Internet, and after manual labeling, they are directly used to train VLLM to obtain a trained VLLM.

[0037] For example, a positive sample image A1 and a negative sample image A2 are obtained from the Internet. Assume that the positive sample image A1 is manually labeled as a pornographic image, which includes entities B1, B2, B3, and B4; the negative sample image is manually labeled as a non-pornographic image, which includes entities B3, B4, B5, and B6, and that in the positive sample image A1, entity B1 is the only key entity that causes the positive sample image A1 to be labeled as a pornographic image, that is, entities B2, B3, and B4 are not the key entities that cause the positive sample image A1 to be labeled as a pornographic image. However, when the VLLM is trained using the positive sample image A1 and the negative sample image A2, since only entities B3, B4, B5, and B6 exist in the negative sample image B1, but not entity B2, the VLLM can only learn that entities B3, B4, B5, and B6 are not key entities that can be used to infer that the target image is a pornographic image, and cannot determine which of entity B1 and entity B2 is the key entity that can be used to infer that the target image is a pornographic image. Therefore, it is possible that both entity B1 and entity B2 are regarded as key entities that can be used to infer that the target image is a pornographic image. For example, there is a target image A3, which includes entity B2, entity B7, entity B8, and entity B9. Then, when the trained VLLM is used to analyze the target image A3, based on the existence of entity B2, the trained VLLM is likely to obtain an image analysis result that characterizes the target image as a pornographic image, but this is obviously inaccurate.

[0038] In summary, at present, in the process of training VLLM, positive and negative sample pairs (including positive sample images and negative sample images) cannot provide clear inferential causal basis for VLLM, which will cause the trained VLLM to easily regard non-critical entities as inferential causal basis when being used to perform causal reasoning tasks for target images, thereby reducing the accuracy of image analysis results.

[0039] In view of the above problems, the present disclosure provides a VLLM training method, which can be applied to electronic devices. The electronic device can be a server or a terminal device. Here, the terminal device can be a workbench, a mainframe computer, a conventional computer (e.g., a desktop computer, a laptop computer, a car computer, etc.), a personal digital processing unit or other similar computing devices. Figure 1 The flowchart diagram shown in the figure illustrates a VLLM training method provided by an embodiment of the present disclosure. It should be noted that although the logical order is shown in the flowchart diagram, in some cases, the steps shown or described in the flowchart can also be performed in other orders.

[0040] Step S101 : determining, from a positive sample image, a key entity that has a causal relationship with an image subject of the positive sample image.

[0041] Among them, the positive sample images can be obtained from the Internet. For example, when the trained VLLM needs to be used to perform network risk assessment tasks in the field of risk control and security, the positive sample images can be network risk images, specifically pornographic images, illegal images, abnormal behavior images, contraband images, etc.; for another example, when the trained VLLM needs to be used to perform environmental monitoring tasks in the field of risk control and security, the positive sample images can be problem environment images, specifically natural disaster images, environmental pollution images, etc.; for another example, when the trained VLLM needs to be used to perform festival recognition tasks in the field of visual content analysis, the positive sample images can be specific festival images, specifically Spring Festival images, Lantern Festival images, Dragon Boat Festival images, Mid-Autumn Festival images, Christmas images, etc.

[0042] In addition, in the disclosed embodiments, the image theme of the positive sample image can be used to describe the main content and core idea of ​​the positive sample image to summarize the most important representational meaning in the positive sample image, and the key entity can be an entity that can infer that the positive sample image has the image theme. For example, when the positive sample image is a Lantern Festival image, its image theme can be "Happy Reunion Celebrating Lantern Festival", and the key entity that has a causal relationship with the image theme can be Lantern Festival; for another example, when the positive sample image is a Mid-Autumn Festival image, its image theme can be "Mid-Autumn Reunion Night", and the key entity that has a causal relationship with the image theme can be moon cake; for another example, when the positive sample image is a Christmas image, its image theme can be "Happy Celebration of Christmas", and the key entity that has a causal relationship with the image theme can be Christmas tree.

[0043] Step S102: Use the target entity to replace the key entity in the positive sample image to obtain a negative sample image.

[0044] The target entity may be another entity that has no causal relationship with the image theme of the positive sample image. For example, if the positive sample image is a Lantern Festival image, and its image theme is "Happy Reunion Celebrating Lantern Festival", the target entity may be dumplings that have no causal relationship with the image theme; for another example, if the positive sample image is a Mid-Autumn Festival image, and its image theme is "Mid-Autumn Reunion Night", the target entity may be rice dumplings that have no causal relationship with the image theme; for another example, if the positive sample image is a Christmas image, and its image theme is "Happy Celebration of Christmas", the target entity may be a palm tree that has no causal relationship with the image theme.

[0045] Step S103: train the VLLM using the positive sample images and the negative sample images to obtain a trained VLLM.

[0046] The VLLM training method provided by the embodiment of the present disclosure can determine the key entity that has a causal relationship with the image subject of the positive sample image from the positive sample image, and use the target entity that has no causal relationship with the image subject of the positive sample image to replace the key entity in the positive sample image to obtain a negative sample image, and then use the positive sample image and the negative sample image to train the VLLM to obtain a trained VLLM. That is to say, in the embodiment of the present disclosure, when training the VLLM, the positive and negative sample pairs (including the positive sample image and the negative sample image) can provide the VLLM with clear inferential causal basis, that is, the VLLM can clearly learn that only the key entity is the inferential causal basis, so that when the trained VLLM is used to perform causal reasoning tasks for the target image, it will no longer regard non-key entities as inferential causal basis, thereby improving the accuracy of the image analysis results.

[0047] In addition, it should be noted that in the embodiment of the present disclosure, step S101, step S102 and step S103 can be executed in a loop, that is, in the embodiment of the present disclosure, multiple positive and negative sample pairs can be used to train the VLLM to ensure that the trained VLLM has excellent image analysis capabilities.

[0048] In some optional implementations, step S101, i.e., “determining from the positive sample image a key entity that has a causal relationship with the image subject of the positive sample image” may include:

[0049] Determining a plurality of first candidate entities from the positive sample images;

[0050] Obtaining basic information of each first candidate entity among multiple first candidate entities;

[0051] Based on basic information of each first candidate entity among the plurality of first candidate entities, a key entity having a causal relationship with an image subject of the positive sample image is determined from the plurality of first candidate entities.

[0052] The first candidate entity may be an object in a positive sample image that has clear shape, color and texture features and can be identified and located by a computer vision algorithm; the basic information of the first candidate entity may include the entity name, entity description, entity attributes, etc. Here, the entity attributes may include size, position, color, etc.

[0053] In one example, after determining multiple first candidate entities from a positive sample image and obtaining basic information of each of the multiple first candidate entities, each of the multiple first candidate entities can be used as a first entity to be evaluated, and the matching degree between the basic information of the first entity to be evaluated and the image theme of the positive sample image can be obtained as a first evaluation reference value of the first entity to be evaluated. For example, a large language model (LLM) can be used to obtain the matching degree between the basic information of the first entity to be evaluated and the image theme of the positive sample image as the first evaluation reference value of the first entity to be evaluated. The matching degree can be used to characterize the correlation between the basic information of the first entity to be evaluated and the image theme of the positive sample image. Specifically, the higher the matching degree between the basic information of the first entity to be evaluated and the image theme of the positive sample image, the more relevant the basic information of the first entity to be evaluated is to the image theme of the positive sample image.

[0054] After obtaining the first evaluation reference value of each of the multiple first candidate entities, the first candidate entity with the largest first evaluation reference value can be selected from the multiple first candidate entities as the key entity that has a causal relationship with the image subject of the positive sample image.

[0055] In the above manner, in the embodiment of the present disclosure, the key entity that has a causal relationship with the image theme of the positive sample image can be determined from the multiple first candidate entities based on the basic information of each first candidate entity in the multiple first candidate entities. In this way, it is possible to deeply analyze whether there is a causal relationship between each first candidate entity in the multiple first candidate entities and the image theme of the positive sample image based on the basic information of each first candidate entity in the multiple first candidate entities, so as to ensure that the key entity that has a causal relationship with the image theme of the positive sample image can be accurately determined from the multiple first candidate entities. Moreover, in this process, each first candidate entity in the multiple first candidate entities is used as the first entity to be evaluated, and the matching degree between the basic information of the first entity to be evaluated and the image theme of the positive sample image is obtained as the first evaluation reference value of the first entity to be evaluated, and then the first candidate entity with the largest first evaluation reference value is selected from the multiple first candidate entities as the key entity that has a causal relationship with the image theme of the positive sample image. Since the acquisition of the matching degree does not involve a complex data processing process, it is possible to improve the efficiency of determining the key entity, thereby improving the training efficiency of the VLLM.

[0056] In some optional implementations, step S102, that is, “using the target entity to replace the key entity in the positive sample image to obtain the negative sample image” may include:

[0057] Masking the key entities in the positive sample image to obtain an intermediate sample image including the masked area;

[0058] The target entity is fused into the masked area in the intermediate sample image to obtain a negative sample image.

[0059] In one example, the visual application model can be used to mask the key entities in the positive sample image to obtain an intermediate sample image including the masked area. The visual application model can include a zero-sample detector (GroundingDINO) and a zero-sample segmentation model (Segment Anything Model, SAM), that is, the visual application model can be a zero-sample visual application model, also known as a Grounded-SAM model.

[0060] In another example, the target entity can be fused into the mask area in the intermediate sample image through a diffusion model (Stable Diffusion) to obtain a negative sample image.

[0061] Through the above method, in the embodiment of the present disclosure, the key entity in the positive sample image can be masked to obtain an intermediate sample image including the masked area, and the target entity can be fused into the masked area in the intermediate sample image to obtain a negative sample image. This process can realize the automated production of negative sample images, eliminating the manual production process, and at the same time, does not involve complex data processing processes, so it can improve the efficiency of obtaining negative sample images, thereby further improving the training efficiency of VLLM.

[0062] Furthermore, in the embodiment of the present disclosure, the target entity can be obtained by:

[0063] Determine the entity data source;

[0064] Obtaining basic information of each second candidate entity among the plurality of second candidate entities;

[0065] Based on the basic information of each second candidate entity in the plurality of second candidate entities, a target entity is selected from the plurality of second candidate entities.

[0066] Among them, multiple second candidate entities can be provided by an entity data source. Here, the entity data source can be a pre-created entity database; the second candidate entity can be an object provided by the entity data source that has clear shape, color and texture features and can be recognized and located by a computer vision algorithm; the basic information of the second candidate entity can include the entity name, entity description, entity attributes, etc. of the second candidate entity. Here, the entity attributes can include size, position, color, etc.

[0067] In one example, after determining an entity data source for providing a plurality of second candidate entities and obtaining basic information of each of the plurality of second candidate entities, each of the plurality of second candidate entities can be used as a second entity to be evaluated, and the similarity between the basic information of the second entity to be evaluated and the basic information of the key entity can be obtained as a second evaluation reference value of the second entity to be evaluated. For example, the cosine similarity between the basic information of the second entity to be evaluated and the basic information of the key entity can be obtained as the second evaluation reference value of the second entity to be evaluated.

[0068] After obtaining the second evaluation reference value of each second candidate entity in the plurality of second candidate entities, the second candidate entity with the largest second evaluation reference value and no causal relationship with the image subject of the positive sample image can be selected from the plurality of second candidate entities as the target entity. For example, LLM can be used to select the second candidate entity with the largest second evaluation reference value and no causal relationship with the image subject of the positive sample image from the plurality of second candidate entities, and after manual verification, the second candidate entity can be used as the target entity.

[0069] Through the above manner, in the embodiment of the present disclosure, the target entity can be selected from the multiple second candidate entities based on the basic information of each second candidate entity in the multiple second candidate entities, thereby ensuring that the target entity can be accurately selected from the multiple second candidate entities. Moreover, in this process, each second candidate entity in the multiple second candidate entities is specifically used as the second entity to be evaluated, and the similarity between the basic information of the second entity to be evaluated and the basic information of the key entity is obtained as the second evaluation reference value of the second entity to be evaluated, and then the second candidate entity with the largest second evaluation reference value and no causal relationship with the image subject of the positive sample image is selected from the multiple second candidate entities as the target entity. That is to say, in the embodiment of the present disclosure, the target entity has a high similarity with the key entity. In this way, the ability to determine the inferential causal basis of the VLLM can be further improved, so that when the trained VLLM is used to perform the causal reasoning task for the target image, it will no longer regard the non-key entity as the inferential causal basis, thereby further improving the accuracy of the image analysis results.

[0070] In some optional implementations, step S103, that is, “training the VLLM using the positive sample images and the negative sample images to obtain a trained VLLM” may include:

[0071] The positive sample image and the negative sample image are respectively input into the VLLM as the target sample images, so as to obtain the inference image analysis result for the target sample image by using the VLLM;

[0072] Based on the inferential image analysis result, the VLLM is trained to obtain a trained VLLM.

[0073] For the inferential image analysis results, when the trained VLLM needs to be used to perform network risk assessment tasks in the field of risk control and security, it can be used to characterize whether the target sample image is a network risk image; when the trained VLLM needs to be used to perform environmental monitoring tasks in the field of risk control and security, it can be used to characterize whether the target sample image is a problem environment image; when the trained VLLM needs to be used to perform festival recognition tasks in the field of visual content analysis, it can be used to characterize whether the target sample image is a specific festival image.

[0074] After obtaining the inferential image analysis results for the target sample image, the VLLM can be trained based on the inferential image analysis results to obtain a trained VLLM. For example, based on the first loss value between the inferential image analysis results and the original image label of the target sample image, the parameters of the VLLM can be adjusted to obtain a trained VLLM. Among them, when the target sample image is a positive sample image and the trained VLLM needs to be used to perform network risk assessment tasks in the fields of risk control and security, its original image label is used to characterize the target sample image as a network risk image; when the target sample image is a negative sample image and the trained VLLM needs to be used to perform network risk assessment tasks in the fields of risk control and security, its original image label is used to characterize that the target sample image is not a network risk image; when the target sample image is a positive sample image and the trained VLLM needs to be used to perform environmental monitoring tasks in the fields of risk control and security, its original image label is used to characterize that the target sample image is a problem environment image. Image; when the target sample image is a negative sample image and the trained VLLM needs to be used to perform environmental monitoring tasks in the field of risk control and security, its original image label is used to characterize that the target sample image is not a problem environment image; when the target sample image is a positive sample image and the trained VLLM needs to be used to perform a festival recognition task in the field of visual content analysis, its original image label is used to characterize that the target sample image is a specific festival image; when the target sample image is a negative sample image and the trained VLLM needs to be used to perform a festival recognition task in the field of visual content analysis, its original image label is used to characterize that the target sample image is not a specific festival image.

[0075] Through the above method, in the embodiment of the present disclosure, the positive sample image and the negative sample image can be respectively input into the VLLM as the target sample images, so as to use the VLLM to obtain the inferential image analysis results for the target sample images, and based on the inferential image analysis results, the VLLM is trained to obtain a trained VLLM, thereby further improving the training efficiency of the VLLM.

[0076] In some other optional implementations, step S103, that is, "training the VLLM using the positive sample images and the negative sample images to obtain a trained VLLM" may include:

[0077] The positive sample image and the negative sample image are respectively used as the target sample image, and the reference description information for the target sample image is obtained by using the generative model;

[0078] Inputting the reference description information and the target sample image into the VLLM, so that the VLLM takes the reference description information as a learning object, obtains the inference description information for the target sample image, and obtains the inference image analysis result for the target sample image based on the inference description information;

[0079] Based on the inferential image analysis result, the VLLM is trained to obtain a trained VLLM.

[0080] Among them, the generative model can be a multimodal large language model (MLLM) with excellent performance, such as the fourth-generation generative pre-trained transformer (GPT-4), the third-generation large language model (LLaMA-3), etc.

[0081] After using the generative model to obtain reference description information for the target sample image, the reference description information and the target sample image can be input into the VLLM so that the VLLM takes the reference description information as a learning object to obtain inferential description information for the target sample image, and based on the inferential description information, obtains inferential image analysis results for the target sample image, and then based on the inferential image analysis results, the VLLM is trained to obtain a trained VLLM.

[0082] The reference description information is used to describe the specific content displayed in the target sample image, including multiple entities and basic information of each of the multiple entities, such as entity name, entity description, entity attributes, etc. The inference description information is also used to describe the specific content displayed in the target sample image, including multiple entities and basic information of each of the multiple entities, such as entity name, entity description, entity attributes, etc. Here, entity attributes may include size, position, color, etc.

[0083] In addition, as mentioned above, for the inferential image analysis results, when the trained VLLM needs to be used to perform network risk assessment tasks in the field of risk control and security, it can be used to characterize whether the target sample image is a network risk image; when the trained VLLM needs to be used to perform environmental monitoring tasks in the field of risk control and security, it can be used to characterize whether the target sample image is a problem environment image; when the trained VLLM needs to be used to perform festival recognition tasks in the field of visual content analysis, it can be used to characterize whether the target sample image is a specific festival image.

[0084] After obtaining the inferential image analysis results for the target sample image, the VLLM can be trained based on the inferential image analysis results to obtain a trained VLLM. For example, based on the first loss value between the inferential image analysis results and the original image label of the target sample image, the parameters of the VLLM can be adjusted to obtain a trained VLLM. Among them, when the target sample image is a positive sample image and the trained VLLM needs to be used to perform network risk assessment tasks in the fields of risk control and security, its original image label is used to characterize the target sample image as a network risk image; when the target sample image is a negative sample image and the trained VLLM needs to be used to perform network risk assessment tasks in the fields of risk control and security, its original image label is used to characterize that the target sample image is not a network risk image; when the target sample image is a positive sample image and the trained VLLM needs to be used to perform environmental monitoring tasks in the fields of risk control and security, its original image label is used to characterize that the target sample image is a problem environment image. Image; when the target sample image is a negative sample image and the trained VLLM needs to be used to perform environmental monitoring tasks in the field of risk control and security, its original image label is used to characterize that the target sample image is not a problem environment image; when the target sample image is a positive sample image and the trained VLLM needs to be used to perform a festival recognition task in the field of visual content analysis, its original image label is used to characterize that the target sample image is a specific festival image; when the target sample image is a negative sample image and the trained VLLM needs to be used to perform a festival recognition task in the field of visual content analysis, its original image label is used to characterize that the target sample image is not a specific festival image.

[0085] Through the above method, in the embodiment of the present disclosure, the positive sample image and the negative sample image can be used as the target sample image respectively, and the reference description information for the target sample image can be obtained by using the generative model, and the reference description information and the target sample image can be input into the VLLM, so that the VLLM uses the reference description information as a learning object to obtain the inferential description information for the target sample image, and based on the inferential description information, the inferential image analysis result for the target sample image is obtained, and then based on the inferential image analysis result, the VLLM is trained to obtain a trained VLLM. In this way, when the trained VLLM is used to perform causal reasoning tasks for the target image, it can better capture the description information of the target image, thereby further improving the accuracy of the image analysis result.

[0086] A VLLM training method provided by an embodiment of the present disclosure is described below through a specific example. In this example, a VLLM used to perform a festival recognition task in the field of visual content analysis needs to be trained.

[0087] Please combine Figure 2 First, a positive sample image C1, specifically a Christmas image, can be obtained from the Internet. Thereafter, the key entity, the Christmas tree, which has a causal relationship with the image theme of the positive sample image C1, "Christmas Joy Celebration", can be determined from the positive sample image C1, and the key entity in the positive sample image C1 is replaced with the target entity (palm tree) which has no causal relationship with the image theme of the positive sample image C1, "Christmas Joy Celebration", to obtain the negative sample image C2. Among them, the positive sample image C1 retains the causal relationship of "Christmas tree -> specific holiday image (specifically Christmas image)"; the negative sample image C2 destroys the causal relationship of "Christmas tree -> specific holiday image (specifically Christmas image)".

[0088] After the positive sample image C1 and the negative sample image C2 are obtained, the VLLM is trained using the positive sample image C1 and the negative sample image C2 to obtain a trained VLLM.

[0089] The present disclosure also provides an image analysis method, which can be applied to an electronic device. The electronic device can be a server or a terminal device. The terminal device can be a workbench, a mainframe computer, a conventional computer (e.g., a desktop computer, a laptop computer, a car computer, etc.), a personal digital processing unit or other similar computing device. Figure 3 The flowchart diagram shown in the figure illustrates an image analysis method provided by an embodiment of the present disclosure. It should be noted that although the logical order is shown in the flowchart diagram, in some cases, the steps shown or described in the flowchart can also be performed in other orders.

[0090] Step S301, acquiring a target image.

[0091] The target image can be any image that needs to be analyzed.

[0092] Step S302: input the target image into the trained VLLM to obtain an image analysis result for the target image using the trained VLLM.

[0093] The trained VLLM may be obtained by training using the aforementioned VLLM training method.

[0094] In one example, after inputting the target image into the trained VLLM, the trained VLLM can obtain description information for the target image, and based on the description information, obtain the image analysis result for the target image. The description information is used to describe the specific content displayed in the target image, including multiple entities, and basic information of each entity in the multiple entities, such as entity name, entity description, entity attributes, etc. Here, the entity attributes may include size, position, color, etc.

[0095] For the image analysis results, when the trained VLLM needs to be used to perform network risk assessment tasks in the field of risk control and security, it can be used to characterize whether the target image is a network risk image; when the trained VLLM needs to be used to perform environmental monitoring tasks in the field of risk control and security, it can be used to characterize whether the target image is a problem environment image; when the trained VLLM needs to be used to perform festival recognition tasks in the field of visual content analysis, it can be used to characterize whether the target image is a specific festival image.

[0096] The image analysis method provided by the embodiment of the present disclosure can be used to obtain a target image, and the target image can be input into a trained VLLM, so as to obtain an image analysis result for the target image using the trained VLLM. Since the trained VLLM is obtained after training using the aforementioned VLLM training method, when the trained VLLM is used to perform a causal reasoning task for the target image, non-critical entities will no longer be considered as inferential causal evidence, thereby improving the accuracy of the image analysis result.

[0097] See also Figure 4 , is a schematic diagram of an application scenario of a VLLM training method provided in an embodiment of the present disclosure.

[0098] The VLLM training method provided in the embodiments of the present disclosure is applied to an electronic device. The electronic device may be a server or a terminal device. The terminal device may be a workbench, a mainframe computer, a conventional computer (e.g., a desktop computer, a laptop computer, a vehicle-mounted computer, etc.), a personal digital processing unit or other similar computing devices.

[0099] Here, electronic devices are used to:

[0100] Determining from the positive image a key entity that has a causal relationship with the image subject of the positive image;

[0101] Using the target entity, the key entity in the positive sample image is replaced to obtain a negative sample image; wherein there is no causal relationship between the target entity and the image subject of the positive sample image;

[0102] The VLLM is trained using the positive sample images and the negative sample images to obtain a trained VLLM.

[0103] It should be noted that in the embodiments of the present disclosure, Figure 4 The application scenario diagram shown is only illustrative and not restrictive. Those skilled in the art can Figure 4 Various obvious changes and / or substitutions are made to the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0104] See also Figure 5 , which is a schematic diagram of an application scenario of an image analysis method provided in an embodiment of the present disclosure.

[0105] The image analysis method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device may be a server or a terminal device. The terminal device may be a workbench, a mainframe computer, a conventional computer (e.g., a desktop computer, a laptop computer, a vehicle-mounted computer, etc.), a personal digital processing device or other similar computing device.

[0106] Here, electronic devices are used to:

[0107] Get the target image;

[0108] The target image is input into the trained VLLM to obtain an image analysis result for the target image using the trained VLLM.

[0109] It should be noted that in the embodiments of the present disclosure, Figure 5 The application scenario diagram shown is only illustrative and not restrictive. Those skilled in the art can Figure 5 Various obvious changes and / or substitutions are made to the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0110] In order to better implement the aforementioned VLLM training method, the disclosed embodiment also provides a VLLM training device, which can be integrated into an electronic device. The electronic device can be a server or a terminal device. Here, the terminal device can be a workbench, a mainframe computer, a conventional computer (e.g., a desktop computer, a laptop computer, a car computer, etc.), a personal digital processing unit or other similar computing devices. Figure 6 The schematic structural block diagram shown illustrates a VLLM training device 600 provided in the disclosed embodiment.

[0111] The VLLM training device 600 comprises:

[0112] A key entity determination unit 601 is used to determine a key entity having a causal relationship with an image subject of the positive sample image from the positive sample image;

[0113] An entity replacement unit 602 is used to replace a key entity in a positive sample image with a target entity to obtain a negative sample image; wherein there is no causal relationship between the target entity and the image subject of the positive sample image;

[0114] The model training unit 603 is used to train the VLLM using the positive sample images and the negative sample images to obtain a trained VLLM.

[0115] In some optional implementations, the key entity determination unit 601 is used to:

[0116] Determining a plurality of first candidate entities from the positive sample images;

[0117] Obtaining basic information of each first candidate entity among multiple first candidate entities;

[0118] Based on basic information of each first candidate entity among the plurality of first candidate entities, a key entity having a causal relationship with an image subject of the positive sample image is determined from the plurality of first candidate entities.

[0119] In some optional implementations, the key entity determination unit 601 is used to:

[0120] Obtaining a degree of match between basic information of a first entity to be evaluated and an image subject of a positive sample image as a first evaluation reference value of the first entity to be evaluated; wherein the first entity to be evaluated is each first candidate entity among a plurality of first candidate entities;

[0121] A first candidate entity with a maximum first evaluation reference value is selected from the multiple first candidate entities as a key entity having a causal relationship with the image subject of the positive sample image.

[0122] In some optional implementations, the VLLM training device 600 further includes a target entity selection unit, which is used to:

[0123] Determine an entity data source; wherein the entity data source is used to provide a plurality of second candidate entities;

[0124] Obtaining basic information of each second candidate entity among the plurality of second candidate entities;

[0125] Based on the basic information of each second candidate entity in the plurality of second candidate entities, a target entity is selected from the plurality of second candidate entities.

[0126] In some optional implementations, the target entity selection unit is used to:

[0127] Obtaining the similarity between the basic information of the second entity to be evaluated and the basic information of the key entity as a second evaluation reference value of the second entity to be evaluated; wherein the second entity to be evaluated is each second candidate entity in the plurality of second candidate entities;

[0128] A second candidate entity having the largest second evaluation reference value and having no causal relationship with the image subject of the positive sample image is selected from the multiple second candidate entities as the target entity.

[0129] In some optional implementations, the entity replacement unit 602 is configured to:

[0130] Masking the key entities in the positive sample image to obtain an intermediate sample image including the masked area;

[0131] The target entity is fused into the masked area in the intermediate sample image to obtain a negative sample image.

[0132] In some optional implementations, the model training unit 603 is used to:

[0133] The positive sample image and the negative sample image are respectively used as the target sample image, and the reference description information for the target sample image is obtained by using the generative model;

[0134] Inputting the reference description information and the target sample image into the VLLM, so that the VLLM takes the reference description information as a learning object, obtains the inference description information for the target sample image, and obtains the inference image analysis result for the target sample image based on the inference description information;

[0135] Based on the inferential image analysis result, the VLLM is trained to obtain a trained VLLM.

[0136] In the embodiment of the present disclosure, the specific functions and examples of each unit in the VLLM training device 600 can refer to the relevant descriptions of the corresponding steps in the aforementioned VLLM training method embodiment, which will not be repeated here.

[0137] In order to better implement the aforementioned image analysis method, the embodiment of the present disclosure further provides an image analysis device, which can be integrated into an electronic device. The electronic device can be a server or a terminal device. Here, the terminal device can be a workbench, a mainframe computer, a conventional computer (e.g., a desktop computer, a laptop computer, a car computer, etc.), a personal digital processing unit or other similar computing devices. Figure 7 The schematic structural block diagram shown is used to illustrate an image analysis device 700 provided in the disclosed embodiment.

[0138] The image analysis device 700 comprises:

[0139] An image acquisition unit 701 is used to acquire a target image;

[0140] The image analysis unit 702 is used to input the target image into the trained VLLM to obtain an image analysis result for the target image using the trained VLLM.

[0141] In the embodiment of the present disclosure, the specific functions and examples of each unit in the image analysis device 700 can refer to the relevant descriptions of the corresponding steps in the embodiment of the image analysis method, which will not be repeated here.

[0142] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0143] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0144] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device 800 is intended to represent various forms of digital computers, such as vehicle-mounted computing devices, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 800 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0145] like Figure 8As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 to a random access memory (RAM) 803. In RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0146] A number of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of renderers, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0147] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the training method and / or image analysis method of the VLLM. For example, in some embodiments, the training method and / or image analysis method of the VLLM may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps in the VLLM training method and / or image analysis method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured as the VLLM training method and / or image analysis method in any other appropriate manner (e.g., by means of firmware).

[0148] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0149] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data optimization device so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed completely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine or completely on a remote machine or server.

[0150] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM) or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] To provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a rendering device (e.g., a cathode ray tube (CRT) renderer or a liquid crystal display (LCD)) for rendering information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices are also used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0152] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: Local Area Network (LAN), Wide Area Network (WAN), and the Internet.

[0153] The computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0154] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute a VLLM training method and / or an image analysis method.

[0155] The embodiments of the present disclosure also provide a computer program product, including a computer program, which implements a VLLM training method and / or an image analysis method when executed by a processor.

[0156] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and this document does not limit this. In addition, in the present disclosure, relational terms such as "first", "second", "third", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. In addition, "multiple" in the present disclosure can be understood as at least two.

[0157] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for training a large visual language model, comprising: Determining, from the positive sample image, a key entity that has a causal relationship with an image subject of the positive sample image; Using the target entity, the key entity in the positive sample image is replaced to obtain a negative sample image; wherein there is no causal relationship between the target entity and the image subject of the positive sample image; The visual large language model is trained using the positive sample images and the negative sample images to obtain a trained visual large language model.

2. The method according to claim 1, wherein: The determining, from the positive sample image, a key entity having a causal relationship with the image subject of the positive sample image comprises: Determining a plurality of first candidate entities from the positive sample images; Acquire basic information of each first candidate entity among the multiple first candidate entities; Based on the basic information of each first candidate entity among the multiple first candidate entities, a key entity having a causal relationship with the image subject of the positive sample image is determined from the multiple first candidate entities.

3. The method according to claim 2, wherein: The determining, based on the basic information of each of the plurality of first candidate entities, a key entity having a causal relationship with the image subject of the positive sample image from the plurality of first candidate entities comprises: Obtaining a degree of match between basic information of a first entity to be evaluated and an image subject of the positive sample image as a first evaluation reference value of the first entity to be evaluated; wherein the first entity to be evaluated is each first candidate entity among the multiple first candidate entities; A first candidate entity with a maximum first evaluation reference value is selected from the multiple first candidate entities as a key entity having a causal relationship with the image subject of the positive sample image.

4. The method according to claim 2, further comprising: Determine an entity data source; wherein the entity data source is used to provide a plurality of second candidate entities; Acquire basic information of each second candidate entity among the plurality of second candidate entities; The target entity is selected from the plurality of second candidate entities based on basic information of each second candidate entity in the plurality of second candidate entities.

5. The method according to claim 4, wherein: The selecting the target entity from the plurality of second candidate entities based on the basic information of each second candidate entity in the plurality of second candidate entities comprises: Acquire the similarity between the basic information of the second entity to be evaluated and the basic information of the key entity as a second evaluation reference value of the second entity to be evaluated; wherein the second entity to be evaluated is each second candidate entity in the plurality of second candidate entities; A second candidate entity having the largest second evaluation reference value and having no causal relationship with the image subject of the positive sample image is selected from the multiple second candidate entities as the target entity.

6. The method according to any one of claims 1 to 5, wherein: The step of replacing the key entity in the positive sample image with the target entity to obtain a negative sample image includes: Masking the key entity in the positive sample image to obtain an intermediate sample image including the masked area; The target entity is fused into the mask region in the intermediate sample image to obtain the negative sample image.

7. The method according to any one of claims 1 to 5, wherein: The using the positive sample image and the negative sample image to train the visual large language model to obtain a trained visual large language model includes: The positive sample image and the negative sample image are respectively used as target sample images, and reference description information for the target sample images is obtained by using a generation model; Inputting the reference description information and the target sample image into the visual large language model, so that the visual large language model takes the reference description information as a learning object, obtains inference description information for the target sample image, and obtains inference image analysis results for the target sample image based on the inference description information; Based on the inferential image analysis result, the visual large language model is trained to obtain a trained visual large language model.

8. An image analysis method, comprising: Get the target image; The target image is input into a trained visual large language model to obtain an image analysis result for the target image using the trained visual large language model.

9. A training device for a visual large language model, comprising: A key entity determination unit, used to determine a key entity having a causal relationship with an image subject of the positive sample image from the positive sample image; An entity replacement unit, used to replace the key entity in the positive sample image with a target entity to obtain a negative sample image; wherein there is no causal relationship between the target entity and the image subject of the positive sample image; The model training unit is used to train the visual large language model using the positive sample images and the negative sample images to obtain a trained visual large language model.

10. The device according to claim 9, wherein: The key entity determination unit is used for: Determining a plurality of first candidate entities from the positive sample images; Acquire basic information of each first candidate entity among the multiple first candidate entities; Based on the basic information of each first candidate entity among the multiple first candidate entities, a key entity having a causal relationship with the image subject of the positive sample image is determined from the multiple first candidate entities.

11. The device according to claim 10, wherein: The key entity determination unit is used for: Obtaining a degree of match between basic information of a first entity to be evaluated and an image subject of the positive sample image as a first evaluation reference value of the first entity to be evaluated; wherein the first entity to be evaluated is each first candidate entity among the multiple first candidate entities; A first candidate entity with a maximum first evaluation reference value is selected from the multiple first candidate entities as a key entity having a causal relationship with the image subject of the positive sample image.

12. The device according to claim 10, further comprising a target entity selection unit, configured to: Determine the entity data source; where, The entity data source is used to provide a plurality of second candidate entities; Acquire basic information of each second candidate entity among the plurality of second candidate entities; The target entity is selected from the plurality of second candidate entities based on basic information of each second candidate entity in the plurality of second candidate entities.

13. The device according to claim 12, wherein: The target entity selection unit is used for: Acquire the similarity between the basic information of the second entity to be evaluated and the basic information of the key entity as a second evaluation reference value of the second entity to be evaluated; wherein the second entity to be evaluated is each second candidate entity in the plurality of second candidate entities; A second candidate entity having the largest second evaluation reference value and having no causal relationship with the image subject of the positive sample image is selected from the multiple second candidate entities as the target entity.

14. The device according to any one of claims 9 to 13, wherein: The entity replacement unit is used for: Masking the key entity in the positive sample image to obtain an intermediate sample image including the masked area; The target entity is fused into the mask region in the intermediate sample image to obtain the negative sample image.

15. The device according to any one of claims 9 to 13, wherein: The model training unit is used to: The positive sample image and the negative sample image are respectively used as target sample images, and reference description information for the target sample images is obtained by using a generation model; Inputting the reference description information and the target sample image into the visual large language model, so that the visual large language model takes the reference description information as a learning object, obtains inference description information for the target sample image, and obtains inference image analysis results for the target sample image based on the inference description information; Based on the inferential image analysis result, the visual large language model is trained to obtain a trained visual large language model.

16. An image analysis device, comprising: An image acquisition unit, used for acquiring a target image; The image analysis unit is used to input the target image into a trained visual large language model to obtain an image analysis result for the target image using the trained visual large language model.

17. An electronic device comprising: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

19. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Training method of multimedia resource generation model and multimedia resource generation method

    CN120494016A