Image understanding method and apparatus, and device and medium
By combining the image understanding model with the first and second fine-tuning datasets, the problem of low object recognition accuracy in image understanding is solved, achieving higher recognition and understanding accuracy and fine-grained perception.
Patent Information
- Application Number
- PCT/CN2025/082018
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2025-03-12
- Publication Date
- 2025-10-30
AI Technical Summary
Existing image understanding technologies have low accuracy in object recognition and comprehension, making it difficult to effectively identify and understand objects in images.
An image understanding model is employed. By acquiring the target object image, the target test image, and the first target text, the model is trained using a first fine-tuning dataset and/or a second fine-tuning dataset to improve the accuracy of object recognition and understanding. The first fine-tuning dataset is obtained by cropping and adjusting the image of the object's location from the base dataset, while the second fine-tuning dataset is obtained by adjusting the video dataset, thereby enhancing the model's understanding ability and fine-grained perception.
It improves the accuracy and fine-grained perception of image understanding, enabling better identification and understanding of complex visual information and enhancing the accuracy of object recognition.
Smart Images

Figure CN2025082018_30102025_PF_FP_ABST
Abstract
Description
An image understanding method, apparatus, device, and medium
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410509275.1, filed on April 25, 2024, entitled "An Image Understanding Method, Apparatus, Device and Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of image processing technology, and in particular to an image understanding method, apparatus, device, and medium. Background Technology
[0004] With the continuous development of image processing technology, image understanding is being applied more and more widely. Due to the excellent performance of multimodal large models in visual perception, reasoning, and multimodal knowledge, they can be applied to image understanding. Summary of the Invention
[0005] To address the aforementioned technical problems, this disclosure provides an image understanding method, apparatus, device, and medium.
[0006] This disclosure provides an image understanding method, the method comprising:
[0007] Obtain the target object image, the target test image, and the first target text;
[0008] The target object image, the target test image, and the first target text are input into the image understanding model to identify and analyze the target object based on the target object image, thereby obtaining the second target text corresponding to the first target text.
[0009] Wherein, the first target text indicates that a target operation is performed on the target object in the target test graph, and the second target text indicates the execution result of the target operation;
[0010] The image understanding model is trained based on a first fine-tuning dataset and / or a second fine-tuning dataset. The first fine-tuning dataset is obtained by adjusting the image of the object region in the base dataset, and the second fine-tuning dataset is obtained by adjusting the video dataset.
[0011] This disclosure also provides an image understanding apparatus, the apparatus comprising:
[0012] The acquisition module is used to acquire the target object image, the target test image, and the first target text.
[0013] The understanding module is used to input the target object image, the target test image, and the first target text into the image understanding model, and to identify and analyze the target object in the target test image based on the target object image to obtain the second target text corresponding to the first target text.
[0014] Wherein, the first target text indicates that a target operation is performed on the target object in the target test graph, and the second target text indicates the execution result of the target operation;
[0015] The image understanding model is trained based on a first fine-tuning dataset and / or a second fine-tuning dataset. The first fine-tuning dataset is obtained by adjusting the image of the object region in the base dataset, and the second fine-tuning dataset is obtained by adjusting the video dataset.
[0016] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the image understanding method provided in this disclosure.
[0017] This disclosure also provides a computer-readable storage medium storing a computer program for performing the image understanding method provided in this disclosure.
[0018] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: The image understanding scheme provided in this disclosure acquires a target object image, a target test image, and a first target text; the target object image, the target test image, and the first target text are input into an image understanding model to identify and analyze the target object in the target test image based on the target object image, thereby obtaining a second target text corresponding to the first target text; wherein, the first target text represents the target operation performed on the target object in the target test image, and the second target text represents the execution result of the target operation; the image understanding model is trained based on a first fine-tuning dataset and / or a second fine-tuning dataset, the first fine-tuning dataset being obtained by adjusting the image of the area where the object is located by cropping from the basic dataset, and the second fine-tuning dataset being obtained by adjusting the video dataset. Using the above technical solution, the image understanding model identifies and analyzes the target object in the target test image based on the target object image, performs a target operation based on the first target text, and obtains the second target text corresponding to the first target text, i.e., the execution of the target operation. Since the image understanding model is trained based on the first fine-tuning dataset and / or the second fine-tuning dataset, the first fine-tuning dataset is obtained by cropping and adjusting the image of the area where the object is located. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 is a schematic flowchart of an image understanding method provided in an embodiment of this disclosure;
[0021] Figure 2 is a schematic diagram of a target object provided in an embodiment of this disclosure;
[0022] Figure 3 is a schematic diagram of a target test pattern provided in an embodiment of this disclosure;
[0023] Figure 4 is a flowchart illustrating another image understanding method provided in an embodiment of this disclosure;
[0024] Figure 5 is a schematic diagram of the construction process of a second fine-tuning data provided in an embodiment of this disclosure;
[0025] Figure 6 is a schematic diagram of the structure of an image understanding device provided in an embodiment of this disclosure;
[0026] Figure 7 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0033] Most related technologies focus on perceiving visual categories, resulting in low accuracy and poor performance in object recognition and understanding within images. To address the issue of accuracy in object recognition within images in related technologies, this disclosure provides an image understanding method, which will be described below with reference to specific embodiments.
[0034] Figure 1 is a flowchart illustrating an image understanding method according to an embodiment of this disclosure. This method can be executed by an image understanding device, which can be implemented in software and / or hardware and is generally integrated into an electronic device. As shown in Figure 1, the method includes:
[0035] Step 101: Obtain the target object image, the target test image, and the first target text.
[0036] The target object image can be a single image containing only the target object, serving as the foundational image for object recognition and understanding. The target test image can be an image that requires recognition and understanding of the same target object based on the target image. There can be one or more target test images. The specific sources of the target object image and target test images are not limited; they can be from the device's local storage or the internet. The target object can be the object targeted for subsequent recognition and understanding. An object refers to an individual with personalized characteristics, which can include people, animals, objects, and buildings, etc.
[0037] The first target text indicates that a target operation is performed on a target object in a target test image. The first target text includes the name of the target object image and text instructions for performing the target operation on the target object. The target operation includes at least one of the following: matching, localization, question answering, and description. The first target text includes at least one of the following: instructions for matching the target object in the target test image based on the target object image; instructions for localization; questions for question answering; and instructions for description. Matching can be matching an image containing a specific object from multiple other images based on an object image of that specific object. Localization can be locating an object in another image containing multiple similar interfering objects based on an object image. Question answering can identify objects in one or more images and answer questions related to those objects. Description can identify objects in one or more images and provide a description of those objects.
[0038] For example, Figure 2 is a schematic diagram of a target object diagram provided by an embodiment of the present disclosure. As shown in Figure 2, the figure shows an exemplary target object diagram, in which the target object includes a white lamb. For example, Figure 3 is a schematic diagram of a target test diagram provided by an embodiment of the present disclosure. As shown in Figure 3, the figure shows an exemplary target test diagram, which may include four objects, each of which is a white lamb, one of which is the target object of Figure 2.
[0039] Step 102: Input the target object image, the target test image, and the first target text into the image understanding model to identify and analyze the target object based on the target object image in the target test image, and obtain the second target text corresponding to the first target text; the image understanding model is trained based on the first fine-tuning dataset and / or the second fine-tuning dataset. The first fine-tuning dataset is obtained by adjusting the image of the area where the object is located by cropping from the basic dataset, and the second fine-tuning dataset is obtained by adjusting the film and television dataset.
[0040] The image understanding model can be a model used for object recognition, analysis, and understanding of images. The image understanding model may have one or more functions, which may vary depending on the training data. For example, the image understanding model may perform at least one of matching, localization, question answering, and description. This image understanding model can be trained based on a multimodal large model, which can understand and process information from multiple data modalities, such as text, images, audio, and video. The image understanding model in this embodiment may include a visual encoder, a mapping layer composed of a cross-attention mechanism, and a large oracle model; these are merely examples.
[0041] The second target text represents the execution result of the target operation. Since the target operation includes at least one of the following: matching, locating, question answering, and description, the execution result of the target operation includes at least one of the following: matching result, locating result, question answering result, and description result.
[0042] Image understanding models can be trained on instruction-based fine-tuning datasets. These datasets serve as training data for the model and can be obtained after pre-training, through further adjustments to a specific dataset to enhance the model's understanding capabilities and its accuracy and controllability during task execution. In this embodiment, the instruction-based fine-tuning data may include a first fine-tuning dataset and / or a second fine-tuning dataset. The first fine-tuning dataset can be an adaptive dataset obtained by cropping the image of the object's region from the base dataset. Cropping the object image from the original reduces the difficulty of image understanding and allows the image understanding model to be trained to model multi-image relationships. The second fine-tuning dataset can be obtained by adjusting a film and television dataset, where different characters have relationships, which is beneficial for training the image understanding model's context learning and fine-grained perception capabilities.
[0043] Specifically, the image understanding device can input the target object image, the target test image, and the first target text into the image understanding model to identify the target object based on the target object image. After identifying the target object, a target operation is performed on the target object, and the second target text corresponding to the first target text is obtained, which is the execution result of the target operation.
[0044] The image understanding scheme provided in this disclosure acquires a target object image, a target test image, and a first target text; inputs the target object image, the target test image, and the first target text into an image understanding model to identify and analyze the target object in the target test image based on the target object image, and obtains a second target text corresponding to the first target text; wherein, the first target text represents a target operation performed on the target object in the target test image, and the second target text represents the execution result of the target operation; the image understanding model is trained based on a first fine-tuning dataset and / or a second fine-tuning dataset, the first fine-tuning dataset being obtained by adjusting the image of the area where the object is located by cropping from the basic dataset, and the second fine-tuning dataset being obtained by adjusting a film and television dataset. By employing the above technical solution, an image understanding model is used to identify and analyze target objects in a target test image based on a target object image. Target operations are performed based on a first target text to obtain a second target text corresponding to the first target text, i.e., the execution of the target operation. Since the image understanding model is trained on a first fine-tuning dataset and / or a second fine-tuning dataset, the first fine-tuning dataset is obtained by cropping and adjusting the image of the region where the object is located, without processing any redundant areas of the image, thus improving the accuracy of image understanding. The second fine-tuning dataset is obtained by adjusting a film and television dataset, in which objects have relationships, improving the fine-grained perception capability of image understanding. The image understanding model is able to understand more complex visual information.
[0045] For example, Figure 4 is a schematic flowchart of another image understanding method provided in this embodiment of the disclosure. As shown in Figure 4, the image understanding model of this embodiment of the disclosure is trained in the following manner:
[0046] Step 401: Construct the first fine-tuning dataset, and / or, construct the second fine-tuning dataset.
[0047] The first fine-tuned dataset was obtained by adjusting the image of the cropped object region in the base dataset, and the second fine-tuned dataset was obtained by adjusting the film and television dataset.
[0048] In some embodiments, constructing a first fine-tuning dataset may include: obtaining a base dataset; cropping the region where an object is located in a first image in the base dataset to obtain a first object image, and using the first image as a first test image; obtaining the target name of the first object image, and substituting the target name into the first text data in the base dataset to obtain second text data; and combining the first object image, the first test image, and the second text data with a fine-tuned data format to obtain the first fine-tuning dataset.
[0049] The base dataset can be an existing dataset used by a multimodal large model. In this embodiment, the base dataset may include at least one of a question-answering dataset, a localization dataset, and a description dataset. The question-answering dataset can be a dataset including questions and answers. In this embodiment, the question-answering dataset can be a large-scale dataset of visual commonsense reasoning, which may include a large number of images and related natural language questions and reasons. These questions typically require combining image content and common sense knowledge to arrive at the correct answer. The questions are about specific objects in the image, such as a person or object, and the annotation information includes the location information of the specific object in the image. The localization dataset can be a dataset for visual localization based on descriptions, including a large number of images and corresponding descriptive text pointing to specific objects. The localization of specific objects in the image is performed based on the descriptive text, and the annotation information includes the location of the specific object. The description dataset can be a dataset of image descriptions, which may include a large number of images and corresponding text describing specific objects in the images, and the annotation information includes the location of the specific objects in the images.
[0050] The first image can be an existing image in the base dataset. The first object image can be a sub-image obtained by cropping the region where the object is located in the first image. This sub-image serves as the object image, which can be a reference image for image understanding of the object. The first object image is the object image determined for the first fine-tuning dataset. The first test image can be a test image for which object recognition needs to be performed. The target name can be a specific name obtained by custom naming the first object image. For example, the target name can be determined based on the object characteristics of the first object image, and can be an object name or a person's name. The first text data can be text data included in the base dataset. For example, when the base dataset includes a question-and-answer dataset, the first text data can include questions and answers. When the base dataset includes a description dataset, the first text data can include text describing the object. The second text data can be text data with the target name obtained by replacing the text representing a specific object in the first text data with the target name. For example, the first text data can be "What is the car doing in the image?", where the specific object is the car. The target name for the first object image of the car can be "0", and the second text data can be "What is 0 doing in the image?". This is just an example.
[0051] Specifically, when constructing the first fine-tuning dataset, the image processing device can first obtain a basic dataset, crop the area where the object is located in the first image based on the object location information in the basic dataset to obtain a first object image, and use the first image as a first test image; then, it can obtain the user's target name for the first object image, and input the target name into the first text data of the basic dataset to construct the second text data; finally, it combines the first object image, the first test image, and the second text data according to the fine-tuning data format to obtain the first fine-tuning dataset. The input data of the first fine-tuning dataset may include the first object image, the first test image, and the question in the second text data, and the output data may include the object location information, answer, description, etc. in the second text data.
[0052] Fine-tuning the data format can be set according to the needs of the model and the characteristics of the task. Fine-tuning the data format can include the encoding format of text data, sequence length, file format, dataset size, multilingual settings, etc.
[0053] For example, when the base dataset includes a question-and-answer dataset, a location dataset, and a description dataset, for the question-and-answer dataset, the object to be queried can be cropped into a sub-image based on its location information and assigned a target name. This target name is then substituted into the question-and-answer data to construct the question-and-answer data. For the location dataset, the object to be located is cropped into a sub-image based on its location information and assigned a target name. This name is then substituted into the corresponding text data. For the description dataset, the object to be identified is cropped into a sub-image based on its location information and assigned a target name. This name is then substituted into the corresponding text data describing the object. Finally, the first fine-tuning dataset is obtained. Cropping the object graph from the original image reduces the difficulty of object recognition, enabling the model to perform multi-graph relationship modeling and to identify and provide location information in the test image using the object graph for location, question-and-answer, and description. Optionally, the first fine-tuning dataset can also incrementally include training data for a lightweight multimodal large model, allowing the image understanding model to avoid forgetting its original knowledge during object recognition training.
[0054] Optionally, after obtaining the first object image, the image processing device may also perform data cleaning, data enhancement, and other operations on the first object image. Data cleaning may include, for example, filtering out images with inappropriate size ratios or filtering out images with only one similar object. Data enhancement may include flipping, adding noise, etc.
[0055] In some embodiments, constructing a second fine-tuning dataset may include: acquiring a film and television dataset; extracting images from the film and television dataset that include only preset characters as a second object image, and extracting images from the film and television dataset that include preset characters and other characters as a second test image; generating third text data based on the second test image and the character information of the multiple characters included in the second test image; and combining the second object image, the second test image, and the third text data with a fine-tuned data format to obtain the second fine-tuning dataset.
[0056] The film and television dataset can be a dataset used for film and television understanding, which may include a large amount of multimodal data, such as trailers, stills, and plot descriptions, as well as different types of annotation information, such as character names and character locations. In this embodiment, the film and television dataset may be a movie dataset. The second object graph can be an object graph where the object is determined to be a character based on the film and television dataset; it is a reference image for the image understanding object determined for the second fine-tuning dataset. The second test image can be a test image for object recognition that includes multiple characters, determined based on the film and television dataset. The third text data can be text data generated for the object recognition task that requires text processing and text generation. For example, for question-and-answer purposes, the third text data may include questions and answers; for description purposes, the third text data may include descriptive text.
[0057] Specifically, when constructing the second fine-tuning dataset, the image processing device can first acquire a film and television dataset, traverse all images in the dataset, and determine images containing only preset characters as the second object images. The preset characters can be specific objects to be identified, and the specific settings are based on the actual situation. Images including the aforementioned preset characters and at least one other character are extracted as the second test images. Based on the second test images and the character information of the multiple characters included in the second test images, third text data can be generated. For example, a generative pre-trained model can be used to generate the third text data. Finally, the second object images, the second test images, and the third text data are combined according to the fine-tuning data format to obtain the second fine-tuning dataset. This second fine-tuning data can include input data and output data. The input data can include questions from the second object images, the second test images, and the third text data, while the output data can include object location information, answers, descriptions, etc., from the third text data.
[0058] Optionally, the character information includes the character name and the character location. The third text data is generated based on the second test image and the character information of the multiple characters included in the second test image. This can include: inputting the second test image, the character names and locations of the characters included in the second test image, and preset prompt words into the text generation model to obtain text information including the character names, and determining the text information as the third text data.
[0059] The preset prompts can be set according to the specific task of the model and can be instructions for generating third-party text data. The text generation model can be a model used to generate third-party text data, and its functions can be set according to the task of the image understanding model. Specifically, when generating third-party text data, the image understanding model can input the second test image, the names and positions of the characters included in the second test image, and the preset prompts into the text generation model. After analysis by the text generation model, text information including the character names can be obtained, and this text information is identified as the third-party text data.
[0060] For example, Figure 5 is a schematic diagram of the construction process of a second fine-tuning data provided in an embodiment of this disclosure. As shown in Figure 5, the construction of the second fine-tuning data can specifically include: obtaining a second object graph and a second test graph by traversing and extracting from a film and television dataset; inputting the second test graph, the character names and positions of each character included in the second test graph, and preset prompt words into a text generation model to obtain third text data; and combining the second object graph, the second test graph, and the third text data according to the fine-tuning data format to obtain the second fine-tuning dataset.
[0061] Step 402: Train the multimodal large model based on the first fine-tuning dataset and / or the second fine-tuning dataset to obtain the image understanding model.
[0062] After constructing a first fine-tuning dataset and / or a second fine-tuning dataset, the image understanding device can train a multimodal large model using the first fine-tuning dataset and / or the second fine-tuning dataset to obtain an image understanding model. When training using the first fine-tuning dataset and the second fine-tuning dataset, the multimodal large model can be trained first using the first fine-tuning dataset to obtain a first model, and then the first model can be trained using the second fine-tuning dataset to obtain the image understanding model.
[0063] In the above scheme, an instruction fine-tuning dataset is constructed, which includes a first fine-tuning dataset and / or a second fine-tuning dataset. An image understanding model is trained using this instruction fine-tuning dataset. The instruction fine-tuning dataset is used to stimulate the ability of multiple large models to recognize objects. When faced with complex visual inputs of multiple objects, the image understanding model can remember the detailed features of specific objects, improve recognition accuracy, and further associate objects in different images. It can accurately understand the semantic information of multiple image inputs, has the ability to learn context and perceive fine-grained information, and realize the memory and recognition of objects in context, and understand more complex visual information.
[0064] In some embodiments, after step 402, the image understanding method may further include: evaluating the accuracy of the image understanding model to obtain an accuracy score, and outputting the image understanding model when the accuracy score is greater than a preset score.
[0065] The accuracy score is a quantitative evaluation of the object recognition ability and accuracy of the image understanding model. A higher accuracy score indicates a better image understanding model. The preset score can be set according to actual needs; when a higher level of image understanding is required, the preset score can be set higher.
[0066] This disclosure embodiment can construct a test dataset. First, test images meeting certain conditions can be collected. Since the recognition ability of the image understanding model needs to be tested, the test images need to contain multiple objects of the same category, and each object should have specific characteristics, such as different states or attributes. The test dataset mainly comes from sources such as movie stills obtained from film datasets, anime images with download links collected from the internet, and images of buildings, vehicles, and animals collected from object datasets. For each object in a sample, there is a corresponding individual image and one or more test images of multiple objects together. After obtaining the benchmark images, questions, instructions, and standard answers are prepared. For each subtask, a text generation model can be used to generate, such as matching instructions, instructions for generating descriptions, instructions for location, question-and-answer angles, and questions. In question-and-answer tasks, questions can be asked about the state, attributes, location, and relationships of specific objects. Then, the questions for each sample are manually labeled, and further refinement and balancing are performed to increase the diversity of question expression and grammatical structure. Finally, the correct answers are manually labeled. For example, for descriptive tasks, a correct and detailed description is given; for question-and-answer tasks, the correct answer is given, while the state labels of other objects are used to distract the answer as candidate options for multiple-choice questions.
[0067] The accuracy of image understanding models can be evaluated through four tasks. The first task is matching, where given an image of an object, such as an animal or a building, the model is required to select one image from four candidate images that contains the same object. This is the simplest task because each image contains only one object. The second task is locating objects in an image and providing their positional information (i.e., bounding box coordinates), which is a visual localization task based on the image of the object. In the test images, there will be many similar distracting objects. The third task is question answering, and the fourth task is generating descriptions. These tasks require the model to identify objects in one or more query images and answer related questions or provide an overall description. This is the most difficult level, where the model not only needs to recognize each object in the image but also needs to provide appropriate answers or descriptions based on their names, states, actions, positions, and other information.
[0068] In the above scheme, the recognition ability and accuracy of the image understanding model can be obtained by evaluating the accuracy of the image understanding model. Object recognition of the image is only performed when the recognition ability and accuracy meet the preset requirements, which ensures that the image understanding model can understand complex visual input and further improves the accuracy of image understanding.
[0069] Figure 6 is a schematic diagram of an image understanding device provided in an embodiment of this disclosure. This device can be implemented by software and / or hardware and is generally integrated into an electronic device. As shown in Figure 6, the device includes:
[0070] The acquisition module 601 is used to acquire the target object image, the target test image, and the first target text;
[0071] Understanding module 602 is used to input the target object image, the target test image, and the first target text into the image understanding model, so as to identify and analyze the target object based on the target object image to obtain the second target text corresponding to the first target text;
[0072] Wherein, the first target text indicates that a target operation is performed on the target object in the target test graph, and the second target text indicates the execution result of the target operation;
[0073] The image understanding model is trained based on a first fine-tuning dataset and / or a second fine-tuning dataset. The first fine-tuning dataset is obtained by adjusting the image of the object region in the base dataset, and the second fine-tuning dataset is obtained by adjusting the video dataset.
[0074] Optionally, the device further includes a model building module, comprising:
[0075] Data set unit, used to construct the first fine-tuning dataset, and / or, to construct the second fine-tuning dataset;
[0076] The training unit is used to train the multimodal large model based on the first fine-tuning dataset and / or the second fine-tuning dataset to obtain the image understanding model.
[0077] Optionally, the dataset unit includes a first sub-unit for:
[0078] Obtain the basic dataset;
[0079] The region where the object is located in the first image in the basic dataset is cropped to obtain the first object image, and the first image is used as the first test image.
[0080] Obtain the target name of the first object graph, and substitute the target name into the first text data of the basic dataset to obtain the second text data;
[0081] The first fine-tuned dataset is obtained by combining the first object graph, the first test graph, and the second text data with a fine-tuned data format.
[0082] Optionally, the base dataset includes at least one of a question-and-answer dataset, a location dataset, and a description dataset.
[0083] Optionally, the dataset unit includes a second subunit for:
[0084] Obtain film and television datasets;
[0085] Images containing only the preset characters are extracted from the film and television dataset as the second object image, and images containing the preset characters and other characters are extracted from the film and television dataset as the second test image;
[0086] The third text data is generated based on the second test map and the character information of the multiple characters included in the second test map;
[0087] The second fine-tuned dataset is obtained by combining the second object graph, the second test graph, and the third text data with fine-tuned data formats.
[0088] Optionally, the role information includes the role name and role location, and the second sub-unit is specifically used for:
[0089] The second test image, the character names and positions of each character included in the second test image, and preset prompts are input into the text generation model to obtain text information including the character names, and this text information is determined as the third text data.
[0090] Optionally, the target operation includes at least one of the following: matching, locating, question answering, and description, and the execution result of the target operation includes at least one of the following: matching result, locating result, question answering result, and description result.
[0091] The image understanding apparatus provided in this disclosure can execute the image understanding method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0092] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the image understanding method provided in any embodiment of this disclosure.
[0093] Figure 7 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0094] Referring specifically to Figure 7, which illustrates a structural schematic suitable for implementing the electronic device 700 in the embodiments of this disclosure, the electronic device 700 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 7 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this disclosure.
[0095] As shown in Figure 7, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0096] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 shows electronic device 700 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0097] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the image understanding method of embodiments of this disclosure.
[0098] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0099] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0100] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0101] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire a target object image, a target test image, and a first target text; input the target object image, the target test image, and the first target text into an image understanding model to identify and analyze the target object based on the target object image in the target test image, thereby obtaining a second target text corresponding to the first target text; wherein the first target text represents a target operation performed on the target object in the target test image, and the second target text represents the execution result of the target operation; the image understanding model is trained based on a first fine-tuning dataset and / or a second fine-tuning dataset, wherein the first fine-tuning dataset is obtained by adjusting the image of the area where the object is located by cropping from the base dataset, and the second fine-tuning dataset is obtained by adjusting a film and television dataset.
[0102] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0104] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0105] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0106] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0107] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0108] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0109] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0110] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image understanding method, comprising: Obtain the target object image, the target test image, and the first target text; The target object image, the target test image, and the first target text are input into the image understanding model to identify and analyze the target object based on the target object image, thereby obtaining the second target text corresponding to the first target text. Wherein, the first target text indicates that a target operation is performed on the target object in the target test graph, and the second target text indicates the execution result of the target operation; The image understanding model is trained based on a first fine-tuning dataset and / or a second fine-tuning dataset. The first fine-tuning dataset is obtained by adjusting the image of the object region in the base dataset, and the second fine-tuning dataset is obtained by adjusting the video dataset.
2. The method according to claim 1, wherein the image understanding model is trained in the following manner: Construct the first fine-tuning dataset, and / or construct the second fine-tuning dataset; The image understanding model is obtained by training a multimodal large model based on the first fine-tuning dataset and / or the second fine-tuning dataset.
3. The method according to claim 2, wherein constructing the first fine-tuning dataset comprises: Obtain the basic dataset; The region where the object is located in the first image in the basic dataset is cropped to obtain the first object image, and the first image is used as the first test image. Obtain the target name of the first object graph, and substitute the target name into the first text data of the basic dataset to obtain the second text data; The first fine-tuned dataset is obtained by combining the first object graph, the first test graph, and the second text data with a fine-tuned data format.
4. The method according to claim 3, wherein the basic dataset includes at least one of a question-answer dataset, a location dataset, and a description dataset.
5. The method of claim 2, wherein constructing the second fine-tuning dataset comprises: Obtain film and television datasets; Images containing only the preset characters are extracted from the film and television dataset as the second object image, and images containing the preset characters and other characters are extracted from the film and television dataset as the second test image; The third text data is generated based on the second test map and the character information of the multiple characters included in the second test map; The second fine-tuned dataset is obtained by combining the second object graph, the second test graph, and the third text data with fine-tuned data formats.
6. The method according to claim 5, wherein the character information includes character name and character location, and generating third text data based on the second test map and the character information of multiple characters included in the second test map, includes: The second test image, the character names and positions of each character included in the second test image, and preset prompts are input into the text generation model to obtain text information including the character names, and this text information is determined as the third text data.
7. The method according to any one of claims 1-6, wherein the target operation includes at least one of the following: matching, locating, question answering, and description, and the execution result of the target operation includes at least one of the following: matching result, locating result, question answering result, and description result.
8. An image understanding device, comprising: The acquisition module is used to acquire the target object image, the target test image, and the first target text. The understanding module is used to input the target object image, the target test image, and the first target text into the image understanding model, and to identify and analyze the target object in the target test image based on the target object image to obtain the second target text corresponding to the first target text. Wherein, the first target text indicates that a target operation is performed on the target object in the target test graph, and the second target text indicates the execution result of the target operation; The image understanding model is trained based on a first fine-tuning dataset and / or a second fine-tuning dataset. The first fine-tuning dataset is obtained by adjusting the image of the object region in the base dataset, and the second fine-tuning dataset is obtained by adjusting the video dataset.
9. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the image understanding method according to any one of claims 1-7.
10. A computer-readable storage medium storing a computer program for performing the image understanding method according to any one of claims 1-7.
Citation Information
Patent Citations
Power transformation equipment image defect identification method and system
CN116596895A
Video detection method and system, storage medium and electronic equipment
CN116778426A
Image processing method and device, electronic equipment and computer readable storage medium
CN116958320A
Text visual question and answer method and device, computer equipment and storage medium
CN117033609A