Image processing method and apparatus
By generating initial structured description information using a large language model and combining it with object location, the problem of multimodal models understanding complex image descriptions is solved, achieving efficient structured processing of image information and simplifying information extraction for downstream tasks.
Patent Information
- Application Number
- PCT/CN2025/100832
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-05
- Filing Date
- 2025-06-13
- Publication Date
- 2026-01-08
AI Technical Summary
Existing multimodal models struggle to understand complex image descriptions, leading to difficulties in downstream tasks.
An initial structured description information is generated using a large language model. Combined with object location information, structured image description information is generated. The image information extraction model is trained using a pre-set structured description template and the ICL method, simplifying the output format of the multimodal model.
It improves the richness and comprehensibility of image descriptions, simplifies information extraction for downstream tasks, and enhances the accuracy and efficiency of image processing.
Smart Images

Figure CN2025100832_08012026_PF_FP_ABST
Abstract
Description
Image processing method and device
[0001] The present disclosure claims priority to Chinese Patent Application No. 202410904423.X, filed on July 5, 2024, with the Chinese Patent Office, entitled "Image processing method and device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to the technical field of computer, in particular to an image processing method. BACKGROUND
[0003] A large number of images and corresponding text data are needed to train a text-image multimodal model. Generally speaking, high-quality text descriptions can significantly help the multimodal downstream task to understand the image content. However, due to the richness of picture information, high-quality text descriptions are usually long and complex. The downstream multimodal model lacking large language understanding ability is usually difficult to understand such complex text information. Therefore, there is an urgent need for a method that can provide rich but easy-to-understand picture description information for downstream multimodal models. SUMMARY
[0004] Therefore, one embodiment of the present disclosure provides an image processing method. One or more embodiments of the present disclosure also provide an image processing device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects in the prior art.
[0005] According to a first aspect of an embodiment of the present disclosure, an image processing method is provided, comprising:
[0006] obtaining a to-be-processed image;
[0007] inputting the to-be-processed image into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object;
[0008] determining object position information corresponding to each object according to the to-be-processed image and the initial structured description information;
[0009] generating structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object.
[0010] According to a second aspect of an embodiment of the present disclosure, an image processing method is provided, applied to a cloud-side device, comprising:
[0011] receiving a to-be-processed image sent by an end-side device;
[0012] input the image to be processed into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information comprises at least one object and object description information of each object;
[0013] determine object position information corresponding to each object according to the image to be processed and the initial structured description information;
[0014] generate structured image description information corresponding to the image to be processed according to the initial structured description information and the object position information corresponding to each object;
[0015] send the structured image description information to the terminal device.
[0016] According to a third aspect of the embodiments of the present disclosure, an image processing apparatus is provided, comprising:
[0017] an acquisition module configured to acquire an image to be processed;
[0018] an information extraction module configured to input the image to be processed into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information comprises at least one object and object description information of each object;
[0019] a determination module configured to determine object position information corresponding to each object according to the image to be processed and the initial structured description information;
[0020] a generation module configured to generate structured image description information corresponding to the image to be processed according to the initial structured description information and the object position information corresponding to each object.
[0021] According to a fourth aspect of the embodiments of the present disclosure, a computing device is provided, comprising:
[0022] a memory and a processor;
[0023] the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which realize the steps of the image processing method described above when executed by the processor.
[0024] According to a fifth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores computer programs / instructions, which realize the steps of the image processing method described above when executed by the processor.
[0025] According to a sixth aspect of the embodiments of the present disclosure, a computer program product is provided, comprising computer programs / instructions which, when executed by a processor, implement the steps of the image processing method described above.
[0026] According to an image processing method provided by one embodiment of the present disclosure, an image to be processed is obtained; the image to be processed is input into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information comprises at least one object and object description information of each object; object position information corresponding to each object is determined according to the image to be processed and the initial structured description information; and structured image description information corresponding to the image to be processed is generated according to the initial structured description information and the object position information corresponding to each object.
[0027] According to the image processing method provided by the embodiments of the present disclosure, the understanding ability of a large language model is used to analyze the image to be processed, and initial structured description information corresponding to the object to be processed is generated according to the preset structured information, wherein the initial structured description information comprises at least one object and object description information corresponding to each object, which is used to describe the image to be processed, facilitating the subsequent use of information in the image to be processed. In addition, each object in the image to be processed is positioned by combining the initial structured description information and the image to be processed, and object position information of each object is obtained, so that the structured image description information of the image to be processed is further enriched, and more accurate object and object description information can be obtained from the image to be processed according to the requirements of a downstream task. BRIEF DESCRIPTION OF DRAWINGS
[0028] FIG. 1 is a flowchart of an image processing method according to one embodiment of the present disclosure;
[0029] FIG. 2 is a schematic diagram of the composition of initial structured description information according to one embodiment of the present disclosure;
[0030] FIG. 3 is a schematic diagram of generating sample structured description information according to one embodiment of the present disclosure;
[0031] FIG. 4 is a schematic diagram of object positioning according to one embodiment of the present disclosure;
[0032] FIG. 5 is a processing framework diagram of an image processing method according to one embodiment of the present disclosure;
[0033] FIG. 6 is a flowchart of an image processing method applied to a cloud-side device according to one embodiment of the present disclosure;
[0034] FIG. 7 is a structural schematic diagram of an image processing apparatus according to one embodiment of the present disclosure;
[0035] FIG. 8 is a structural schematic diagram of an image processing device applied to a cloud-side device according to an embodiment of the present disclosure;
[0036] FIG. 9 is an architecture diagram of an image processing system according to an embodiment of the present disclosure;
[0037] FIG. 10 is a structural block diagram of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0038] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, the present disclosure can be practiced without the specific details, which are not necessary for the understanding of the present disclosure. In other instances, well-known methods, procedures, and components have not been described in detail so as not to unnecessarily obscure aspects of the present disclosure.
[0039] The terminology used in this disclosure, including the specific embodiments described herein, is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0040] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used solely to distinguish one from another only. For example, without departing from the scope of one or more embodiments of the present disclosure, first can be termed second, and similarly, second can be termed first. The term "if' as used herein, can be interpreted as meaning "when" or "in response to determining" depending on the context.
[0041] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards in relevant regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0042] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, billions, tens of billions, hundreds of billions, or even tens of billions of model parameters. The large model can also be referred to as a foundation model. Through large-scale unlabeled corpus pre-training, a pre-trained model with hundreds of millions of parameters is output. Such a model can adapt to a wide range of downstream tasks and has good generalization ability. For example, a large language model (LLM) and a multi-modal pre-training model.
[0043] In actual application, the large model only needs a small amount of samples to fine-tune the pre-trained model and can be applied to different tasks. The large model can be widely applied to natural language processing (NLP) and computer vision fields. Specifically, it can be applied to computer vision field tasks such as visual question answering (VQA), image captioning (IC), and image generation, and natural language processing field tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the large model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0044] First, the technical terms related to one or more embodiments of the present disclosure are explained.
[0045] Multi-modal small model: a model that needs to input multiple modal information (such as image and text information) at the same time, but lacks the ability to understand text information like a large language model.
[0046] Text-image multi-modal model needs an image and the corresponding text of the image as input. Usually, an image contains a large amount of information, which needs a large amount of text to accurately describe the image. The more accurate and rich the description information of the image is, the more beneficial it is to complete the multi-modal downstream task. However, most multi-modal small models lack the text understanding ability of a large language model, and even if the text information is rich, the multi-modal small model cannot be applied. Therefore, a method is needed to provide image description information that is rich in content and easy for multi-modal small models to understand and use.
[0047] Based on this, in the present disclosure, an image processing method is provided, and the present disclosure also relates to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.
[0048] Referring to FIG. 1, FIG. 1 shows a flowchart of an image processing method according to an embodiment of the present disclosure, which specifically includes the following steps.
[0049] Step 102: Obtain a to-be-processed image.
[0050] The to-be-processed image can be understood as an image that needs to be described in text. In the method provided in the embodiments of the present disclosure, the purpose is to describe the to-be-processed image in text and obtain image description information corresponding to the to-be-processed image.
[0051] In the method provided in the embodiments of the present disclosure, it is necessary to generate image description information corresponding to the to-be-processed image. It is necessary to first obtain the to-be-processed image. The way to obtain the to-be-processed image can be to obtain the to-be-processed image from a specified storage location, or the user can upload the to-be-processed image on a specified interactive interface. In the method provided in the embodiments of the present disclosure, the specific implementation manner of obtaining the to-be-processed image is not limited, and the actual application is used as the criterion.
[0052] Obtaining the to-be-processed image provides a data basis for subsequent image description of the to-be-processed image.
[0053] Step 104: Input the to-be-processed image into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object.
[0054] The image information extraction model can be understood as a model that understands the to-be-processed image and generates corresponding description information. In one or more embodiments of the present disclosure, the image information extraction model can be understood as a multi-modal large language model, which has the ability to understand and analyze multi-modal data.
[0055] The initial structured description information can be understood as the extraction result of the image information extraction model. The initial structured description information is not the final description information of the to-be-processed image. The initial structured description information is structured description information, and the initial structured description information includes at least one object and object description information of each object. The object can be understood as an object in the to-be-processed image, for example, a street lamp, a person, a car, a sky, and other elements in the to-be-processed image, which can be understood as objects in the to-be-processed image. The object description information can be understood as the description content of the object.
[0056] For example, taking a picture including a sky, two buildings, a street between the buildings, and two cars on the street as an example. The initial structured description information can specifically include:
[0057] “sky (non-living object; background; the sky presents a blue gradient and is interspersed with clouds. Color information: blue hue.);
[0058] building 1 (non-living object; background; the building has a unique angular design and internally illuminated windows. Color information: black and yellow light.);
[0059] building 2 (non-living object; background; the building is very high, has many windows, some of which are lighted, and has a cylindrical shape. Color information: white and yellow lighted windows.);
[0060] street (non-living object; foreground / background; the street shows multiple lanes, vehicles, and wet ground reflecting light. Color information: dark gray asphalt, white road markings.);
[0061] car 1 (non-living object; foreground; a car with headlights on is driving on the street, and the car body type is a sedan. Color information: black.);
[0062] car 2 (non-living object; foreground; another car with headlights on is following car 1, and looks like a hatchback. Color information: silver.)”。
[0063] Among them, the sky, building 1, building 2, street, car 1, and car 2 are six objects, and the content in the brackets after each object is the object description information.
[0064] It should be noted that the objects and object description information in the initial structured description information provided in the embodiments of the present disclosure can be displayed in the form of an object list. Through the list form of objects and object description information, the user can more intuitively and accurately understand the objects and object description information contained in the image to be processed.
[0065] In a specific implementation provided in the present disclosure, the initial structured description information output by the image information extraction model includes:
[0066] The initial graph structure description information output by the image information extraction model, wherein the initial graph structure description information includes at least one object node and object description information corresponding to each object node.
[0067] In the embodiment, the image information extraction model outputs initial graph structure description information according to the input to-be-processed image, that is, the initial structured description information can be the initial graph structure description information. The graph structure is a nonlinear data structure composed of nodes and edges. In the method provided in the embodiments of the present disclosure, the initial graph structure description information includes at least one object node and object description information corresponding to each object node.
[0068] Further, the initial structured description information further includes image overall description information of the to-be-processed image and the association relationship between objects.
[0069] In actual application, the initial structured description information further includes image overall description information of the to-be-processed image, which is used to reflect the content of the to-be-processed image as a whole. The initial structured description information further includes the association relationship between objects.
[0070] For example, still taking the to-be-processed image including the sky, two buildings, a street between the buildings, and two cars on the street as an example. The image overall description information can include:
[0071] “Style: This picture is a photo with a realistic style.
[0072] Theme: The theme of the picture is the urban traffic at dusk.
[0073] Background description: The background is composed of a cloudy sky, which presents a gradient from deep blue at the top to light blue at the bottom. On the left side is a large modern building named “***”, which is characterized by its unique angular design and internally illuminated windows. On the right side is a high-rise building with many windows, some of which emit warm light. Further away, there are more urban buildings, including one with a green glass facade. The ambient light indicates that it is either dawn or dusk, and artificial light is starting to have a significant impact on the scene.
[0074] Foreground depiction: The foreground shows a city street scene with multiple lanes of traffic, the street is busy, and the vehicles all have their headlights on, indicating poor light conditions. The vehicles vary in size and shape, with a mix of private and commercial vehicles. The sidewalk along the street is particularly clean under the wet light, which is a typical feature of an urban environment at night.
[0075] In addition to the above image overall description information and object list, the initial structured description information further includes the association relationship between objects, specifically:
[0076] “The sky covers the buildings;
[0077] Building 1 is located on the left side of the street;
[0078] The building 2 is located on the right side of the street;
[0079] The car 1 is driving on the street;
[0080] The car 2 is driving on the street following the car 1.
[0081] The above is the introduction of the specific content of the initial structured description information. In a specific implementation provided by an embodiment of the present disclosure, the initial structured description information at least includes an object list, which can be understood as at least one object and object description information of each pair of objects. Furthermore, the initial structured description information can also include image overall description information of the to-be-processed image and / or an association relationship between objects.
[0082] Referring to FIG. 2, FIG. 2 shows a composition diagram of initial structured description information provided by an embodiment of the present disclosure. As shown in FIG. 2, after the to-be-processed image is processed by the image information extraction model, the initial structured description information is generated, which includes three parts. The first part is image overall description information, the second part is an object list, and the third part is an association relationship between objects.
[0083] The first part is the image overall description information, which includes overall description information of the to-be-processed image, such as image style, image theme, image background description, image foreground description and the like.
[0084] The second part is the object list, which includes at least one object in the to-be-processed image and object description information of each object.
[0085] The third part is the association relationship, which includes the association relationship between objects in the to-be-processed image.
[0086] The initial structured description information is a special format of graph structure description information. In order to facilitate the model output, the content analyzed by the model can be displayed in the form of the initial structured description information. The object and object description information in the initial structured description information can be understood as a node in the graph structure and description information of the node.
[0087] In the method provided by an embodiment of the present disclosure, the image information extraction model is a pre-trained large language model, which is trained to output initial structured description information in a fixed format according to the input image. In a specific implementation provided by the present disclosure, the image information extraction model is obtained by the following steps:
[0088] Obtain a sample image and sample structured description information corresponding to the sample image, wherein the sample structured description information is generated by a preset structured description template.
[0089] inputting the sample image into an image information extraction model to obtain predicted structured description information output by the image information extraction model.
[0090] calculating a model loss value according to the sample structured description information and the predicted structured description information.
[0091] adjusting model parameters of the image information extraction model according to the model loss value, and continuing to train the image information extraction model until a model training stop condition is reached.
[0092] Specifically, the training method of the image information extraction model provided in the embodiments of the present disclosure uses supervised training, which includes a plurality of training sample pairs. A certain training sample pair includes a sample image and sample structured description information corresponding to the sample image. The sample image can be understood as a sample used to train the image information extraction model. The sample structured description information can be understood as structured description information corresponding to the sample image, which includes at least one sample object in the sample image and sample object description information corresponding to each sample object.
[0093] In actual applications, although the multi-modal large language model has strong capabilities, it cannot directly output the structured description information we hope. Therefore, in the method provided in the embodiments of the present disclosure, a template is provided for the image information extraction model based on the preset structured description template, which informs the image information extraction model to output corresponding structured description information according to the format of the preset structured description template.
[0094] Specifically, in a specific implementation provided in the present disclosure, obtaining a sample image and sample structured description information corresponding to the sample image includes:
[0095] obtaining a sample image and a preset structured description template;
[0096] inputting the sample image and the preset structured description template into a multi-modal large language model to obtain sample structured description information output by the multi-modal large language model.
[0097] The preset structured description template can be understood as a template pre-set in one or more embodiments of the present disclosure, which is used to inform the multi-modal large language model to output corresponding structured description information according to the format of the preset structured description template.
[0098] In actual applications, the preset structured description template includes at least one type of preset identification character, and the preset identification character represents a type of information in the preset structured description template. For example, the preset identification character can be "%%", "&&", "<>", "()", ";", "[]", and the like. Each preset identification character identifies a type of information in the preset structured template, for example, "%%" is used to distinguish a main title, "&&" is used to distinguish a sub-title, "<>" is used to identify a noun, "()" is used to identify an object attribute, ";" is used to separate object attributes, and "[]" is used to identify a relationship, and the like.
[0099] By using various types of preset identifiers in the preset structured description template, the preset structured description template can be divided into pre-specified information parts, so as to specify the preset character string format of the preset structured description template.
[0100] After the preset character string format is designed, in order to enable the multi-modal large language model to output structured description information according to the preset structured description template and the preset character string format, in one or more embodiments provided by the present disclosure, a target information type instance corresponding to a target information type can also be included in the preset structured description template. The In-context learning method (ICL, a method that can complete a natural language processing task by providing only a small number of examples) is used to provide examples for the multi-modal large language model, so that the multi-modal large language model outputs corresponding structured description information according to the preset character format.
[0101] The traditional ICL method needs to provide multiple examples, which are input to the multi-modal large language model in the form of a dialogue. However, in the method provided by the embodiments of the present disclosure, the content of the preset character string format is relatively large, and the cost of providing examples for all information is relatively high. Therefore, in the method provided by the embodiments of the present disclosure, a corresponding target information type instance is provided for a target information type. The target information type can be understood as a pre-set information type, for example, the target information type can be a main title, a sub-title, numbering different individuals of the same category, an output format of object information, a relationship between objects, and the like. Providing a corresponding target information type instance for the target information type can better help the multi-modal large language model to output according to the given preset character string format.
[0102] In the method provided by the embodiments of the present disclosure, the ICL method is used, and only a target information type instance is provided at the position of the target information type, so that the multi-modal large language model outputs corresponding structured description information according to the specified format. The length of the text input to the multi-modal large language model is shortened, and only a single round of dialogue is used, so that the multi-modal large language model outputs the structured description information.
[0103] Referring to FIG. 3, FIG. 3 shows a schematic diagram of generating sample structured description information according to an embodiment of the present disclosure. As shown in FIG. 3, the sample image and the preset structured description template are input into the multi-modal large language model. After understanding and analyzing the sample image, the multi-modal large language model can output corresponding sample structured description information according to the preset structured description template. The sample structured description information is generated according to the pre-set string format and examples.
[0104] As shown in FIG. 3, the sample structured description information includes three parts. The first part is the overall description part, which further includes style, theme, background description, foreground description and the like. The second part is the object list, which further includes each object in the image and the object description information of each object. The third part is the association relationship, which further includes the association relationship between each object in the image.
[0105] The sample structured description information is a pre-set structured format of image description information provided by an embodiment of the present disclosure. After understanding and analyzing the sample image, the multi-modal large language model outputs corresponding sample structured description information according to the pre-set string format. The sample image is described through a specific structured format, so as to train the image information extraction model subsequently. The trained image information extraction model can output initial structured description information in a specified format.
[0106] After obtaining the sample structured description information corresponding to the sample image through the above steps, the image information extraction model can be trained according to the sample image and the sample structured description information. Specifically, the sample image is input into the image information extraction model. At this time, the image information extraction model can be understood as the multi-modal large language model described above, which has the ability to output the preset format according to the sample image. The sample image is input into the image information extraction model for processing. The image information extraction model can output the predicted structured description information according to the sample image.
[0107] The predicted structured description information can be understood as the description information output by the untrained image information extraction model according to the input sample image. The data structure thereof corresponds to the pre-set character format.
[0108] At this time, the image information extraction model is an untrained model. There is still a difference between the predicted structured description information output by the image information extraction model and the sample structured description information. The model loss value needs to be calculated according to the predicted structured description information and the sample structured description information. In the method provided by an embodiment of the present disclosure, there are many methods for calculating the model loss value, such as cross-entropy loss function, maximum loss function, average value loss function and the like. In the embodiments provided by the present disclosure, the specific method of the loss function is not limited, and the actual application is used as the criterion.
[0109] After obtaining the model loss value, the image information extraction model can be back propagated according to the model loss value to adjust the model parameters in the image information extraction model. Then the above steps can be repeated to continue training the image information extraction model until the model training stopping condition is reached. In actual application, the model training stopping condition of the initial image processing model includes:
[0110] The model loss value is less than a preset threshold, and / or the training round reaches a preset training round.
[0111] Specifically, during the training of the image information extraction model, the training stopping condition of the model can be set as the model loss value being less than a preset threshold, that is, when the model loss value is less than the preset threshold, there is no need to adjust the model parameters of the image information extraction model.
[0112] The model training stopping condition of the image information extraction model can also be set as the training round reaching a preset training round, for example, the preset training round is 10 rounds, and when the training round of the model reaches 10 rounds, that is, the model training stopping condition is reached.
[0113] In the method provided in the present disclosure, the model training stopping condition is not limited. When the image information extraction model reaches the model training stopping condition, it means that the image information extraction model training is completed, and the final image information extraction model is obtained.
[0114] Step 106: determining the object position information corresponding to each object according to the to-be-processed image and the initial structured description information.
[0115] The object position information can be understood as the position information of the object in the to-be-processed image. Although the image information extraction model has certain object positioning ability, the positioning effect is poor. In order to facilitate the downstream task to extract the corresponding image description information from the to-be-processed image, the more accurate object position information of each object needs to be further obtained.
[0116] In the present embodiment, after obtaining the initial structured description information, the object position information corresponding to each object in the to-be-processed image can be further determined according to at least one object in the initial structured description information and the object description information of each object. Specifically, the object position information of each object in the to-be-processed image can be marked in the form of a marking box.
[0117] In a specific embodiment provided in the present disclosure, the object position information corresponding to each object is determined according to the to-be-processed image and the initial structured description information, including:
[0118] S1062, acquire the to-be-processed object in the initial structured description information and to-be-processed object description information corresponding to the to-be-processed object, wherein the to-be-processed object is any one of the at least one object.
[0119] In actual application, the initial structured description information includes at least one object and object description information corresponding to each object. In this embodiment, one of the objects is taken as an example for explanation. The to-be-processed object is the object to be operated and processed, and the to-be-processed object description information is the object description information corresponding to the to-be-processed object. Specifically, the to-be-processed object and the to-be-processed object description information can be selected from the object list.
[0120] The to-be-processed object can be any one of the objects, and the processing between the objects can be parallel or sequential. In the embodiments provided in this disclosure, the processing order between the objects is not limited.
[0121] S1064, determine at least one candidate object bounding box according to the to-be-processed object in the to-be-processed image.
[0122] After the to-be-processed object is determined, at least one candidate object bounding box can be determined according to the to-be-processed object in the to-be-processed image. The candidate object bounding box is a bounding box generated in a detection positioning task for an object in an image, which can mark the position information of the object in the image.
[0123] In a specific embodiment provided in this embodiment, the candidate object bounding box can be generated by a positioning model. Determining at least one candidate object bounding box according to the to-be-processed object in the to-be-processed image includes:
[0124] The to-be-processed object and the to-be-processed image are input into the positioning model to obtain at least one candidate object bounding box determined by the positioning model according to the to-be-processed object in the to-be-processed image.
[0125] The positioning model can be understood as a model for positioning and marking in an image according to input information. The positioning model can mark the object corresponding to the information in the to-be-processed image according to the input information. For example, taking the to-be-processed object "man" as an example, the to-be-processed object and the to-be-processed image are input into the positioning model, and the positioning model marks the candidate object bounding box related to "man" in the to-be-processed image.
[0126] At this time, the positioning model can only mark at least one candidate bounding box in the to-be-processed image according to the input to-be-processed object, but the positioning model is difficult to distinguish different individuals of the same category, for example, when the to-be-processed object is "man", the positioning model can only mark multiple man objects in the to-be-processed image, but cannot distinguish the differences between the man. Therefore, the method provided in the embodiments of the present disclosure needs to further screen the candidate object bounding box determined by the positioning model to obtain the final target object bounding box.
[0127] In a specific embodiment provided by the present disclosure, determining at least one candidate object bounding box in the to-be-processed image according to the to-be-processed object comprises:
[0128] inputting the to-be-processed object and the to-be-processed image into a positioning model to obtain at least one initial candidate object bounding box determined by the positioning model in the to-be-processed image according to the to-be-processed object;
[0129] inputting the to-be-processed object and each initial candidate object bounding box into a multi-modal large language model to obtain at least one candidate object bounding box output by the multi-modal large language model, wherein the multi-modal large language model screens each initial candidate object bounding box according to the to-be-processed object.
[0130] In actual application, there may be a large gap between the candidate object bounding box and the to-be-processed object. In order to improve the processing efficiency in the subsequent processing process, the multi-modal large language model can also be used to screen the candidate object bounding box.
[0131] Specifically, the to-be-processed object and the to-be-processed image are first input into the positioning model, and the positioning model outputs at least one initial candidate object bounding box. At this time, due to the accuracy problem of the positioning model, there may be a positioning error bounding box. Therefore, the to-be-processed object and each initial candidate object bounding box can be input into the multi-modal large language model for screening. The initial candidate bounding box that is obviously inconsistent with the to-be-processed object is deleted, and the candidate bounding box that matches the to-be-processed object is retained.
[0132] In one or more specific embodiments provided by the present disclosure, the to-be-processed object and each initial candidate bounding box can be input into the multi-modal large language model for screening, and the multi-modal large language model outputs the candidate bounding box that matches the to-be-processed object. Alternatively, one initial candidate bounding box can be selected from multiple initial candidate object bounding boxes, and the selected initial candidate bounding box and the to-be-processed object are input into the multi-modal large language model for screening. The multi-modal large language model determines the matching degree between the initial candidate bounding box and the to-be-processed object, and the initial candidate bounding box that meets the preset threshold is taken as the candidate bounding box.
[0133] S1066, determine a target object marking box from the at least one candidate object marking box according to the to-be-processed object description information, and determine marking box position information of the target object marking box.
[0134] In the method provided in the embodiments of the present disclosure, the to-be-processed object description information of the to-be-processed object is used to select a target object marking box from at least one candidate object marking box. The to-be-processed object description information is a detailed description of the to-be-processed object. Through the to-be-processed object description information, the target object marking box can be selected from the candidate object marking box, and the marking box position information of the target object marking box can be determined.
[0135] Specifically, the target object marking box is determined from the at least one candidate object marking box according to the to-be-processed object description information, including:
[0136] calculating a matching similarity between the to-be-processed object description information and each candidate object marking box;
[0137] determining a candidate object marking box whose matching similarity meets a preset condition as the target object marking box.
[0138] In one or more embodiments provided in the present disclosure, the matching similarity between the to-be-processed object description information and each candidate object marking box is calculated. The matching similarity represents the similarity between the to-be-processed object description information and the candidate object marking box. The higher the matching similarity, the more matched the to-be-processed object and the candidate object marking box, and vice versa.
[0139] After the matching similarity of each candidate object marking box is calculated, the candidate object marking box corresponding to the matching similarity meeting the preset condition is determined as the target object marking box. The preset condition can be understood as a condition for selecting the target object marking box from multiple candidate object marking boxes. For example, the candidate object marking box with the highest matching similarity can be the target object marking box. After the target object marking box is determined, the corresponding marking box position information can be determined according to the target object marking box.
[0140] S1068, determine the marking box position information as object position information corresponding to the to-be-processed object.
[0141] After the marking box position information is determined, the marking box position information is determined as the object position information corresponding to the to-be-processed object.
[0142] Referring to FIG. 4, FIG. 4 shows a schematic diagram of object positioning provided by an embodiment of the present application. As shown in FIG. 4, the image to be processed is input into the image information extraction model for extraction, and initial structured description information output by the image information extraction model is obtained. The initial structured description information includes an object list, and the object list includes at least one object included in the image to be processed and object description information of each object.
[0143] The object to be processed "person" is selected from the object list, and the corresponding object description information to be processed is "wearing a hat". Since the positioning model can only process object names, the object to be processed "person" is input into the positioning model. The positioning model performs positioning on the object to be processed "person" in the image to be processed, and generates four initial candidate marking boxes.
[0144] The four initial candidate marking boxes and the object to be processed "person" are input into the multi-modal large language model for screening. The initial candidate marking boxes that are obviously not matched with the object to be processed "person" are deleted, and three candidate marking boxes are determined.
[0145] The three candidate marking boxes are matched with the object description information to be processed "wearing a hat" respectively. The candidate marking box with the highest matching information is selected as the target marking box, and the target marking box is marked in the image to be processed.
[0146] Step 108: generating structured image description information corresponding to the image to be processed according to the initial structured description information and object position information corresponding to each object.
[0147] After determining the object position information corresponding to each object, the object position information corresponding to each object can be added to the initial structured description information, and each object is one-to-one corresponding. The final structured image description information generated not only includes each object and object description information of each object, but also includes object position information corresponding to each object.
[0148] The structured image description information to which the object position information is added can describe the image to be processed from multiple dimensions, so that the downstream task can select corresponding sub-information from the structured image description information according to different task requirements, and perform subsequent downstream tasks. Specifically, in another specific embodiment provided by the present disclosure, the method further includes:
[0149] receiving a data request instruction of a downstream task, wherein the data request instruction carries data request information;
[0150] extracting image description sub-information from the structured image description information in response to the data request information, and sending the image description sub-information to the downstream task.
[0151] The downstream task can be understood as another image processing task that needs to rely on the structured image description information. Different downstream tasks can perform different operations on the to-be-processed image, and thus different downstream tasks obtain different information from the to-be-processed image. After receiving the data request instruction of the downstream task, the data request information in the data request instruction can be used to filter the structured image description information, extract image description sub-information corresponding to the data request information, and send the image description sub-information to the downstream task, so that the multi-modal model of the downstream task can perform a corresponding downstream task according to the image description sub-information.
[0152] The image processing method provided by the embodiment of the present disclosure includes obtaining a to-be-processed image; inputting the to-be-processed image into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object; determining object position information corresponding to each object according to the to-be-processed image and the initial structured description information; and generating structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object.
[0153] The image information extraction model generates initial structured description information corresponding to the to-be-processed image according to a preset format. According to at least one object and object description information corresponding to each object in the initial structured description information, object position information of each object in the to-be-processed image can be determined, and final structured image description information can be determined according to the initial structured description information and the object position information of each object. The content of the image is described in a specific structured description image manner, which facilitates the extraction of related information from the structured image description information.
[0154] Secondly, the image information extraction model is trained based on a specific structured description image manner. The ICL method is used to provide corresponding target information type instances only for target information types, so that the image information extraction model can well learn the preset string format, and thus output initial structured description information according to the preset string format. The training cost is reduced, and the learning process is simplified.
[0155] Finally, the objects are positioned in the to-be-processed image by using the initial structured description information and a positioning model to determine position information corresponding to each object. Thus, each object in the initial structured description information can be better and more detailedly depicted. The convenience of obtaining information from the structured image description information by a subsequent downstream task is further improved.
[0156] FIG. 5 shows a processing framework diagram of an image processing method provided by an embodiment of the present disclosure. As shown in FIG. 5, in the method provided by the embodiment of the present disclosure, a preset structured description template is designed by a preset string format. A sample structured description information is generated by a multi-modal large language model using the preset structured description template. A sample structured description information and a sample image information extraction model are trained, so that the image information extraction model has the ability to output structured description information according to input image.
[0157] After the model training is completed, the to-be-processed image is input into the trained image information extraction model to obtain initial structured description information output by the image information extraction model. The initial structured description information includes at least one object and object description information corresponding to each object, but does not include position information of each object in the to-be-processed image.
[0158] Based on the object and the object description information, the position information of each object in the to-be-processed image is obtained through positioning processing of the positioning model. The position information of each object is added to the initial structured description information to generate final structured image description information. The structured image description information includes at least one object, object description information and object position information of each object.
[0159] FIG. 6 shows a flowchart of an image processing method applied to a cloud-side device provided by an embodiment of the present disclosure. The method is applied to a cloud-side device and specifically includes:
[0160] Step 602: receiving a to-be-processed image sent by an end-side device.
[0161] Step 604: inputting the to-be-processed image into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object.
[0162] Step 606: determining object position information corresponding to each object according to the to-be-processed image and the initial structured description information.
[0163] Step 608: generating structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object.
[0164] Step 610: sending the structured image description information to the end-side device.
[0165] In a specific embodiment provided by the present disclosure, the method further includes:
[0166] receiving an information processing instruction sent by the end-side device for the structured image description information;
[0167] obtaining image description adjustment information in response to the information processing instruction processing the structured image description information;
[0168] sending the image description adjustment information to the terminal device.
[0169] In actual application, since the image information extraction model is a multi-modal large language model, it needs better computing resources at runtime, and the terminal device may not have corresponding processing capability. Therefore, the image information extraction model can be deployed on the cloud side device, that is, the image processing process provided by the embodiment of the disclosure is implemented on the cloud side device. After obtaining the structured image description information generated by the image information extraction model, the cloud side device can also send the structured image description information to the terminal device.
[0170] The image processing method provided by the embodiment of the disclosure generates initial structured description information corresponding to the to-be-processed image in a preset format through the image information extraction model, and according to at least one object in the initial structured description information and object description information corresponding to each object, the object position information of each object in the to-be-processed image can be determined, and the final structured image description information is determined according to the initial structured description information and the object position information of each object. Through a specific structured description image, the content of the image is described, which facilitates the extraction of related information from the structured image description information in the subsequent process.
[0171] Secondly, the image information extraction model is trained based on a specific structured description image, and the ICL method is used to provide only corresponding target information type instances for the target information type, so that the image information extraction model can well learn the preset string format, and thus output the initial structured description information according to the preset string format. The training cost is reduced, and the learning process is simplified.
[0172] Finally, the initial structured description information and the positioning model are used to position each object in the to-be-processed image, and the position information corresponding to each object is determined. Thus, each object in the initial structured description information is better described in more detail. Further improve the convenience of subsequent downstream tasks to obtain information from the structured image description information.
[0173] Corresponding to the above method embodiments, the disclosure also provides image processing device embodiments. FIG. 7 shows a structural schematic diagram of an image processing device according to an embodiment of the disclosure. As shown in FIG. 7, the device includes:
[0174] The acquisition module 702 is configured to acquire a to-be-processed image.
[0175] The information extraction module 704 is configured to input the to-be-processed image into an image information extraction model, and obtain initial structured description information output by the image information extraction model, where the initial structured description information includes at least one object and object description information of each object.
[0176] The determination module 706 is configured to determine object position information corresponding to each object according to the to-be-processed image and the initial structured description information.
[0177] The generation module 708 is configured to generate structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object.
[0178] Optionally, the information extraction module 704 is configured to:
[0179] obtain initial graph structure description information output by the image information extraction model, where the initial graph structure description information includes at least one object node and object description information corresponding to each object node.
[0180] Optionally, the apparatus further includes a training module configured to:
[0181] obtain a sample image and sample structured description information corresponding to the sample image, where the sample structured description information is generated by using a preset structured description template;
[0182] input the sample image into an image information extraction model, and obtain predicted structured description information output by the image information extraction model;
[0183] calculate a model loss value according to the sample structured description information and the predicted structured description information;
[0184] adjust model parameters of the image information extraction model according to the model loss value, and continue to train the image information extraction model until a model training stop condition is reached.
[0185] Optionally, the training module is further configured to:
[0186] obtain a sample image and a preset structured description template;
[0187] input the sample image and the preset structured description template into a multi-modal large language model, and obtain sample structured description information output by the multi-modal large language model.
[0188] Optionally, the preset structured description template includes at least one type of preset identification character, and the preset identification character represents a type of information in the preset structured description template.
[0189] Optionally, the preset structured description template includes a target information type instance corresponding to a target information type.
[0190] Optionally, the determining module 706 is further configured to:
[0191] obtain a to-be-processed object in the initial structured description information and to-be-processed object description information corresponding to the to-be-processed object, wherein the to-be-processed object is any one of the at least one object;
[0192] determine at least one candidate object bounding box of the to-be-processed object in the to-be-processed image;
[0193] determine a target object bounding box in the at least one candidate object bounding box according to the to-be-processed object description information, and determine bounding box position information of the target object bounding box;
[0194] determine the bounding box position information as object position information corresponding to the to-be-processed object.
[0195] Optionally, the determining module 706 is further configured to:
[0196] input the to-be-processed object and the to-be-processed image into a positioning model to obtain at least one candidate object bounding box determined by the positioning model according to the to-be-processed object in the to-be-processed image.
[0197] Optionally, the determining module 706 is further configured to:
[0198] input the to-be-processed object and the to-be-processed image into a positioning model to obtain at least one initial candidate object bounding box determined by the positioning model according to the to-be-processed object in the to-be-processed image.
[0199] input the to-be-processed object and each initial candidate object bounding box into a multi-modal large language model to obtain at least one candidate object bounding box output by the multi-modal large language model, wherein the multi-modal large language model filters each initial candidate object bounding box according to the to-be-processed object.
[0200] Optionally, the determining module 706 is further configured to:
[0201] calculate a matching similarity between the to-be-processed object description information and each candidate object bounding box;
[0202] determine a candidate object bounding box that satisfies a preset condition in the matching similarity as a target object bounding box.
[0203] Optionally, the initial structured description information further comprises image overall description information of the to-be-processed image and a correlation between objects.
[0204] Optionally, the apparatus further comprises an information extraction module configured to:
[0205] receive a data request instruction of a downstream task, wherein the data request instruction carries data request information;
[0206] extract image description sub-information from the structured image description information in response to the data request information, and send the image description sub-information to the downstream task.
[0207] The image processing apparatus provided by the embodiments of the present disclosure comprises: obtaining a to-be-processed image; inputting the to-be-processed image into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information comprises at least one object and object description information of each object; determining object position information corresponding to each object according to the to-be-processed image and the initial structured description information; and generating structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object.
[0208] The image information extraction model generates initial structured description information corresponding to the to-be-processed image according to a preset format, and according to at least one object and object description information of each object in the initial structured description information, object position information of each object can be determined in the to-be-processed image, and final structured image description information can be determined according to the initial structured description information and the object position information of each object. By using a specific structured description image, the content of the image is described, which facilitates the extraction of related information from the structured image description information in the subsequent process.
[0209] Secondly, the image information extraction model is trained based on a specific structured description image, and an ICL method is used to provide corresponding target information type instances only for target information types, so that the image information extraction model can well learn the preset string format, and thus output initial structured description information according to the preset string format. The training cost is reduced, and the learning process is simplified.
[0210] Finally, each object is positioned in the to-be-processed image by using the initial structured description information and a positioning model to determine position information corresponding to each object, so that each object in the initial structured description information can be better and more detailedly depicted. The convenience of obtaining information from the structured image description information in the subsequent downstream task is further improved.
[0211] The above is a schematic scheme of the image processing device of the embodiment. It should be noted that the technical scheme of the image processing device belongs to the same concept as the technical scheme of the image processing method described above, and the details of the technical scheme of the image processing device that are not described in detail can be referred to the description of the technical scheme of the image processing method.
[0212] Corresponding to the method embodiments described above, the disclosure also provides image processing device embodiments. FIG. 8 shows a structural schematic diagram of an image processing device applied to a cloud-side device according to an embodiment of the disclosure. As shown in FIG. 8, the device applied to the cloud-side device includes:
[0213] The receiving module 802 is configured to receive the to-be-processed image sent by the end-side device;
[0214] The information extraction module 804 is configured to input the to-be-processed image into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object;
[0215] The determination module 806 is configured to determine object position information corresponding to each object according to the to-be-processed image and the initial structured description information;
[0216] The generation module 808 is configured to generate structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object;
[0217] The sending module 810 is configured to send the structured image description information to the end-side device.
[0218] Optionally, the device further includes an adjustment module configured to:
[0219] Receive the information processing instruction sent by the end-side device for the structured image description information;
[0220] Process the structured image description information in response to the information processing instruction to obtain image description adjustment information;
[0221] Send the image description adjustment information to the end-side device.
[0222] The image processing apparatus provided in the embodiment of the present disclosure generates initial structured description information corresponding to a to-be-processed image in a preset format through an image information extraction model, determines object position information of each object in the to-be-processed image according to at least one object in the initial structured description information and object description information corresponding to each object, and determines final structured image description information according to the initial structured description information and the object position information of each object. The content of the image is described in a specific structured description image manner, which facilitates subsequent extraction of related information from the structured image description information.
[0223] Secondly, the image information extraction model is trained based on the specific structured description image manner, and the ICL method is used to provide corresponding target information type instances only for target information types, so that the image information extraction model can well learn the preset string format, and thus output the initial structured description information according to the preset string format. The training cost is reduced, and the learning process is simplified.
[0224] Finally, each object is positioned in the to-be-processed image through the initial structured description information and a positioning model, and the position information corresponding to each object is determined. Thus, each object in the initial structured description information is better described in more detail. The convenience of subsequent downstream tasks for obtaining information from the structured image description information is further improved.
[0225] The above is a schematic scheme of the image processing apparatus of the embodiment. It should be noted that the technical scheme of the image processing apparatus belongs to the same concept as the technical scheme of the image processing method described above, and the details of the technical scheme of the image processing apparatus that are not described in detail can be referred to the description of the technical scheme of the image processing method.
[0226] Referring to FIG. 9, FIG. 9 shows an architecture diagram of an image processing system according to an embodiment of the present disclosure. The image processing system can include a client 100 and a server 200.
[0227] The client 100 is configured to send a to-be-processed image to the server 200.
[0228] The server 200 is configured to input the to-be-processed image into an image information extraction model, obtain initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object; determine object position information corresponding to each object according to the to-be-processed image and the initial structured description information; generate structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object; and send the structured image description information to the client 100.
[0229] The client 100 is also configured to receive the structured image description information sent by the server 200.
[0230] The image processing system can include a plurality of clients 100 and a server 200, where the client 100 can be referred to as an end-side device, and the server 200 can be referred to as a cloud-side device. The plurality of clients 100 can establish a communication connection through the server 200. In the image processing scenario, the server 200 is used to provide image processing services between the plurality of clients 100. The plurality of clients 100 can respectively act as a sending end or a receiving end to implement communication through the server 200.
[0231] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100 or send data to other clients 100, and the like. In the image processing scenario, the user can publish a data stream to the server 200 through the client 100. The server 200 generates structured image description information according to the data stream and pushes the structured image description information to other clients that establish a communication connection.
[0232] The client 100 and the server 200 establish a connection through a network. The network provides a medium for a communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or optical fiber cables, and the like. The data transmitted by the client 100 can need to be processed through encoding, transcoding, compression, and the like before being published to the server 200.
[0233] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also referred to as a small program, a lightweight application), or a cloud application, and the like. The client 100 can be developed based on a software development kit (SDK) provided by the server 200 for a corresponding service, such as an RTC (Real Time Communication) SDK. The client 100 can be deployed in an electronic device and needs to depend on the device or some APP in the device to run, and the like. The electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, and the like. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant communication tools, mailbox clients, social platform software, and the like.
[0234] The server 200 can include a server providing various services, for example, a server providing a communication service for a plurality of clients, for example, a server for background training providing support for a model used on a client, for example, a server processing data sent by a client, and the like. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of a cloud service, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, and the like basic cloud computing services, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0235] It should be noted that the image processing method provided in the embodiments of the present disclosure is generally executed by the server, but in other embodiments of the present disclosure, the client can also have similar functions as the server, so as to execute the image processing method provided in the embodiments of the present disclosure. In other embodiments, the image processing method provided in the embodiments of the present disclosure can also be executed by the client and the server together.
[0236] FIG. 10 shows a structural block diagram of a computing device 1000 according to an embodiment of the present disclosure. The components of the computing device 1000 include, but are not limited to, a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 through a bus 1030, and a database 1050 is used to save data.
[0237] The computing device 1000 also includes an access device 1040 that enables the computing device 1000 to communicate via one or more networks 1060. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or combinations of such networks, such as the Internet. The access device 1040 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0238] In one embodiment of the present disclosure, the above-mentioned components of the computing device 1000 and other components not shown in FIG. 10 can also be connected to each other, for example, through a bus. It should be understood that the computing device structure block diagram shown in FIG. 10 is only for the purpose of example, and is not a limitation on the scope of the present disclosure. Other components can be added or replaced by those skilled in the art as needed.
[0239] The computing device 1000 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1000 can also be a mobile or stationary server.
[0240] The processor 1020 is configured to execute computer program / instructions that implement the steps of the above image processing method when the computer program / instructions are executed by the processor.
[0241] The various embodiments in the present disclosure are described in a progressive manner, and the same or similar parts among the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the computing device embodiment is basically similar to the image processing method embodiment, and thus the description is relatively simple, and the relevant parts can be referred to the description of the image processing method embodiment.
[0242] An embodiment of the present disclosure further provides a computer readable storage medium, which stores computer programs / instructions, and the computer programs / instructions are executed by a processor to realize the steps of the image processing method.
[0243] The various embodiments in the present disclosure are described in a progressive manner, and the same or similar parts among the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the computer readable storage medium embodiment is basically similar to the image processing method embodiment, and thus the description is relatively simple, and the relevant parts can be referred to the description of the image processing method embodiment.
[0244] An embodiment of the present disclosure further provides a computer program product, which includes computer programs / instructions, and the computer programs / instructions are executed by a processor to realize the steps of the image processing method.
[0245] The above is a schematic scheme of the computer program product of the embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the image processing method belong to the same concept, and the details of the technical scheme of the computer program product which are not described in detail can be referred to the description of the technical scheme of the image processing method.
[0246] The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0247] The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or subtractions according to the requirements of patent practice. For example, according to the patent practice in some regions, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0248] It should be noted that the above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve the desired results. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous. Secondly, those skilled in the art should know that the embodiments described in the present disclosure are preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of the present disclosure.
[0249] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0250] The preferred embodiments of the present disclosure disclosed above are only used to help explain the present disclosure. The alternative embodiments do not describe all the details and do not limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of the present disclosure, many modifications and changes can be made. The present disclosure selects and describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can well understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. An image processing method, comprising: obtaining a to-be-processed image; inputting the to-be-processed image into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information comprises at least one object and object description information of each object; determining object position information corresponding to each object according to the to-be-processed image and the initial structured description information; and generating structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object.
2. The method of claim 1, wherein obtaining the initial structured description information output by the image information extraction model comprises: obtaining initial graph structure description information output by the image information extraction model, wherein the initial graph structure description information comprises at least one object node and object description information corresponding to each object node.
3. The method of claim 1 or 2, wherein the image information extraction model is trained by the following steps: obtaining a sample image and sample structured description information corresponding to the sample image, wherein, generating sample structured description information by using a preset structured description template; inputting a sample image into the image information extraction model to obtain predicted structured description information output by the image information extraction model; calculating a model loss value according to the sample structured description information and the predicted structured description information; adjusting model parameters of the image information extraction model according to the model loss value, and continuing to train the image information extraction model until a model training stop condition is reached.
4. The method of claim 3, wherein obtaining a sample image and sample structured description information corresponding to the sample image comprises: obtaining a sample image and a preset structured description template; inputting the sample image and the preset structured description template into a multi-modal large language model to obtain sample structured description information output by the multi-modal large language model.
5. The method of claim 3 or 4, wherein the preset structured description template comprises at least one type of preset identification character, and the preset identification character represents a type of information in the preset structured description template.
6. The method of any one of claims 3-5, wherein the preset structured description template comprises a target information type instance corresponding to a target information type.
7. The method of any one of claims 1-6, wherein determining object position information corresponding to each object according to the to-be-processed image and the initial structured description information comprises: obtaining a to-be-processed object in the initial structured description information and to-be-processed object description information corresponding to the to-be-processed object, wherein the to-be-processed object is any one of the at least one object; determining at least one candidate object bounding box according to the to-be-processed object in the to-be-processed image; determining a target object bounding box in the at least one candidate object bounding box according to the to-be-processed object description information, and determining bounding box position information of the target object bounding box; determining the bounding box position information as the object position information corresponding to the to-be-processed object.
8. The method of claim 7, wherein determining at least one candidate object bounding box of the to-be-processed object in the to-be-processed image comprises: inputting the to-be-processed object and the to-be-processed image into a positioning model to obtain at least one candidate object bounding box determined by the positioning model according to the to-be-processed object in the to-be-processed image.
9. The method of claim 7, wherein determining at least one candidate object bounding box of the to-be-processed object in the to-be-processed image comprises: inputting the to-be-processed object and the to-be-processed image into a positioning model to obtain at least one initial candidate object bounding box determined by the positioning model according to the to-be-processed object in the to-be-processed image; inputting the to-be-processed object and each initial candidate object bounding box into a multi-modal large language model to obtain at least one candidate object bounding box output by the multi-modal large language model, wherein the multi-modal large language model filters each initial candidate object bounding box according to the to-be-processed object.
10. The method of any one of claims 7-9, wherein determining a target object bounding box of the to-be-processed object description information in at least one candidate object bounding box comprises: calculating a matching similarity between the to-be-processed object description information and each candidate object bounding box; determining a candidate object bounding box whose matching similarity meets a preset condition as the target object bounding box.
11. The method of any one of claims 1-10, wherein the initial structured description information further comprises image overall description information of the to-be-processed image and a correlation between objects.
12. The method of any one of claims 1-11, further comprising: receiving a data request instruction of a downstream task, wherein the data request instruction carries data request information; extracting image description sub-information from the structured image description information in response to the data request information, and sending the image description sub-information to the downstream task.
13. An image processing method applied to a cloud-side device, comprising: receiving a to-be-processed image sent by an end-side device; inputting the to-be-processed image into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information comprises at least one object and object description information of each object; determining object position information corresponding to each object according to the to-be-processed image and the initial structured description information; generating structured image description information corresponding to the to-be-processed image according to the initial structured description information and the object position information corresponding to each object; sending the structured image description information to the end-side device.
14. The method of claim 13, further comprising: receiving an information processing instruction sent by the end-side device for the structured image description information; processing the structured image description information in response to the information processing instruction to obtain image description adjustment information; sending the image description adjustment information to the end-side device.
15. An image processing apparatus, comprising: an acquisition module configured to acquire a to-be-processed image; An information extraction module configured to input the image to be processed into an image information extraction model, and obtain initial structured description information output by the image information extraction model, wherein the initial structured description information comprises at least one object and object description information of each object; A determination module configured to determine object position information corresponding to each object according to the image to be processed and the initial structured description information; A generation module configured to generate structured image description information corresponding to the image to be processed according to the initial structured description information and the object position information corresponding to each object.
16. A computing device comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, and the computer programs / instructions, when executed by the processor, implement the steps of the method of any one of claims 1 to 14.
17. A computer readable storage medium storing computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method of any one of claims 1 to 14.
18. A computer program product comprising computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method of any one of claims 1 to 14.
Citation Information
Patent Citations
Image figure behavior description generation method based on multi-stage image context coding and decoding
CN113449801A
Image description generation method and device, equipment, medium and product
CN114627353A
Image text information generation method and deep learning model training method
CN115359323A
Method and device for determining picture description information generation model, medium and equipment
CN116050496A
Image processing method and device
CN118968081A