Image description method and system, storage medium and electronic equipment
Through object detection and panoramic segmentation combined with large language model and feature extraction technology, the problem that the image description model accuracy is limited by the training data volume is solved, high-precision image description is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510566618.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-26
AI Technical Summary
The accuracy of existing deep learning-based image description models is limited by the number of training data and cannot meet the practical application requirements.
The CO-Detr object detector, Mask DINO segmentation model, Oneformer segmentation model or mask2former segmentation model is used for object detection and panoramic segmentation, combined with CLIP text encoder and image encoder to extract features, use large language models such as chatGPT, GPT-3 and GPT-4 to generate description text, and select the best description through cosine similarity.
It breaks through the limit on the amount of training data, realizes accurate image description, and improves the reliability and user experience of the model.
Smart Images

Figure CN120544196A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and in particular relates to an image description method, system, storage medium and electronic device. Background Art
[0002] Image description takes an image as input and uses mathematical models and calculations to enable the computer to output natural language description text corresponding to the image, giving the computer the ability to "speak by looking at the image". It is another new task in the field of image processing after image recognition, image segmentation and target tracking.
[0003] The essence of image captioning is to convert computer-extracted visual features of an image into high-level semantic information, enabling the computer to generate a textual description of the image that is similar to what the human brain understands. This allows for image classification, retrieval, analysis, and other processing tasks.
[0004] With the continuous development of deep learning technology, neural networks have been widely used in computer vision and natural language processing. Inspired by the encoder-decoder model in machine translation, image description can directly achieve the mapping between images and description sentences through end-to-end learning methods, transforming the image description process into a "translation" process from image to description. Deep learning methods can directly learn the mapping from images to description sentences from large amounts of data, generating more accurate descriptions with performance far exceeding that of traditional methods.
[0005] However, the accuracy of existing image description models based on deep learning methods is limited by the amount of training data and cannot meet the needs of practical applications. Summary of the Invention
[0006] In view of the above-mentioned shortcomings of the prior art, an object of the present invention is to provide an image description method, system, storage medium and electronic device, which can achieve accurate image description and are not limited by the amount of training data.
[0007] In a first aspect, the present invention provides an image description method, comprising the following steps: performing target detection on an image to obtain a target object in the image; performing panoramic segmentation on the image to obtain a scene in the image; obtaining an input text prompt, wherein the input text prompt is used to indicate a description of the image based on the target object and the scene; repeatedly inputting the input text prompt into at least one large language model to obtain a description text of the image; extracting text features of the description text; extracting image features of the image; calculating the similarity between the text features and the image features; and selecting the description text with the greatest similarity as the optimal description text of the image.
[0008] In an implementation manner of the first aspect, object detection is performed on an image based on a CO-Detr object detector, a YOLOV8 model, or a YOLOX model.
[0009] In an implementation manner of the first aspect, panoptic segmentation is performed on the image based on a Mask DINO segmentation model, a Oneformer segmentation model, or a mask2former segmentation model.
[0010] In an implementation of the first aspect, the at least one large language model includes one or more combinations of chatGPT, GPT-3, and GPT-4.
[0011] In an implementation of the first aspect, text features of the description text are extracted based on a CLIP text encoder.
[0012] In an implementation manner of the first aspect, image features of the image are extracted based on a CLIP image encoder.
[0013] In an implementation manner of the first aspect, the similarity adopts cosine similarity.
[0014] In a second aspect, the present invention provides an image description system, the system comprising a detection module, a segmentation module, an acquisition module, a processing module, a first extraction module, a second extraction module, a calculation module, and a description module;
[0015] The detection module is used to perform target detection on the image to obtain the target object in the image;
[0016] The segmentation module is used to perform panoramic segmentation on the image to obtain the scene in the image;
[0017] The acquisition module is used to acquire an input text prompt, where the input text prompt is used to indicate a description of the image according to the target object and the scene;
[0018] The processing module is used to repeatedly input the input text prompt into at least one large language model to obtain a description text of the image;
[0019] The first extraction module is used to extract text features of the description text;
[0020] The second extraction module is used to extract image features of the image;
[0021] The calculation module is used to calculate the similarity between the text feature and the image feature;
[0022] The description module is used to select the description text with the greatest similarity as the best description text for the image.
[0023] In a third aspect, the present invention provides an electronic device, comprising: a processor and a memory;
[0024] The memory is used to store computer programs;
[0025] The processor is configured to execute the computer program stored in the memory, so as to enable the electronic device to perform the above-mentioned image description method.
[0026] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned image description method when executed by an electronic device.
[0027] As described above, the image description method, system, storage medium, and electronic device of the present invention have the following beneficial effects:
[0028] (1) Ability to accurately describe images;
[0029] (2) It breaks through the limitation of the amount of training data and ensures the reliability of the image description model;
[0030] (3) High degree of intelligence, which greatly improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Shown is a schematic diagram of a scene of an electronic device in an embodiment of the present invention;
[0032] Figure 2 Shown is a flow chart of an image description method according to an embodiment of the present invention;
[0033] Figure 3 Shown is a schematic diagram of the architecture of an image description method according to an embodiment of the present invention;
[0034] Figure 4 Shown is a schematic structural diagram of an image description system according to an embodiment of the present invention;
[0035] Figure 5 FIG. 1 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.
[0037] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0038] The following embodiments of the present invention provide an image description method, which can be applied to Figure 1 The electronic device shown. The electronic device described in the present invention may include a mobile phone 11 with a wireless charging function, a tablet computer 12, a laptop computer 13, a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiment of the present invention does not impose any restrictions on the specific type of electronic device.
[0039] For example, the electronic device may be a station (STAION, ST) in a WLAN with a wireless charging function, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with a wireless charging function, a computing device or other processing device, a computer, a laptop computer, a handheld communication device, a handheld computing device, and / or other devices for communicating on a wireless system and a next-generation communication system, such as a mobile terminal in a 5G network, a mobile terminal in a future-evolved Public Land Mobile Network (PLMN), or a mobile terminal in a future-evolved Non-terrestrial Network (NTN).
[0040] For example, the electronic device can communicate with a network and other devices via wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS can include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS) and / or Satellite Based Augmentation Systems (SBAS).
[0041] The technical solutions in the embodiments of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0042] like Figure 2 and Figure 3 As shown, in one embodiment, the image description method of the present invention includes steps S1 to S8.
[0043] Step S1: perform target detection on an image to obtain a target object in the image.
[0044] Specifically, target detection is performed on images based on the CO-Detr target detector, the YOLOv8 model, or the YOLOX model. The CO-Detr target detector is an end-to-end target detector based on Transformer, which combines one-to-many label assignment and one-to-one set matching to improve the training efficiency and effectiveness of the encoder and decoder. The YOLOv8 model is a state-of-the-art target detection and instance segmentation model that provides multiple scales and resolutions based on improvements to YOLOv5. The YOLOX model is an open-source target detection model developed by Megvii that outperforms YOLOv5.
[0045] For example, by performing target detection on the image to be described, we can obtain the target objects that may appear in the image, [object 1, object 2, object 3...], such as [person, dog, car...].
[0046] Step S2: performing panorama segmentation on the image to obtain the scene in the image.
[0047] Specifically, panoptic segmentation is performed on the image based on a Mask DINO segmentation model, a Oneformer segmentation model, or a Mask2former segmentation model. The Mask DINO segmentation model is a unified Transformer-based object detection and segmentation framework. Building on the content-based query embedding, DINO has two branches for box prediction and label prediction. These boxes are dynamically updated and used to guide deformable attention in each Transformer decoder. Mask DINO adds another branch for mask prediction and minimally expands several key components of detection to accommodate segmentation. The Oneformer segmentation model is the first Transformer-based multi-task general image segmentation framework that achieves new state-of-the-art results on semantic, instance, and panoptic segmentation tasks with only a single training run. OneFormer uses task tokens, a task-conditioned joint training strategy, and a query-text contrastive loss to achieve inter-task and inter-class discrimination, and uses ConvNeXt and DiNAT backbones for further performance improvements. The Mask2former segmentation model is a universal image segmentation model that draws inspiration from the designs of Detr and DeformableDetr to unify semantic and instance segmentation tasks.
[0048] For example, by performing panoramic segmentation on the image to be described, possible scenes can be obtained, [scene 1, scene 2, scene 3...], such as [sky, grass, road...].
[0049] Step S3: Obtain an input text prompt, where the input text prompt is used to indicate a description of the image based on the target object and the scene.
[0050] Specifically, users provide input via prompts. A prompt is a piece of text or instruction that provides input to a model to guide it to produce a specific output. It is a user-provided paragraph of text when interacting with a model, describing the information, answer, or text they want from the model. The purpose of a prompt is to guide the model to produce a desired response, allowing for greater control over the generated output.
[0051] In one embodiment, the prompt reads: "I am an intelligent image description robot. I believe there are {N} people in the image, including {scene 1}, {scene 2}, {scene 3}, and so on. I also believe there may be {object 1}, {object 2}, and {object 3} in this image. I can generate a creative, short caption to describe this image." It should be noted that if no people are detected during the object detection process, the sentence "I believe there are {N} people in the image" is deleted. If so, the number of people is counted, assuming it is N.
[0052] Step S4: repeatedly input the input text prompt into at least one large language model to obtain a description text of the image.
[0053] Specifically, Large Language Models (LLMs) are AI models designed to understand and generate human language. They are trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. LLMs are characterized by their massive size, containing billions of parameters, which helps them learn complex patterns in language data. These models are often based on deep learning architectures such as Transformers, which contributes to their impressive performance on various NLP tasks.
[0054] In the present invention, the at least one large language model includes one or more combinations of chatGPT, GPT-3 and GPT-4. Among them, chatGPT is an artificial intelligence program based on a language model developed by OpenAI, which can interact with humans in natural language. It is built based on GPT (Generative Pre-trained Transformer) technology. The GPT model is an autoregressive language model based on a neural network. The model uses an architecture called "Transformer", which is a new sequence-to-sequence model that can avoid the gradient disappearance problem in traditional recurrent neural networks (RNN) when processing long sequence data. The key components of the Transformer architecture include multi-head attention mechanisms and residual connections. GPT uses the decoder part of the Transformer. GPT-3 is a language model that can learn from a small number of samples, and the number of model parameters is 175 billion. GPT-4 can accept image and text inputs and achieve "human level" performance on various professional and academic benchmarks.
[0055] Therefore, by inputting the input text prompt into the large language model, a description text of the image can be obtained. Repeating the input text prompt into each large language model N times will produce N description texts. If M large language models are provided, a total of N*M description texts can be obtained.
[0056] Step S5: extract text features of the description text.
[0057] Specifically, the text features of the descriptive text are extracted based on the CLIP text encoder. CLIP (Contrastive Language-Image Pre-training) is a pre-training model based on contrasting text-image pairs. The CLIP model consists of two parts, namely the text encoder (Text Encoder) and the image encoder (Image Encoder). The Text Encoder uses the Text Transformer model; the Image Encoder uses two models, one is the CNN-based ResNet (comparing ResNets with different numbers of layers), and the other is the Transformer-based ViT.
[0058] Step S6: extracting image features of the image.
[0059] Specifically, image features of the image are extracted based on a CLIP image encoder.
[0060] Step S7: Calculate the similarity between the text feature and the image feature.
[0061] Specifically, the cosine similarity between the text feature and the image feature is calculated.
[0062] Step S8: Select the description text with the greatest similarity as the best description text for the image.
[0063] Therefore, the above similarity selection method effectively improves the accuracy of image description and overcomes the impact of the amount of training data.
[0064] The protection scope of the image description method described in the embodiment of the present invention is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, reducing, or replacing steps in the prior art based on the principles of the present invention are included in the protection scope of the present invention.
[0065] An embodiment of the present invention further provides an image description system, which can implement the image description method described in the present invention. However, the implementation device of the image description system described in the present invention includes but is not limited to the structure of the image description system listed in this embodiment. Any structural deformation and replacement of the existing technology made according to the principles of the present invention are included in the protection scope of the present invention.
[0066] like Figure 4 As shown, in one embodiment, the image description system of the present invention includes a detection module 41 , a segmentation module 42 , an acquisition module 43 , a processing module 44 , a first extraction module 45 , a second extraction module 46 , a calculation module 47 and a description module 48 .
[0067] The detection module 41 is used to perform target detection on an image to obtain a target object in the image.
[0068] The segmentation module 42 is used to perform panoramic segmentation on the image to obtain the scene in the image.
[0069] The acquisition module 43 is connected to the detection module 41 and the segmentation module 42 and is used to acquire an input text prompt, where the input text prompt is used to indicate a description of the image according to the target object and the scene.
[0070] The processing module 44 is connected to the acquisition module 43 and is configured to repeatedly input the input text prompt into at least one large language model to acquire the description text of the image.
[0071] The first extraction module 45 is connected to the processing module 44 and is used to extract text features of the description text.
[0072] The second extraction module 46 is used to extract image features of the image.
[0073] The calculation module 47 is connected to the first extraction module 45 and the second extraction module 46 and is used to calculate the similarity between the text feature and the image feature.
[0074] The description module 48 is connected to the calculation module 47 and is used to select the description text with the greatest similarity as the best description text of the image.
[0075] Among them, the structures and principles of the detection module 41, the segmentation module 42, the acquisition module 43, the processing module 44, the first extraction module 45, the second extraction module 46, the calculation module 47 and the description module 48 correspond one-to-one to the steps in the above-mentioned image description method, so they are not repeated here.
[0076] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices or methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules / units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules or units, which can be electrical, mechanical or other forms.
[0077] Modules / units described as separate components may or may not be physically separate, and components displayed as modules / units may or may not be physical modules, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules / units may be selected based on actual needs to achieve the objectives of the embodiments of the present invention. For example, the functional modules / units in various embodiments of the present invention may be integrated into a single processing module, each module / unit may exist physically separately, or two or more modules / units may be integrated into a single module / unit.
[0078] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0079] The embodiment of the present invention also provides a computer-readable storage medium. A person skilled in the art will understand that all or part of the steps in the method for implementing the above embodiment can be completed by instructing a processor through a program, and the program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as a random access memory, a read-only memory, a flash memory, a hard disk, a solid-state drive, a magnetic tape, a floppy disk, an optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid-state drive (SSD)), etc.
[0080] An embodiment of the present invention further provides an electronic device comprising a processor and a memory.
[0081] The memory is used to store computer programs.
[0082] The memory includes various media that can store program codes, such as ROM, RAM, magnetic disk, USB flash drive, memory card or optical disk.
[0083] The processor is connected to the memory and is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned image description method.
[0084] Preferably, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0085] like Figure 5As shown, the electronic device of the present invention is in the form of a general-purpose computing device. Components of the electronic device may include, but are not limited to: one or more processors or processing units 51, a memory 52, and a bus 53 connecting different system components (including the memory 52 and the processing unit 51).
[0086] Bus 53 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0087] Electronic devices typically include a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, removable and non-removable media.
[0088] The memory 52 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 521 and / or cache memory 522. The electronic device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 523 may be used to read and write non-removable, non-volatile magnetic media ( Figure 5 Not shown, often called a "hard drive"). Although Figure 5 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 53 via one or more data medium interfaces. Memory 52 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.
[0089] A program / utility 524 having a set (at least one) of program modules 5241 may be stored, for example, in memory 52. Such program modules 5241 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 5241 generally implement the functions and / or methods of the embodiments described herein.
[0090] The electronic device may also communicate with one or more external devices (e.g., keyboards, pointing devices, displays, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via input / output (I / O) interface 54. Furthermore, the electronic device may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via network adapter 55. Figure 5 As shown, the network adapter 55 communicates with other modules of the electronic device via the bus 53. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0091] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.
Claims
1. An image description method, characterized in that: The method comprises the following steps: Performing target detection on the image to obtain the target object in the image; Performing panorama segmentation on the image to obtain a scene in the image; Obtaining an input text prompt, wherein the input text prompt is used to indicate a description of the image according to the target object and the scene; repeatedly inputting the input text prompt into at least one large language model to obtain a description text of the image; Extracting text features of the description text; extracting image features of the image; Calculating the similarity between the text feature and the image feature; The description text with the greatest similarity is selected as the best description text for the image.
2. The image description method according to claim 1, wherein: Perform object detection on images based on the CO-Detr object detector, YOLOV8 model, or YOLOX model.
3. The image description method according to claim 1, wherein: Perform panoptic segmentation on the image based on a Mask DINO segmentation model, a Oneformer segmentation model, or a mask2former segmentation model.
4. The image description method according to claim 1, wherein: The at least one large language model includes one or more combinations of chatGPT, GPT-3, and GPT-4.
5. The image description method according to claim 1, wherein: The text features of the description text are extracted based on the CLIP text encoder.
6. The image description method according to claim 1, wherein: Image features of the image are extracted based on a CLIP image encoder.
7. The image description method according to claim 1, wherein: The similarity adopts cosine similarity.
8. An image description system, characterized in that: The system includes a detection module, a segmentation module, an acquisition module, a processing module, a first extraction module, a second extraction module, a calculation module and a description module; The detection module is used to perform target detection on the image to obtain the target object in the image; The segmentation module is used to perform panoramic segmentation on the image to obtain the scene in the image; The acquisition module is used to acquire an input text prompt, where the input text prompt is used to indicate a description of the image according to the target object and the scene; The processing module is used to repeatedly input the input text prompt into at least one large language model to obtain a description text of the image; The first extraction module is used to extract text features of the description text; The second extraction module is used to extract image features of the image; The calculation module is used to calculate the similarity between the text feature and the image feature; The description module is used to select the description text with the greatest similarity as the best description text for the image.
9. An electronic device, characterized in that: The electronic device includes: a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the computer program stored in the memory, so as to enable the electronic device to perform the image description method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by an electronic device, the image description method according to any one of claims 1 to 7 is implemented.