Method and device for generating image text description
By building an image description generation model based on Transformer and multi-layer perceptron, combining visual packet networks and GPT networks, the problem of insufficient feature details and semantic information in image text descriptions is solved, and the quality of text descriptions is improved.
Patent Information
- Application Number
- CN202510193712.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the image text description method has problems such as insufficient description of feature details and insufficient carrying semantic information.
Semantic feature extraction branches are built using Transformer and multi-layer perception mechanism, and visual packet network is built using similarity calculation formulas and attention mechanisms. Visual packet network is built using random vector generation network, Transformer, visual packet network, average pooling layer and multi-layer perception mechanism. The image description generation model is built using linear layer, semantic feature extraction branches, visual packet branches, multi-layer perceptrons and GPT networks, and the image description generation model is built and trained to generate target text descriptions.
The quality of generated text descriptions is improved, and the problems of insufficient description of feature details and insufficient semantic information are solved.
Smart Images

Figure CN120355968A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method and apparatus for generating image text descriptions. Background Art
[0002] Currently, there are two mainstream methods for Image Captioning in Chinese (ICC). One is a method where the encoder uses a convolutional neural network and a FastText word embedding model, and the decoder uses a recurrent neural network. To a certain extent, the model of this method improves the description quality, but there are also problems. For example, during the encoding process, the model processes the image as a fixed-length vector, resulting in insufficient image information carried during the decoding process, thus causing a certain deviation between the generated description sentences and the actual content of the image.
[0003] Another model that combines visual attention and text attention extracts the visual attention and topic feature vectors of the input image by using a CNN. Among them, visual attention is used to reduce the deviation between the description sentences and visual content, and the topic features are used to maintain sentence accuracy and ensure description diversity. To a certain extent, the model of this method solves the problem of information loss, but it still does not extract sufficient detailed information in the image. Summary of the Invention
[0004] In view of this, embodiments of the present disclosure provide a method, apparatus, electronic device, and computer-readable storage medium for generating image text descriptions to solve the problems of insufficient description of feature detail information and insufficient semantic information carried in the prior art.
[0005] In a first aspect of the embodiments of the present disclosure, a method for generating an image text description is provided, including: constructing a semantic feature extraction branch by using a Transformer and a multi-layer perceptron; constructing a visual grouping network by using a similarity calculation formula and an attention mechanism, and constructing a visual grouping branch by using a random vector generation network, a Transformer, the visual grouping network, a Transformer, an average pooling layer, and a multi-layer perceptron; constructing an image description generation model by using a linear layer, the semantic feature extraction branch, the visual grouping branch, a multi-layer perceptron, and a GPT network; training the image description generation model by using training images, and generating a target text description of a target image by using the trained image description generation model.
[0006] In a second aspect of the embodiments of the present disclosure, there is provided an apparatus for generating an image text description, including: a first construction module configured to construct a semantic feature extraction branch by using a Transformer and a multi-layer perceptron; a second construction module configured to construct a visual grouping network by using a similarity calculation formula and an attention mechanism; a third construction module configured to construct a visual grouping branch by using a random vector generation network, a Transformer, the visual grouping network, a Transformer, an average pooling layer and a multi-layer perceptron; a fourth construction module configured to construct an image description generation model by using a linear layer, the semantic feature extraction branch, the visual grouping branch, a multi-layer perceptron and a GPT network; a training module configured to train the image description generation model by using training images, and generate a target text description of a target image by using the trained image description generation model.
[0007] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps of the above method when executing the computer program.
[0008] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, where the computer program implements the steps of the above method when executed by a processor.
[0009] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: constructing a semantic feature extraction branch by using a Transformer and a multi-layer perceptron; constructing a visual grouping network by using a similarity calculation formula and an attention mechanism; constructing a visual grouping branch by using a random vector generation network, a Transformer, the visual grouping network, a Transformer, an average pooling layer and a multi-layer perceptron; constructing an image description generation model by using a linear layer, the semantic feature extraction branch, the visual grouping branch, a multi-layer perceptron and a GPT network; training the image description generation model by using training images, and generating a target text description of a target image by using the trained image description generation model. By adopting the above technical means, the problems of insufficient description of feature detail information and insufficient semantic information carried in the prior art can be solved, and further the quality of the generated text description can be improved. Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0011] Figure 1It is a schematic flowchart of a method for generating an image text description provided by an embodiment of the present disclosure;
[0012] Figure 2 It is a schematic flowchart of another method for generating an image text description provided by an embodiment of the present disclosure;
[0013] Figure 3 It is a schematic structural diagram of an apparatus for generating an image text description provided by an embodiment of the present disclosure;
[0014] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0015] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art should understand that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present disclosure with unnecessary details.
[0016] Next, a method and an apparatus for generating an image text description according to an embodiment of the present disclosure will be described in detail with reference to the accompanying drawings.
[0017] Figure 1 It is a schematic flowchart of a method for generating an image text description provided by an embodiment of the present disclosure. Figure 1 The method for generating an image text description may be executed by a computer or a server, or software on a computer or a server. As Figure 1 shown, the method for generating an image text description includes:
[0018] S101, constructing a semantic feature extraction branch by using a Transformer and a multi-layer perceptron;
[0019] S102, constructing a visual grouping network by using a similarity calculation formula and an attention mechanism;
[0020] S103, constructing a visual grouping branch by using a random vector generation network, a Transformer, a visual grouping network, a Transformer, an average pooling layer, and a multi-layer perceptron;
[0021] S104, constructing an image description generation model by using a linear layer, a semantic feature extraction branch, a visual grouping branch, a multi-layer perceptron, and a GPT network;
[0022] S105, training the image description generation model by using training images, and generating a target text description of a target image by using the trained image description generation model.
[0023] Connect the Transformer and the multi-layer perceptron in sequence to obtain the semantic feature extraction branch. Construct a visual grouping network using a similarity calculation formula and an attention mechanism. The visual grouping network calculates the similarity between features using cosine similarity and performs matching through a self-attention mechanism. Connect the random vector generation network, Transformer, visual grouping network, Transformer, average pooling layer, and multi-layer perceptron in sequence to obtain the visual grouping branch. The random vector generation network is used to generate randomly initialized grouping vectors. Connect the linear layer, semantic feature extraction branch, visual grouping branch, multi-layer perceptron, and GPT network in sequence to obtain an image description generation model. Finally, use the training images to train the image description generation model. After training, use the image description generation model to generate the target text description of the target image. The target image is the image for which the text description is to be generated.
[0024] GPT (Generative Pre-trained Transformer) is a large language model developed by OpenAI, based on the Transformer architecture. The target text description is the text description of the content of the target image.
[0025] According to the technical solution provided by the embodiments of the present application, construct a semantic feature extraction branch using the Transformer and the multi-layer perceptron; construct a visual grouping network using a similarity calculation formula and an attention mechanism, and construct a visual grouping branch using the random vector generation network, Transformer, visual grouping network, Transformer, average pooling layer, and multi-layer perceptron; construct an image description generation model using the linear layer, semantic feature extraction branch, visual grouping branch, multi-layer perceptron, and GPT network; use the training images to train the image description generation model, and use the trained image description generation model to generate the target text description of the target image. By adopting the above technical means, the problems of insufficient description of feature detail information and insufficient semantic information carried in the prior art can be solved, thereby improving the quality of the generated text description.
[0026] Further, using the training images to train the image description generation model includes: inputting the training images into the image description generation model: extracting the image features of the training images through the linear layer; processing the image features through the semantic feature extraction branch to obtain image semantic features; processing the image features through the visual grouping branch to obtain image visual features; processing the image semantic features and image visual features through the multi-layer perceptron to obtain image mapping features; generating the training text description of the training images through the GPT network based on the image mapping features; calculating the loss between the training text description and the labels of the training images; and optimizing the model parameters of the image description generation model according to the loss.
[0027] Inside the image description generation model: The linear layer extracts the image features of the training images. The semantic feature extraction branch processes the image features to obtain the image semantic features. The visual grouping branch processes the image features to obtain the image visual features. The multi-layer perceptron processes the image semantic features and the image visual features to obtain the image mapping features. The GPT network generates the training text description of the training images based on the image mapping features. The cross-entropy loss function can be used to calculate the loss between the training text description and the labels of the training images. The labels of the training images are the text descriptions of the training images pre-annotated. The model parameters of the image description generation model are optimized according to the loss, and the training of the image description generation model is completed.
[0028] Furthermore, by processing the image features through the semantic feature extraction branch, the image semantic features are obtained, including: Inside the semantic feature extraction branch: The image features are processed by the Transformer to obtain the first feature; The first feature is processed by the multi-layer perceptron to obtain the image semantic features.
[0029] Inside the semantic feature extraction branch: The Transformer processes the image features to obtain the first feature. The multi-layer perceptron processes the first feature to obtain the image semantic features.
[0030] Furthermore, by processing the image features through the visual grouping branch, the image visual features are obtained, including: Inside the visual grouping branch: The random vector generation network generates the randomly initialized grouping vectors; The first Transformer processes the randomly initialized grouping vectors and the image features respectively to obtain the random features and the second feature; The visual grouping network processes the random features and the second feature to obtain the third feature; The second Transformer processes the third feature to obtain the fourth feature; The fourth feature is processed by the average pooling layer to obtain the fifth feature; The fifth feature is processed by the multi-layer perceptron to obtain the image visual features.
[0031] Inside the visual grouping branch: The random vector generation network generates the randomly initialized grouping vectors. The first Transformer in the visual grouping branch processes the randomly initialized grouping vectors to obtain the random features and processes the image features to obtain the second feature. The visual grouping network processes the random features and the second feature to obtain the third feature. The similarity between the random features and the second feature is calculated in the visual grouping network, and then the random features and the second feature are allocated according to the similarity (this part is completed by the attention mechanism). It should be noted that both the random features and the second feature have multiple. The second Transformer in the visual grouping branch processes the third feature to obtain the fourth feature. The average pooling layer processes the fourth feature to obtain the fifth feature. The multi-layer perceptron processes the fifth feature to obtain the image visual features.
[0032] Further, the image semantic features and the image visual features are processed by a multi-layer perceptron to obtain image mapping features, including: adding the image semantic features and the image visual features to obtain a sixth feature; processing the sixth feature by a multi-layer perceptron to obtain image mapping features.
[0033] The sixth feature is obtained by adding the image semantic features and the image visual features, and the image mapping features are obtained by processing the sixth feature by a multi-layer perceptron. The multi-layer perceptron realizes the function of feature mapping.
[0034] Figure 2 It is a schematic flowchart of another method for generating an image text description provided by an embodiment of the present disclosure. As Figure 2 shown, the method includes:
[0035] S201, inputting a target image into an image description generation model:
[0036] S202, extracting target features of the target image through a linear layer;
[0037] S203, processing the target features through a semantic feature extraction branch to obtain target semantic features;
[0038] S204, processing the target features through a visual grouping branch to obtain target visual features;
[0039] S205, processing the target semantic features and the target visual features through a multi-layer perceptron to obtain target mapping features;
[0040] S206, generating a target text description of the target image based on the target mapping features through a GPT network.
[0041] Inside the image description generation model: The linear layer extracts the target features of the target image. The semantic feature extraction branch processes the target features to obtain target semantic features. The visual grouping branch processes the target features to obtain target visual features. The multi-layer perceptron processes the target semantic features and the target visual features to obtain target mapping features. The GPT network generates a target text description of the target based on the target mapping features.
[0042] Further, processing the target features through a semantic feature extraction branch to obtain target semantic features, including: Inside the semantic feature extraction branch: processing the target features through a Transformer to obtain a seventh feature; processing the seventh feature through a multi-layer perceptron to obtain target semantic features.
[0043] Further, the target visual features are processed through the visual grouping branch, including: inside the visual grouping branch: a randomly initialized grouping vector is generated through a random vector generation network; the randomly initialized grouping vector and the target features are respectively processed through a first Transformer to obtain random features and eighth features; the random features and the eighth features are processed through a visual grouping network to obtain ninth features; the ninth features are processed through a second Transformer to obtain tenth features; the tenth features are processed through an average pooling layer to obtain eleventh features; the eleventh features are processed through a multi-layer perceptron to obtain the target visual features.
[0044] Further, the target mapping features are obtained by processing the target semantic features and the target visual features through a multi-layer perceptron, including: adding the target semantic features and the target visual features to obtain twelfth features; the twelfth features are processed through a multi-layer perceptron to obtain the target mapping features.
[0045] Further, an image description generation model is constructed using a linear layer, a semantic feature extraction branch, a visual grouping branch, a feature fusion network, a multi-layer perceptron, and a GPT network.
[0046] The feature fusion network (Feature Fusion Network, FFN) is used to fuse semantic features and visual features. The multi-layer perceptron processes the fusion result to obtain mapping features.
[0047] All the above optional technical solutions can be combined arbitrarily to form alternative embodiments of the present application, which will not be elaborated here one by one.
[0048] The following is an embodiment of the apparatus of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For details not disclosed in the embodiment of the apparatus of the present disclosure, please refer to the embodiment of the method of the present disclosure.
[0049] Figure 3 It is a schematic diagram of an image text description generation apparatus provided by an embodiment of the present disclosure. As Figure 3 shown, the image text description generation apparatus includes:
[0050] A first construction module 301, configured to construct a semantic feature extraction branch using a Transformer and a multi-layer perceptron;
[0051] A second construction module 302, configured to construct a visual grouping network using a similarity calculation formula and an attention mechanism,
[0052] A third construction module 303, configured to construct a visual grouping branch using a random vector generation network, a Transformer, a visual grouping network, a Transformer, an average pooling layer, and a multi-layer perceptron;
[0053] The fourth building block 304 is configured to construct an image caption generation model by using a linear layer, a semantic feature extraction branch, a visual grouping branch, a multi-layer perceptron, and a GPT network;
[0054] The training module 305 is configured to train the image caption generation model by using training images, and generate a target text description of the target image by using the trained image caption generation model.
[0055] Connect the Transformer and the multi-layer perceptron in sequence to obtain the semantic feature extraction branch. Construct a visual grouping network by using a similarity calculation formula and an attention mechanism. The visual grouping network calculates the similarity between features by using cosine similarity and performs matching through a self-attention mechanism. Connect the random vector generation network, the Transformer, the visual grouping network, the Transformer, the average pooling layer, and the multi-layer perceptron in sequence to obtain the visual grouping branch. The random vector generation network is used to generate randomly initialized grouping vectors. Connect the linear layer, the semantic feature extraction branch, the visual grouping branch, the multi-layer perceptron, and the GPT network in sequence to obtain the image caption generation model. Finally, use the training images to train the image caption generation model. After training, use the image caption generation model to generate a target text description of the target image. The target image is the image for which the text description is to be generated.
[0056] GPT (Generative Pre-trained Transformer) is a large language model developed by OpenAI, based on the Transformer architecture. The target text description is the text description of the content of the target image.
[0057] According to the technical solution provided by the embodiments of the present application, construct a semantic feature extraction branch by using a Transformer and a multi-layer perceptron; construct a visual grouping network by using a similarity calculation formula and an attention mechanism, and construct a visual grouping branch by using a random vector generation network, a Transformer, a visual grouping network, a Transformer, an average pooling layer, and a multi-layer perceptron; construct an image caption generation model by using a linear layer, a semantic feature extraction branch, a visual grouping branch, a multi-layer perceptron, and a GPT network; train the image caption generation model by using training images, and generate a target text description of the target image by using the trained image caption generation model. By adopting the above technical means, the problems of insufficient description of feature detail information and insufficient semantic information carried in the prior art can be solved, and thus the quality of the generated text description can be improved.
[0058] In some embodiments, the training module 305 is further configured to input the training image into an image description generation model: extract the image features of the training image through a linear layer; process the image features through a semantic feature extraction branch to obtain image semantic features; process the image features through a visual grouping branch to obtain image visual features; process the image semantic features and the image visual features through a multi-layer perceptron to obtain image mapping features; generate a training text description of the training image based on the image mapping features through a GPT network; calculate the loss between the training text description and the label of the training image; and optimize the model parameters of the image description generation model according to the loss.
[0059] Inside the image description generation model: The linear layer extracts the image features of the training image. The semantic feature extraction branch processes the image features to obtain image semantic features. The visual grouping branch processes the image features to obtain image visual features. The multi-layer perceptron processes the image semantic features and the image visual features to obtain image mapping features. The GPT network generates a training text description of the training image based on the image mapping features. The loss between the training text description and the label of the training image can be calculated using the cross-entropy loss function. The label of the training image is the pre-annotated text description of the training image. The model parameters of the image description generation model are optimized according to the loss, and the training of the image description generation model is completed.
[0060] In some embodiments, the training module 305 is further configured to, inside the semantic feature extraction branch: process the image features through a Transformer to obtain first features; process the first features through a multi-layer perceptron to obtain image semantic features.
[0061] Inside the semantic feature extraction branch: The Transformer processes the image features to obtain first features. The multi-layer perceptron processes the first features to obtain image semantic features.
[0062] In some embodiments, the training module 305 is further configured to, inside the visual grouping branch: generate a randomly initialized grouping vector through a random vector generation network; process the randomly initialized grouping vector and the image features respectively through a first Transformer to obtain random features and second features; process the random features and the second features through a visual grouping network to obtain third features; process the third features through a second Transformer to obtain fourth features; process the fourth features through an average pooling layer to obtain fifth features; process the fifth features through a multi-layer perceptron to obtain image visual features.
[0063] Inside the visual grouping branch: The random vector generation network generates randomly initialized grouping vectors. The first Transformer in the visual grouping branch processes the randomly initialized grouping vectors to obtain random features and processes the image features to obtain second features. The visual grouping network processes the random features and the second features to obtain third features. The visual grouping network calculates the similarity between the random features and the second features, and then assigns the random features and the second features according to the similarity (this part is completed by the attention mechanism). It should be noted that there are multiple random features and second features. The second Transformer in the visual grouping branch processes the third features to obtain fourth features. The average pooling layer processes the fourth features to obtain fifth features. The multi-layer perceptron processes the fifth features to obtain the image visual features.
[0064] In some embodiments, the training module 305 is further configured to add the image semantic features and the image visual features to obtain a sixth feature; and process the sixth feature through a multi-layer perceptron to obtain the image mapping feature.
[0065] The image semantic features and the image visual features are added to obtain a sixth feature, and the multi-layer perceptron processes the sixth feature to obtain the image mapping feature. This multi-layer perceptron realizes the function of feature mapping.
[0066] In some embodiments, the training module 305 is further configured to input the target image into the image description generation model: extract the target features of the target image through a linear layer; process the target features through the semantic feature extraction branch to obtain the target semantic features; process the target features through the visual grouping branch to obtain the target visual features; process the target semantic features and the target visual features through a multi-layer perceptron to obtain the target mapping feature; and generate the target text description of the target image based on the target mapping feature through the GPT network.
[0067] Inside the image description generation model: The linear layer extracts the target features of the target image. The semantic feature extraction branch processes the target features to obtain the target semantic features. The visual grouping branch processes the target features to obtain the target visual features. The multi-layer perceptron processes the target semantic features and the target visual features to obtain the target mapping feature. The GPT network generates the target text description of the target based on the target mapping feature.
[0068] In some embodiments, the training module 305 is further configured to, inside the semantic feature extraction branch: process the target features through a Transformer to obtain a seventh feature; and process the seventh feature through a multi-layer perceptron to obtain the target semantic features.
[0069] In some embodiments, the training module 305 is further configured to, within the visual grouping branch: generate a randomly initialized grouping vector through a random vector generation network; process the randomly initialized grouping vector and the target feature respectively through a first Transformer to obtain a random feature and an eighth feature; process the random feature and the eighth feature through a visual grouping network to obtain a ninth feature; process the ninth feature through a second Transformer to obtain a tenth feature; process the tenth feature through an average pooling layer to obtain an eleventh feature; and process the eleventh feature through a multi-layer perceptron to obtain a target visual feature.
[0070] In some embodiments, the training module 305 is further configured to add the target semantic feature and the target visual feature to obtain a twelfth feature; and process the twelfth feature through a multi-layer perceptron to obtain a target mapping feature.
[0071] In some embodiments, the training module 305 is further configured to construct an image description generation model by using a linear layer, a semantic feature extraction branch, a visual grouping branch, a feature fusion network, a multi-layer perceptron, and a GPT network.
[0072] The feature fusion network (Feature Fusion Network, FFN) is used to fuse semantic features and visual features. The multi-layer perceptron processes the fusion result to obtain a mapping feature.
[0073] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0074] Figure 4 It is a schematic diagram of the electronic device 4 provided by the embodiments of the present disclosure. As Figure 4 shown, the electronic device 4 in this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above device embodiments are implemented.
[0075] The electronic device 4 may be a desktop computer, a notebook, a palm computer, a cloud server, and other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 this is only an example of the electronic device 4 and does not constitute a limitation to the electronic device 4. It may include more or fewer components than shown in the figure, or different components.
[0076] The processor 401 may be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0077] The memory 402 may be an internal storage unit of the electronic device 4. For example, the hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4. For example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 4. The memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.
[0078] Those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, only the above-mentioned functional units and modules are divided for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. In the embodiments, each functional unit and module can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in hardware form or in the form of software functional units.
[0079] When an integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0080] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included within the protection scope of the present disclosure.
Claims
1. A method for generating an image text description, characterized in that, Including: Construct a semantic feature extraction branch using a Transformer and a multi-layer perceptron; Construct a visual grouping network using a similarity calculation formula and an attention mechanism; Construct a visual grouping branch using a random vector generation network, a Transformer, a visual grouping network, a Transformer, an average pooling layer, and a multi-layer perceptron; Construct an image description generation model using a linear layer, a semantic feature extraction branch, a visual grouping branch, a multi-layer perceptron, and a GPT network; Train the image description generation model using training images, and use the trained image description generation model to generate a target text description of the target image.
2. The method according to claim 1, wherein Training the image description generation model using training images includes: Input the training images into the image description generation model: Extract the image features of the training images through the linear layer; Process the image features through the semantic feature extraction branch to obtain image semantic features; Process the image features through the visual grouping branch to obtain image visual features; Process the image semantic features and the image visual features through the multi-layer perceptron to obtain image mapping features; Generate a training text description of the training images through the GPT network based on the image mapping features; Calculate the loss between the training text description and the labels of the training images; Optimize the model parameters of the image description generation model according to the loss.
3. The method according to claim 2, wherein Processing the image features through the semantic feature extraction branch to obtain image semantic features includes: Inside the semantic feature extraction branch: Process the image features through a Transformer to obtain a first feature; Process the first feature through a multi-layer perceptron to obtain the image semantic features.
4. The method according to claim 2, wherein Processing the image features through the visual grouping branch to obtain image visual features includes: Inside the visual grouping branch: Generate a randomly initialized grouping vector through a random vector generation network; Process the randomly initialized grouping vector and the image features respectively through the first Transformer to obtain a random feature and a second feature; Process the random feature and the second feature through the visual grouping network to obtain a third feature; Process the third feature through the second Transformer to obtain a fourth feature; Process the fourth feature through an average pooling layer to obtain a fifth feature; Process the fifth feature through a multi-layer perceptron to obtain the image visual features.
5. The method according to claim 2, wherein Processing the image semantic features and the image visual features through the multi-layer perceptron to obtain image mapping features includes: Add the image semantic features and the image visual features to obtain a sixth feature; Process the sixth feature through the multi-layer perceptron to obtain the image mapping features.
6. The method according to claim 1, characterized in that Generating a target text description of the target image using the trained image description generation model includes: Input the target image into the image description generation model: Extract the target features of the target image through the linear layer; Process the target features through the semantic feature extraction branch to obtain target semantic features; Process the target feature through the visual grouping branch to obtain a target visual feature; Process the target semantic feature and the target visual feature through the multi-layer perceptron to obtain a target mapping feature; Generate a target text description of the target image through the GPT network based on the target mapping feature.
7. The method according to claim 1, wherein Construct an image description generation model using a linear layer, a semantic feature extraction branch, a visual grouping branch, a feature fusion network, a multi-layer perceptron, and a GPT network.
8. An apparatus for generating an image text description, characterized in that, Including: A first construction module configured to construct a semantic feature extraction branch using a Transformer and a multi-layer perceptron; A second construction module configured to construct a visual grouping network using a similarity calculation formula and an attention mechanism; A third construction module configured to construct a visual grouping branch using a random vector generation network, a Transformer, a visual grouping network, a Transformer, an average pooling layer, and a multi-layer perceptron; A fourth construction module configured to construct an image description generation model using a linear layer, a semantic feature extraction branch, a visual grouping branch, a multi-layer perceptron, and a GPT network; A training module configured to train the image description generation model using training images and generate a target text description of a target image using the trained image description generation model.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.