A method for determining image description text and related equipment
By extracting visual feature and integrating information representation of images, image description text is generated, and the technical difficulties of image description text are solved, and high-accurate image description is achieved.
Patent Information
- Application Number
- CN202111294586.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-03
AI Technical Summary
How to accurately determine the explanatory text of an image data to solve the technical challenges that the problem in the prior art has not been effectively solved.
By extracting visual features of the image to be understood, visual information representation data and reference information representation data are determined respectively, and text generation is performed after fusion processing to generate image description text.
Accurate description of image data is achieved, the accuracy of image description text is improved, and the information carried by the image can be more accurately represented.
Smart Images

Figure CN114021646B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method for determining image description text and related equipment. Background Art
[0002] In some image application fields, there may be the following image processing requirements: using a paragraph of text to describe an image data. For example, you can use the paragraph "There are relatively small waves in the sea, and there is an empty chair next to the sea" to describe Figure 1 The image data shown.
[0003] However, how to determine the caption of an image data is still a technical problem that needs to be solved urgently. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides an image description text determination method and related equipment, which can accurately determine the description text of an image data.
[0005] In order to achieve the above purpose, the technical solutions provided by the embodiments of the present application are as follows:
[0006] The present application provides a method for determining image description text, the method comprising:
[0007] After acquiring the image to be understood, performing visual feature extraction processing on the image to be understood to obtain the visual features of the image to be understood;
[0008] Determining visual information representation data of the image to be understood and reference information representation data of the image to be understood respectively according to the visual features;
[0009] fusing the visual information representation data and the reference information representation data to obtain image information representation data of the image to be understood;
[0010] The image information representation data is processed for text generation to obtain an image description text of the image to be understood.
[0011] In a possible implementation, the reference information representation data includes at least one of image classification result representation data, image target detection result representation data, and image segmentation result representation data; wherein the image classification result representation data is used to represent the image classification result of the image to be understood; the image target detection result representation data is used to represent the image target detection result of the image to be understood; and the image segmentation result representation data is used to represent the image segmentation result of the image to be understood.
[0012] In a possible implementation, the process of determining the image classification result representation data includes:
[0013] Inputting the visual features into a pre-built classification representation model to obtain image classification result representation data of the image to be understood output by the classification representation model;
[0014] The process of determining the image target detection result representation data includes:
[0015] Inputting the visual features into a pre-built target detection representation model to obtain image target detection result representation data of the image to be understood output by the target detection representation model;
[0016] The process of determining the image segmentation result representation data includes:
[0017] The visual features are input into a pre-built segmentation representation model to obtain image segmentation result representation data of the image to be understood output by the segmentation representation model.
[0018] In a possible implementation, the process of determining the visual information representation data includes:
[0019] The visual features are input into a pre-built visual information representation model to obtain visual information representation data of the image to be understood output by the visual information representation model.
[0020] In a possible implementation manner, the performing text generation processing on the image information representation data to obtain the image description text of the image to be understood includes:
[0021] The image information representation data is input into a pre-built text generation model to obtain an image description text of the image to be understood output by the text generation model.
[0022] In a possible implementation, the process of constructing the text generation model includes:
[0023] Obtaining a sample image and actual description text of the sample image;
[0024] Inputting the sample image into the model to be trained to obtain the predicted description text of the sample image output by the model to be trained;
[0025] According to the predicted description text of the sample image and the actual description text of the sample image, the model to be trained is updated, and the step of inputting the sample image into the model to be trained is continued until a preset stop condition is reached, and the text generation model is determined according to the model to be trained.
[0026] In a possible implementation manner, the process of obtaining the actual description text includes:
[0027] According to the label information of the sample image, an actual description text of the sample image is determined.
[0028] In a possible implementation, the label information includes a to-be-used label, the to-be-used label includes an image classification label, an image object detection label, or an image segmentation label, and determining the actual description text of the sample image according to the label information of the sample image includes:
[0029] The to-be-used label of the sample image is subjected to text conversion processing to obtain an actual description text of the sample image.
[0030] In a possible implementation, the label information includes a plurality of reference labels, the plurality of reference labels include at least two of an image classification label, an image object detection label, and an image segmentation label, and determining the actual description text of the sample image according to the label information of the sample image includes:
[0031] Performing text conversion processing on each of the reference tags respectively to obtain description text of each of the reference tags; determining the actual description text of the sample image according to the description texts of the multiple reference tags and a preset text generation template.
[0032] In a possible implementation, the label information includes an image description text label and at least one reference label, the at least one reference label includes at least one of an image classification label, an image object detection label, and an image segmentation label, and determining the actual description text of the sample image according to the label information of the sample image includes:
[0033] Perform text conversion processing on each of the reference tags to obtain description text of each of the reference tags; determine the actual description text of the sample image based on the image description text tag, the description text of at least one reference tag, and a pre-set text generation template.
[0034] In a possible implementation manner, performing text conversion processing on the to-be-used label of the sample image to obtain the actual description text of the sample image includes:
[0035] According to the text conversion rule corresponding to the tag to be used, the tag to be used of the sample image is subjected to text conversion processing to obtain the actual description text of the sample image; wherein the text conversion rule corresponding to the tag to be used is determined according to the tag type of the tag to be used.
[0036] In a possible implementation, the model to be trained includes a visual feature extraction layer, a visual information representation layer, a reference information representation layer, a fusion layer, and a text generation layer; the input data of the text generation layer includes the output data of the fusion layer; the input data of the fusion layer includes the output data of the reference information representation layer and the output data of the visual information representation layer; the input data of the reference information representation layer includes the output data of the visual feature extraction layer; the input data of the visual information representation layer includes the output data of the visual feature extraction layer;
[0037] Determining the text generation model according to the model to be trained includes:
[0038] The text generation model is determined according to the text generation layer in the model to be trained.
[0039] In a possible implementation, the step of inputting the sample image into a model to be trained to obtain a predicted description text of the sample image output by the model to be trained includes:
[0040] The sample image and the label type representation data of the sample image are input into the model to be trained to obtain the predicted description text of the sample image output by the model to be trained; wherein the label type representation data of the sample image is determined based on the label information of the sample image.
[0041] In a possible implementation manner, the performing visual feature extraction processing on the image to be understood to obtain the visual features of the image to be understood includes:
[0042] Performing image feature extraction processing on the image to be understood to obtain image features of the image to be understood;
[0043] Perform visual encoding processing on the image features to obtain visual features of the image to be understood.
[0044] The present application also provides a device for determining image description text, including:
[0045] an extraction unit, configured to, after acquiring the image to be understood, perform visual feature extraction processing on the image to be understood to obtain the visual features of the image to be understood;
[0046] a determining unit, configured to determine, according to the visual features, visual information representation data of the image to be understood and reference information representation data of the image to be understood;
[0047] A fusion unit, configured to fuse the visual information representation data and the reference information representation data to obtain image information representation data of the image to be understood;
[0048] The generating unit is used to perform text generation processing on the image information representation data to obtain the image description text of the image to be understood.
[0049] The present application also provides a device, the device comprising a processor and a memory:
[0050] The memory is used to store computer programs;
[0051] The processor is used to execute any implementation of the image description text determination method provided in the embodiments of the present application according to the computer program.
[0052] The embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute any implementation of the image description text determination method provided in the embodiment of the present application.
[0053] The embodiment of the present application also provides a computer program product. When the computer program product is executed on a terminal device, the terminal device executes any implementation of the method for determining image description text provided in the embodiment of the present application.
[0054] Compared with the prior art, the embodiments of the present application have at least the following advantages:
[0055] In the technical solution provided by the embodiment of the present application, after obtaining the image to be understood, the image to be understood is first subjected to visual feature extraction processing to obtain the visual features of the image to be understood; then, based on the visual features, the visual information representation data of the image to be understood and the reference information representation data of the image to be understood are determined respectively; then, the visual information representation data and the reference information representation data are fused to obtain the image information representation data of the image to be understood; finally, the image information representation data is subjected to text generation processing to obtain the image description text of the image to be understood, so that the text content in the image description text can accurately describe the image to be understood, thereby achieving the purpose of accurately determining the description text of an image data.
[0056] In addition, since the image information representation data of the image to be understood is determined by comprehensively combining the visual information representation data of the image to be understood and the reference information representation data of the image to be understood (for example, at least one of the image classification result representation data, the image target detection result representation data, and the image segmentation result representation data), the image information representation data can more accurately represent the image information carried by the image to be understood, so that the text content in the image description text generated based on the image information representation data can more accurately describe the image to be understood, which is conducive to improving the accuracy of the description text of an image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0058] Figure 1 A schematic diagram of image data provided in an embodiment of the present application;
[0059] Figure 2 A flowchart of a method for determining image description text provided in an embodiment of the present application;
[0060] Figure 3 A schematic diagram of a first conversion rule provided in an embodiment of the present application;
[0061] Figure 4 A schematic diagram of a second conversion rule provided in an embodiment of the present application;
[0062] Figure 5 A schematic diagram of a third conversion rule provided in an embodiment of the present application;
[0063] Figure 6 A schematic diagram of the structure of a model to be trained provided in an embodiment of the present application;
[0064] Figure 7 A schematic diagram of another model to be trained provided in an embodiment of the present application;
[0065] Figure 8 A schematic diagram of the structure of an image description text determination device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] In order to solve the technical problems in the background technology part, an embodiment of the present application provides a method for determining image description text, which includes: after obtaining an image to be understood, first performing visual feature extraction processing on the image to be understood to obtain the visual features of the image to be understood; then, based on the visual features, respectively determining the visual information representation data of the image to be understood and the reference information representation data of the image to be understood; then, performing fusion processing on the visual information representation data and the reference information representation data to obtain the image information representation data of the image to be understood; finally, performing text generation processing on the image information representation data to obtain the image description text of the image to be understood, so that the text content recorded in the image description text can accurately describe the image to be understood, thereby achieving the purpose of accurately determining the description text of an image data.
[0067] In addition, since the image information representation data of the image to be understood is determined by comprehensively combining the visual information representation data of the image to be understood and the reference information representation data of the image to be understood (for example, at least one of the image classification result representation data, the image target detection result representation data, and the image segmentation result representation data), the image information representation data can more accurately represent the image information carried by the image to be understood, so that the text content in the image description text generated based on the image information representation data can more accurately describe the image to be understood, which is conducive to improving the accuracy of the description text of an image data.
[0068] In addition, the embodiments of the present application do not limit the execution subject of the method for determining the image description text. For example, the method for determining the image description text provided in the embodiments of the present application can be applied to data processing devices such as terminal devices or servers. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA) or a tablet computer. The server can be an independent server, a cluster server or a cloud server.
[0069] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0070] Method Example 1
[0071] See also Figure 2 , which is a flowchart of a method for determining image description text provided in an embodiment of the present application.
[0072] The method for determining image description text provided in the embodiment of the present application includes S1-S5:
[0073] S1: After acquiring the image to be understood, performing visual feature extraction processing on the image to be understood to obtain visual features of the image to be understood.
[0074] The image to be understood refers to image data that needs to be processed by image description text extraction; and the embodiment of the present application does not limit the image to be understood, for example, it can be Figure 1 The image data shown.
[0075] The above-mentioned "visual feature extraction process" is used to extract visual features from an image data; and the embodiments of the present application do not limit the implementation method of the "visual feature extraction process", and can be implemented by any existing or future method that can perform visual feature extraction process on an image data. For another example, it can be implemented with the help of a pre-built visual feature extraction model.
[0076] It should be noted that the above-mentioned "visual feature extraction model" is used to perform visual feature extraction processing on the input data of the visual feature extraction model; and the embodiments of the present application do not limit the "visual feature extraction model", for example, it can be any machine learning model. In addition, the embodiments of the present application do not limit the construction process of the "visual feature extraction model", for example, it can be implemented by any existing or future model construction method. For example, the following method can be used square Method Example 2 The model building method shown is implemented.
[0077] The above-mentioned “visual features of the image to be understood” are used to represent the image visual information carried by the image to be understood.
[0078] In addition, in order to further improve the accuracy of visual features, the embodiment of the present application also provides a possible implementation method for determining the “visual features of the image to be understood”, which may specifically include steps 11 and 12:
[0079] Step 11: Perform image feature extraction processing on the image to be understood to obtain image features of the image to be understood.
[0080] The above-mentioned “image feature extraction process” is used to extract image features from an image data; and the embodiments of the present application do not limit the implementation method of the “image feature extraction process”. For example, it can be implemented by using any existing or future image feature extraction method. For another example, it can be implemented with the help of a pre-built image feature extraction model.
[0081] It should be noted that the above-mentioned "image feature extraction model" is used to perform image feature extraction processing on the input data of the image feature extraction model; and the embodiment of the present application does not limit the "image feature extraction model", for example, it can be any machine learning model (for example, a deep learning model based on convolutional neural networks (CNN)). In addition, the embodiment of the present application does not limit the construction process of the "image feature extraction model", for example, it can be implemented by any existing or future model construction method. For example, the following method can be used Method Implementation Example 2 The model building method shown is implemented.
[0082] The above-mentioned “image features of the image to be understood” are used to represent the image information carried by the image to be understood; and the embodiment of the present application does not limit the acquisition process of the “image features of the image to be understood”, for example, it may specifically include: inputting the image to be understood into a pre-built image feature extraction model, and obtaining the image features of the image to be understood output by the image feature extraction model.
[0083] Step 12: Perform visual encoding processing on the image features of the image to be understood to obtain the visual features of the image to be understood.
[0084] The embodiment of the present application does not limit the implementation method of step 12. For example, it can specifically be: inputting the image features of the image to be understood into a pre-built visual coding model, and obtaining the visual features of the image to be understood output by the visual coding model, so that the visual features can accurately represent the image visual information carried by the image to be understood.
[0085] It should be noted that the above-mentioned "visual coding model" is used to perform visual coding processing on the input data of the visual coding model; and the embodiment of the present application does not limit the "visual coding model", for example, it can be any machine learning model (for example, a model based on convolutional neural networks (CNN), a model based on residual networks (Recognition Networks, ResNet), etc.). In addition, the embodiment of the present application does not limit the construction process of the "visual coding model", for example, it can be implemented by any existing or future model construction method.
[0086] For example, we can use the following Method Example 2 The model building method shown is implemented.
[0087] Based on the relevant content of S1 above, it can be known that after obtaining the image to be understood, visual feature extraction processing is performed on the image to be understood to obtain the visual features of the image to be understood, so that the visual features can accurately represent the image visual information carried by the image to be understood, thereby enabling the image information representation data of the image to be understood to be determined based on the visual features.
[0088] S2: Determine visual information representation data of the image to be understood according to the visual features of the image to be understood.
[0089] The above-mentioned "visual information representation data of the image to be understood" is used to represent the image visual information carried by the image to be understood; and the embodiment of the present application does not limit the determination process of the "visual information representation data of the image to be understood", for example, the visual features of the image to be understood can be directly determined as the visual information representation data of the image to be understood.
[0090] In addition, in order to improve the information representation effect of the above-mentioned "visual information representation data of the image to be understood", the implementation of the present application also provides another possible implementation method for determining the "visual information representation data of the image to be understood", which may specifically include: inputting the visual features of the image to be understood into a pre-constructed visual information representation model, and obtaining the visual information representation data of the image to be understood output by the visual information representation model.
[0091] It should be noted that the above-mentioned "visual information representation model" is used to perform visual information representation processing on the input data of the visual information representation model; and the embodiments of the present application do not limit the "visual information representation model", for example, it can be any machine learning model (for example, a model based on the patch attention mechanism (Patches Attention)). In addition, the embodiments of the present application do not limit the construction process of the "visual information representation model", for example, it can be implemented by any existing or future model construction method. For example, the following method can be used Method Example 2 The model building method shown is implemented.
[0092] S3: Determine reference information representation data of the image to be understood according to the visual features of the image to be understood.
[0093] The above-mentioned "reference information representation data of the image to be understood" is used to represent at least one reference information of the image to be understood (for example, at least one of the image classification result, image segmentation result, and image target detection result); and the embodiment of the present application does not limit the "reference information representation data of the image to be understood", for example, it may specifically include: image classification result representation data of the image to be understood, image target detection result representation data of the image to be understood, and at least one of the image segmentation result representation data of the image to be understood.
[0094] The above-mentioned "image classification result representation data of the image to be understood" is used to represent the image classification result of the image to be understood; and the embodiment of the present application does not limit the determination process of the "image classification result representation data of the image to be understood", for example, it may specifically include: inputting the visual features of the image to be understood into a pre-constructed classification representation model, and obtaining the image classification result representation data of the image to be understood output by the classification representation model.
[0095] It should be noted that the above-mentioned "classification representation model" is used to perform image classification result representation processing on the input data of the classification representation model; and the embodiment of the present application does not limit the model structure of the "classification representation model". For example, it can be implemented using other model structure parts except the classification result output layer (that is, the last layer in the image classification model) in any existing or future image classification model. In addition, the embodiment of the present application does not limit the construction process of the "classification representation model". For example, it can be implemented using any existing or future model construction method. For example, the following method can be used. Method Example 2 The model building method shown is implemented.
[0096] The above-mentioned "image target detection result representation data of the image to be understood" is used to represent the image target detection result of the image to be understood; and the embodiment of the present application does not limit the determination process of the "image target detection result representation data of the image to be understood", for example, it may specifically include: inputting the visual features of the image to be understood into a pre-constructed target detection representation model, and obtaining the image target detection result representation data of the image to be understood output by the target detection representation model.
[0097] It should be noted that the above-mentioned "target detection representation model" is used to perform image target detection result representation processing on the input data of the target detection representation model; and the embodiment of the present application does not limit the model structure of the "target detection representation model". For example, it can be implemented using other model structure parts except the target detection result output layer (that is, the last layer in the image target detection model) in any existing or future image target detection model. In addition, the embodiment of the present application does not limit the construction process of the "target detection representation model". For example, it can be implemented using any existing or future model construction method. For example, the following method can be used. Method Example 2 The model building method shown is implemented.
[0098] The above-mentioned "image segmentation result representation data of the image to be understood" is used to represent the image segmentation result of the image to be understood; and the embodiment of the present application does not limit the determination process of the "image segmentation result representation data of the image to be understood", for example, it may specifically include: inputting the visual features of the image to be understood into a pre-constructed segmentation representation model, and obtaining the image segmentation result representation data of the image to be understood output by the segmentation representation model.
[0099] It should be noted that the above-mentioned "segmentation representation model" is used to perform image segmentation result representation processing on the input data of the segmentation representation model; and the embodiment of the present application does not limit the model structure of the "segmentation representation model". For example, it can be implemented using other model structure parts except the segmentation result output layer (that is, the last layer in the image segmentation model) in any existing or future image segmentation model. In addition, the embodiment of the present application does not limit the construction process of the "segmentation representation model". For example, it can be implemented using any existing or future model construction method. For example, the following method can be used. Method Example 2 The model building method shown is implemented.
[0100] In addition, the embodiments of the present application do not limit the implementation methods of S3. For example, in one possible implementation method, S3 may specifically include: inputting the visual features of the image to be understood into a pre-built reference information representation model, and obtaining reference information representation data of the image to be understood output by the reference information representation model.
[0101] It should be noted that the above-mentioned "reference information representation model" is used to perform reference information representation processing on the input data of the reference information representation model; and the embodiments of the present application do not limit the "reference information representation model", for example, it can be any machine learning model. In addition, the embodiments of the present application do not limit the construction process of the "reference information representation model", for example, it can be implemented by any existing or future model construction method. For example, the following can be used square Method Example 2 The model building method shown is implemented.
[0102] S4: fusing the visual information representation data of the image to be understood and the reference information representation data of the image to be understood to obtain the image information representation data of the image to be understood.
[0103] The fusion process is used to fuse two data; and the embodiment of the present application does not limit the implementation method of the fusion process. For example, any existing or future method that can fuse two data (for example, splicing, addition, etc.) can be used for implementation. In addition, it can also be implemented with the help of a pre-built fusion model.
[0104] It should be noted that the above-mentioned "fusion model" is used to perform fusion processing on the input data of the fusion model; and the embodiment of the present application does not limit the "fusion model", for example, it can be any machine learning model (for example, a model based on the attention mechanism). In addition, the embodiment of the present application does not limit the construction process of the "fusion model", for example, it can be implemented by any existing or future model construction method. For example, the following Method Example 2 The model building method shown is implemented.
[0105] The above-mentioned "image information representation data of the image to be understood" is used to represent the image information carried by the image to be understood; and the embodiment of the present application does not limit the determination process of the "image information representation data of the image to be understood", for example, it can be specifically: inputting the visual information representation data of the image to be understood and the reference information representation data of the image to be understood into a pre-constructed fusion model to obtain the image information representation data of the image to be understood output by the fusion model.
[0106] S5: Perform text generation processing on the image information representation data of the image to be understood to obtain an image description text of the image to be understood.
[0107] Among them, the text generation processing is used to extract the explanatory text of the image data from the image information representation data of the image data; and the embodiment of the present application does not limit the implementation method of the text generation processing. For example, it can be implemented with the help of a pre-built text generation model.
[0108] It should be noted that the above-mentioned "text generation model" is used to perform text generation processing on the input data of the text generation model; and the embodiments of the present application do not limit the "text generation model", for example, it can be any machine learning model (for example, conditional generation model (Generative Pre-Training, GPT)). In addition, the embodiments of the present application do not limit the construction process of the "text generation model", for example, it can be implemented by any existing or future model construction method. For another example, the following method can be used Method Example 2 The model building method shown is implemented.
[0109] The above-mentioned "image description text of the image to be understood" is used to provide a textual description of the image to be understood; and the embodiment of the present application does not limit the determination process of the "image description text of the image to be understood", for example, it may specifically include: inputting the image information representation data of the image to be understood into a pre-constructed text generation model, and obtaining the image description text of the image to be understood output by the text generation model.
[0110] Based on the relevant contents of S1 to S5 above, it can be known that for the image description text determination method provided in the embodiment of the present application, after obtaining the image to be understood, the image to be understood is first subjected to visual feature extraction processing to obtain the visual features of the image to be understood; then, based on the visual features, the visual information representation data of the image to be understood and the reference information representation data of the image to be understood are determined respectively; then, the visual information representation data and the reference information representation data are fused to obtain the image information representation data of the image to be understood; finally, the image information representation data is subjected to text generation processing to obtain the image description text of the image to be understood, so that the text content in the image description text can accurately describe the image to be understood, thereby achieving the purpose of accurately determining the description text of an image data.
[0111] In addition, since the image information representation data of the image to be understood is determined by comprehensively combining the visual information representation data of the image to be understood and the reference information representation data of the image to be understood (for example, at least one of the image classification result representation data, the image target detection result representation data, and the image segmentation result representation data), the image information representation data can more accurately represent the image information carried by the image to be understood, so that the text content in the image description text generated based on the image information representation data can more accurately describe the image to be understood, which is conducive to improving the accuracy of the description text of an image data.
[0112] Method Example 2
[0113] In order to further improve the data processing effect of the above models (for example, visual feature extraction model, image feature extraction model, visual encoding model, visual information representation model, classification representation model, target detection representation model, segmentation representation model, reference information representation model, fusion model, text generation model, etc.), the embodiment of the present application also provides a model construction method, which may specifically include steps 21 to 25:
[0114] Step 21: Obtain a sample image and actual description text of the sample image.
[0115] The sample image refers to the image data required for use in the model building process; and the embodiment of the present application does not limit the number of the sample images.
[0116] The above-mentioned “actual description text of the sample image” refers to the actual explanatory text of the sample image, so that the “actual description text of the sample image” is used to represent the image information actually carried by the sample image (for example, existing objects, and the relationship between different objects, etc.).
[0117] In addition, the embodiment of the present application does not limit the method for obtaining the above-mentioned "actual description text of the sample image". For example, it may specifically include: determining the actual description text of the sample image according to the label information of the sample image.
[0118] The above-mentioned “label information of the sample image” refers to the annotation information of the sample image; and the embodiment of the present application does not limit the “label information of the sample image”, for example, it may include at least one of the image description text label of the sample image, the image classification label of the sample image, the image target detection label of the sample image, and the image segmentation label of the sample image. Among them, the above-mentioned “image description text label of the sample image” refers to the image description text annotation of the sample image, so that the image description text label is used to represent the actual description text of the sample image. The above-mentioned “image classification label of the sample image” refers to the image classification annotation of the sample image, so that the image classification label is used to represent the actual classification information of the sample image. The above-mentioned “image target detection label of the sample image” refers to the image target detection annotation of the sample image, so that the image target detection label is used to represent the actual target object information of the sample image. The above-mentioned “image segmentation label of the sample image” refers to the image segmentation annotation of the sample image, so that the image segmentation label is used to represent the actual segmentation information of the sample image.
[0119] In order to facilitate understanding of the method of obtaining the above-mentioned “actual description text of the sample image”, four cases are described below.
[0120] Case 1: when the above-mentioned “label information of the sample image” includes the image description text label of the sample image, the image description text label can be directly determined as the actual description text of the sample image.
[0121] Case 2: When the above-mentioned "label information of the sample image" includes the label to be used of the sample image, and the label to be used can be the image classification label of the sample image, the image target detection label of the sample image, or the image segmentation label of the sample image, the label to be used can be converted into text to obtain the actual description text of the sample image.
[0122] The above-mentioned "text conversion processing" is used to extract text content that can describe the label data from a label data (for example, an image classification label, an image target detection label, or an image segmentation label, etc.); and the embodiment of the present application does not limit the "text conversion processing", for example, it can be pre-set.
[0123] In order to further improve the text conversion effect of the above-mentioned "text conversion processing", different text conversion rules can be set for different label types. Based on this, when the above-mentioned "label information of the sample image" includes the label to be used of the sample image, the text conversion rule corresponding to the label to be used can be determined according to the label type of the label to be used; then, according to the text conversion rule corresponding to the label to be used, the label to be used is subjected to text conversion processing to obtain the actual description text of the sample image.
[0124] The above-mentioned “label type of the label to be used” is used to indicate the type of the label to be used; and the embodiment of the present application does not limit the “label type”, for example, it can be an image classification type, an image target detection type, or an image segmentation type.
[0125] The above-mentioned "text conversion rules corresponding to the label to be used" refer to the text conversion rules required to be followed when performing text conversion processing on the label to be used; and the embodiment of the present application does not limit the determination process of the "text conversion rules corresponding to the label to be used", for example, it can be specifically: searching for the text conversion rules corresponding to the label type of the label to be used from the pre-constructed mapping relationship to be used, and determining them as the text conversion rules corresponding to the label to be used.
[0126] The above-mentioned "mapping relationship to be used" is used to record the text conversion rules corresponding to each label type; and the embodiment of the present application does not limit the "mapping relationship to be used", for example, it may specifically include: the correspondence between the image classification type and the first conversion rule, the correspondence between the image target detection type and the second conversion rule, and the correspondence between the image segmentation type and the third conversion rule.
[0127] The above-mentioned “first conversion rule” is used to convert an image classification label of an image data into a text data, so that the text data can describe the image classification label of the image data in a text description manner; and the embodiment of the present application does not limit the “first conversion rule”, for example, it can be adopted Figure 3 The first conversion rule shown is implemented; and based on Figure 3 The first conversion rule shown in FIG. 1 shows that when the sample image is Figure 1 , and the image classification label of the sample image includes the sea and the chair, the image classification label can be converted into the text data "the sample image includes the elements [sea] and [chair]" according to the first conversion rule. Figure 3 For example, Q is a positive integer.
[0128] The above-mentioned “second conversion rule” is used to convert an image target detection label of an image data into a text data, so that the text data can describe the image target detection label of the image data in a text description manner; and the embodiment of the present application does not limit the “second conversion rule”, for example, it can be adopted Figure 4 The second conversion rule shown is implemented; and based on Figure 4 The second conversion rule shown in FIG. 1 shows that when the sample image is Figure 1 The image data shown in FIG. 1 , and the image target detection label of the sample image includes [sea, X 1 , Y 1 , X 2 , Y 2 ] and [chair, X 3 , Y 3 , X 4 , Y 4 ], the image target detection label can be converted into the text data "a [sea] is in the upper left, upper right, middle, top, left and right positions of the sample image, and a [chair] is in the lower right position of the sample image" according to the second conversion rule.
[0129] It should be noted that the above “X 1 , Y 1 , X 2 , Y 2 "Used to describe the sea in Figure 1 The actual position in the image data shown; the above “X 3 , Y 3 , X 4 , Y 4 "Used to describe the chair in Figure 1 The actual position in the image data shown. Figure 4 For example, Q is a positive integer.
[0130] It should also be noted that the above-mentioned "upper left", "upper right", "middle", "directly above", "directly left", "directly right", and "lower right" can be determined by a preset image position conversion process; and the image position conversion process can be specifically: firstly divide an image data (for example, a sample image) into a 3×3 square grid area (similar to a nine-square grid) to obtain nine directions (for example, upper left, lower right, lower left, upper right, middle, directly above, directly below, directly left, directly right and other nine directions) and the area range of the nine directions; then the object position recorded by an object detection tag (for example, X 1 , Y 1 , X 2 , Y 2) is compared with the area range of each direction to obtain the location of the object identifier (for example, [sea], etc.) recorded by the target detection tag.
[0131] The above-mentioned “third conversion rule” is used to convert an image segmentation label of an image data into a text data, so that the text data can describe the image segmentation label of the image data in a text description manner; and the embodiment of the present application does not limit the “third conversion rule”, for example, it can be used Figure 5 The third conversion rule shown is implemented; and based on this Figure 5 The third conversion rule shown in FIG. 1 shows that when the sample image is Figure 1 The image data shown in FIG. 1 , and the image segmentation labels of the sample image include [sea, X 1 , Y 1 , X 2 , Y 2 , X 3 , Y 3 , X 4 , Y 4 , X 5 , Y 5 ] and [chair, X 6 , Y 6 , X 7 , Y 7 , X 8 , Y 8 , X 9 , Y 9 , X 10 , Y 10 ], the image segmentation label can be converted into the text data "a [sea] is in the upper left, upper right, middle, top, left and right positions of the sample image, and a [chair] is in the lower right position of the sample image" according to the third conversion rule.
[0132] It should be noted that the above “X 1 , Y 1 , X 2 , Y 2 , X 3 , Y 3 , X 4 , Y 4 , X 5 , Y 5 "Used to describe the sea in Figure 1 The actual position of the region in the image data shown; the above “X 6 , Y 6 , X 7 , Y 7 , X 8 , Y 8 , X 9 , Y9 , X 10 , Y 10 "Used to describe the chair in Figure 1 The actual position of the area in the image data shown. Figure 5 For example, Q is a positive integer.
[0133] Based on the relevant content of the above situation 2, it can be known that when the label information of the sample image only includes the image classification label of the sample image, the image target detection label of the sample image, or the image segmentation label of the sample image, the label information can be directly converted into text according to the pre-set text conversion rules to obtain the actual description text of the sample image, so that the actual description text can explain the label information in the form of text description.
[0134] Case 3, when the above-mentioned “label information of the sample image” includes R reference labels of the sample image, and the R reference labels include at least two of the image classification label of the sample image, the image object detection label of the sample image, and the image segmentation label of the sample image, the determination process of the above-mentioned “actual description text of the sample image” may include steps 31-32:
[0135] Step 31: Perform text conversion processing on the r-th reference label of the sample image to obtain the description text of the r-th reference label, where r is a positive integer, r≤R, R is a positive integer, R≥2.
[0136] It should be noted that step 31 can be implemented using any implementation method of the process for determining the "actual description text of the sample image" shown in Case 2 above. It only requires replacing the "label to be used" with the "rth reference label" and replacing the "actual description text of the sample image" with the "description text of the rth reference label" in any implementation method of the process for determining the "actual description text of the sample image" shown in Case 2 above.
[0137] Step 32: Determine the actual description text of the sample image based on the description texts of the R reference tags and a pre-set text generation template. The text generation template refers to a template required to generate a text data from multiple text data; and the text generation template can be pre-set according to the application scenario.
[0138] In an embodiment of the present application, after obtaining the description text of the 1st reference label to the description text of the Rth reference label, the R description texts can be processed for text generation according to a pre-set text generation template to obtain the actual description text of the sample image, so that the actual description text can explain the R reference labels of the sample image in the form of text description.
[0139] Based on the relevant content of the above situation 3, it can be known that when the above-mentioned "label information of the sample image" includes at least two of the image classification label of the sample image, the image target detection label of the sample image, and the image segmentation label of the sample image, text conversion processing can be performed on each label data in the label information to obtain the description text of each label data; and then the actual description text of the sample image can be extracted from the description text of these label data according to a pre-set text generation template, so that the actual description text can explain these reference labels in the form of text description.
[0140] Case 4, when the above-mentioned “label information of the sample image” includes the image description text label of the sample image and T reference labels, and the “T reference labels” include at least one of the image classification label of the sample image, the image object detection label of the sample image, and the image segmentation label of the sample image, the determination process of the above-mentioned “actual description text of the sample image” may include steps 41-42:
[0141] Step 41: Perform text conversion processing on the t-th reference label of the sample image to obtain the description text of the t-th reference label, where t is a positive integer, t≤T, T is a positive integer, T≥1.
[0142] It should be noted that step 41 can be implemented using any implementation method of the process for determining the "actual description text of the sample image" shown in Case 2 above. It only requires replacing the "label to be used" with the "tth reference label" and replacing the "actual description text of the sample image" with the "description text of the tth reference label" in any implementation method of the process for determining the "actual description text of the sample image" shown in Case 2 above.
[0143] Step 42: Determine the actual description text of the sample image according to the image description text label of the sample image, the description texts of T reference labels, and a preset text generation template.
[0144] In an embodiment of the present application, after obtaining the description text of the 1st reference label to the description text of the Tth reference label, the T description texts and the image description text label of the sample image can be processed for text generation according to a pre-set text generation template to obtain the actual description text of the sample image, so that the actual description text can explain the T reference labels and the image description text labels of the sample image in the form of text description.
[0145] Based on the relevant content of the above situation 4, it can be known that when the above-mentioned "label information of the sample image" includes at least one of the image classification label of the sample image, the image target detection label of the sample image, and the image segmentation label of the sample image, and the image description text label of the sample image, text conversion processing can be first performed on each label data in the label information except the image description text label to obtain the description text of each label data; and then the actual description text of the sample image is extracted from the description texts of these label data and the image description text label according to a pre-set text generation template, so that the actual description text can explain all the label data in the label information in the form of text description.
[0146] Based on the relevant content of the above step 21, it can be known that after obtaining the sample image and the label information of the sample image, the actual description text of the sample image can be determined according to the label information of the sample image, so that the actual description text can explain the label information in a textual description form, so that the sample image and the actual description text of the sample image can be used to train the model to be trained subsequently.
[0147] Step 22: Input the sample image into the model to be trained, and obtain the predicted description text of the sample image output by the model to be trained.
[0148] The above-mentioned “model to be trained” is used to perform image caption text extraction processing on the input data of the model to be trained; and the embodiment of the present application does not limit the “model to be trained”, for example, it can be any machine learning model.
[0149] In addition, the embodiment of the present application does not limit the model structure of the above-mentioned "model to be trained", for example, Figure 6 As shown, the model to be trained 600 may specifically include: a visual feature extraction layer 601, a visual information representation layer 602, a reference information representation layer 603, a fusion layer 604, and a text generation layer 605. Among them, the input data of the text generation layer 605 includes the output data of the fusion layer 604; the input data of the fusion layer 604 includes the output data of the reference information representation layer 603 and the output data of the visual information representation layer 602; the input data of the reference information representation layer 603 includes the output data of the visual feature extraction layer 601; the input data of the visual information representation layer 602 includes the output data of the visual feature extraction layer 601.
[0150] The visual feature extraction layer 601 is used to perform visual feature extraction processing on an image data; and the embodiment of the present application does not limit the network structure of the visual feature extraction layer 601. For example, it can be implemented using the model structure of the above "visual feature extraction model". For example, in some possible implementations, Figure 7As shown, the visual feature extraction layer 601 may include an image feature extraction layer (Image Feature Extractor) and a visual encoding layer (Visual Embedding), and the input data of the visual encoding layer includes the output data of the image feature extraction layer.
[0151] It should be noted that the image feature extraction layer is used to perform image feature extraction processing on an image data; and the embodiment of the present application does not limit the network structure of the "image feature extraction layer", for example, it can be implemented using the model structure of the above "image feature extraction model". The visual coding layer is used to perform visual coding processing on the input data of the visual coding layer; and the embodiment of the present application does not limit the network structure of the "visual coding layer", for example, it can be implemented using the model structure of the above "visual coding model".
[0152] The visual information representation layer 602 is used to perform visual information representation processing on the input data of the visual information representation layer 602; and the embodiment of the present application does not limit the network structure of the visual information representation layer 602. For example, it can be implemented using the model structure of the "visual information representation model" above. For another example, it can be implemented using Figure 7 The patch attention layer shown is implemented.
[0153] The reference information representation layer 603 is used to perform reference information representation processing on the input data of the reference information representation layer 603; and the embodiment of the present application does not limit the network structure of the reference information representation layer 603. For example, it can be implemented using the model structure of the "reference information representation model" above. For another example, in some possible implementations, it can include a classification representation layer, a target detection representation layer, and a segmentation representation layer.
[0154] It should be noted that the above-mentioned "classification representation layer" is used to perform image classification result representation processing on the input data of the classification representation layer (that is, the above-mentioned "input data of the reference information representation layer 603"); and the embodiment of the present application does not limit the network structure of the "classification representation layer", for example, it can be implemented using the model structure of the above-mentioned "classification representation model". The above-mentioned "target detection representation layer" is used to perform image target detection result representation processing on the input data of the target detection representation layer (that is, the above-mentioned "input data of the reference information representation layer 603"); and the embodiment of the present application does not limit the network structure of the "target detection representation layer", for example, it can be implemented using the model structure of the above-mentioned "target detection representation model". The above-mentioned "segmentation representation layer" is used to perform image segmentation result representation processing on the input data of the segmentation representation layer (that is, the above-mentioned "input data of the reference information representation layer 603"); and the embodiment of the present application does not limit the network structure of the "segmentation representation layer", for example, it can be implemented using the model structure of the above-mentioned "segmentation representation model".
[0155] It should also be noted that for Figure 7 For example, the “reference information representation data” is obtained by performing reference information representation processing on the output data of the visual coding layer by the above-mentioned “reference information representation layer 603”; “Det” is used to represent the image target detection result representation data obtained by performing image target detection result representation processing on the output data of the visual coding layer by the above-mentioned “target detection representation layer”; “Cls” is used to represent the image classification result representation data obtained by performing image classification result representation processing on the output data of the visual coding layer by the above-mentioned “classification representation layer”; “Seg” is used to represent the image segmentation result representation data obtained by performing image segmentation result representation processing on the output data of the visual coding layer by the above-mentioned “segmentation representation layer”.
[0156] The fusion layer 604 is used to perform fusion processing on the input data of the fusion layer 604; and the embodiment of the present application does not limit the network structure of the fusion layer 604. For example, it can be implemented using the model structure of the above "fusion model".
[0157] The text generation layer 605 is used to perform text generation processing on the input data of the text generation layer 605; and the embodiment of the present application does not limit the network structure of the "text generation layer 605", for example, it can be implemented using the model structure of the above "text generation model". It should be noted that for Figure 7 For the “text generation model” shown in the figure, “S 1 ”, “S 2 ”、……“S n " respectively represent the output results at different time states when the text generation model performs text generation processing.
[0158] The above-mentioned “prediction description text of the sample image” is used to indicate the prediction description text of the sample image.
[0159] Based on the relevant content of the above step 22, it can be known that after obtaining the sample image, the sample image can be input into the model to be trained so that the model to be trained can perform image caption text extraction processing on the sample image, obtain and output the predicted description text of the sample image, so that the predicted description text can be referenced later to determine the image caption text extraction performance of the model to be trained.
[0160] Step 23: Determine whether the preset stop condition is reached, if so, execute step 25; if not, execute step 24.
[0161] Among them, the preset stop condition can be set in advance; and the embodiment of the present application does not limit the preset stop condition. For example, it can be that the model loss value of the model to be trained is lower than the first threshold, or the rate of change of the model loss value of the model to be trained is lower than the second threshold, or the number of updates of the model to be trained reaches the third threshold.
[0162] The above-mentioned “model loss value of the model to be trained” is used to represent the image description text extraction performance of the model to be trained; and the embodiment of the present application does not limit the determination process of the “model loss value of the model to be trained”, for example, Figure 7 As shown, it can be implemented with the help of Maximum Likelihood Estimate (MLE).
[0163] Step 24: Update the model to be trained according to the predicted description text of the sample image and the actual description text of the sample image, and return to execute step 22.
[0164] In an embodiment of the present application, after determining that the current round of the model to be trained has not yet reached the preset stop condition, it can be determined that the image description text extraction performance of the model to be trained is still relatively poor, so the model to be trained can be updated based on the difference between the predicted description text of the sample image and the actual description text of the sample image, so that the updated model to be trained has better image description text extraction performance, so that step 22 and its subsequent steps can be continued based on the updated model to be trained to achieve the next round of training process for the model to be trained.
[0165] Step 25: According to the model to be trained, determine at least one of a visual feature extraction model, an image feature extraction model, a visual encoding model, a visual information representation model, a classification representation model, a target detection representation model, a segmentation representation model, a reference information representation model, a fusion model, and a text generation model.
[0166] In an embodiment of the present application, after determining that the model to be trained in the current round reaches the preset stop condition, it can be determined that the model to be trained has good image description text extraction performance, so at least one of the visual feature extraction model, the image feature extraction model, the visual coding model, the visual information representation model, the classification representation model, the target detection representation model, the segmentation representation model, the reference information representation model, the fusion model, and the text generation model can be determined according to the model to be trained; and the determination process can specifically include: determining the visual feature extraction model according to the visual feature extraction layer in the model to be trained (for example, determining the visual feature extraction layer in the model to be trained as the visual feature extraction model); determining the image feature extraction model according to the image feature extraction layer in the model to be trained (for example, determining the image feature extraction layer in the model to be trained as the image feature extraction model); determining the visual coding model according to the visual coding layer in the model to be trained (for example, determining the visual coding layer in the model to be trained as the visual coding model); determining the visual information representation model according to the visual information representation model in the model to be trained According to the representation layer of the to-be-trained model, a visual information representation model is determined (for example, the visual information representation layer in the to-be-trained model is determined as the visual information representation model); according to the classification representation layer in the to-be-trained model, a classification representation model is determined (for example, the classification representation layer in the to-be-trained model is determined as the classification representation model); according to the target detection representation layer in the to-be-trained model, a target detection representation model is determined (for example, the target detection representation layer in the to-be-trained model is determined as the target detection representation model); according to the segmentation representation layer in the to-be-trained model, a segmentation representation model is determined (for example, the segmentation representation layer in the to-be-trained model is determined as the segmentation representation model); according to the reference information representation layer in the to-be-trained model, a reference information representation model is determined (for example, the reference information representation layer in the to-be-trained model is determined as the reference information representation model); according to the fusion layer in the to-be-trained model, a fusion model is determined (for example, the fusion layer in the to-be-trained model is determined as the fusion model); according to the text generation layer in the to-be-trained model, a text generation model is determined (for example, the text generation layer in the to-be-trained model is determined as the text generation model).
[0167] Based on the relevant contents of the above steps 21 to 25, it can be known that after obtaining the sample image and at least one of the image description text label, image classification label, image target detection label, and image segmentation label of the sample image, the actual description text of the sample image can be determined based on these labels, so that the actual description text is used to record the actual description text of the sample image; then, with reference to the sample image and the actual description text of the sample image, the model to be trained is trained to obtain the trained model to be trained, so that the trained model to be trained has better image description text extraction performance, so that subsequently, the entire network or part of the network in the trained model to be trained can be directly used to perform image description text extraction processing on an image data, and the image description text of the image data can be obtained and output, so that the image description text can accurately describe the image data.
[0168] In addition, in some application scenarios (for example, when the text generation layer in the above-mentioned “model to be trained” is implemented using GPT), in order to further improve the image description text extraction performance of the model to be trained, the embodiment of the present application also provides another possible implementation of the model construction method. In this implementation, in addition to the above-mentioned step 21 and steps 23-25, the model construction method may also include steps 26-27:
[0169] Step 26: Determine the label type representation data of the sample image according to the label information of the sample image.
[0170] The above-mentioned “label type characterization data of the sample image” is used to indicate the type of label information of the sample image; and the embodiment of the present application does not limit the determination process of the “label type characterization data of the sample image”, for example, it may specifically include step 51-step 52:
[0171] Step 51: Determine the label type identifier of the sample image according to the label information of the sample image.
[0172] The above-mentioned “label type identifier of the sample image” is used to identify the type of label information of the sample image; and the embodiment of the present application does not limit the representation method of the “label type identifier of the sample image”, for example, Figure 7 As shown, it can be used with the help of t 1 ,t 2 ,t 3 , and t 4 If the above “label information of the sample image” includes the image description text label of the sample image, True is assigned to t 1 ; If the above “label information of sample image” does not include the image description text label of the sample image, then assign False to t1 ; If the above “label information of sample image” includes the image classification label of the sample image, True is assigned to t 2 ; If the above “label information of sample image” does not include the image classification label of the sample image, False is assigned to t 2 ; If the above “label information of the sample image” includes the image target detection label of the sample image, True is assigned to t 3 ; If the above “label information of the sample image” does not include the image target detection label of the sample image, assign False to t 3 ; If the above “label information of the sample image” includes the image segmentation label of the sample image, True is assigned to t 4 ; If the above “label information of sample image” does not include the image segmentation label of the sample image, False is assigned to t 4 .
[0173] Step 52: Perform character encoding processing on the label type identifier of the sample image to obtain label type representation data of the sample image.
[0174] The above-mentioned "character encoding processing" is used to perform encoding processing on a character data; and the embodiment of the present application does not limit the "character encoding processing". For example, any existing or future method that can perform encoding processing on a character data (for example, Word2Vec, etc.) can be used for implementation.
[0175] Based on the relevant content of the above step 26, it can be known that after obtaining the label information of the sample image, the label type representation data of the sample image can be determined according to the label information, so that the label type representation data can represent the type of the label information of the sample image, so that the text generation layer in the subsequent model to be trained can use the label type representation data as a reference condition to perform text generation processing.
[0176] Step 27: Input the sample image and the label type representation data of the sample image into the model to be trained, and obtain the predicted description text of the sample image output by the model to be trained.
[0177] It should be noted that for the relevant content of the "model to be trained", please refer to step 22 above.
[0178] Based on the relevant contents of steps 26 to 27 above, it can be known that when the text generation layer in the above-mentioned "model to be trained" is implemented using GPT, the label type representation data of the sample image can be first determined based on the label information of the sample image; then the GPT uses the label type representation data and the image information representation data of the sample image as reference conditions to perform text generation processing, and obtains and outputs the predicted description text of the sample image, which is conducive to improving the accuracy of the predicted description text.
[0179] Based on the image description text determination method provided by the above method embodiment, the embodiment of the present application also provides an image description text determination device, which is explained and illustrated below in conjunction with the accompanying drawings.
[0180] Device Embodiment
[0181] For technical details of the image description text determination device provided by the device embodiment, please refer to the above method embodiment.
[0182] See also Figure 8 , which is a structural schematic diagram of an image description text determination device provided in an embodiment of the present application.
[0183] The image description text determination device 800 provided in the embodiment of the present application includes:
[0184] The extraction unit 801 is used for performing visual feature extraction processing on the image to be understood after acquiring the image to be understood, so as to obtain the visual features of the image to be understood;
[0185] A determining unit 802, configured to determine visual information representation data of the image to be understood and reference information representation data of the image to be understood, respectively, according to the visual features;
[0186] A fusion unit 803 is used to fuse the visual information representation data and the reference information representation data to obtain image information representation data of the image to be understood;
[0187] The generating unit 804 is used to perform text generation processing on the image information representation data to obtain an image description text of the image to be understood.
[0188] In a possible implementation, the reference information representation data includes at least one of image classification result representation data, image target detection result representation data, and image segmentation result representation data; wherein the image classification result representation data is used to represent the image classification result of the image to be understood; the image target detection result representation data is used to represent the image target detection result of the image to be understood; and the image segmentation result representation data is used to represent the image segmentation result of the image to be understood.
[0189] In a possible implementation manner, the determining unit 802 includes:
[0190] A classification determination subunit, used for inputting the visual features into a pre-built classification representation model to obtain image classification result representation data of the image to be understood output by the classification representation model;
[0191] The target detection subunit is used to input the visual features into a pre-built target detection representation model to obtain image target detection result representation data of the image to be understood output by the target detection representation model;
[0192] The segmentation determination subunit is used to input the visual features into a pre-built segmentation representation model to obtain image segmentation result representation data of the image to be understood output by the segmentation representation model.
[0193] In a possible implementation manner, the determining unit 802 includes:
[0194] The visual determination subunit is used to input the visual features into a pre-built visual information representation model to obtain visual information representation data of the image to be understood output by the visual information representation model.
[0195] In a possible implementation, the generating unit 804 is specifically configured to: input the image information representation data into a pre-built text generation model to obtain an image description text of the image to be understood output by the text generation model.
[0196] In a possible implementation manner, the image description text determination device 800 further includes:
[0197] A data acquisition subunit, used for acquiring a sample image and actual description text of the sample image;
[0198] The model training subunit is used to input the sample image into the model to be trained to obtain the predicted description text of the sample image output by the model to be trained; update the model to be trained according to the predicted description text of the sample image and the actual description text of the sample image, and continue to execute the step of inputting the sample image into the model to be trained until a preset stop condition is reached, and then determine the text generation model according to the model to be trained.
[0199] In a possible implementation, the data acquisition subunit includes:
[0200] The text determination subunit is used to determine the actual description text of the sample image according to the label information of the sample image.
[0201] In a possible implementation, the label information includes labels to be used, and the labels to be used include image classification labels, image target detection labels, or image segmentation labels, and the text determination subunit is specifically used to: perform text conversion processing on the labels to be used of the sample image to obtain actual description text of the sample image.
[0202] In a possible implementation, the label information includes multiple reference labels, the multiple reference labels include at least two of an image classification label, an image target detection label, and an image segmentation label, and the text determination subunit is specifically used to: perform text conversion processing on each of the reference labels to obtain a description text of each of the reference labels; and determine the actual description text of the sample image based on the description texts of the multiple reference labels and a pre-set text generation template.
[0203] In a possible implementation, the label information includes an image description text label and at least one reference label, the at least one reference label includes at least one of an image classification label, an image target detection label, and an image segmentation label, and the text determination subunit is specifically used to: perform text conversion processing on each of the reference labels to obtain a description text of each of the reference labels; determine the actual description text of the sample image based on the image description text label, the description text of the at least one reference label, and a pre-set text generation template.
[0204] In a possible implementation, the text determination subunit is specifically used to: perform text conversion processing on the label to be used of the sample image according to the text conversion rule corresponding to the label to be used, so as to obtain the actual description text of the sample image; wherein the text conversion rule corresponding to the label to be used is determined according to the label type of the label to be used.
[0205] In a possible implementation, the model to be trained includes a visual feature extraction layer, a visual information representation layer, a reference information representation layer, a fusion layer, and a text generation layer; the input data of the text generation layer includes the output data of the fusion layer; the input data of the fusion layer includes the output data of the reference information representation layer and the output data of the visual information representation layer; the input data of the reference information representation layer includes the output data of the visual feature extraction layer; the input data of the visual information representation layer includes the output data of the visual feature extraction layer;
[0206] The model training subunit comprises:
[0207] The model determination subunit is used to determine the text generation model according to the text generation layer in the model to be trained.
[0208] In a possible implementation, the model training subunit includes:
[0209] The text prediction subunit is used to input the sample image and the label type representation data of the sample image into the model to be trained, and obtain the predicted description text of the sample image output by the model to be trained; wherein the label type representation data of the sample image is determined based on the label information of the sample image.
[0210] In a possible implementation, the extraction unit 801 is specifically used to: after acquiring the image to be understood, perform image feature extraction processing on the image to be understood to obtain image features of the image to be understood; perform visual encoding processing on the image features to obtain visual features of the image to be understood.
[0211] Based on the relevant contents of the above-mentioned image description text determination device 800, it can be known that for the image description text determination device 800, after obtaining the image to be understood, the image to be understood is first subjected to visual feature extraction processing to obtain the visual features of the image to be understood; then, based on the visual features, the visual information representation data of the image to be understood and the reference information representation data of the image to be understood are determined respectively; then, the visual information representation data and the reference information representation data are fused to obtain the image information representation data of the image to be understood; finally, the image information representation data is subjected to text generation processing to obtain the image description text of the image to be understood, so that the text content in the image description text can accurately describe the image to be understood, thereby achieving the purpose of accurately determining the description text of an image data.
[0212] In addition, since the image information representation data of the image to be understood is determined by comprehensively combining the visual information representation data of the image to be understood and the reference information representation data of the image to be understood (for example, at least one of the image classification result representation data, the image target detection result representation data, and the image segmentation result representation data), the image information representation data can more accurately represent the image information carried by the image to be understood, so that the text content in the image description text generated based on the image information representation data can more accurately describe the image to be understood, which is conducive to improving the accuracy of the description text of an image data.
[0213] Furthermore, an embodiment of the present application also provides a device, the device comprising a processor and a memory:
[0214] The memory is used to store computer programs;
[0215] The processor is used to execute any implementation of the image description text determination method provided in the embodiments of the present application according to the computer program.
[0216] Furthermore, the embodiments of the present application also provide a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute any implementation of the image description text determination method provided in the embodiments of the present application.
[0217] Furthermore, the embodiments of the present application also provide a computer program product. When the computer program product is executed on a terminal device, the terminal device executes any implementation of the image description text determination method provided in the embodiments of the present application.
[0218] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0219] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as above in the preferred embodiment, it is not used to limit the present invention. Any technician familiar with the art can make many possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the scope of the technical solution of the present invention, or modify it into an equivalent embodiment of equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.
Claims
1. A method for determining image description text, It is characterized in that The method comprises: After acquiring the image to be understood, performing visual feature extraction processing on the image to be understood to obtain the visual features of the image to be understood; Determine, according to the visual features, visual information representation data of the image to be understood and reference information representation data of the image to be understood, respectively, wherein the reference information representation data includes part or all of image classification result representation data, image target detection result representation data, and image segmentation result representation data; fusing the visual information representation data and the reference information representation data to obtain image information representation data of the image to be understood; The image information representation data is subjected to text generation processing to obtain an image description text of the image to be understood. The text generation processing is implemented by using a text generation model. The text generation model is determined based on a trained model to be trained. The training process of the model to be trained includes updating the model to be trained based on the predicted description text of the sample image and the actual description text of the sample image. The predicted description text is obtained by the model to be trained by processing the sample image and the label type representation data of the sample image. The label type representation data is used to describe the type of label information of the sample image. The actual description text is used to describe the label information in a text description. The label information includes part or all of an image description text label, an image classification label, an image target detection label and an image segmentation label. The type of the image description text label, the type of the image classification label, the type of the image target detection label and the type of the image segmentation label are all different.
2. The method according to claim 1, It is characterized in that The process of determining the image classification result representation data includes: Inputting the visual features into a pre-built classification representation model to obtain image classification result representation data of the image to be understood output by the classification representation model; The process of determining the image target detection result representation data includes: Inputting the visual features into a pre-built target detection representation model to obtain image target detection result representation data of the image to be understood output by the target detection representation model; The process of determining the image segmentation result representation data includes: The visual features are input into a pre-built segmentation representation model to obtain image segmentation result representation data of the image to be understood output by the segmentation representation model.
3. The method according to claim 1, It is characterized in that The process of determining the visual information representation data includes: The visual features are input into a pre-built visual information representation model to obtain visual information representation data of the image to be understood output by the visual information representation model.
4. The method according to claim 1, It is characterized in that The step of performing text generation processing on the image information representation data to obtain an image description text of the image to be understood includes: The image information representation data is input into a pre-built text generation model to obtain an image description text of the image to be understood output by the text generation model.
5. The method according to claim 4, It is characterized in that The construction process of the text generation model includes: Obtaining a sample image and actual description text of the sample image; Inputting the sample image into the model to be trained to obtain the predicted description text of the sample image output by the model to be trained; According to the predicted description text of the sample image and the actual description text of the sample image, the model to be trained is updated, and the step of inputting the sample image into the model to be trained is continued until a preset stop condition is reached, and the text generation model is determined according to the model to be trained.
6. The method according to claim 5, It is characterized in that The process of obtaining the actual description text includes: According to the label information of the sample image, an actual description text of the sample image is determined.
7. The method according to claim 6, It is characterized in that The label information includes a to-be-used label, wherein the to-be-used label includes an image classification label, an image target detection label, or an image segmentation label, and determining the actual description text of the sample image according to the label information of the sample image includes: Performing text conversion processing on the to-be-used label of the sample image to obtain actual description text of the sample image; or, The label information includes a plurality of reference labels, the plurality of reference labels include at least two of an image classification label, an image object detection label, and an image segmentation label, and determining the actual description text of the sample image according to the label information of the sample image includes: Performing text conversion processing on each of the reference labels to obtain description texts of each of the reference labels; determining the actual description text of the sample image according to the description texts of the multiple reference labels and a preset text generation template; or, The label information includes an image description text label and at least one reference label, wherein the at least one reference label includes at least one of an image classification label, an image object detection label, and an image segmentation label, and determining the actual description text of the sample image according to the label information of the sample image includes: Perform text conversion processing on each of the reference tags to obtain description text of each of the reference tags; determine the actual description text of the sample image based on the image description text tag, the description text of at least one reference tag, and a pre-set text generation template.
8. The method according to claim 7, It is characterized in that The performing text conversion processing on the to-be-used label of the sample image to obtain the actual description text of the sample image includes: According to the text conversion rule corresponding to the tag to be used, the tag to be used of the sample image is subjected to text conversion processing to obtain the actual description text of the sample image; wherein the text conversion rule corresponding to the tag to be used is determined according to the tag type of the tag to be used.
9. The method according to claim 6, It is characterized in that The model to be trained includes a visual feature extraction layer, a visual information representation layer, a reference information representation layer, a fusion layer, and a text generation layer; the input data of the text generation layer includes the output data of the fusion layer; The input data of the fusion layer includes the output data of the reference information representation layer and the output data of the visual information representation layer; The input data of the reference information representation layer includes the output data of the visual feature extraction layer; the input data of the visual information representation layer includes the output data of the visual feature extraction layer; Determining the text generation model according to the model to be trained includes: The text generation model is determined according to the text generation layer in the model to be trained.
10. The method according to claim 5, It is characterized in that The step of inputting the sample image into the model to be trained to obtain the predicted description text of the sample image output by the model to be trained includes: The sample image and the label type representation data of the sample image are input into the model to be trained to obtain the predicted description text of the sample image output by the model to be trained; wherein the label type representation data of the sample image is determined based on the label information of the sample image.
11. The method according to claim 1, It is characterized in that The performing visual feature extraction processing on the image to be understood to obtain the visual features of the image to be understood includes: Performing image feature extraction processing on the image to be understood to obtain image features of the image to be understood; Perform visual encoding processing on the image features to obtain visual features of the image to be understood.
12. A device for determining image description text, It is characterized in that include: an extraction unit, configured to, after acquiring the image to be understood, perform visual feature extraction processing on the image to be understood to obtain the visual features of the image to be understood; a determination unit, configured to determine, according to the visual features, visual information representation data of the image to be understood and reference information representation data of the image to be understood, respectively, wherein the reference information representation data includes part or all of image classification result representation data, image target detection result representation data, and image segmentation result representation data; A fusion unit, configured to fuse the visual information representation data and the reference information representation data to obtain image information representation data of the image to be understood; A generation unit is used to perform text generation processing on the image information representation data to obtain image description text of the image to be understood, wherein the text generation processing is implemented by using a text generation model, and the text generation model is determined based on a trained model to be trained. The training process of the model to be trained includes updating the model to be trained based on the predicted description text of the sample image and the actual description text of the sample image. The predicted description text is obtained by the model to be trained processing the sample image and the label type representation data of the sample image. The label type representation data is used to describe the type of label information of the sample image, and the actual description text is used to explain the label information in a text description manner. The label information includes part or all of an image description text label, an image classification label, an image target detection label, and an image segmentation label, and the type of the image description text label, the type of the image classification label, the type of the image target detection label, and the type of the image segmentation label are all different.
13. A device, It is characterized in that The device comprises a processor and a memory: The memory is used to store computer programs; The processor is configured to execute the method of any one of claims 1 to 11 according to the computer program.
14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 11.
15. A computer program product, It is characterized in that When the computer program product is executed on a terminal device, the terminal device executes the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Training method of medical image report generation model and image report generation method
CN112992308A