Graphical text dialogue method, model training method, device and electronic device

By performing object detection and visual feature extraction on the target image, generating prompt text and inputting into a large model, the problem of low accuracy of dialogue text in the prior art is solved, and a more accurate graphic and text dialogue method is realized.

CN117648419BActive Publication Date: 2025-07-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311588690.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-07-18
Estimated Expiration
2043-11-24

AI Technical Summary

Technical Problem

The accuracy of dialogue text in existing graphic and text dialogue methods is low, especially when the object detection results include object categories, entity description errors are prone to occur.

Method used

By performing object detection on the target image, the object detection results are obtained, and prompt text is generated based on the detection results and visual features, input it into the big model to generate dialogue text, and combining feature extraction and feature fusion techniques to improve text accuracy.

Benefits of technology

Improve the accuracy of dialogue text, especially when detecting object categories, avoid entity description errors, and enhance the fine-grained dialogue text and the accuracy of spatial positional relationship description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117648419B_ABST
    Figure CN117648419B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for graphic conversation, a method for model training, a device, and an electronic device, which relate to the field of artificial intelligence technology, particularly to the fields of computer vision, deep learning, and large model technology, and can be applied to scenarios such as content generation in artificial intelligence. The method includes: obtaining a target image to be conversed; performing object detection on the target image to obtain an object detection result of the target image; obtaining a prompt text based on the object detection result; and inputting the prompt text into a large model, and outputting a conversation text for the target image by the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of computer vision, deep learning, and large model technology, and can be applied to scenarios such as artificial intelligence content generation, and in particular to a text-graphic dialogue method, a model training method, an apparatus, an electronic device, a storage medium, and a computer program product. Background Art

[0002] At present, with the continuous development of artificial intelligence technology, large models have advantages such as good generalization and have been widely used in information extraction, text credibility assessment, machine translation and other fields. For example, in the field of image-text dialogue, dialogue text for images can be generated based on large models. However, the image-text dialogue method in related technologies has the problem of low accuracy of dialogue text. Summary of the invention

[0003] The present disclosure proposes a text-image dialogue method, a model training method, a device, an electronic device, a storage medium and a computer program product.

[0004] According to the first aspect of the present disclosure, a method for image-text dialogue is proposed, comprising: obtaining a target image to be communicated; performing target detection on the target image to obtain a target detection result of the target image; obtaining a prompt text based on the target detection result; inputting the prompt text into a large model, and having the large model output a dialogue text for the target image.

[0005] According to a second aspect of the present disclosure, a model training method is proposed, comprising: acquiring a sample image and a sample conversation text for the sample image; performing target detection on the sample image to obtain a sample target detection result of the sample image; obtaining a sample prompt text based on the sample target detection result; inputting the sample prompt text into a large model, and having the large model output a predicted conversation text for the sample image; and training the large model based on the predicted conversation text and the sample conversation text.

[0006] According to a third aspect of the present disclosure, a graphic-text dialogue device is proposed, comprising: a first acquisition module, used to acquire a target image to be communicated with; a detection module, used to perform target detection on the target image to obtain a target detection result of the target image; a second acquisition module, used to obtain a prompt text based on the target detection result; and a third acquisition module, used to input the prompt text into a large model, and the large model outputs a dialogue text for the target image.

[0007] According to a fourth aspect of the present disclosure, a model training device is provided, including: a first acquisition module configured to acquire a sample image and sample dialogue text for the sample image; a detection module configured to perform object detection on the sample image to obtain a sample object detection result of the sample image; a second acquisition module configured to obtain sample prompt text based on the sample object detection result; a third acquisition module configured to input the sample prompt text into a large model, and the large model outputs predicted dialogue text for the sample image; and a training module configured to train the large model based on the predicted dialogue text and the sample dialogue text.

[0008] According to a fifth aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the graphic-text dialogue method proposed in the first aspect and the model training method proposed in the second aspect above.

[0009] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the graphic-text dialogue method proposed in the first aspect and the model training method proposed in the second aspect above.

[0010] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the graphic-text dialogue method proposed in the first aspect and the model training method proposed in the second aspect above are implemented.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0013] Figure 1 is a schematic flowchart of a graphic-text dialogue method according to an embodiment of the present disclosure;

[0014] Figure 2 is a schematic diagram of a graphic-text dialogue method according to an embodiment of the present disclosure;

[0015] Figure 3 is a schematic flowchart of a graphic-text dialogue method according to another embodiment of the present disclosure;

[0016] Figure 4Schematic flowchart of the graphic conversation method according to another embodiment of the present disclosure;

[0017] Figure 5 Schematic flowchart of the graphic conversation method according to another embodiment of the present disclosure;

[0018] Figure 6 Schematic flowchart of the model training method according to an embodiment of the present disclosure;

[0019] Figure 7 Schematic structural diagram of the graphic conversation device according to an embodiment of the present disclosure;

[0020] Figure 8 Schematic structural diagram of the model training device according to an embodiment of the present disclosure;

[0021] Figure 9 Schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.

[0023] AI (Artificial Intelligence) is a technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Currently, AI technology has the advantages of high automation, high precision, and low cost, and has been widely applied.

[0024] Computer Vision refers to machine vision that uses cameras and computers to replace human eyes to identify, track, and measure targets, and further performs graphics processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. Computer Vision is a comprehensive discipline, including computer science and engineering, signal processing, physics, applied mathematics and statistics, neurophysiology, and cognitive science, etc.

[0025] DL (Deep Learning) is a new research direction in the field of ML (Machine Learning). It is to learn the internal laws and representation levels of sample data, enabling the machine to have the ability of analysis and learning like humans, and being able to recognize data such as text, images, and sounds. It is widely applied in speech and image recognition.

[0026] A large model refers to a machine learning model with a huge parameter scale and complexity, which requires a large amount of computing resources and storage space for training and storage, and often needs to perform distributed computing and special hardware acceleration technologies. Large models have stronger generalization ability and expression ability. Large models include LLM (Large Language Model). A large language model refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. Large language models can handle various natural language tasks, such as text classification, question answering, dialogue, etc., and are an important way to artificial intelligence.

[0027] Figure 1 The flowchart of the graphic conversation method according to an embodiment of the present disclosure is shown as Figure 1 shown, and the method includes:

[0028] S101, obtaining a target image to be conversed.

[0029] It should be noted that the execution subject of the graphic conversation method in the embodiment of the present disclosure can be a hardware device with data information processing ability and / or the necessary software for driving the hardware device to work. Optionally, the execution subject may include a workstation, a server, a computer, a user terminal, and other intelligent devices. Among them, the user terminal includes but is not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, etc.

[0030] It should be noted that there are no excessive limitations on the target image. For example, it may include two-dimensional images, three-dimensional images, etc.

[0031] In one implementation, obtaining the target image to be conversed includes receiving the target image sent by the client. For example, taking the server as the execution subject of the graphic conversation method, the client can obtain the target image based on the operation information of the user operating the client (such as the text input by the user, the icon clicked, the voice information of the user, etc.), and send the target image to the server. Correspondingly, the server can receive the target image sent by the client.

[0032] In one implementation, obtaining the target image to be conversed includes receiving a conversation request sent by the client and determining the target image based on the conversation request.

[0033] For example, if the conversation request carries the identifier of the target image, the image identified by the identifier of the target image can be used as the target image. It should be noted that there are no excessive limitations on the identifier. For example, it may include a name, a storage path, etc.

[0034] For example, if the conversation request carries the target image, the target image can be extracted from the conversation request.

[0035] S102. Perform object detection on the target image to obtain the object detection result of the target image.

[0036] It should be noted that for object detection of the target image, any object detection method in related technologies can be used to implement it, and no excessive limitations are imposed here. For example, the object detection method can include object detection methods based on traditional machine learning, object detection methods based on deep learning, object detection methods based on single-stage detection, object detection methods based on multi-object detection, etc.

[0037] It should be noted that no excessive limitations are imposed on the object detection result. For example, it can include the category of the object in the target image and / or the detection box information of the object. Among them, the object can include a human body, an animal, an object, etc., and the detection box information can include the position and size of the detection box. Among them, the position of the detection box can include the position of the center point of the detection box, the position of the vertex of the detection box, etc., and the size of the detection box can include the width, height, radius, etc. of the detection box.

[0038] In one implementation, performing object detection on the target image to obtain the object detection result of the target image includes inputting the target image into an object detection model, and the object detection model outputs the object detection result. For example, the object detection model includes an OVD (Open-Vocabulary Detector) model. The OVD model is not limited to the few categories labeled during the training process and can perform object detection of any number and any category, with good generalization.

[0039] S103. Obtain a prompt text based on the object detection result.

[0040] It should be noted that the prompt text refers to the text input into the large model, and the prompt text is used to obtain the dialogue text of the target image. For example, the prompt text can be in vector format.

[0041] In one implementation, obtaining a prompt text based on the object detection result includes extracting features from the object detection result to obtain object detection features, and obtaining the prompt text based on the object detection features. Thus, considering the object detection features, the prompt text can be obtained. Compared with directly obtaining the prompt text based on the object detection result, the object detection result can be further processed to obtain object detection features, which has the advantages of simplifying data and reducing the amount of calculation.

[0042] It should be noted that for feature extraction from the object detection result, any text feature extraction method in related technologies can be used to implement it, and no excessive limitations are imposed here. For example, the object detection result can be input into a text encoder, and the text encoder outputs the object detection features.

[0043] In one embodiment, based on the object detection result, a prompt text is obtained, including obtaining a question text, and based on the object detection result and the question text, the prompt text is obtained. Thus, the prompt text can be obtained by comprehensively considering the object detection result and the question text, making the prompt text more comprehensive.

[0044] It should be noted that the question text is not overly limited. For example, it can include text composed of any language such as Chinese, English, etc.

[0045] For example, taking the application scenario of image captioning as an example, the question text can include "Please give a specific description of this image".

[0046] For example, taking the application scenario of visual question answering as an example, the question text can include "What objects are there in this image".

[0047] For example, taking the application scenario of visual localization as an example, the question text can include "Please give the position of the objects in this image", "Where is the position of the vehicle in this image".

[0048] For example, taking the application scenario of document analysis as an example, the target image can include an image of a document (such as a scanned image, a photographed image, etc.), and the question text can include "Please give an abstract of this document", "What are the titles of this document".

[0049] For example, taking the application scenario of image review as an example, the question text can include "Is this image illegal", "Please give the risk level of this image", etc.

[0050] For example, taking the application scenario of medical diagnosis as an example, the question text can include "Based on this image, please give a description of the patient's condition".

[0051] In some examples, obtaining the question text includes receiving the question text sent by the client.

[0052] In some examples, obtaining the question text includes receiving a conversation request sent by the client and extracting the question text from the conversation request.

[0053] In some examples, obtaining the question text includes performing text recognition on the target image to obtain the question text. It should be noted that for performing text recognition on the target image, any image text recognition method in related technologies can be used to implement it, and no excessive limitation is made here. For example, the image text recognition method can include OCR (Optical Character Recognition), etc.

[0054] S104, input the prompt text into the large model, and the large model outputs the conversation text for the target image.

[0055] It should be noted that the large model can be implemented using any large model in related technologies, and no excessive limitations are imposed here. For example, it can be a Transformer model. It should be noted that the Transformer model is a neural network model based on the self-attention mechanism. For example, the large model can be a large language model.

[0056] It should be noted that no excessive limitations are imposed on the dialogue text.

[0057] For example, taking the application scenario of image captioning as an example, the dialogue text may include the descriptive text of the target image.

[0058] For example, taking the application scenario of visual question answering as an example, the dialogue text may include "The object xx is included in this image".

[0059] For example, taking the application scenario of visual localization as an example, the dialogue text may include "The position of the object xx in this image is xx".

[0060] For example, taking the application scenario of document analysis as an example, the dialogue text may include "The summary of this document is xx" and "The title of this document includes xx".

[0061] For example, taking the application scenario of image review as an example, the dialogue text may include "This image is not in violation" and "The risk level of this image is xx".

[0062] For example, taking the application scenario of medical diagnosis as an example, the dialogue text may include "Based on this image, the patient's condition description is xx".

[0063] The proposed graphic-text dialogue method of the present disclosure obtains the target image to be dialogued, performs object detection on the target image to obtain the object detection result of the target image, obtains the prompt text based on the object detection result, inputs the prompt text into the large model, and the large model outputs the dialogue text for the target image. Thus, considering the object detection result of the target image, the prompt text can be obtained and input into the large model, and the large model can use the object detection result to generate the dialogue text, improving the accuracy of the dialogue text. Especially when the object detection result includes the category of the object, the large model can consider the category of the object to generate the dialogue text, avoiding the problem of incorrect entity description in the dialogue text.

[0064] Based on any of the above embodiments, regarding obtaining the prompt text based on the object detection result in step S103, it can be combined with Figure 2 For further understanding, Figure 2 is a schematic flowchart of the graphic-text dialogue method according to another embodiment of the present disclosure. As Figure 2 shown, the method includes:

[0065] S201, obtaining the target image to be dialogued.

[0066] S202. Perform object detection on the target image to obtain the object detection result of the target image.

[0067] For the relevant content of steps S201 - S202, refer to the above embodiments and will not be elaborated here.

[0068] S203. Extract features from the target image to obtain the visual features of the target image.

[0069] It should be noted that for feature extraction of the target image, any image feature extraction method in related technologies can be used to implement it, and no excessive limitation is made here.

[0070] In one implementation, the visual features include visual features at the object granularity. It should be noted that the visual features at the object granularity refer to the visual features of the objects in the target image, which are the visual features of a certain region of the target image. Compared with the visual features at the image granularity, the visual features at the object granularity are fine-grained visual features. Thus, when generating prompt texts, the visual features at the object granularity can be considered, and the large model can use the visual features at the object granularity to generate dialogue texts, helping the large model obtain fine-grained object perception ability and improving the accuracy of the description of the spatial position relationship in the dialogue text.

[0071] In one implementation, the visual features include visual features at the image granularity. It should be noted that the visual features at the image granularity refer to the visual features of the entire target image. Compared with the visual features at the object granularity, the visual features at the image granularity are coarse-grained visual features.

[0072] As Figure 3 shown, when extracting features from the target image to obtain the visual features of the target image, it includes inputting the visual features at the image granularity into the encoder in the object detection model, and the encoder in the object detection model outputs the visual features at the object granularity. Thus, through the encoder in the object detection model, the visual features at the image granularity can be processed to obtain the visual features at the object granularity.

[0073] In one implementation, the visual features include visual features at the image granularity. As Figure 3 shown, when extracting features from the target image to obtain the visual features of the target image, it includes inputting the target image into the visual encoder, and the visual encoder outputs the visual features at the image granularity. Thus, the visual encoder can be used to extract features from the target image to obtain the visual features at the image granularity. For example, the visual encoder includes ViT (Vision Transformer) in BLIP-2. It should be noted that BLIP-2 is a visual model and ViT is a visual encoder.

[0074] S204. Obtain a prompt text based on the object detection result and the visual feature.

[0075] In one implementation, obtaining a prompt text based on the object detection result and the visual feature includes concatenating the object detection result and the visual feature in a set order to obtain the prompt text.

[0076] In one implementation, the visual feature includes the visual feature at the object granularity and the visual feature at the image granularity. Obtaining a prompt text based on the object detection result and the visual feature includes concatenating the object detection result, the visual feature at the object granularity, and the visual feature at the image granularity in a set order to obtain the prompt text.

[0077] S205. Input the prompt text into a large model, and the large model outputs a dialogue text for the target image.

[0078] For the relevant content of step S205, reference can be made to the above embodiments, which will not be elaborated here.

[0079] The proposed graphic-text dialogue method in the present disclosure extracts features from the target image to obtain the visual feature of the target image, and obtains a prompt text based on the object detection result and the visual feature. Thus, the object detection result and the visual feature can be comprehensively considered to obtain the prompt text, making the prompt text more comprehensive, and further improving the accuracy of the dialogue text.

[0080] Based on any of the above embodiments, as Figure 3 shown, perform object detection on the target image to obtain the object detection result of the target image, including inputting the visual feature at the object granularity into the decoder in the object detection model, and the decoder outputs the object detection result to achieve the acquisition of the object detection result.

[0081] In the above embodiments, regarding obtaining a prompt text based on the object detection result and the visual feature in step S204, it can be further understood in combination with Figure 4 For Figure 4 is a schematic flowchart of the graphic-text dialogue method according to another embodiment of the present disclosure. As Figure 4 shown, the method includes:

[0082] S401. Obtain a target image to be dialogued.

[0083] S402. Perform object detection on the target image to obtain the object detection result of the target image.

[0084] S403. Extract features from the target image to obtain the visual feature of the target image.

[0085] S404. Obtain a question text.

[0086] For the relevant content of steps S401 - S404, reference can be made to the above - mentioned embodiments, which will not be elaborated here.

[0087] S405. Based on the object detection result, visual feature, and problem text, obtain a prompt text.

[0088] In one implementation, based on the object detection result, visual feature, and problem text, obtaining a prompt text includes concatenating the object detection result, visual feature, and problem text in a set order to obtain the prompt text.

[0089] In one implementation, based on the object detection result, visual feature, and problem text, obtaining a prompt text includes extracting features from the problem text to obtain text features, and based on the object detection result, visual feature, and text features, obtaining the prompt text. Thus, the prompt text can be obtained by comprehensively considering the object detection result, visual feature, and text features. Compared with directly obtaining the prompt text based on the problem text, text features can be further processed from the problem text, which has the advantages of simplifying data and reducing the amount of calculation.

[0090] It should be noted that for feature extraction of the problem text, any text feature extraction method in related technologies can be used to implement it, and no excessive limitation is made here. For example, the problem text can be input into a text encoder, and the text encoder outputs text features.

[0091] In some examples, based on the object detection result, visual feature, and text features, obtaining a prompt text includes concatenating the object detection result, visual feature, and text features in a set order to obtain the prompt text.

[0092] In some examples, the visual feature includes visual features at the object granularity and visual features at the image granularity. Based on the object detection result, visual feature, and text features, obtaining a prompt text includes extracting features from the object detection result to obtain object detection features, and concatenating the visual features at the image granularity, visual features at the object granularity, object detection features, and text features in the order of visual features at the image granularity, visual features at the object granularity, object detection features, and text features to obtain the prompt text.

[0093] S406. Input the prompt text into a large - model, and the large - model outputs a dialogue text for the target image.

[0094] For the relevant content of step S406, reference can be made to the above - mentioned embodiments, which will not be elaborated here.

[0095] The graphic conversation method proposed by the present disclosure obtains a question text and obtains a prompt text based on the object detection result, visual feature, and question text. Thus, the prompt text can be obtained by comprehensively considering the object detection result, visual feature, and question text, making the prompt text more comprehensive, and further improving the accuracy of the conversation text.

[0096] In the above embodiment, the visual feature includes visual features of at least one granularity. For example, the visual feature includes at least one of the visual feature of the object granularity and the visual feature of the image granularity.

[0097] Regarding obtaining the prompt text based on the object detection result, visual feature, and text feature, it can be combined with Figure 5 For further understanding, Figure 5 It is a schematic flowchart of the graphic conversation method according to another embodiment of the present disclosure. As Figure 5 shown, the method includes:

[0098] S501, obtain the target image to be conversed.

[0099] S502, perform object detection on the target image to obtain the object detection result of the target image.

[0100] S503, perform feature extraction on the target image to obtain the visual feature of the target image.

[0101] S504, obtain the question text.

[0102] S505, perform feature extraction on the question text to obtain the text feature.

[0103] For the relevant content of steps S501 - S505, reference can be made to the above embodiment, which will not be elaborated here.

[0104] S506, perform feature fusion on the visual feature of at least one granularity and the text feature to obtain a fusion feature.

[0105] It should be noted that for performing feature fusion on the visual feature of at least one granularity and the text feature, any feature fusion method in the related art can be used to implement it, and no excessive limitation is made here.

[0106] In the embodiment of the present disclosure, for performing feature fusion on the visual feature of at least one granularity and the text feature to obtain a fusion feature, the following several possible implementation manners may be included:

[0107] Method 1: Concatenate the visual feature of at least one granularity and the text feature to obtain a fusion feature.

[0108] In some examples, the visual features include visual features at the image granularity and visual features at the object granularity. Concatenating the visual features of at least one granularity and the text features to obtain a fused feature, including concatenating the visual features at the image granularity, the visual features at the object granularity, and the text features in the order of the visual features at the image granularity, the visual features at the object granularity, and the text features to obtain a fused feature.

[0109] Method 2: Perform feature mapping on the visual features of any granularity to obtain the visual mapping features of any granularity of the visual features of any granularity in the feature space where the text features are located, and perform feature fusion on the visual mapping features of at least one granularity and the text features to obtain a fused feature.

[0110] It should be noted that performing feature mapping on the visual features of any granularity can be implemented by using any feature mapping method in related technologies, and no excessive limitation is made here. For example, the feature mapping method may include a dimensionality reduction method, a dimensionality increase method, etc.

[0111] Method 3: Perform feature mapping on the text features to obtain the text mapping features of the text features in the feature space where the visual features are located, and perform feature fusion on the visual features of at least one granularity and the text mapping features to obtain a fused feature.

[0112] Method 4: Input the visual features of at least one granularity and the text features into a lightweight neural network model, and output a fused feature by the lightweight neural network model.

[0113] Thus, through the lightweight neural network model, the visual features of at least one granularity and the text features can be subjected to feature fusion to obtain a fused feature.

[0114] For example, as Figure 3 shown, the visual features of at least one granularity and the text features can be input into a lightweight neural network model, and a fused feature is output by the lightweight neural network model.

[0115] It should be noted that the lightweight neural network model may include Q-former (Querying Transformer) in BILP-2. It should be noted that Q-former is a lightweight Transformer model.

[0116] In some examples, as Figure 3As shown, input visual features and text features of at least one granularity into a lightweight neural network model, and the lightweight neural network model outputs fused features, including obtaining a set of Query vectors, inputting visual features, text features of at least one granularity, and the set of Query vectors into the lightweight neural network model, and the lightweight neural network model outputs fused features. Thus, the lightweight neural network model can utilize the set of Query vectors to perform feature fusion on visual features and text features of at least one granularity to obtain fused features.

[0117] It should be noted that the set of Query vectors is a set of learnable Query vectors. The set of Query vectors can be preset or updated dynamically. For example, the set of Query vectors can be used as the model parameters of the lightweight neural network model and updated during the training process of the lightweight neural network model.

[0118] In some examples, the visual features include visual features at the object granularity. Before outputting the fused features, it also includes obtaining the sum value of the dimension of the set of Query vectors and the dimension of the visual features at the object granularity as the target dimension, and compressing the dimension of the fused features to the target dimension. Thus, through the lightweight neural network model, feature compression can be performed on the fused features to reduce the dimension of the fused features, which has the advantages of simplifying data and reducing the amount of calculation, and the dimension of the compressed fused features is the sum value of the dimension of the set of Query vectors and the dimension of the visual features at the object granularity.

[0119] It should be noted that compressing the dimension of the fused features to the target dimension can be achieved by any feature compression method in related technologies, and no more limitations are made here. For example, dimensionality reduction processing can be performed on the fused features, or downsampling can be performed on the fused features.

[0120] S507, obtain a prompt text based on the object detection result and the fused features.

[0121] In one implementation, obtaining a prompt text based on the object detection result and the fused features includes concatenating the object detection result and the fused features in a set order to obtain the prompt text.

[0122] In one implementation, obtaining a prompt text based on the object detection result and the fused features includes extracting features from the object detection result to obtain object detection features, and concatenating the object detection features and the fused features in a set order to obtain the prompt text.

[0123] S508, input the prompt text into a large model, and the large model outputs a dialogue text for the target image.

[0124] The relevant content of step S508 can be referred to the above embodiments and will not be elaborated here.

[0125] The graphic conversation method proposed by the present disclosure, where the visual features include visual features of at least one granularity, perform feature fusion on the visual features of at least one granularity and the text features to obtain fused features, and based on the object detection result and the fused features, obtain a prompt text. Thus, feature fusion can be performed on the visual features of at least one granularity and the text features to obtain fused features, and by comprehensively considering the object detection result and the fused features, a prompt text can be obtained, which helps to improve the accuracy, robustness, and generalization of the graphic conversation method.

[0126] Figure 6 It is a schematic flowchart of a model training method according to an embodiment of the present disclosure. As Figure 6 shown, the method includes:

[0127] S601, obtain a sample image and a sample conversation text for the sample image.

[0128] S602, perform object detection on the sample image to obtain a sample object detection result of the sample image.

[0129] S603, obtain a sample prompt text based on the sample object detection result.

[0130] S604, input the sample prompt text into a large model, and the large model outputs a predicted conversation text for the sample image.

[0131] The relevant content of steps S601 - S604 can be referred to the above embodiments and will not be elaborated here.

[0132] It should be noted that the execution subject of the model training method in the embodiments of the present disclosure can be a hardware device with data information processing capabilities and / or the necessary software to drive the hardware device to work. Optionally, the execution subject may include workstations, servers, computers, user terminals, and other intelligent devices. Among them, the user terminal includes but is not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle terminals, etc.

[0133] It should be noted that for the relevant content of the sample image, reference can be made to the relevant content of the target image in the above embodiments; for the relevant content of the sample object detection result, reference can be made to the relevant content of the object detection result of the target image in the above embodiments; for the relevant content of the sample prompt text, reference can be made to the relevant content of the prompt text in the above embodiments; for the relevant content of the predicted conversation text and the sample conversation text, reference can be made to the relevant content of the conversation text of the target image in the above embodiments, and will not be elaborated here.

[0134] In one implementation, obtaining a sample prompt text based on the sample object detection result includes extracting features from the sample image to obtain sample visual features of the sample image, and obtaining the sample prompt text based on the sample object detection result and the sample visual features.

[0135] In some examples, the sample visual features include the sample visual features of the object granularity.

[0136] In some examples, based on the sample object detection result and the sample visual features, a sample prompt text is obtained, including obtaining a sample question text, and based on the sample object detection result, the sample visual features, and the sample question text, the sample prompt text is obtained.

[0137] S605, train the large model based on the predicted dialogue text and the sample dialogue text.

[0138] It should be noted that training the large model based on the predicted dialogue text and the sample dialogue text can be implemented by any model training method in related technologies, which will not be elaborated here. For example, the model training method may include LoRA (Low-Rank Adaptation), Instruction Fine-Tuning method. It should be noted that LoRA is a method for efficient fine-tuning of model parameters.

[0139] In one implementation, training the large model based on the predicted dialogue text and the sample dialogue text includes obtaining the loss function of the large model based on the predicted dialogue text and the sample dialogue text, and training the large model based on the loss function. It should be noted that the loss function is not overly limited. For example, it may include CE (Cross Entropy), MSE (Mean-Square Error), KL (Kullback-Leibler) divergence, contrast loss function, etc.

[0140] In one implementation, before training the large model based on the predicted dialogue text and the sample dialogue text, it further includes pre-training the large model based on the sample texts of multiple knowledge domains. Thus, the large model can be pre-trained based on the sample texts of multiple knowledge domains, so that the large language model can learn the sample texts of multiple knowledge domains during the pre-training process, enabling the large model to generate general text data.

[0141] It should be noted that the knowledge domains are not overly limited. For example, they may include medicine, meteorology, literature, advertising, etc.

[0142] The model training method proposed by the present disclosure obtains a sample image and sample dialogue text for the sample image, performs object detection on the sample image to obtain the sample object detection result of the sample image, obtains a sample prompt text based on the sample object detection result, inputs the sample prompt text into a large model, and the large model outputs a predicted dialogue text for the sample image. Based on the predicted dialogue text and the sample dialogue text, the large model is trained. Thus, during the training process, the large model can learn the relationship between the sample prompt text and the sample dialogue text, and the trained large model can generate dialogue text based on the prompt text for use in a graphic-text dialogue method.

[0143] In the technical solution of the present disclosure, the processing of collecting, storing, using, processing, transmitting, providing, and disclosing the user's personal information involved all complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0144] According to an embodiment of the present disclosure, the present disclosure also provides a graphic-text dialogue device for implementing the above graphic-text dialogue method.

[0145] Figure 7 It is a block diagram of a graphic-text dialogue device according to an embodiment of the present disclosure.

[0146] As Figure 7 shown, the graphic-text dialogue device 700 includes: a first acquisition module 701, a detection module 702, a second acquisition module 703, and a third acquisition module 704.

[0147] The first acquisition module 701 is used to acquire a target image to be dialogued.

[0148] The detection module 702 is used to perform object detection on the target image to obtain the target detection result of the target image.

[0149] The second acquisition module 703 is used to obtain a prompt text based on the target detection result.

[0150] The third acquisition module 704 is used to input the prompt text into a large model, and the large model outputs a dialogue text for the target image.

[0151] In an embodiment of the present disclosure, the second acquisition module 703 is further used to: extract features from the target image to obtain the visual features of the target image; and obtain the prompt text based on the target detection result and the visual features.

[0152] In an embodiment of the present disclosure, the visual features include visual features at the object granularity.

[0153] In one embodiment of the present disclosure, the visual feature further includes a visual feature of image granularity, and the second acquisition module 703 is further configured to: input the visual feature of image granularity into an encoder in the target detection model, and output the visual feature of object granularity by the encoder in the target detection model.

[0154] In one embodiment of the present disclosure, the detection module 702 is further configured to: input the visual feature of object granularity into a decoder in the target detection model, and output the target detection result by the decoder.

[0155] In one embodiment of the present disclosure, the visual feature includes a visual feature of image granularity, and the second acquisition module 703 is further configured to: input the target image into a visual encoder, and output the visual feature of image granularity by the visual encoder.

[0156] In one embodiment of the present disclosure, the second acquisition module 703 is further configured to: acquire a question text; obtain the prompt text based on the target detection result, the visual feature, and the question text.

[0157] In one embodiment of the present disclosure, the second acquisition module 703 is further configured to: extract features from the question text to obtain a text feature; obtain the prompt text based on the target detection result, the visual feature, and the text feature.

[0158] In one embodiment of the present disclosure, the visual feature includes visual features of at least one granularity, and the second acquisition module 703 is further configured to: perform feature fusion on the visual features of at least one granularity and the text feature to obtain a fusion feature; obtain the prompt text based on the target detection result and the fusion feature.

[0159] In one embodiment of the present disclosure, the second acquisition module 703 is further configured to: input the visual features of at least one granularity and the text feature into a lightweight neural network model, and output the fusion feature by the lightweight neural network model.

[0160] In one embodiment of the present disclosure, the second acquisition module 703 is further configured to: acquire a set of query vectors; input the visual features of at least one granularity, the text feature, and the set of query vectors into the lightweight neural network model, and output the fusion feature by the lightweight neural network model.

[0161] In one embodiment of the present disclosure, the visual feature includes the visual feature of the object granularity. Before outputting the fused feature, the second acquisition module 703 is further configured to: obtain the sum value of the dimension of the query vector set and the dimension of the visual feature of the object granularity as the target dimension; compress the dimension of the fused feature to the target dimension.

[0162] In one embodiment of the present disclosure, the target detection result includes the category of the object in the target image and / or the detection box information of the object.

[0163] The proposed image-text dialogue device of the present disclosure performs target detection on a target image to obtain a target detection result of the target image, obtains a prompt text based on the target detection result, inputs the prompt text into a large model, and the large model outputs a dialogue text for the target image. Thus, considering the target detection result of the target image, a prompt text can be obtained and input into the large model. The large model can use the target detection result to generate a dialogue text, improving the accuracy of the dialogue text. Especially when the target detection result includes the category of the object, the large model can consider the category of the object to generate a dialogue text, avoiding the problem of incorrect entity description in the dialogue text.

[0164] According to an embodiment of the present disclosure, the present disclosure also provides a model training device for implementing the above model training method.

[0165] Figure 8 It is a block diagram of a model training device according to an embodiment of the present disclosure.

[0166] As Figure 8 shown, the model training device 800 includes: a first acquisition module 801, a detection module 802, a second acquisition module 803, a third acquisition module 804, and a training module 805.

[0167] The first acquisition module 801 is configured to acquire a sample image and a sample dialogue text for the sample image;

[0168] The detection module 802 is configured to perform target detection on the sample image to obtain a sample target detection result of the sample image;

[0169] The second acquisition module 803 is configured to obtain a sample prompt text based on the sample target detection result;

[0170] The third acquisition module 804 is configured to input the sample prompt text into a large model, and the large model outputs a predicted dialogue text for the sample image;

[0171] The training module 805 is configured to train the large model based on the predicted dialogue text and the sample dialogue text.

[0172] The model training device proposed by the present disclosure obtains a sample image, performs object detection on the sample image to obtain the sample object detection result of the sample image, obtains a sample prompt text based on the sample object detection result, inputs the sample prompt text into a large model, and the large model outputs a predicted dialogue text for the sample image. Based on the predicted dialogue text and the sample dialogue text, the large model is trained. Thus, during the training process, the large model can learn the relationship between the sample prompt text and the sample dialogue text. Therefore, the trained large model can generate dialogue text based on the prompt text and be used in the image-text dialogue method.

[0173] According to an embodiment of the present disclosure, the present disclosure also proposes an electronic device, a readable storage medium, and a computer program product.

[0174] Figure 9 A schematic block diagram of an example electronic device that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0175] As Figure 9 shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0176] A plurality of components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 906, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0177] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the graphic-text dialogue method and the model training method. For example, in some embodiments, the graphic-text dialogue method and the model training method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the graphic-text dialogue method and the model training method described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the graphic-text dialogue method and the model training method by any other suitable means (e.g., by means of firmware).

[0178] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0179] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0180] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0181] To present an interaction with a user account, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user account can provide input to the computer. Other kinds of devices can also be used to present an interaction with the user account; for example, the feedback presented to the user account can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user account can be received in any form (including acoustic input, speech input, or tactile input).

[0182] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user account computer having a graphical user account interface or a web browser through which the user account can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0183] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0184] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, it implements the steps of the graphic conversation method and the model training method described in the above embodiments of the present disclosure.

[0185] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is made herein.

[0186] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for graphic conversation, comprising: Obtaining a target image to be conversed based on an identifier carried in a conversation request; Determining a target detection result corresponding to the target image through an OVD model in a target detection model; Performing feature extraction on the target image to obtain visual features at the image granularity, inputting the visual features at the image granularity into an encoder in the target detection model, outputting visual features at the object granularity by the encoder, obtaining question text from the conversation request or the target image, and obtaining a prompt text based on the target detection result, the visual features at the image granularity, the visual features at the object granularity, and the question text, where the prompt text is in vector format; Inputting the prompt text into a trained large model, and outputting conversation text for the target image by the trained large model.

2. The method according to claim 1, wherein The method for determining the target detection result further comprises: Inputting the visual features at the object granularity into a decoder in the target detection model, and outputting the target detection result by the decoder.

3. The method according to claim 1, wherein Performing feature extraction on the target image to obtain visual features at the image granularity, including: Inputting the target image into a visual encoder, and outputting the visual features at the image granularity by the visual encoder.

4. The method according to claim 1, wherein The obtaining the prompt text based on the target detection result, the visual features, and the question text includes: Performing feature extraction on the question text to obtain text features; Obtaining the prompt text based on the target detection result, the visual features, and the text features.

5. The method according to claim 4, wherein, Obtaining the prompt text based on the target detection result, the visual features at the image granularity, the visual features at the object granularity, and the question text includes: Performing feature fusion on visual features at at least one granularity and the text features to obtain fusion features; Obtaining the prompt text based on the target detection result and the fusion features.

6. The method according to claim 5, wherein The performing feature fusion on visual features at at least one granularity and the text features to obtain fusion features includes: Inputting visual features at at least one granularity and the text features into a lightweight neural network model, and outputting the fusion features by the lightweight neural network model.

7. The method according to claim 6, wherein, The inputting visual features at at least one granularity and the text features into a lightweight neural network model, and outputting the fusion features by the lightweight neural network model includes: Obtaining a query vector set; Inputting visual features at at least one granularity, the text features, and the query vector set into the lightweight neural network model, and outputting the fusion features by the lightweight neural network model.

8. The method according to claim 7, wherein, Before outputting the fusion features, it further comprises: Obtaining a sum value of the dimension of the query vector set and the dimension of the visual features at the object granularity as a target dimension; Compressing the dimension of the fusion features to the target dimension.

9. The method according to any one of claims 1-3, wherein The target detection result includes the category of the object in the target image, and / or, the detection box information of the object.

10. A model training method, comprising: Obtaining a sample image and sample conversation text for the sample image; Perform object detection on the sample image to obtain the sample object detection result of the sample image; Based on the sample object detection result, obtain a sample prompt text; Input the sample prompt text into a large model, and the large model outputs a predicted dialogue text for the sample image; Based on the predicted dialogue text and the sample dialogue text, train the large model to obtain a trained large model, and the trained large model determines the dialogue text of the target image as described in any one of claims 1-8.

11. A graphic-text dialogue device, comprising: A first acquisition module, configured to acquire a target image to be dialogued based on an identifier carried in a dialogue request; A detection module, configured to perform object detection on the target image corresponding to the target image through an OVD model in an object detection model to obtain the object detection result of the target image; A second acquisition module, configured to extract features from the target image to obtain visual features at the image granularity, input the visual features at the image granularity into an encoder in the object detection model, and the encoder outputs visual features at the object granularity, obtain question text from the dialogue request or the target image, and based on the object detection result, the visual features at the image granularity, the visual features at the object granularity, and the question text, obtain a prompt text, where the prompt text is in vector format; A third acquisition module, configured to input the prompt text into a trained large model, and the trained large model outputs a dialogue text for the target image.

12. The apparatus according to claim 11, wherein, The detection module is further configured to: Input the visual features at the object granularity into a decoder in the object detection model, and the decoder outputs the object detection result.

13. The apparatus according to claim 11, wherein, The second acquisition module is further configured to: Input the target image into a visual encoder, and the visual encoder outputs the visual features at the image granularity.

14. The apparatus according to claim 11, wherein, The second acquisition module is further configured to: Extract features from the question text to obtain text features; Based on the object detection result, the visual features, and the text features, obtain the prompt text.

15. The apparatus according to claim 14, wherein, The second acquisition module is further configured to: Perform feature fusion on visual features at at least one granularity and the text features to obtain fusion features; Based on the object detection result and the fusion features, obtain the prompt text.

16. The apparatus according to claim 15, wherein, The second acquisition module is further configured to: Input visual features at at least one granularity and the text features into a lightweight neural network model, and the lightweight neural network model outputs the fusion features.

17. The apparatus according to claim 16, wherein, The second acquisition module is further configured to: Obtain a query vector set; Input visual features at at least one granularity, the text features, and the query vector set into the lightweight neural network model, and the lightweight neural network model outputs the fusion features.

18. The apparatus according to claim 17, wherein Before outputting the fusion features, the second acquisition module is further configured to: Obtain the sum value of the dimension of the query vector set and the dimension of the visual features at the object granularity as the target dimension; Compress the dimension of the fusion features to the target dimension.

19. The apparatus according to any one of claims 11-13, wherein, The target detection result includes the category of the object in the target image, and / or, the detection box information of the object.

20. A model training device, comprising: A first acquisition module, configured to acquire a sample image and a sample dialogue text for the sample image; A detection module, configured to perform target detection on the sample image to obtain a sample target detection result of the sample image; A second acquisition module, configured to obtain a sample prompt text based on the sample target detection result; A third acquisition module, configured to input the sample prompt text into a large model, and output a predicted dialogue text for the sample image by the large model; A training module, configured to train the large model based on the predicted dialogue text and the sample dialogue text to obtain a trained large model, and determine the dialogue text of the target image according to any one of claims 1-9 by the trained large model.

21. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.

23. A computer program product, comprising a computer program, wherein the computer program implements the method according to any one of claims 1-10 when executed by a processor.

Citation Information

Patent Citations

  • Visual question and answer method and device, electronic equipment and storage medium

    CN114707017A

  • Question and answer method and device

    CN114722178A