Dialogue method, device, equipment and storage medium based on multimodal large model
Through the multimodal large model combined with visual information and pop-up window detection, the existing question-and-answer system lacks understanding of user intentions in professional software, and achieves more accurate and efficient operation guidance.
Patent Information
- Application Number
- CN202510918795.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-03
AI Technical Summary
The existing question-and-answer system lacks understanding of user visual information in the professional software operation interface, which makes it difficult to accurately capture user intentions, and lacks automated mechanisms to identify and utilize pop-up information, resulting in insufficient operation guidance.
The dialogue method based on multimodal large model is adopted, combining user text input and visual information on the operation interface, and rewriting query text using the visual language model, combining pop-up window detection and vector database retrieval to generate accurate answer text.
By combining visual information and pop-up window detection, the system can have a deeper understanding of user intentions, provide more accurate and efficient operation guidance, and improve the intelligent assistance level of professional software.
Smart Images

Figure CN120448507B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to technical fields such as deep learning, large language models, question-answering systems, and home decoration design software. Background Art
[0002] In recent years, artificial intelligence technologies, represented by large language models (LLMs), have experienced rapid development and are widely used in building various intelligent question-answering systems. To overcome the inherent problems of large language models, such as knowledge staleness and the tendency to generate hallucinations, Retrieval Augmented Generation (RAG) technology has emerged. This technology combines large language models with external knowledge bases, retrieving relevant information from the knowledge base before generating answers. This significantly improves the accuracy and reliability of question-answering systems in specific domains.
[0003] Following this trend, many professional software tools, especially those in the fields of computer-aided design (CAD) and home decoration design, have also begun to integrate built-in intelligent question-and-answer systems, aiming to provide users with real-time operational guidance and problem-solving solutions to improve user experience and learning curve. Summary of the Invention
[0004] The present disclosure provides a multimodal large model-based dialogue method, apparatus, device, and storage medium to solve or alleviate one or more technical problems in the prior art.
[0005] In a first aspect, the present disclosure provides a multimodal large model-based dialogue method, comprising:
[0006] In response to a dialog operation on the front-end interface, an input original query text and an interface image of the front-end interface are obtained; wherein the original query text includes inquiry information about the target event;
[0007] Using the visual language model, the original query text is rewritten according to the visual information contained in the interface image to obtain the enhanced query text;
[0008] Retrieving external knowledge results from a vector database based on the original query text and the enhanced query text; wherein the external knowledge results are obtained by reordering the recall results of the original query text and the recall results of the enhanced query text;
[0009] The enhanced query text and external knowledge results are input into the question-answering model to obtain the answer text; wherein the answer text contains guidance information or operation process information for calling the tool module in the front-end interface to solve the target event.
[0010] In a second aspect, the present disclosure provides a conversation device based on a multimodal large model, comprising:
[0011] An acquisition module, configured to acquire the input original query text and the interface image of the front-end interface in response to the dialogue operation of the front-end interface; wherein the original query text includes the query information about the target event;
[0012] The rewriting module is used to rewrite the original query text based on the visual information contained in the interface image using the visual language model to obtain the enhanced query text;
[0013] A retrieval module is used to retrieve external knowledge results from the vector database based on the original query text and the enhanced query text; wherein the external knowledge results are obtained by re-ranking the recall results of the original query text and the recall results of the enhanced query text;
[0014] A generation module is used to input the enhanced query text and external knowledge results into the question-answering model to obtain the answer text; wherein the answer text contains guidance information or operation process information for calling the tool module in the front-end interface to solve the target event.
[0015] According to a third aspect, an electronic device is provided, including:
[0016] at least one processor; and
[0017] a memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.
[0019] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0020] In a fifth aspect, a computer program product is provided, comprising a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.
[0021] The beneficial effects of the technical solution provided by the present disclosure include at least: it can combine the user's text input and real-time visual information of the operation interface to more deeply understand the user's intention, thereby providing more accurate and effective answers.
[0022] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided in accordance with the present disclosure and should not be regarded as limiting the scope of the present disclosure.
[0024] Figure 1 1 is a flow chart of a multimodal large model-based dialogue method according to an embodiment of the present disclosure;
[0025] Figure 2 This is a flow chart of pop-up window detection in a multimodal large model-based conversation method according to an embodiment of the present disclosure;
[0026] Figure 3 1 is a schematic structural diagram of a multimodal large model-based conversation device according to an embodiment of the present disclosure;
[0027] Figure 4 It is a block diagram of an electronic device used to implement the multimodal large model-based dialogue method of an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] The present disclosure will be described in further detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0029] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, circuits, etc. well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.
[0030] In the related art, current question-answering systems, especially in scenarios that are closely integrated with specific software operating interfaces, still have some shortcomings in the retrieval link of external knowledge bases. First, many question-answering systems mainly or only use the original text entered by the user as the retrieval basis when performing knowledge retrieval. This approach ignores other modal information that the user may also provide, such as screenshots of the software operating interface. When the user's question content is unclear, such as "How do I handle this?" or "How do I achieve this effect?", the system will find it difficult to accurately capture the user's true intentions due to the lack of understanding of the visual information of the interface.
[0031] In order to integrate visual information, general multimodal large models have emerged in the existing technology, which can describe the content of images. However, these general models show certain limitations when applied to auxiliary question-answering scenarios of professional software. On the one hand, when describing images in specific fields (such as home decoration design renderings or software graphical user interfaces (UI)), the language generated by general models is relatively colloquial and usually does not contain professional terms in this vertical field, resulting in low subsequent retrieval accuracy based on this description. On the other hand, when faced with complex software interface screenshots, when general models identify graphic elements that are not clearly labeled, their recognition order and focus are relatively random, and they lack internal knowledge of the recognition priority of key elements in specific software application scenarios.
[0032] Furthermore, when using professional software, users often encounter various system pop-ups, such as error reports, warnings, or feature prompts. Users often need to manually read, understand, and enter keywords in the pop-up window to search for solutions in the help center, a cumbersome process. Existing Q&A processes generally lack an automated mechanism to proactively identify pop-up information contained in user screenshots and directly use its content for knowledge retrieval and question answering, thus failing to provide the most efficient and direct guidance.
[0033] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, the embodiments of the present disclosure provide a dialogue method based on a multimodal large model. By utilizing the technical solutions of the embodiments of the present disclosure, the user's text input and real-time visual information of the operation interface can be combined to more deeply understand the user's intention, thereby providing more accurate and effective answers.
[0034] Example 1
[0035] The embodiments of the present disclosure provide a multimodal large model-based dialogue method, which can be applied to professional computer-aided design software, especially home decoration design software. Figure 1 FIG is a flow chart of a multimodal large model-based dialogue method according to an embodiment of the present disclosure. Figure 1 As shown, the method at least includes:
[0036] S110: In response to the dialogue operation on the front-end interface, the input original query text and the interface image of the front-end interface are obtained, wherein the original query text includes query information about the target event.
[0037] In a specific application scenario, a user is designing an interior using the front-end interface (the user interface) of a home design software. When the user encounters a question, they can initiate a conversation through the software's built-in dialog box. In response to this action, the system simultaneously captures the original query text entered by the user in the dialog box and automatically captures the current front-end interface as an image. The original query text typically contains the user's inquiry about a target event.
[0038] S120: Using the visual language model, the original query text is rewritten according to the visual information contained in the interface image to obtain an enhanced query text.
[0039] In the disclosed embodiments, the Visual Large Language Model (VLM) is an artificial intelligence model capable of understanding and processing both image and textual information. The VLM analyzes the interface image acquired in S110, extracts the visual information contained therein, and combines it with the original query text to generate a rewritten enhanced query text. The enhanced query text uses visual information to enhance understanding of the user's intent and / or target event.
[0040] S130: Retrieve external knowledge results from the vector database based on the original query text and the enhanced query text, wherein the external knowledge results are obtained by re-ranking the recall results of the original query text and the recall results of the enhanced query text.
[0041] Using the original query text obtained in S110 and the enhanced query text generated in S120, a search is performed in a pre-built vector database. A vector database is a data system specifically used to store and query high-dimensional vectors, in which the vectors are converted from a large number of knowledge documents. The purpose of this step is to find the help information most relevant to the user's question and form external knowledge results. The original query text and the enhanced query text can be searched in the vector database separately to obtain two parts of recall results. The two parts of the recall results can be re-sorted based on and finally obtain the external knowledge results.
[0042] S140: Input the enhanced query text and the external knowledge results into the question-answering model to obtain an answer text, wherein the answer text includes guidance information or operation process information for calling the tool module in the front-end interface to solve the target event.
[0043] The external knowledge results obtained in S130 and the enhanced query text in S120 are input into a question-answering model (typically a large language model). Alternatively, the original query text, the enhanced query text, and the external knowledge results are all input into the question-answering model simultaneously. Based on this input information, the question-answering model generates a natural language response text. This response text content is highly instructive, for example, containing detailed guidance or operational process information for invoking specific tool modules (such as "lofting" and "modeling" functional components) in the front-end interface to resolve the user's target event.
[0044] According to the solution of the embodiment of the present disclosure, it is possible to combine the user's text input and real-time visual information of the operation interface to more deeply understand the user's intention, thereby providing more accurate and effective operation guidance, and improving the intelligent assistance level of the design software.
[0045] Example 2
[0046] This embodiment is intended to specifically explain the target event first mentioned in Example 1. The target event is the core problem that the user hopes to solve through the query information. In the embodiment of this disclosure, it can be specifically divided into the following two types:
[0047] 1. Operation goal: The operation goal refers to the user's desire to use the software's functions to complete a specific design task or achieve a certain visual effect.
[0048] Example 1: When designing a living room, a user wants to create an irregular, curved ceiling. In this case, the original query text may be "How to make a curved ceiling?", and the target event is to complete the operation goal of "making a curved ceiling".
[0049] Example 2: A user sees a reference image and hopes to replicate a wall effect in his own design. He may ask, "How can I make the wall have a micro-cement texture?" The target event is the action goal of "achieving a micro-cement wall effect."
[0050] 2. Abnormal state: Abnormal state refers to the unexpected software state encountered by users during operation, which usually hinders their normal work process, such as various pop-up windows.
[0051] Example 3: When the user is performing a rendering operation, the software pops up a dialog box prompting "Rendering failed, error code: 1053". The user does not understand its meaning and solution, and may ask "What does this pop-up window mean?" The target event is to resolve the abnormal state of "Error code 1053".
[0052] Example 3
[0053] This embodiment aims to elaborate on step S120 (rewriting the original query text) in embodiment 1.
[0054] In one possible implementation, S120 uses a visual language model to rewrite the original query text according to the visual information contained in the interface image to obtain an enhanced query text, further including:
[0055] S121. Using the visual language model, identify at least one of the following visual information from the interface image: scene design subject information, style information, and pattern information.
[0056] S122: Concatenate the visual information and the original query text with prompt words to obtain an enhanced query text, so as to enhance the ability to understand the user's intention and / or target event through the visual information.
[0057] In the disclosed embodiment, when analyzing interface images, the VLM will focus on identifying the following types of visual information that are helpful in understanding the design intent:
[0058] Scene design subject information: Identify the main components of the current design space, such as the "ceiling of the living room", "the wall of the bedroom" or "the cabinets of the kitchen".
[0059] Style information: Determine the overall style of the current design based on the furniture, color scheme, lines and other elements in the interface, such as "modern minimalist style" or "new Chinese style".
[0060] Style information: Identify the visual style of the specific object that the user is currently editing or selecting, such as a "curved wall outline" or a "European headboard with carvings".
[0061] After identifying the above visual information, a prompt concatenation operation will be performed to merge this structured visual information with the user's original, more colloquial query text to generate an enhanced query text with richer information and clearer intent.
[0062] Example: Continuing with Example 1 in the above embodiment, the user's original query text is "How to make a curved ceiling?".
[0063] VLM analyzes the interface image and identifies the following visual information: {scene subject: "living room ceiling", style: "modern simplicity", pattern: "irregular closed curve in the central area of the ceiling"}.
[0064] The system splices prompt words and rewrites the original query text into an enhanced query text: "In a modern minimalist living room, how can I use modeling tools to create an irregular curved ceiling based on the curved outline drawn on the ceiling?"
[0065] Compared with the original query text, this enhanced query text contains scenes, styles and specific styles, which greatly enhances the ability to understand the user's target events and lays the foundation for subsequent accurate retrieval.
[0066] According to the solution of the embodiment of the present disclosure, the ability to understand the user's intention and / or the target event is enhanced through visual information.
[0067] Example 4
[0068] This embodiment adds a set of parallel pop-up event processing procedures based on embodiment 1. Figure 2 As shown, the process includes:
[0069] S150: Use the target detection model to detect pop-up windows. When the system obtains the interface image, it will simultaneously input the image into an object detection model. This model has been specially trained to identify whether the interface image contains pop-up windows. Specifically, it includes:
[0070] S151: Input the interface image into a model pre-trained based on the YOLO (You Only Look Once) series of algorithms.
[0071] S152: After model processing, if a pop-up window is detected, the bounding box coordinates of the pop-up window event (i.e., the exact location and size of the pop-up window in the image) and type label information (e.g., error, warning, info, representing error, warning, and prompt pop-ups, respectively) are output.
[0072] S160: Identify pop-up window text information. If the detection result of S150 is that there is a pop-up window event, the system will perform the following steps:
[0073] S161: Accurately crop the original interface image according to the bounding box coordinates to obtain a pop-up window image containing only the pop-up window content.
[0074] S162: Performing optical character recognition (OCR) on the pop-up window image to extract text information from the image to form structured pop-up window text information.
[0075] Example: Continuing with Example 3 in Example 2, an error box pops up on the user interface. When the object detection model identifies the pop-up window, it outputs its location coordinates and the error type label. The system then crops the pop-up window image based on the coordinates. Using OCR technology, it identifies the text within the pop-up window as: "Rendering failed, ray tracing engine initialization error." This process allows the system to automatically and accurately capture and understand software anomalies without requiring the user to manually paraphrase.
[0076] Example 5
[0077] This embodiment aims to elaborate on various implementations of step S130 (searching for external knowledge results in the vector database) in embodiment 1.
[0078] Method 1: Standard multi-channel recall and reordering process
[0079] S131: Convert the original query text into a first query vector. The system uses a first text vectorization model (such as BERT or Sentence-BERT) to convert the original query text (such as "How to make a curved ceiling?") into a high-dimensional first query vector.
[0080] S132: Convert the enhanced query text into a second query vector. Similarly, the system uses a second text vectorization model (which can be the same as or different from the first model) to convert the enhanced query text (e.g., "In a modern minimalist living room...") into a second query vector.
[0081] S133: Retrieve preliminary candidate documents from the vector database. The system uses the first query vector and the second query vector to perform a similarity search (e.g., cosine similarity calculation) in the vector database, respectively. The system retrieves a set of documents that are most similar to the query vectors, forming a first set of candidate documents and a second set of candidate documents, respectively. These two sets together constitute the preliminary set of candidate documents.
[0082] S134: Rearrange and merge candidate documents. To further improve the accuracy of the retrieval results, the system will perform rearrangement and fusion operations on the preliminary candidate document set.
[0083] In a possible implementation, S134 reorders and merges the candidate documents, further including:
[0084] S134a (Score): The system uses a re-ranking model (eg, Cross-Encoder) to score the relevance of each document in the first and second candidate document sets to the corresponding original / enhanced query text.
[0085] The Cross-Encoder inputs the query, document pair as a whole into the model. Specifically, it concatenates the query and document text (for example, formatted as [CLS] query text [SEP] document text [SEP]) and then feeds it into a deep neural network (such as the Transformer model like BERT). This structure allows the model to perform deep token-level interactions and attention calculations between the query and document within each layer, capturing richer and more fine-grained semantic connections than vector cosine similarity.
[0086] S134b (Reorder): Reorder the documents in the two collections based on the scores obtained in S134a. Specifically, within each collection, all documents are sorted in descending order based on their relevance scores obtained in S134a. This step optimizes the internal order of both document lists, with documents ranked higher in the rankings indicating greater relevance to the query.
[0087] S134c (Fusion): Merge the two re-ranked document lists using a specific strategy (such as weighted fusion, cross-ranking, etc.) to generate a high-quality, highly relevant final external knowledge result.
[0088] After re-sorting, two high-quality document lists are obtained. The purpose of this step is to intelligently merge these two lists into a final, optimal document list as the final external knowledge result. The disclosed embodiments can adopt a variety of fusion strategies, such as weighted fusion and inverse sort fusion.
[0089] In one example, weighted fusion assigns different weights to the scores of documents from different paths. For example, if the documents recalled by the enhanced query text are considered to be of higher quality, their scores can be given a higher weight. Specifically, for each document, its final fusion score, FinalScore, can be calculated using the following formula:
[0090] FinalScore = w1 * score_1 + w2 * score_2
[0091] Where score_1 and score_2 are the scores of the document in the first and second lists, respectively (if a document exists in only one list, the score in the other list is 0), and w1 and w2 are preset weights (for example, w1=0.4, w2=0.6). Finally, all documents are sorted by FinalScore.
[0092] In another example, the inverse ranking fusion algorithm does not depend on the specific score value but only on the ranking of the document in the list. The specific process can be:
[0093] For each document d, its RRF score is calculated using the formula RRF_Score(d) = Σ (1 / (k + rank_i(d))) , where rank_i(d) is the rank of document d in the i-th list, and k is a constant (e.g., 60) that reduces the influence of lower-ranked documents. The system calculates the RRF score for all documents and sorts them from highest to lowest.
[0094] Regardless of the fusion strategy used, step S134c ultimately produces a unique, optimally ranked list of documents. This list is the final external knowledge result provided to the question-answering model in S140, which is far superior to the initial, unranked result in terms of relevance, accuracy, and comprehensiveness.
[0095] Method 2: Retrieval process including pop-up information
[0096] When the system detects a pop-up event through the process of Example 4, the retrieval step S130 will adopt this enhanced method.
[0097] S131': Convert the original query vector. (Same as S131)
[0098] S132': Generate a multi-way query vector. In this step, the second text vectorization model integrates multiple pieces of information to generate a multi-way query vector with richer information dimensions. Specifically, the model combines the enhanced query text, the pop-up text information identified in Example 4 (e.g., "Rendering failed, ray tracing engine initialization error"), and the type label information of the pop-up event (e.g., error).
[0099] S133': Use a multi-way query vector for search. The system uses the original query vector and the multi-way query vector to search the vector database. When using a multi-way query vector for search, the type tag information of the pop-up event can be used for optimization. For example, if the type tag is "error," the system will prioritize searching the "Error Codes and Solutions" section of the help document because that section matches the type tag. The external knowledge results obtained in this way will include software help documents that are highly relevant to the pop-up event.
[0100] Through the above two methods, the embodiment of the present disclosure realizes a flexible and powerful retrieval capability, which can not only process common operation target inquiries, but also efficiently solve specific abnormal status problems.
[0101] The first text vectorization model and the second text vectorization model are two independent text vectorization models. They can use specially optimized models to process text inputs with different characteristics (length, information density, and structure), thereby maximizing the vectorization quality of each input and ultimately improving the overall effect of knowledge retrieval.
[0102] The entire process can be viewed as two parallel text processing pipelines:
[0103] Pipeline A: Raw query processing (short text)
[0104] Processing object: The original query text entered by the user. This type of text is usually very short and colloquial, but contains the user's most direct intention. For example: "How do I make this curved roof?"
[0105] Model used: First Text Vectorization Model. This model is optimized to capture the core semantics and keyword information in short sentences.
[0106] Output: A first query vector that accurately reflects the user's original intent.
[0107] Pipeline B: Multi-channel information processing (long text)
[0108] Processing object: A long, information-rich text pieced together by a program. It usually contains:
[0109] Enhanced query text: More detailed descriptive text generated by the Visual Language Model (VLM) based on the interface screenshot.
[0110] Pop-up text information: The specific text content recognized by OCR from the pop-up window on the interface (if any).
[0111] Pop-up type tag: metadata that identifies the nature of the pop-up, such as "error", "warning", etc. (if available).
[0112] Model used: Second Text Vectorization Model. This model is designed to process longer and more complex text paragraphs and can effectively extract and integrate information from different sources.
[0113] Output: A multi-way query vector that fully reflects the problem context (including vision, system state, etc.).
[0114] According to the solution of the embodiment of the present disclosure, the model optimized for short texts can better understand the subtle differences in word order and synonyms, while the model optimized for long texts is better at grasping themes and key information from large paragraphs of text. Processing the two separately can ensure that the query vectors produced by the two paths have the highest signal-to-noise ratio and the most accurate representation ability, laying a solid foundation for subsequent retrieval. On the other hand, the two paths form a complementary and fault-tolerant mechanism. The original query path can be regarded as a stable "benchmark" that is not affected by the possible understanding deviations of the VLM. In the event that the VLM's interpretation of the image is not accurate enough, resulting in a deviation in the direction of the enhanced query text, the original query path of the benchmark can still ensure that the system recalls the guaranteed relevant results. This design greatly enhances the stability and reliability of the entire question-answering system.
[0115] Example 6
[0116] This embodiment aims to provide a detailed description of the external knowledge document, which is the construction basis of the vector database mentioned in Embodiment 1.
[0117] The data in the vector database comes from the results of vectorization processing of a large number of diverse external knowledge documents. In the embodiment of the present disclosure, the external knowledge documents include at least one or more of the following:
[0118] Software help documentation: This is the core source of knowledge and can be the official guidance material provided by the software developer.
[0119] Knowledge documents in specific fields: For example, in the home decoration design scenario, professional documents such as interior design theory, materials knowledge, color matching principles, and ergonomic standards can be introduced.
[0120] Professional terminology corpus: covers professional vocabulary and their explanations in fields such as home decoration, construction, and software operation, used to improve the model's ability to understand terminology.
[0121] Software help documents are typically structured as a collection of documents with a pre-set structure to facilitate retrieval and understanding. Each document may contain one or more of the following types of information:
[0122] Module-Functional Description: This document details the functionality of a specific tool module within the software. For example, a document might specifically describe the purpose of the "Stakeout Tool," the meaning of all parameters, and a simple use case.
[0123] Event-Solution Mapping Information: This type of document typically has a title or keywords that directly correspond to a specific operational goal, and the content contains detailed steps to achieve that goal. For example, a document titled "How to Make a Curved Ceiling" might contain a step-by-step illustrated tutorial, from drawing lines to modeling the finished product.
[0124] Status-Cause Analysis Information: This type of document specifically explains various abnormal conditions. For example, a document titled "Detailed Explanation of Rendering Error Codes" would list various error codes like "1053" and "1077," and explain their possible causes and recommended solutions.
[0125] Example 7: Training Example of Visual Language Model (VLM)
[0126] This embodiment describes in detail the specific training process of the visual language model (VLM) described in step S120 of embodiment 1. This training aims to inject professional home improvement domain knowledge into the model to improve its ability to understand home improvement scenes and its response accuracy.
[0127] S210: Construct multimodal image-text pair training data.
[0128] In order to enable the VLM to understand specific concepts and terms in the field of home improvement, the embodiment of the present disclosure constructs a batch of image-text pair data with different purposes from the help center documents based on a large amount of usage experience and accumulated examples.
[0129] Example 1 (Injecting Recognition Priority): Create a data entry containing the text "How do I create a curved ceiling?" This data is used to train the model to prioritize and focus on the visual information of the ceiling area in the interface image when receiving inquiries related to ceilings.
[0130] Example 2 (Aligning Professional Terminology): We construct a data set containing the text "How to draw a partition wall." This data allows the model to learn and align professional terminology from the home improvement field and specific design software (such as Cool Home). This allows the model to use colloquial descriptions such as "partition wall" rather than "a wall in the room" when generating augmented query text.
[0131] Example 3 (Combined Use): Create a data item containing the text "How do I draw piano corner lines?" This data can simultaneously train the model to prioritize the specific element "corner lines" and learn the term "piano corner lines."
[0132] S220: Fine-tune the large visual language model. Based on a pre-trained general large visual language model, this disclosed embodiment utilizes the multimodal image-text data constructed in S210 to perform targeted fine-tuning on the model. This process expands the model's capabilities, enabling it to incorporate specialized knowledge in the home improvement field, thereby generating more accurate and professional enhanced query text when performing query rewriting tasks.
[0133] Example 8: Training Example of Target Detection Model
[0134] This embodiment describes in detail the specific training process of the target detection model described in step S150 of embodiment 4.
[0135] S310: Constructing pop-up detection training data. Considering the high cost of collecting and manually labeling images with pop-up windows from real user data, the embodiment of the present disclosure adopts an efficient data synthesis method to construct training data.
[0136] S311 (background preparation): Collect a large number of interface screenshots of design software. These screenshots have diverse content and can be used as rich background images.
[0137] S312 (foreground preparation): Since the types and content quantity of pop-up windows in the software are relatively fixed, these fixed pop-up window images can be used as foreground materials.
[0138] S313 (data synthesis): The program randomly pastes the pop-up window foreground in S312 onto the screenshot background in S311, thereby automatically and batch-building a large amount of training data with precise annotations (pop-up window location and type).
[0139] S320: Training the object detection and classification model. This disclosed embodiment uses advanced object detection models such as YOLOv10 for training on the synthetic dataset constructed in S310. The goal of this training is to enable the model to not only accurately detect the bounding box of pop-up windows within the interface image, but also accurately classify the pop-up window type (e.g., error, warning, prompt). This type information can be directly mapped to the partition of the help center document, enabling more accurate knowledge retrieval.
[0140] Example 9: Training Example of Rearrangement Model
[0141] This embodiment describes in detail the specific training process of the rearrangement model described in step S134a of embodiment 5.
[0142] S410: Construct positive and negative sample training data. In order to train the model to accurately determine the correlation between the query and the document, the embodiment of the present disclosure selects materials from the help center documents and constructs positive and negative sample pairs for training. The training data includes:
[0143] Query: Select the title of a help document as the query.
[0144] Positive sample: The document body content corresponding to the title is selected as the positive sample.
[0145] Negative samples: We select a document from other documents whose content is related to the query topic but not an exact match as a negative sample. Selecting relevant content as negative samples increases the training difficulty and forces the model to learn more fine-grained semantic differences.
[0146] Example:
[0147] Query: "How to make a sloped ceiling".
[0148] Positive sample document content: Contains detailed steps for making a sloped ceiling, such as "First, we enter the whole house hard decoration tools from the industry library in the lower left corner of the tool page", "Use the straight line tool to connect the four corners of the upper and lower inner lines to create a slope."
[0149] Negative sample document content: Contains content about making grille wall panels or other types of ceilings, such as "There are three drawing methods: Method 1... Use the rectangle tool to draw the ceiling profile," and "Method 2... Adsorb the grille wall panels and place them to create the effect."
[0150] S420: Training the ranking model. The disclosed embodiment uses a model with a single-tower cross-encoder architecture for training. During training, a (query, document) pair is input into the model at the same time, and the model outputs a correlation score representing the degree of match between the two. By training on the dataset constructed in S410, the model is optimized to output high scores for (query, positive sample) pairs and low scores for (query, negative sample) pairs. After the training is completed, the model can accurately score and rank the initially recalled document set (i.e., "fine ranking") in the retrieval process, thereby significantly improving the relevance and accuracy of the final retrieval results.
[0151] Finally, it should be noted that the above description is merely a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art will appreciate that they may modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents.
[0152] It should be emphasized that all details not elaborated in this specification, such as the training process, data preparation, specific programming language implementation, the operating system environment relied upon, or the basic deep learning framework configuration, are conventional technical means or known technologies well known to those skilled in the art. Based on the technical solution disclosed in this disclosure, those skilled in the art are fully capable of implementing this solution, and any omissions should not affect the establishment and implementation of the technical solution disclosed in this disclosure.
[0153] Figure 3 Schematic diagram of a multimodal large model-based dialogue device according to an embodiment of the present disclosure. Figure 3 As shown, the device includes:
[0154] The acquisition module 301 is used to acquire the input original query text and the interface image of the front-end interface in response to the dialogue operation of the front-end interface, wherein the original query text includes the query information about the target event.
[0155] The rewriting module 302 is used to rewrite the original query text using the visual language model according to the visual information contained in the interface image to obtain an enhanced query text.
[0156] The retrieval module 303 is used to retrieve external knowledge results in the vector database according to the original query text and the enhanced query text.
[0157] The generation module 304 is used to input the enhanced query text and the external knowledge results into the question-answering model to obtain the answer text. The answer text includes guidance information or operation process information for calling the tool module in the front-end interface to solve the target event.
[0158] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0159] Figure 4 FIG. 1 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 4 As shown, the electronic device includes: a memory 410 and a processor 420. The memory 410 stores a computer program that can be executed on the processor 420. The number of memory 410 and processor 420 can be one or more. The memory 410 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device performs the method provided by the above method embodiment. The electronic device may also include: a communication interface 430 for communicating with external devices and performing data exchange.
[0160] If the memory 410, processor 420, and communication interface 430 are implemented independently, the memory 410, processor 420, and communication interface 430 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0161] Optionally, in a specific implementation, if the memory 410, the processor 420 and the communication interface 430 are integrated on a chip, the memory 410, the processor 420 and the communication interface 430 can communicate with each other through an internal interface.
[0162] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.
[0163] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM).
[0164] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present disclosure is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, data subscriber line (DSL)) or wireless (e.g., infrared, Bluetooth, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)). It is worth noting that the computer-readable storage medium mentioned in the present disclosure may be a non-volatile storage medium, in other words, a non-transient storage medium.
[0165] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0166] In the description of the embodiments of the present disclosure, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0167] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means or. For example, A / B can mean A or B. "And / or" in this document is only a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0168] In the description of the embodiments of the present disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "plurality" means two or more.
[0169] The above description is merely an exemplary embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the scope of protection of the present disclosure.
Claims
1. A multimodal large model-based dialogue method, comprising: In response to a dialogue operation on the front-end interface, an input original query text and an interface image of the front-end interface are obtained; wherein the original query text includes inquiry information about the target event; Using a visual language model, the original query text is rewritten according to the visual information contained in the interface image to obtain an enhanced query text; According to the original query text and the enhanced query text, an external knowledge result is retrieved from a vector database; wherein the external knowledge result is obtained by re-ranking the recall result of the original query text and the recall result of the enhanced query text; Inputting the enhanced query text and the external knowledge result into a question-answering model to obtain an answer text; wherein the answer text includes guidance information or operation process information for calling a tool module in the front-end interface to resolve the target event; wherein the target event includes at least one of the following: Operation goal refers to the user's intention to use one or more tool modules in the front-end interface to achieve a specific design effect or complete a specific operation task; Abnormal status refers to an unexpected status displayed on the front-end interface that hinders normal user operations, including error pop-ups or warning prompts; The step of retrieving external knowledge results from a vector database based on the original query text and the enhanced query text includes: Converting the original query text into a first query vector using a first text vectorization model; Converting the enhanced query text into a second query vector using a second text vectorization model; Using the first query vector and the second query vector respectively, retrieve a preliminary candidate document set from a vector database; Rearranging and fusing the candidate document set to generate a final external knowledge result; The step of rearranging and fusing the candidate document set to generate a final external knowledge result includes: Scoring the relevance of a plurality of first candidate documents and a plurality of second candidate documents in the candidate document set, respectively; wherein the plurality of first candidate documents are recalled by the first query vector, and the plurality of second candidate documents are recalled by the second query vector; reordering the plurality of first candidate documents and the plurality of second candidate documents according to the scoring result; The re-ranked results are weighted fused or inversely sorted to generate the final external knowledge results.
2. The method according to claim 1, wherein The method of rewriting the original query text using the visual language model according to the visual information contained in the interface image to obtain the enhanced query text includes: Using a visual language model, identifying at least one of the following visual information from the interface image: scene design subject information, style information, and pattern information; The visual information and the original query text are concatenated with prompt words to obtain an enhanced query text.
3. The method according to claim 1, further comprising: Using the target detection model, perform pop-up detection on the interface image; When the interface image includes a pop-up event, identify the pop-up text information of the pop-up event.
4. The method according to claim 3, wherein: The method of using the target detection model to perform pop-up window detection on the interface image includes: Input the interface image into a target detection model to perform pop-up window detection; wherein the target detection model is pre-trained using the YOLO model; According to the output result of the target detection model, the bounding box coordinates and type label information of the pop-up event are obtained.
5. The method according to claim 4, wherein When the interface image includes a pop-up window event, identifying pop-up window text information of the pop-up window event includes: In the case where the interface image includes a pop-up window event, cropping the interface image according to the bounding box coordinates to obtain a pop-up window image; Perform image recognition on the pop-up window image to obtain pop-up window text information contained in the pop-up window image.
6. The method according to any one of claims 3 to 5, wherein: The step of retrieving external knowledge results from a vector database based on the original query text and the enhanced query text includes: Using a first text vectorization model, converting the original query text into an original query vector; Using a second text vectorization model, a multi-way query vector is obtained according to the enhanced query text, the pop-up window text information, and the type label information of the pop-up window event; Retrieving external knowledge results in a vector database using the original query vector and the multi-way query vector respectively; The external knowledge result includes a software help document associated with the pop-up window event, and the type tag information of the pop-up window event matches the partition of the help document.
7. The method according to claim 1, wherein The vector database is obtained by vectorizing external knowledge documents, and the external knowledge documents include at least one of software help documents, knowledge documents in a specific field, and professional terminology corpus.
8. The method according to claim 7, wherein: The software help document is a collection of documents with a preset structure, each of which contains at least one of the following information: Module-Function Description Information: Function description, parameter configuration instructions or usage examples of various tool modules on the front-end interface; Status-Cause Analysis Information: Cause analysis, explanations, and / or solutions for various pop-up events; Event-solution mapping information: contains titles or keywords associated with a specific target event, as well as step-by-step operational procedures or processing suggestions for resolving the target event.
9. A multimodal large model-based dialogue device, comprising: An acquisition module, configured to acquire an input original query text and an interface image of the front-end interface in response to a dialogue operation on the front-end interface; wherein the original query text includes inquiry information about a target event; A rewriting module, configured to rewrite the original query text using a visual language model according to the visual information contained in the interface image to obtain an enhanced query text; a retrieval module, configured to retrieve external knowledge results from a vector database based on the original query text and the enhanced query text; wherein the external knowledge results are obtained by reordering the recall results of the original query text and the recall results of the enhanced query text; A generation module is configured to input the enhanced query text and the external knowledge result into a question-answering model to obtain an answer text; wherein the answer text includes guidance information or operation process information for calling a tool module in the front-end interface to resolve the target event; wherein the target event includes at least one of the following: Operation goal refers to the user's intention to use one or more tool modules in the front-end interface to achieve a specific design effect or complete a specific operation task; Abnormal status refers to an unexpected status displayed on the front-end interface that hinders normal user operations, including error pop-ups or warning prompts; The step of retrieving external knowledge results from a vector database based on the original query text and the enhanced query text includes: Converting the original query text into a first query vector using a first text vectorization model; Converting the enhanced query text into a second query vector using a second text vectorization model; Using the first query vector and the second query vector respectively, retrieve a preliminary candidate document set from a vector database; Rearranging and fusing the candidate document set to generate a final external knowledge result; The step of rearranging and fusing the candidate document set to generate a final external knowledge result includes: Scoring the relevance of a plurality of first candidate documents and a plurality of second candidate documents in the candidate document set, respectively; wherein the plurality of first candidate documents are recalled by the first query vector, and the plurality of second candidate documents are recalled by the second query vector; reordering the plurality of first candidate documents and the plurality of second candidate documents according to the scoring result; The re-ranked results are weighted fused or inversely sorted to generate the final external knowledge results.
10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Visual question and answer method and platform based on knowledge enhancement
CN117972044A
Visual question and answer model and method based on multi-modal retrieval enhancement
CN119066174A