Knowledge Q&A Control Method, System and Electronic Device Based on Image Data
By preprocessing image data, generating description text in the knowledge question and answer system, and dynamically referring pictures in the original document, the problem that traditional systems are difficult to process image data is solved, and answer generation is achieved with pictures and texts, improving user experience and answer quality.
Patent Information
- Application Number
- CN202510038888.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Traditional knowledge question and answer systems are difficult to effectively process documents containing picture data, and cannot make full use of picture information in answers, resulting in limited answer integrity and accuracy.
The semantic results of the image data are obtained through preprocessing steps and description text are generated. A vector database is built using markup language, and pictures in the original document are referenced dynamically, and answers are generated with pictures and text.
Improves user experience and answer quality, and can make full use of image data during the knowledge Q&A process to provide more complete and intuitive answers.
Smart Images

Figure CN119441523B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a knowledge - based question - answering control method, system, and electronic device based on picture data. Background Art
[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have achieved remarkable results in the field of natural language processing. In knowledge - based question - answering related systems, LLMs combined with Retrieval - Augmented Generation (RAG) technology can provide efficient and intelligent question - answering services. Traditional knowledge - based question - answering systems mainly process text information, but have limited capabilities in processing document data containing pictures, charts, etc. However, in many fields, such as education, medical care, industrial design, etc., processing picture data in documents is crucial for understanding and answering questions. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a knowledge - based question - answering control method, system, and electronic device based on picture data. This method can make full use of picture data in knowledge - based question - answering related systems, dynamically reference pictures in the original document during the knowledge - answering process, and provide answers with pictures and texts, improving user experience and answer quality.
[0004] In a first aspect, an embodiment of the present invention provides a knowledge - based question - answering control method based on picture data, the method comprising:
[0005] A pre - processing step: obtaining a document to be processed containing picture data, determining the semantic result of the picture data according to the document level of the document to be processed, generating a description text corresponding to the picture data based on the semantic result, and determining a markup language corresponding to the picture data according to the description text;
[0006] A vector - based interaction step: constructing a vector database corresponding to the document to be processed based on the markup language, and obtaining a corresponding query text block from the vector database based on a knowledge - based question - answering query instruction;
[0007] An answer - generating step: obtaining context data corresponding to the query text block according to a preset context assembly strategy, and generating a knowledge - based question - answering result corresponding to the knowledge - based question - answering query instruction by using the markup language included in the context data.
[0008] Optionally, the pre - processing step includes:
[0009] A document structure analysis step: obtaining the document level corresponding to the document to be processed, and determining the context information of the picture data based on the document level;
[0010] Steps for picture semantic understanding: Determine the semantic result corresponding to the picture data based on the context information of the picture data, and use the semantic result to determine the visual analysis result corresponding to the picture data;
[0011] Steps for picture description generation: Obtain the description text corresponding to the picture data based on the visual analysis result and the context information;
[0012] Steps for picture link conversion: Use the markup format parameters corresponding to the markup language to convert the picture data and the description text into the markup language.
[0013] Optionally, the document structure analysis steps include:
[0014] Obtain the document hierarchy after parsing the document structure of the document to be processed;
[0015] Obtain the chapter titles and paragraphs of the document to be processed based on the document hierarchy;
[0016] Obtain the picture data included in the paragraph, and extract the description text corresponding to the picture data;
[0017] Use the chapter title and the description text to determine the context information of the picture data.
[0018] Optionally, the picture semantic understanding steps include:
[0019] Obtain the picture type data and the picture function data of the picture data based on the context information, and use the picture type data and the picture function data to determine the first description data corresponding to the picture data;
[0020] Obtain the product name data, the operation step data, the structure description data, the annotation prompt data, and the warning information data of the picture data based on the above context information, and use the product name data, the operation step data, the structure description data, the annotation prompt data, and the warning information data to determine the second description data corresponding to the picture data;
[0021] Use the first description data and the second description data to determine the semantic result, and after inputting the semantic result into the trained multi-modal model, obtain the visual analysis result output by the multi-modal model.
[0022] Optionally, the picture description generation steps include:
[0023] Obtain the position structure information of the picture data in the document to be processed based on the context information;
[0024] Obtain the identification annotation information of the picture data, and use the identification annotation information, the position structure information, and the visual analysis result to generate the description information of the picture data;
[0025] After inputting the description information into the multi-modal model, obtain the description text of the output of the multi-modal model.
[0026] Optionally, the picture link conversion step includes:
[0027] Determine the markup format parameters corresponding to the markup language;
[0028] Based on the markup format parameters, convert the picture data into a format link, and based on the markup format parameters, convert the description text into a format text;
[0029] Generate the markup language corresponding to the picture data using the format link and the format text.
[0030] Optionally, the vectorization interaction step includes:
[0031] After updating the markup language to the document to be processed, divide the document to be processed into multiple text blocks based on the semantic result of the markup language;
[0032] Use the text blocks to construct a vector database, and generate a vector index for the text blocks based on the vector database;
[0033] Determine the query vector corresponding to the knowledge Q&A question instruction, and obtain the vector index corresponding to the query vector;
[0034] Determine the query text block according to the text blocks corresponding to the vector index.
[0035] Optionally, the response generation step includes:
[0036] Determine the prompt data corresponding to the query text block, and after assembling the prompt data and the query text block using the context assembly strategy, obtain the context data;
[0037] After inputting the context data into a preset language model, obtain the output result of the language model;
[0038] Parse and obtain the markup language included in the output result, and generate a knowledge Q&A answer result using the markup language.
[0039] In a second aspect, the present invention provides a knowledge Q&A control system based on picture data, and the system includes:
[0040] A preprocessing module: used to obtain a document to be processed containing picture data, determine the semantic result of the picture data according to the document level of the document to be processed, generate a description text corresponding to the picture data based on the semantic result, and determine the markup language corresponding to the picture data according to the description text;
[0041] A vectorization interaction module: used to construct a vector database corresponding to the document to be processed based on the markup language, and obtain the corresponding query text block from the vector database based on the knowledge Q&A question instruction;
[0042] Answer generation module: configured to obtain context data corresponding to a query text block according to a preset context assembly strategy, and generate a knowledge Q&A answer result corresponding to a knowledge Q&A question instruction by using the markup language included in the context data.
[0043] In a third aspect, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor. The processor executes the computer-executable instructions to implement the steps of the knowledge Q&A control method based on picture data provided in the first aspect.
[0044] In a fourth aspect, an embodiment of the present invention further provides a storage medium. The storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the steps of the knowledge Q&A control method based on picture data provided in the first aspect.
[0045] A knowledge Q&A control method, system and electronic device provided by an embodiment of the present invention, in the process of knowledge Q&A control, first obtains a document to be processed containing picture data, determines the semantic result of the picture data according to the document level of the document to be processed, generates a description text corresponding to the picture data based on the semantic result, and determines the markup language corresponding to the picture data according to the description text; then constructs a vector database corresponding to the document to be processed based on the markup language, and obtains a corresponding query text block from the vector database based on a knowledge Q&A question instruction; finally, obtains context data corresponding to the query text block according to a preset context assembly strategy, and generates a knowledge Q&A answer result corresponding to the knowledge Q&A question instruction by using the markup language included in the context data. This method can make full use of picture data in a knowledge Q&A related system, can dynamically reference pictures in the original document during the knowledge answering process, and provide answers with both pictures and texts, improving the user experience and the quality of the answers.
[0046] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.
[0047] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0048] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0049] Figure 1 Flowchart of a method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0050] Figure 2 Flowchart of step S101 in a method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0051] Figure 3 Flowchart of step S201 in a method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0052] Figure 4 Flowchart of step S202 in a method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0053] Figure 5 Flowchart of step S203 in a method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0054] Figure 6 Flowchart of step S204 in a method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0055] Figure 7 Flowchart of step S102 in another method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0056] Figure 8 Flowchart of step S103 in another method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0057] Figure 9 Flowchart of another method for controlling knowledge Q&A based on picture data provided by an embodiment of the present invention;
[0058] Figure 10 Schematic diagram of a knowledge Q&A control system based on picture data provided by an embodiment of the present invention;
[0059] Figure 11 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.
[0060] Icon:
[0061] 1010 - Pre - processing module; 1020 - Vectorized interaction module; 1030 - Response generation module;
[0062] 101 - Processor; 102 - Memory; 103 - Bus; 104 - Communication interface. Detailed implementation manners
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0064] With the rapid development of artificial intelligence technology, large language models (LLMs) have achieved remarkable results in the field of natural language processing. In knowledge - based question - answering related systems, LLMs combined with retrieval - augmented generation (RAG) technology can provide efficient and intelligent question - answering services. In many application scenarios, such as education, medical care, industrial design, etc., it is necessary to process image data in documents. However, traditional knowledge - based question - answering systems mainly process text information and have limited capabilities in processing documents containing image data such as pictures and charts.
[0065] Specifically, most knowledge - based question - answering systems in the prior art are mainly based on text retrieval and matching and cannot effectively process image information in documents. Although some systems can identify and describe pictures, they cannot closely integrate image information with text content and context, let alone dynamically reference relevant pictures during the knowledge - based question - answering process.
[0066] It can be seen from this that the knowledge - based question - answering control process in the prior art cannot make full use of image information: there is a lack of effective processing of pictures in documents, and pictures cannot be included in the answers, resulting in limited integrity and accuracy of the answers. In addition, the context relevance of image data is poor. Even if pictures can be described, they cannot be closely integrated with the text context, affecting users' understanding of the answers. In actual scenarios, the results obtained by existing knowledge - based question - answering also lack the form of answers with both pictures and texts, and users are inefficient in understanding complex questions.
[0067] Based on this, the embodiments of the present invention provide a knowledge - based question - answering control method, system, and electronic device based on image data. This method can make full use of image data in knowledge - based question - answering related systems, can dynamically reference pictures in the original document during the knowledge - answering process, and provide answers with both pictures and texts, improving the user experience and the quality of the answers.
[0068] For ease of understanding this embodiment, first, a method for controlling knowledge Q&A based on picture data disclosed in the embodiments of the present invention will be introduced in detail. This method is applied to a navigation system including a map interface. The method is as follows Figure 1 shown, including:
[0069] Preprocessing step S101: Obtain a document to be processed containing picture data. After determining the semantic result of the picture data according to the document hierarchy of the document to be processed, generate a description text corresponding to the picture data based on the semantic result, and determine the markup language corresponding to the picture data according to the description text;
[0070] Vectorized interaction step S102: Construct a vector database corresponding to the document to be processed based on the markup language, and obtain a corresponding query text block from the vector database based on the knowledge Q&A question instruction;
[0071] Answer generation step S103: Obtain the context data corresponding to the query text block according to a preset context assembly strategy, and generate a knowledge Q&A answer result corresponding to the knowledge Q&A question instruction by using the markup language included in the context data.
[0072] This method first executes the preprocessing step. After obtaining the document to be processed, it parses and determines the document hierarchy relationship, and then determines the semantic result of the picture data, and uses the semantic result to obtain the description text of the picture data. When processing the description text, the picture data can be embedded in the form of a markup language such as Markdown data. Using Markdown syntax makes the citation and rendering of pictures more convenient, reduces the complexity of subsequent processing, ensures the correct display of pictures in the answer, and at the same time, the alternative text in the Markdown picture citation can contain the picture description, which is convenient for understanding the content of the picture data.
[0073] By constructing a vector database corresponding to the document to be processed and updating the data in the vector database by using the markup language, it is used as the data source for the retrieval enhancement generation process. The vectorized interaction step realizes the document vectorization and the query process, and uses the knowledge Q&A question instruction to request the vector database, so as to obtain the query text block corresponding to the knowledge Q&A question instruction.
[0074] The answer generation step can be implemented by the answer of a relevant language model. After assembling the query text block through the context assembly strategy corresponding to the language model to obtain the context data, the language model generates a relevant answer result including the picture data based on the context data. By parsing the answer result according to the markup language to obtain Markdown data, and using the Markdown data to generate the knowledge Q&A answer result for output, so that it corresponds to the knowledge Q&A question instruction.
[0075] Optionally, the preprocessing step S101, as Figure 2 shown, includes:
[0076] Document structure analysis step S201: Obtain the document hierarchy corresponding to the document to be processed, and determine the context information of the picture data based on the document hierarchy;
[0077] Picture semantic understanding step S202: Determine the semantic result corresponding to the picture data according to the context information of the picture data, and use the semantic result to determine the visual analysis result corresponding to the picture data;
[0078] Picture description generation step S203: Obtain the description text corresponding to the picture data based on the visual analysis result and the context information;
[0079] Picture link conversion step S204: Use the markup format parameters corresponding to the markup language to convert the picture data and the description text into the markup language.
[0080] In the preprocessing step, it is necessary to perform document structure analysis, picture semantic understanding, picture description generation and picture link conversion processes. Finally, the picture data is converted into Markdown markup language. Specifically, obtain the document hierarchy corresponding to the document to be processed, and determine the context information of the picture data based on the document hierarchy; then use the context information to determine the semantic result of the picture data, and obtain the visual analysis result of the picture data from the semantic result. After obtaining the visual analysis result, combine the context information to determine the description text of the picture data, and then use the Markdown markup format parameters corresponding to the markup language to convert the picture data and the description text, so as to obtain the Markdown markup language.
[0081] Optionally, the document structure analysis step S201, as Figure 3 shown, includes:
[0082] Step S301, obtain the document hierarchy after parsing the document structure of the document to be processed;
[0083] Step S302, obtain the chapter titles and paragraphs of the document to be processed based on the document hierarchy;
[0084] Step S303, obtain the picture data included in the paragraph, and extract the description text corresponding to the picture data;
[0085] Step S304, use the chapter title and the description text to determine the context information of the picture data.
[0086] During the document structure analysis process, first, the document hierarchical structure is parsed, and then the location structure information of the image data in the document is obtained, such as the table of contents location, chapter relationships, etc. Then, the context content around the image data is extracted, and existing image titles, annotations, etc. are obtained as descriptive text, and then the context information of the image data is obtained by combining the chapter titles.
[0087] In the specific implementation process, to parse the document structure, the python-docx library in Python can be used to read and parse the docx document to be processed. The structure in the document is usually organized in the form of paragraphs and headings. Then, the chapter headings are located, and the chapter headings are identified (such as using different levels of heading styles). For example, Heading 1 represents the main title, and Heading 2 represents the secondary title. Subsequently, the relevant information of the image data is extracted, which can be achieved by searching for directly inserted image objects in the paragraphs. Images in docx documents are usually part of a paragraph, rather than a Markdown-style reference. Subsequently, the context information is extracted. When an image is recognized, the paragraphs before it are traced back to obtain the context information, and the descriptive text after the image is recorded. Then, the context information is combined, and the extracted chapter titles, image titles, and their relevant descriptions are combined into complete context information.
[0088] Optionally, for step S202 of image semantic understanding, as Figure 4 shown, it includes:
[0089] Step S401, obtaining the image type data and image function data of the image data based on the context information, and using the image type data and image function data to determine the first description data corresponding to the image data;
[0090] Step S402, obtaining the product name data, operation step data, structure description data, annotation prompt data, and warning information data of the image data based on the above context information, and using the product name data, operation step data, structure description data, annotation prompt data, and warning information data to determine the second description data corresponding to the image data;
[0091] Step S403, using the first description data and the second description data to determine the semantic result, and after inputting the semantic result into the pre-trained multi-modal model, obtaining the visual analysis result output by the multi-modal model.
[0092] The process of semantic understanding of pictures can use a multi-modal large model (such as GPT-4 Vision) to analyze the picture content and generate visual analysis results. The specific execution process involves the basic information and core content description in the construction of prompts. When constructing prompts, it is necessary to consider its basic information and core content. The basic information is used as the first description data, mainly based on the context information to obtain the picture type data (such as operation step diagrams, product structure diagrams, interface schematic diagrams, etc.) and picture function data (the functional positioning of the picture in the document) of the picture data, so as to use the picture type data and picture function data to determine the first description data corresponding to the picture data.
[0093] The core content description is used as the second description data, mainly based on the context information to obtain the product / component name data, operation step data, structure description data, annotation prompt data and warning information data of the picture data. Using the above data, the second description data of the core content can be obtained. In the process of description, the second description data needs to use professional and accurate terms, and maintain the logic and sequence of the description. It is necessary to highlight the key points of the operation or description, and also retain the numbers and marks in the original picture.
[0094] Subsequently, the semantic result is determined using the first description data and the second description data, and it is input into the multi-modal large model to obtain the specific visual analysis result. Examples of visual analysis results are as follows:
[0095] - Basic information;
[0096] - Type: Product structure decomposition diagram;
[0097] - Function: Display the overall structure and component composition relationship of the XW-200 smart bracelet.
[0098] On this basis, the picture description generation step S203, as Figure 5 shown, includes:
[0099] Step S501, obtain the position structure information of the picture data in the document to be processed based on the context information;
[0100] Step S502, obtain the identification annotation information of the picture data, and use the identification annotation information, position structure information and visual analysis result to generate the description information of the picture data;
[0101] Step S503, after inputting the description information into the multi-modal model, obtain the description text output by the multi-modal model.
[0102] During the process of image description, various types of information need to be integrated to obtain the corresponding description information, specifically including information such as the position and structure information of the image in the document, the identification and annotation information of the original image, and the visual analysis result information, etc., so as to generate the description information of the image data. Furthermore, this description information is input into a multi-modal model for further analysis and extraction to generate a concise and information-complete description text. An example of the description text is as follows:
[0103] "Structural diagram of XW-200 smart bracelet | Schematic diagram of core components | Main board, display screen, battery, sensor | FPC connection relationship | Reference for repair and installation".
[0104] Optionally, for the image link conversion step S204, as Figure 6 shown, it includes:
[0105] Step S601, determining the markup format parameters corresponding to the markup language;
[0106] Step S602, converting the image data into a format link based on the markup format parameters, and converting the description text into a format text based on the markup format parameters;
[0107] Step S603, generating the markup language corresponding to the image data by using the format link and the format text.
[0108] During the process of image link conversion, the image data is converted into a Markdown format link based on the markup format parameters, and the description text is converted into a format text based on the markup format parameters. By replacing the original image corresponding to the image data with the Markdown format link, using the description text as the alternative text for the Markdown link, and generating a unique link for the image. An example of the markup language is as follows:
[0109] "".
[0110] Optionally, for the vectorized interaction step S102, as Figure 7 shown, it includes:
[0111] Step S701, after updating the markup language to the document to be processed, dividing the document to be processed into multiple text blocks based on the semantic results of the markup language;
[0112] Step S702, constructing a vector database by using the text blocks and generating a vector index for the text blocks;
[0113] Step S703, determining the query vector corresponding to the knowledge Q&A question instruction and obtaining the vector index corresponding to the query vector;
[0114] Step S704, determine the query text block according to the text block corresponding to the vector index.
[0115] After updating the Markdown markup language to the document to be processed during the vectorized interaction process, the document to be processed is divided into multiple text blocks by semantic segmentation based on the semantic results of the markup language; by the semantic segmentation method, the integrity of the Markdown image links in the text blocks is maintained. Then, a vector database is constructed using the text blocks, the text blocks are converted into vector representations using an embedding model, and vector indexes of the text blocks are generated based on the vector database. After receiving the knowledge Q&A question instruction sent by the user, the query process is converted into a vector representation by obtaining the query vector corresponding to the knowledge Q&A question instruction, and relevant text blocks are retrieved and obtained in the vector index.
[0116] Optionally, the response generation step S103, as Figure 8 shown, includes:
[0117] Step S801, determine the prompt data corresponding to the query text block, and after assembling the prompt data and the query text block using the context assembly strategy, obtain the context data;
[0118] Step S802, input the context data into a preset language model, and obtain the output result of the language model;
[0119] Step S803, parse and obtain the markup language included in the output result, and generate a knowledge Q&A answer result using the markup language.
[0120] The response generation process can be implemented through the responses of relevant language models. After determining the prompt data corresponding to the query text block, the prompt data and the query text block are assembled using the context assembly strategy to obtain the context data, and then the context data is input into a preset language model to obtain the Markdown content output by the language model, and the Markdown content is used to generate a knowledge Q&A answer result. For example:
[0121] "Regarding your question about the internal structure of the product, you can refer to the following schematic diagram to understand its components and connection methods:
[0122] 
[0123] As shown in the figure, the internal structure of the product contains multiple main components, which cooperate with each other through specific connection methods to achieve the functions of the product".
[0124] Another flowchart of a knowledge Q&A control method based on picture data is shown as follows. In the preprocessing stage, the method accurately generates picture descriptions by collecting information such as the positional structure information of the picture in the document, the context content around the picture, the identifier and annotation of the original picture, etc., as well as the visual analysis results (using multimodal vision large models such as GPT-4o), and uses the large model to summarize and generate picture descriptions. By introducing the multimodal model, the system can accurately obtain the content description of the picture and associate the context where the picture is located, laying a foundation for accurately citing the picture in the answer later. Figure 9 In addition, in the input construction, the method uses picture links in Markdown syntax. When preprocessing the text, the picture is embedded in the form of a Markdown picture link, which is convenient for retaining the picture reference when the LLM generates the answer. Using Markdown syntax makes the reference and rendering of pictures more convenient, reduces the complexity of subsequent processing, ensures that the pictures are correctly displayed in the answer, and at the same time, the alternative text in the Markdown picture reference can contain the picture description, which is convenient for the large model to understand.
[0125] The method optimizes the prompt engineering to guide the LLM to cite pictures in the answer. By designing effective prompt words, it guides the LLM to appropriately include picture references when generating the answer and retain the Markdown picture links. Through the carefully designed prompt words, it ensures that the LLM will not ignore or modify the picture links when generating the answer, so that the answer contains both text explanations and relevant pictures.
[0126] The knowledge Q&A control method based on picture data in the above embodiments has the following advantages:
[0127] Integrity: It can dynamically cite pictures in the original document in the answer, provide a form of answer with both pictures and texts, and enhance the integrity and intuitiveness of the answer.
[0128] Relevance: It closely combines the picture with its context content to help users better understand the answer.
[0129] Ease of use: Using Markdown syntax simplifies the process of picture reference and rendering, reducing the complexity of system implementation and maintenance.
[0130] Intelligence: By introducing the multimodal model, the system has the comprehensive understanding ability of pictures and texts, improving the accuracy of the answer.
[0131]
[0132] As can be seen from the method for controlling knowledge Q&A based on picture data mentioned in the above embodiments, this method additionally sets a drawing thread on the basis of the main thread of the navigation system, and uses the drawing thread to generate canvas trajectory data so as to reduce the resource consumption of the main thread; at the same time, this method combines the visible area to streamline the canvas trajectory data, further reducing the amount of data processing, thereby improving the page fluency of the navigation system.
[0133] Corresponding to the method for controlling knowledge Q&A based on picture data provided in the foregoing embodiments, an embodiment of the present invention provides a knowledge Q&A control system based on picture data, as Figure 10 shown, the system includes:
[0134] Preprocessing module 1010: used to obtain a document to be processed containing picture data, determine the semantic result of the picture data according to the document level of the document to be processed, generate a description text corresponding to the picture data based on the semantic result, and determine the markup language corresponding to the picture data according to the description text;
[0135] Vectorized interaction module 1020: used to construct a vector database corresponding to the document to be processed based on the markup language, and obtain a corresponding query text block from the vector database based on the knowledge Q&A question instruction;
[0136] Answer generation module 1030: used to obtain context data corresponding to the query text block according to a preset context assembly strategy, and generate a knowledge Q&A answer result corresponding to the knowledge Q&A question instruction by using the markup language included in the context data.
[0137] As can be seen from the knowledge Q&A control system based on picture data mentioned in the above embodiments, this system can make full use of picture data in a knowledge Q&A related system, can dynamically reference pictures in the original document during the knowledge answering process, and provide answers with both pictures and texts, improving the user experience and the quality of the answers.
[0138] The knowledge Q&A control system based on picture data provided by the embodiments of the present invention has the same implementation principle and the same technical effects as those of the foregoing embodiments of the method for controlling knowledge Q&A based on picture data. For a brief description, for the parts not mentioned in the system embodiment, reference may be made to the corresponding content in the foregoing embodiments of the method for controlling knowledge Q&A based on picture data.
[0139] This embodiment also provides an electronic device, and the structural schematic diagram of this electronic device is as Figure 11 shown, the device includes a processor 101 and a memory 102; wherein, the memory 102 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the steps of the above method for controlling knowledge Q&A based on picture data.
[0140] Figure 11The electronic device shown also includes a bus 103 and a communication interface 104. The processor 101, the communication interface 104, and the memory 102 are connected via the bus 103.
[0141] Among them, the memory 102 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The bus 103 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 11 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0142] The communication interface 104 is used to connect to at least one user terminal and other network units through a network interface, and send the encapsulated IPv4 packet or IPv4 packet to the user terminal through the network interface.
[0143] The processor 101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 101 or by instructions in software form. The above-mentioned processor 101 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 102, and the processor 101 reads the information in the memory 102 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0144] An embodiment of the present invention further provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the knowledge question and answer control method based on picture data in the foregoing embodiments.
[0145] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, equipment, and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces, and the indirect coupling or communication connection of devices or units may be in an electrical, mechanical, or other form.
[0146] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0147] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0148] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0149] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field of the present invention can still modify the technical solutions recorded in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims described.
Claims
1. A knowledge question answering control method based on image data, characterized in that: The method comprises: Preprocessing step: obtaining a document to be processed containing image data, determining a semantic result of the image data according to the document level of the document to be processed, generating a description text corresponding to the image data based on the semantic result, and determining a markup language corresponding to the image data according to the description text; Vectorized interaction step: constructing a vector database corresponding to the document to be processed based on the markup language, and obtaining a corresponding query text block from the vector database based on the knowledge question-answering instruction; Answer generation step: obtaining context data corresponding to the query text block according to a preset context assembly strategy, and generating a knowledge question answer result corresponding to the knowledge question question instruction using the markup language contained in the context data; The vectorized interaction step includes: After the markup language is updated to the document to be processed, the document to be processed is divided into a plurality of text blocks based on the semantic result of the markup language; constructing the vector database using the text block, and generating a vector index of the text block based on the vector database; Determine a query vector corresponding to the knowledge question and answer instruction, and obtain the vector index corresponding to the query vector; The query text block is determined according to the text block corresponding to the vector index.
2. The knowledge question-answering control method based on image data according to claim 1 is characterized in that: The pre-processing step comprises: Document structure analysis step: obtaining the document level corresponding to the document to be processed, and determining the context information of the image data based on the document level; Image semantic understanding step: determining the semantic result corresponding to the image data according to the context information of the image data, and determining the visual analysis result corresponding to the image data using the semantic result; Picture description generating step: obtaining the description text corresponding to the picture data based on the visual analysis result and the context information; Picture link conversion step: using the markup format parameters corresponding to the markup language, converting the picture data and the description text into the markup language.
3. The knowledge question-answering control method based on image data according to claim 2 is characterized in that: The document structure analysis step comprises: After parsing the document structure of the document to be processed, obtain the document level; Acquire the chapter titles and paragraphs of the document to be processed based on the document level; Acquire the image data contained in the paragraph, and extract the description text corresponding to the image data; The context information of the picture data is determined using the chapter title and the description text.
4. The knowledge question-answering control method based on image data according to claim 3 is characterized in that: The picture semantic understanding step includes: Acquire picture type data and picture function data of the picture data based on the context information, and determine first description data corresponding to the picture data by using the picture type data and the picture function data; Acquire product name data, operation step data, structure description data, annotation prompt data, and warning information data of the image data based on the context information, and determine second description data corresponding to the image data using the product name data, the operation step data, the structure description data, the annotation prompt data, and the warning information data; The semantic result is determined using the first description data and the second description data, and after the semantic result is input into a trained multimodal model, the visual analysis result of the output of the multimodal model is obtained.
5. The knowledge question-answering control method based on image data according to claim 4 is characterized in that: The picture description generating step comprises: Acquire the position structure information of the image data in the document to be processed based on the context information; Acquire identification annotation information of the image data, and generate description information of the image data using the identification annotation information, the position structure information, and the visual analysis result; After the description information is input into the multimodal model, the description text output by the multimodal model is obtained.
6. The knowledge question-answering control method based on image data according to claim 2, characterized in that: The picture link conversion step includes: Determine the markup format parameters corresponding to the markup language; Converting the image data into a format link based on the markup format parameters, and converting the description text into a format text based on the markup format parameters; The markup language corresponding to the image data is generated using the format link and the format text.
7. The knowledge question-answering control method based on image data according to claim 1, characterized in that: The response generation step comprises: Determine prompt data corresponding to the query text block, and assemble the prompt data and the query text block using the context assembly strategy to obtain the context data; After inputting the context data into a preset language model, obtaining an output result of the language model; The markup language included in the output result is parsed and obtained, and the knowledge question and answer result is generated using the markup language.
8. A knowledge question-answering control system based on image data, characterized in that: The system comprises: Preprocessing module: used for obtaining a document to be processed containing image data, determining the semantic result of the image data according to the document level of the document to be processed, generating a description text corresponding to the image data based on the semantic result, and determining the markup language corresponding to the image data according to the description text; Vectorization interaction module: used to construct a vector database corresponding to the document to be processed based on the markup language, and obtain a corresponding query text block from the vector database based on the knowledge question and answer instruction; Answer generation module: used for obtaining context data corresponding to the query text block according to a preset context assembly strategy, and generating a knowledge question answer result corresponding to the knowledge question question instruction by using the markup language contained in the context data; The vectorized interaction module is also used to: after updating the markup language to the document to be processed, divide the document to be processed into multiple text blocks based on the semantic results of the markup language; construct the vector database using the text blocks, and generate the vector index of the text block based on the vector database; determine the query vector corresponding to the knowledge question and answer question instruction, and obtain the vector index corresponding to the query vector; determine the query text block according to the text block corresponding to the vector index.
9. An electronic device, characterized in that: It includes a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the steps of the knowledge question and answer control method based on image data as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge base construction method and question and answer dialogue method and system based on generative large language model
CN117056471A
Image-text fusion question and answer method, device and equipment based on large language model
CN118861229A