Question and answer method and device, electronic equipment and computer storage medium

By performing intent recognition and vector matching on user queries, a multimodal knowledge base is generated, which solves the problem that existing systems in the field of communication services cannot handle complex graphic and textual content, and realizes efficient business knowledge support and synchronous output of knowledge point images.

CN121746535APending Publication Date: 2026-03-27CHINA MOBILE COMM GRP SHAANXI CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing knowledge question answering systems based on large language models cannot be directly applied to the field of communication services because communication service knowledge is complex, and traditional document retrieval methods based on keyword search are inefficient and neglect image information processing, resulting in the loss of key knowledge.

Method used

By collecting users' original question information and performing intent recognition, a knowledge base is generated using multiple pre-set sets of image and text data matching. Vector matching is then performed in the knowledge base to output knowledge points and matching images, thereby achieving intelligent understanding and efficient support of text and image content.

Benefits of technology

It effectively avoids knowledge loss and achieves efficient and intelligent business knowledge support in the field of communication services. It can also include relevant images as reference diagrams when outputting text, thereby improving the efficiency of information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746535A_ABST
    Figure CN121746535A_ABST
Patent Text Reader

Abstract

The invention discloses a question and answer method and device, electronic equipment and a computer storage medium, and the method comprises the steps: carrying out the intention recognition of collected original question information of a user, and obtaining an intention recognition result; and performing vector matching in a preset knowledge base according to the intention recognition result, and searching in the preset knowledge base to obtain knowledge points corresponding to the original question information of the user and pictures matched with the knowledge points. And outputting knowledge points corresponding to the original question information of the user and pictures matched with the knowledge points. Through the multi-modal document segmentation method, the picture can be bound with the previous paragraph or the following paragraph after being recognized in the segmentation process, and the picture can be taken out together as a reference picture when the output text relates to the previous or following content, so that knowledge loss is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of artificial intelligence, and particularly relates to a question and answer method and device, electronic equipment and computer storage medium. BACKGROUND

[0002] In recent years, knowledge question and answer systems based on large language models have achieved remarkable results in general fields and have been tried to be applied to various vertical industries. Such systems aim to improve information acquisition efficiency by understanding natural language questions, automatically retrieving and generating answers from knowledge bases. In the field of communication services, there is also a similar demand for intelligent upgrading in order to achieve the goal of "on-site questions answered immediately and company regulations supplemented quickly".

[0003] However, existing systems cannot be directly used in the field of communication services because the knowledge in the field of communication services is complex and daily work involves various complex scenarios. With the explosive growth of business knowledge documents, traditional keyword-based document search methods are inefficient and generally ignore the processing and utilization of image information in business knowledge documents, resulting in the loss of key knowledge. Therefore, there is an urgent need for a business knowledge question and answer system that can intelligently understand text and image content to truly achieve efficient and intelligent business knowledge support. SUMMARY

[0004] The embodiments of the present application provide a question and answer method, device, equipment and computer storage medium, which can output pictures together as reference pictures when the output text involves previous or subsequent content, thereby avoiding the loss of knowledge.

[0005] In a first aspect, the embodiments of the present application provide a question and answer method, which can include: Collecting user original question information; Performing intent recognition on the user original question information to obtain an intent recognition result, the intent recognition result indicating that the user original question information belongs to basic facts that do not need to be proven or knowledge including multiple steps and having an explicit logical relationship; Performing vector matching in a preset knowledge base according to the intent recognition result, searching for a knowledge point corresponding to the user original question information and a picture matched with the knowledge point in the preset knowledge base, the preset knowledge base being obtained by classifying and summarizing a matching set of multiple sets of picture data and multiple sets of text data, the multiple sets of picture data and the multiple sets of text data being obtained by cutting the acquired business knowledge documents; Outputting the knowledge point corresponding to the user original question information and the picture matched with the knowledge point.

[0006] In one of the embodiments, the vector matching in the preset knowledge base according to the intention recognition result involves, before searching for the knowledge point corresponding to the original question information of the user and the picture matched with the knowledge point in the preset knowledge base, further comprising: obtaining a business knowledge document; segmenting the business knowledge document to obtain multiple sets of picture data and multiple sets of text data, the picture data including a picture, a picture caption, a picture preceding paragraph and a picture following paragraph, and the text data being text in a sentence unit; matching the picture preceding paragraph, the picture following paragraph, the multiple sets of text data and the picture caption in the multiple sets of picture data to obtain a matching set, the matching set being used to output the picture corresponding to the matching picture caption as a reference in the case of outputting a target paragraph or target text data, the target paragraph being a paragraph in the picture preceding paragraph and the picture following paragraph that is more semantically close to the picture caption, and the target text data being text data in the multiple sets of text data that is more semantically similar to the picture caption than a preset threshold; classifying and summarizing the matching set to generate a preset knowledge base, the knowledge point in the preset knowledge base being a basic fact that does not need to be proved or knowledge including multiple steps and having an explicit logical relationship.

[0007] In one of the embodiments, the classifying and summarizing the matching set to generate a preset knowledge base involves: classifying and summarizing the matching set to obtain a single-point knowledge set and a logical knowledge set, the single-point knowledge set including multiple single-point knowledge, the single-point knowledge including a knowledge point and a picture matched with the knowledge point, the logical knowledge set including multiple logical knowledge, the logical knowledge including a knowledge point and a picture matched with the knowledge point, the knowledge point being a target paragraph or target text data, the single-point knowledge being a basic fact that does not need to be proved, and the logical knowledge being knowledge including multiple steps and having an explicit logical relationship; generating the preset knowledge base according to the single-point knowledge set and the logical knowledge set.

[0008] In one of the embodiments, the matching according to the picture preceding paragraph, the picture following paragraph, the multiple sets of text data and the picture caption in the multiple sets of picture data to obtain a matching set involves: matching the picture preceding paragraph, the picture following paragraph, the multiple sets of text data and the picture caption in the multiple sets of picture data to obtain a matching result, the matching result being used to indicate the similarity of the semantic of the picture preceding paragraph, the picture following paragraph, the multiple sets of text data and the picture caption; matching the target paragraph or target text data and the picture caption respectively according to the matching result to obtain the matching set.

[0009] In one of the embodiments, the matching according to the picture context paragraph, the picture context paragraph, the multiple sets of text data and the picture title in the multiple sets of picture data, obtaining the matching result, includes: According to the picture context paragraph, the picture context paragraph and the picture title in the multiple sets of picture data, a first matching result is obtained, and according to the multiple sets of text data and the picture title in the multiple sets of picture data, a second matching result is obtained, the first matching result is used to indicate the similarity of the picture context paragraph and the picture context paragraph with the picture title semantics respectively, and the first matching result is used to indicate the similarity of the text data with the picture title semantics; According to the matching result, the target paragraph or the target text data is matched with the picture title respectively, and a matching set is obtained, including: According to the first matching result, the picture corresponding to the picture title is matched with the target paragraph to generate a first picture matching set, and according to the second matching result, the picture corresponding to the picture title is matched with the target text data to generate a second picture matching set, the first picture matching set includes multiple picture paragraph matching results, the picture paragraph matching result includes the picture, the picture title and the target paragraph corresponding to the picture title, the first picture matching set is used to output the picture as a reference at the same time in the case of outputting the target paragraph, and the second picture matching set is used to output the picture as a reference at the same time in the case of outputting the target text data.

[0010] In one of the embodiments, the classification and induction of the matching set to generate the preset knowledge base includes: The first picture matching set and the second picture matching set are classified and induced to generate a knowledge base, the knowledge base includes a single-point knowledge set and a logic knowledge set, the single-point knowledge set includes multiple single-point knowledge, the single-point knowledge includes a knowledge point and a picture matched with the knowledge point, the logic knowledge set includes multiple logic knowledge, the logic knowledge includes a knowledge point and a picture matched with the knowledge point, the knowledge point is the target paragraph or the target text data, the single-point knowledge is a basic fact without proof, and the logic knowledge is a knowledge including multiple steps and having a clear logical relationship.

[0011] In one of the embodiments, before the intention recognition of the user's original question information to obtain the intention recognition result, it further includes: The user's original question information is vectorized to obtain a user input vector; The intention recognition of the user's original question information obtains the intention recognition result, including: The user input vector is subjected to intention recognition to obtain the intention recognition result.

[0012] In one of the embodiments, the knowledge points corresponding to the output user original question information and the pictures matched with the knowledge points involve: According to the user original question information, the accuracy of the question is determined; According to the accuracy of the question, the knowledge points corresponding to the user original question information and the pictures matched with the knowledge points are output.

[0013] In one of the embodiments, the classification and induction of the first picture matching set and the second picture matching set to generate the knowledge base involve: The knowledge type judgment is performed on the first picture matching set and the second picture matching set to generate the vectorized knowledge set and the knowledge graph, the vectorized knowledge set includes the single-point knowledge set, the knowledge graph includes the logic knowledge set, and the knowledge graph is used to indicate the association relationship between the multiple logic knowledge sets; According to the vectorized knowledge set and the knowledge graph, the knowledge base is generated.

[0014] In one of the embodiments, the first matching result is obtained according to the picture context paragraph, the picture context paragraph and the picture title in the multiple sets of picture data, and the second matching result is obtained according to the picture title in the multiple sets of text data and the multiple sets of picture data, which involve: The picture context paragraph and the picture context paragraph in each set of picture data are summarized and extracted to obtain each set of paragraph summary information, and the paragraph summary information is used to indicate the semantics of the picture context paragraph and the picture context paragraph; The paragraph summary information of the multiple sets of picture data and the picture title are semantically matched to obtain the first matching result, and the picture title in the multiple sets of text data and the multiple sets of picture data are semantically matched to obtain the second matching result.

[0015] In a second aspect, the embodiments of the present application provide a question and answer device, which can include: The acquisition module is configured to acquire user original question information; The recognition module is configured to perform intent recognition on the user original question information to obtain an intent recognition result, and the intent recognition result is used to indicate that the user original question information belongs to a basic fact that does not need to be proved or knowledge including multiple steps and having an explicit logical relationship; The matching module is configured to perform vector matching in a preset knowledge base according to the intent recognition result, and search for the knowledge points corresponding to the user original question information and the pictures matched with the knowledge points in the preset knowledge base, the preset knowledge base is obtained by classifying and inducing a matching set obtained by matching multiple sets of picture data and multiple sets of text data, and the multiple sets of picture data and the multiple sets of text data are obtained by cutting a business knowledge document. An output module is configured to output the knowledge point corresponding to the original question information of the user and the picture matched with the knowledge point.

[0016] In a third aspect, an electronic device is provided, and the device includes: a processor; a memory configured to store processor-executable instructions; The processor is configured to execute the instructions to implement the question and answer method as shown in any one of the embodiments of the first aspect.

[0017] In a fourth aspect, a computer storage medium is provided, and a computer program is stored on the computer readable storage medium. The computer program is executed by a processor to implement the question and answer method as shown in any one of the embodiments of the first aspect.

[0018] In a fifth aspect, a computer program product is provided, and the computer program product includes a computer program stored in a readable storage medium. At least one processor of a device reads and executes the computer program from the storage medium, so that the device executes the question and answer method as shown in any one of the embodiments of the first aspect.

[0019] The question and answer method, device, electronic device, and computer storage medium provided by the embodiments of the present application can obtain an intention recognition result by performing intention recognition on the collected original question information of the user. The knowledge point corresponding to the original question information of the user and the picture matched with the knowledge point are searched in the preset knowledge base according to the intention recognition result. The knowledge point corresponding to the original question information of the user and the picture matched with the knowledge point are output.

[0020] By using the multi-modal document segmentation method, the picture can be bound with the previous paragraph or the next paragraph in the segmentation process after the picture is recognized. When the output text involves the previous or next content, the picture can be taken out as a reference picture, so as to avoid knowledge loss. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiments of the present application will be briefly introduced. Those skilled in the art can obtain other drawings according to these drawings without creating any creative labor.

[0022] Figure 1 is a flowchart of a question and answer method provided by the embodiments of the present application; Figure 2 is a flowchart of another question and answer method provided by the embodiments of the present application; Figure 3is a schematic diagram of a broadband service knowledge question and answer system provided by an embodiment of the present application; Figure 4 is a schematic diagram of splitting a service knowledge document provided by an embodiment of the present application; Figure 5 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 6 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 7 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 8 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 9 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 10 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 11 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 12 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 13 is a schematic diagram of a question and answer method provided by an embodiment of the present application; Figure 14 is a schematic diagram of a question and answer device provided by an embodiment of the present application; Figure 15 is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0023] The features and exemplary embodiments of various aspects of the present application will be described in detail below, in order to make the purposes, technical solutions and advantages of the present application more clear and apparent, the present application will be further described in detail below in combination with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, but not to limit the present application. The present application can be implemented without some of these specific details by those skilled in the art. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.

[0024] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0025] As discussed in the background section, in recent years, knowledge-based question-answering systems based on large language models have achieved significant results in general fields and are being attempted to be applied to various vertical industries. These systems aim to improve information retrieval efficiency by understanding natural language questions, automatically retrieving and generating answers from knowledge bases. Similar intelligent upgrade needs exist in the communications business sector, with the goal of achieving "instant answers to on-site questions and rapid supplementation of company specifications."

[0026] However, existing systems cannot be directly applied to the telecommunications business domain because the knowledge in this domain is quite complex, and daily work involves various complex scenarios. With the explosive growth of business knowledge documents, traditional keyword-based document retrieval methods are inefficient and generally neglect the processing and utilization of image information within these documents, resulting in the loss of critical knowledge. Therefore, there is an urgent need for a business knowledge question-and-answer system capable of intelligently understanding text and image content to truly achieve efficient and intelligent business knowledge support.

[0027] To address the problems existing in the prior art, embodiments of this application provide a question-answering method, apparatus, electronic device, and computer storage medium. The method involves performing intent recognition on the collected original user question information to obtain intent recognition results. Based on the intent recognition results, vector matching is performed in a preset knowledge base to search for the knowledge points corresponding to the original user question information and the images matching those knowledge points. The knowledge points corresponding to the original user question information and the images matching those knowledge points are then output. By employing a multimodal document segmentation method, after identifying images during the segmentation process, they can be bound to preceding or following paragraphs. This allows the images to be included as reference images when the output text involves preceding or following content, preventing knowledge loss.

[0028] The question-and-answer method provided in the embodiments of this application will be introduced first below. For example... Figure 1 As shown, the question-and-answer method provided in this application includes the following steps: S101: Collect the user's original question information; S102: Perform intent recognition on the user's original question information to obtain intent recognition results. The intent recognition results are used to indicate that the user's original question information belongs to basic facts that do not require proof or knowledge that includes multiple steps and has a clear logical relationship. S103: Based on the intent recognition result, perform vector matching in the preset knowledge base, and search the preset knowledge base to obtain the knowledge points corresponding to the user's original question information and the images matching the knowledge points. The preset knowledge base is obtained by classifying and summarizing the matching sets obtained by matching multiple sets of image data and multiple sets of text data. The multiple sets of image data and multiple sets of text data are obtained by segmenting the acquired business knowledge documents. S104: Output the knowledge points corresponding to the user's original question information and the images matching the knowledge points.

[0029] The above describes a question-answering method, apparatus, electronic device, and computer storage medium provided in this application. It involves performing intent recognition on the collected original user question information to obtain intent recognition results. Based on the intent recognition results, vector matching is performed in a preset knowledge base to search for the knowledge points corresponding to the original user question information and the matching images. The knowledge points corresponding to the original user question information and the matching images are then output. Using a multimodal document segmentation method, after recognizing the images during segmentation, they can be bound to the preceding or following paragraphs. This allows the images to be included as reference images when the output text involves preceding or following content, preventing knowledge loss.

[0030] In S101, the system collects the user's original question information. In one example, the system can automatically collect the user's original question information throughout the entire interaction between the user and the question-and-answer system. For example, when an operations and maintenance personnel asks "How to deal with the constant red LOS light on the optical modem", the system will record the entire question text as the user's original question information.

[0031] In step S102, intent recognition is performed on the user's original question information to obtain the intent recognition result. The intent recognition result indicates that the user's original question information belongs to basic facts that do not require proof or knowledge involving multiple steps and having a clear logical relationship. In one example, after collecting the user's original question information, the system performs intelligent intent recognition classification on the question. Specifically, a trained FastText classifier can be used to analyze the input text in real time and output a binary classification result. If it is identified as a basic fact-type intent that does not require proof (corresponding to "single-point knowledge"), it indicates that the answer to the question is usually a definite and concise factual statement (such as standard specifications, equipment parameters, etc.); if it is identified as a process-type intent that involves multiple steps and has a clear logical relationship (corresponding to "logical knowledge"), it indicates that the question requires the system to provide structured and sequential operation guidance or solutions (such as troubleshooting procedures, equipment installation steps, etc.). For example, for the question "What is the standard operating voltage of an OLT device?", the system recognizes it as a factual intent and will quickly retrieve the voltage value from the vector library; while for the question "How to configure the IPTV service of a newly installed broadband?", the system recognizes it as a process-oriented intent and will obtain the step-by-step configuration process and related topology diagrams from the knowledge graph.

[0032] In one specific embodiment, to target different parts of the database based on the different content of the user's question, this solution adds a knowledge determination module before the open-source large model. This module determines whether the user's question is a single-point question or an operational question (i.e., a logical question). A single-point question refers to a question whose answer is usually only one sentence. For example, when a user asks about the on-site broadband installation time, the large model usually answers with only one sentence, which is a single-point question. When a user asks about the broadband installation process, the large model usually answers with detailed operational procedures, which is an operational question. Considering the characteristics of this solution—short input text, high response time requirements, and few classification categories—this solution does not utilize the text classification function of the large model. Instead, it uses a FastText-based text classification method in the knowledge matching module.

[0033] In S103, vector matching is performed in a preset knowledge base based on the intent recognition result. The preset knowledge base searches for the knowledge points corresponding to the user's original question information and the images matching the knowledge points. The preset knowledge base is obtained by classifying and summarizing the matching sets obtained by matching multiple sets of image data and multiple sets of text data. The multiple sets of image data and multiple sets of text data are obtained by segmenting the acquired business knowledge documents.

[0034] In one example, based on the intent recognition result, the system initiates a differentiated retrieval mechanism, performing intelligent vector matching within a pre-defined knowledge base. Specifically, when the intent recognition result is a factual intent, the system automatically converts the user's question text into a high-dimensional vector and performs an approximate nearest neighbor search in the vectorized database. It calculates the Manhattan distance to match the text knowledge unit with the closest semantic meaning and simultaneously retrieves the associated image data bound to that unit. When the intent recognition result is a process-oriented intent, the system first matches the question vector with pre-constructed process node vectors in the knowledge graph to locate the relevant process subgraph. Then, it retrieves multiple knowledge nodes with a logical order and their corresponding instruction images along the graph's relational edges.

[0035] The pre-built knowledge base is constructed through the following process: First, business knowledge documents are segmented at multiple granularities to obtain text paragraphs / sentence units and image units; then, semantic matching algorithms (combining graph-head analysis and contextual understanding) are used to establish graph-text association pairs, forming a set of graph-text matching pairs; finally, based on a large model-driven knowledge classifier, this set is divided into a factual graph-text matching subset (stored in a vector library) and a process-oriented graph-text matching subset (used to construct a knowledge graph), thus forming a domain knowledge base with structured hierarchy and multimodal associations. For example, when searching for "fiber optic fusion splicer cleaning steps," the system matches the "equipment maintenance" process chain in the knowledge graph with process intent, returning step-by-step text descriptions and operation diagrams corresponding to each step.

[0036] In one specific embodiment, this solution annotates 3000 text classification data entries, with the knowledge base data sourced from daily question records of broadband maintenance personnel. Through FastText text classification, user input questions can be effectively categorized into single-point knowledge questions and operational questions, preparing for subsequent knowledge filtering. After knowledge matching is completed, the large model vectorizes the original user question content and autonomously searches the corresponding knowledge base based on the classification results, outputting the search answer and achieving a closed loop for installation and maintenance personnel's questions. The semantic matching method used in this solution is Manhattan distance, which is used to match the degree of matching between the user's input semantic vector and the knowledge vector.

[0037] In S104, the system outputs the knowledge points corresponding to the user's original question and the images matching those knowledge points. In one example, based on the search results, the final output is a structured answer integrating multimodal information. Specifically, this includes: 1) text fragments of core knowledge points retrieved from the knowledge base, and 2) metadata of associated images semantically bound to those knowledge points. In a specific embodiment, the system can design specific prompt word templates and input this information along with the user's original question into the finely tuned Qwen-14B large model, instructing it to generate a natural and fluent final answer containing accurate image references. For example, in response to the user's question "How to correctly use a fiber optic fusion splicer for fiber core alignment?", the system outputs the following: "Fiber core alignment of a fiber optic fusion splicer should follow these steps: First, place the stripped fiber into a V-groove and fix it (see attached diagram: Fiber placement diagram); second, adjust the XYZ axes through screen monitoring to initially align the fiber cores (see attached diagram: Axial adjustment interface diagram); finally, use a precision motor for micron-level automatic alignment until the screen displays a core alignment accuracy of over 99% (see attached diagram: Alignment accuracy indicator diagram)."

[0038] like Figure 2 As shown, as an example, before S103, it may also include: S201: Obtain business knowledge documents; S202: Segment the business knowledge document to obtain multiple sets of image data and multiple sets of text data. The image data includes the image, image title, paragraph above the image, and paragraph below the image. The text data consists of sentences. S203: Match the preceding paragraph, following paragraph, multiple sets of text data, and image titles in multiple sets of image data to obtain a matching set. The matching set is used to output the image corresponding to the matching image title as a reference when outputting the target paragraph or target text data. The target paragraph is the paragraph in the preceding paragraph and following paragraph that is semantically closer to the image title. The target text data is the text data in multiple sets of text data whose semantic similarity to the image title is greater than a preset threshold. S204: Classify and summarize the matching set to generate a preset knowledge base. The knowledge points in the preset knowledge base are basic facts that do not require proof or knowledge that includes multiple steps and has a clear logical relationship.

[0039] In one example, such as Figure 3As shown, the broadband service knowledge question-and-answer system 300 designed in this solution includes two parts: a question-and-answer interaction part 310 and a knowledge base 320. In the question-and-answer interaction part 310, the vectorization module 311 acquires the user's original question, vectorizes it, and inputs it into the knowledge determination module 3121 in the large model 312 for intent recognition. Based on the intent recognition result, vector matching is performed in the knowledge base 320. Finally, the matched knowledge points and their corresponding images are output to the user through the output module 313. The knowledge base 320 processes the documents required for broadband service knowledge question-and-answer through document segmentation 323, document cleaning 324, and document deduplication 325, compiling them into a knowledge graph and a vectorized database, serving as the knowledge foundation for broadband service question-and-answer. The knowledge graph 321 is mainly responsible for arranging the logical structure of multiple document blocks involved in complex questions, while the vectorized database 322 is responsible for quickly outputting information for simple questions, improving the response speed for simple questions. The main model used in this solution is the open-source Qwen-14B model. LoRA fine-tuning is employed to adapt the model to industry knowledge corpus.

[0040] In S201, in one example, the acquired business knowledge document can be labeled text classification data, and the data source can be the daily question records of broadband operation and maintenance personnel. In a specific embodiment, in order to accurately extract the intent of network operation and maintenance personnel, this solution collects the daily chat records of 1,000 network operation and maintenance personnel distributed in different cities, performs semantic annotation on the chat records, and fine-tunes the Qwen-14B large model.

[0041] In S202, such as Figure 4 As shown, the business knowledge document 400 is segmented to obtain multiple sets of image data and multiple sets of text data. The image data includes image 403, image title 404, the preceding paragraph (i.e., paragraph A 401), and the following paragraph (i.e., paragraph B 402). The text data consists of sentences. In one example, step S410: Image 403, image title 404, the preceding paragraph (i.e., paragraph A 401), and the following paragraph (i.e., paragraph B 402) are input into a general large model to obtain paragraph summary A 405 and paragraph summary B 406. Step S420: Image title 404 is matched with paragraph summary A 405 and paragraph summary B 406 respectively to obtain similarity score A 407 and similarity score B 408. Step S430: Compare similarity score A 407 and similarity score B 408. If similarity score A 407 is greater than similarity score B 408, match image 403 with the paragraph above the image (i.e., paragraph A 401); if similarity score A 407 is less than similarity score B 408, match image 403 with the paragraph below the image (i.e., paragraph B 402).

[0042] In one example, this solution utilizes document indexing and text segmentation technologies in the document segmentation module. Upon receiving the input knowledge base document, the module first performs initial segmentation using a paragraph-based method. Considering that broadband business knowledge typically contains numerous reference diagrams illustrating corresponding operations for broadband maintenance personnel, and existing document segmentation solutions often handle images within the document infrequently, directly removing images in most cases, thus affecting the completeness of the answers provided in the knowledge-based Q&A section, this solution stores the image, image title, one paragraph preceding the image, and one paragraph following the image as a single data set when segmenting the document at the paragraph level. Furthermore, after completing the paragraph-level segmentation, it continues with sentence-level segmentation for even finer-grained segmentation.

[0043] In S203, a matching set is obtained by matching the preceding paragraph, following paragraph, multiple sets of text data, and image titles from multiple sets of image data. This matching set is used to output the corresponding image title as a reference when outputting the target paragraph or target text data. The target paragraph is the paragraph in the preceding and following paragraphs that is semantically closer to the image title. The target text data is the text data in the multiple sets of text data whose semantic similarity to the image title is greater than a preset threshold. In one example, a general large model is used to summarize and extract the preceding and following paragraphs of the document image. The summarized information is then vectorized and semantically compared and matched with the image title vector. This scheme uses Manhattan distance in the semantic comparison matching process. If the image title is semantically close to the preceding paragraph's summary, the image is matched with the preceding paragraph's content, and the image is included as a reference when the preceding paragraph is output as a result. Similarly, if the image title is semantically close to the following paragraph's summary, the image is matched with the following paragraph, and the image is output as a reference when the following paragraph's content is output. For the sentence-based segmentation results, this solution also summarizes all sentences contained in the preceding and following paragraphs of the image and performs semantic matching with the image title. Sentences with a high degree of semantic matching with the image title are selected, and the image output can also be used as a reference when the user intends to output the knowledge points corresponding to these sentences.

[0044] In one example, after document segmentation is completed, document cleaning and deduplication modules can be used to improve the quality of the segmented document blocks, making it easier for large models to output higher quality answer results.

[0045] In S204, the matching set is categorized and summarized to generate a preset knowledge base. The knowledge points in the preset knowledge base are basic facts that do not require proof or knowledge involving multiple steps and having clear logical relationships. In one example, the matching set is categorized and summarized to generate a vectorized knowledge base and a knowledge graph. The vectorized knowledge base focuses on single-point knowledge, while the knowledge graph focuses on operational knowledge (i.e., logical knowledge).

[0046] like Figure 5 As shown, as an example, S204 may include: S2041: The matching set is summarized and classified to obtain a single-point knowledge set and a logical knowledge set. The single-point knowledge set includes multiple single-point knowledge points, which include knowledge points and images matching the knowledge points. The logical knowledge set includes multiple logical knowledge points, which include knowledge points and images matching the knowledge points. The knowledge points are target paragraphs or target text data. Single-point knowledge points are basic facts that do not require proof. Logical knowledge points are knowledge that includes multiple steps and has a clear logical relationship. S2042: Generate a preset knowledge base based on single-point knowledge sets and logical knowledge sets.

[0047] In S2041, the matching sets are categorized and summarized to obtain single-point knowledge sets and logical knowledge sets. In one example, the matching sets are categorized and summarized to generate a vectorized knowledge base and a knowledge graph. The vectorized knowledge base mainly includes single-point knowledge, while the knowledge graph mainly includes operational knowledge (i.e., logical knowledge). Single-point knowledge in this solution refers to standard knowledge in broadband operation and maintenance scenarios. This type of knowledge usually does not require users to judge usage conditions, and the knowledge content usually does not need to be answered in multiple points. Operational knowledge in this solution refers to knowledge points and execution steps that require users to select based on specific on-site situations. The answer content is usually divided into multiple sentences, with logical relationships between the sentences. An example of single-point knowledge is: Broadband installation usually requires the first contact with the customer within 2 hours. The corresponding question is the duration of the first contact for broadband installation, which is a single-point knowledge point. Examples of operational knowledge points are: 1. Confirm whether the optical power meter is abnormal, restart or check the optical power meter by other means. 2. Use a light pen to detect breakpoints and rework at the breakpoints by fusion splicing. 3. Light leakage at connectors or splices: Re-terminate and inspect. 4. Light leakage in the fiber optic link: Locate the problem and repair the link.

[0048] In S2042, a preset knowledge base is generated based on single-point knowledge sets and logical knowledge sets.

[0049] like Figure 6 As shown, as an example, S203 may include: S2031: Match the preceding paragraph, following paragraph, multiple sets of text data, and image title in multiple sets of image data to obtain matching results. The matching results are used to indicate the semantic similarity between the preceding paragraph, following paragraph, multiple sets of text data, and image title. S2032: Based on the matching results, match the target paragraph or target text data with the image title to obtain a matching set.

[0050] In S2031, matching is performed based on the preceding and following paragraphs of multiple sets of image data, multiple sets of text data, and the image title to obtain matching results. These results indicate the semantic similarity between the preceding and following paragraphs, multiple sets of text data, and the image title. In one example, the title text and context paragraphs of each image in the document are first extracted, and all segmented text units are converted into semantic vectors. Then, the similarity between the title vector and each text vector is calculated using Manhattan distance, generating a matching result with three dimensions: this result indicates both the semantic proximity between the image and its immediate context (preceding / following paragraphs) and identifies the specific paragraph or sentence that best matches the image content from all text units through global comparison. For example, when processing broadband operation and maintenance manuals, the system calculates that an image titled "Optical Power Meter Calibration Interface" has a similarity of 0.92 with the "Equipment Calibration Operation Steps" paragraph, while its similarity with the paragraphs before and after it is less than 0.7. Based on this, it is determined that the image should be linked to the operation steps paragraph, so as to ensure that when users query calibration methods in subsequent Q&A sessions, the system can accurately return the content of that paragraph and attach the corresponding operation interface diagram.

[0051] In S2032, based on the matching results, the target paragraph or target text data is matched with the image title to obtain a matching set. In one example, based on the matching results, the system performs the final binding operation of image-text association. Specifically, the system will filter out target paragraphs or target sentences that meet the preset threshold conditions from the text units to be matched based on the calculated semantic similarity index, and establish a one-to-many mapping relationship with the corresponding image and its title to form a structured image-text matching set.

[0052] like Figure 7 As shown, as an example, S2031 may include: S20311: Based on the image context paragraphs, image captions, and image titles in multiple sets of image data, obtain the first matching result, and based on the image titles in multiple sets of text data and multiple sets of image data, obtain the second matching result. The first matching result is used to indicate the degree of semantic similarity between the image context paragraphs and image captions, respectively. The first matching result is used to indicate the degree of semantic similarity between the text data and the image titles. Accordingly, S2032 may include: S20312: Based on the first matching result, match the image corresponding to the image title with the target paragraph to generate a first image matching set; and based on the second matching result, match the image corresponding to the image title with the target text data to generate a second image matching set. The first image matching set includes multiple sets of image paragraph matching results. The image paragraph matching results include the image, the image title, and the target paragraph corresponding to the image title. The first image matching set is used to output the image as a reference when outputting the target paragraph. The second image matching set is used to output the image as a reference when outputting the target text data.

[0053] In S20311, during the first matching stage, the system focuses on local contextual relationships. It calculates the semantic similarity between the image title and its directly adjacent preceding and following paragraphs to obtain the first matching result, which clearly indicates the closeness of the image to its most direct textual context. In the second matching stage, the system performs a global semantic scan, comparing the image title with all textual data units generated by document segmentation. It calculates the similarity to obtain the second matching result, which reveals the semantic association strength between the image and all potentially related text blocks within the document. For example, for an image titled "Schematic Diagram of Splitter Port Allocation," the first matching result might show a similarity of 0.88 with the following paragraph (describing port configuration rules), while the similarity with the preceding paragraph (introducing the principle of the splitter) is only 0.42. Simultaneously, the second matching result further reveals a high semantic match (0.91 similarity) between the image and a specific clause in the "Resource Allocation Specifications" section of another chapter of the document.

[0054] In S20312, based on the first matching result, the system binds each image to its most semantically similar direct context paragraph (preceding or following paragraph), generating a first image matching set. This set is organized in the form of "image-title-target paragraph" triples, ensuring that when a paragraph is matched by knowledge retrieval, the collaborative output of associated images is automatically triggered. Simultaneously, based on the deep semantic relationships revealed in the second matching result, the system associates each image with the most semantically compatible text data unit (which may be a sentence, clause, or independent point) globally, forming a second image matching set, achieving refined image-text correspondence across chapters and structures.

[0055] like Figure 8 As shown, as an example, S204 may include: S2043: Classify and summarize the first image matching set and the second image matching set to generate a knowledge base. The knowledge base includes a single-point knowledge set and a logical knowledge set. The single-point knowledge set includes multiple single-point knowledge points, each containing a knowledge point and the image matching that knowledge point. The logical knowledge set includes multiple logical knowledge points, each containing a knowledge point and the image matching that knowledge point. The knowledge point is the target paragraph or target text data. Single-point knowledge points are basic facts that do not require proof, while logical knowledge points are knowledge that includes multiple steps and has a clear logical relationship.

[0056] In S2043, the first and second image matching sets are merged and deduplicated to form a complete multimodal knowledge unit library. Each unit contains text content (paragraphs or independent text data), associated images, and their titles. Subsequently, the system uses a large language model fine-tuned with domain data to intelligently classify the text content in each knowledge unit, determining whether it belongs to "basic, unprovable factual knowledge" or "process-oriented knowledge containing multiple steps and clear logical relationships." Based on the classification results, the system assigns these multimodal knowledge units to two different knowledge storage systems: units classified as factual knowledge have their text and image vectorized representations stored in a vectorized database to support efficient approximate semantic retrieval; while units classified as process-oriented knowledge are used to construct a domain knowledge graph.

[0057] In one example, this solution constructs a dual-type knowledge base, categorizing different knowledge types into a vectorized knowledge base and a knowledge graph. During the knowledge base construction process, the document segmentation described above is first used to segment all input documents at the sentence and paragraph levels, and then images are associated. After completion, a general large model is used, employing personalized prompts to summarize all sentence and paragraph-level knowledge points. Since the general large model does not differentiate between different types of knowledge points, the prompts used in this solution are developed through independent debugging and iteration. The independently constructed prompts for this solution are shown below: Logical knowledge points: These knowledge points involve multiple steps or methods, usually requiring several sentences for a detailed description. These descriptions may involve different steps to complete the same operation, or corresponding methods used in different situations. For example, solving a math problem may require first understanding the problem, then choosing the appropriate formula, then performing calculations, and finally verifying the answer. Each step in this process is necessary and has a clear logical relationship.

[0058] Axiom-type knowledge points: These knowledge points are basic facts or principles that do not require proof and can usually be concisely described with a few words. For example, "the shortest distance between two points is a straight line" is an axiom in geometry that does not require further explanation or derivation and is a directly accepted basic fact.

[0059] After classifying different types of knowledge points using a general large model, all segment-level and sentence-level knowledge points will be divided into two main parts. One part consists of single-point knowledge points, including the relevant segment-level and sentence-level knowledge points, and corresponding images. The other part consists of logical knowledge points, including the relevant segment-level and sentence-level knowledge points, and corresponding images.

[0060] Single-point knowledge can be quickly addressed by using a vectorized knowledge base, avoiding the long wait times associated with using a uniform knowledge graph. Meanwhile, the broadband installation and maintenance scenario knowledge base contains a significant amount of operational process-related knowledge, which needs to be summarized using a knowledge graph to avoid missing knowledge points due to direct searching of the vectorized knowledge base, thus preventing disruption to on-site operations for those asking questions.

[0061] like Figure 9 As shown, as an example, S102 may also include: S901: Vectorize the user's original question information to obtain the user input vector; Accordingly, S102 may include: S902: Perform intent recognition on the user input vector to obtain the intent recognition result.

[0062] In the S901, after receiving the user's original query information, the system first starts the vectorization preprocessing module. This module uses a pre-trained semantic embedding model to transform the user's input natural language text into a high-dimensional dense vector, i.e., the user input vector.

[0063] In S902, based on the generated user input vector, the system feeds it into a lightweight FastText classification model for real-time intent recognition. This classification model has been trained using labeled broadband operations and maintenance (O&M) question data and can quickly determine the question type based on the semantic patterns of the vectors.

[0064] like Figure 10 As shown, as an example, S104 may include: S1041: Determine the accuracy of the question based on the user's original question information; S1042: Based on the precision of the question, output the knowledge points corresponding to the user's original question information and the matching images for those knowledge points.

[0065] In S1041, the system can determine the precision of a user's original query by analyzing its syntactic structure and semantic features. Specifically, the system identifies whether the query contains precise identifiers such as specific device models, error codes, or operation step numbers, and whether it uses broad interrogative words (such as "how" or "how to") or specific technical parameter terms. For example, a question like "What is the normal range of optical power received by a certain device's PON port?" is considered a high-precision question because it contains a clear device model and specific technical specifications; while a question like "What should I do if the network is slow?" is considered a low-precision question because it lacks specific context and technical focus.

[0066] In S1042, the system executes differentiated knowledge retrieval and output strategies based on the determined precision of the question. For high-precision questions, the system prioritizes strict semantic matching when matching knowledge points, directly locating knowledge units strongly related to the specific parameters and models in the question, and outputting their precisely corresponding text and image content; for low-precision questions, the system activates semantic generalization and multi-path retrieval mechanisms.

[0067] like Figure 11 As shown, as an example, S2043 may include: S20431: Perform knowledge type judgment on the first image matching set and the second image matching set, and generate a vectorized knowledge set and a knowledge graph. The vectorized knowledge set includes a single-point knowledge set, and the knowledge graph includes a logical knowledge set. The knowledge graph shown is used to indicate the relationship between multiple logical knowledge sets. S20432: Generate a knowledge base based on the vectorized knowledge set and knowledge graph.

[0068] In S20431, based on the integrated image-text units from the first and second image matching sets, the system performs deep discrimination and structured reorganization of knowledge types. Specifically, the system intelligently analyzes the core text of each image-text unit using a fine-tuned domain-specific large language model to determine its knowledge type: if it is independent, definite, and requires no contextual reasoning as single-point knowledge, its complete unit (including text, image, and related information) is encoded into a semantic vector and incorporated into the vectorized knowledge set; if it is logical knowledge containing logical relationships such as steps, conditions, and causality, it is transformed into a knowledge graph.

[0069] In S20432, based on the vectorized knowledge set and knowledge graph after classification processing, the system performs the final fusion construction to generate a unified scheduling multimodal hybrid knowledge base.

[0070] like Figure 12 As shown, as an example, S20311 may include: S203111: Summarize and extract the preceding and following paragraphs of each set of image data to obtain the summary information of each set of paragraphs. The summary information is used to indicate the semantics of the preceding and following paragraphs of the image. S203112: Semantically match the paragraph summary information and image titles of multiple sets of image data to obtain the first matching result, and semantically match the image titles of multiple sets of text data and multiple sets of image data to obtain the second matching result.

[0071] In S203111, the system performs intelligent semantic condensation processing on the context associated with each image in the document. Specifically, for the preceding and following paragraphs of an image, the system calls an optimized text summarization model to extract its core semantic information and generate structured paragraph summary information. This summary information is not a simple excerpt from the original text, but a concise expression of the paragraph's main idea, key entities, and core arguments, aiming to accurately depict the semantic core carried by the paragraph. For example, for an image displaying "fiber optic fusion splicer screen alarm codes," the preceding paragraph might describe the equipment self-test process in detail, which is summarized as "Common Status Descriptions During Equipment Self-Test"; the following paragraph might list the meanings of various alarm codes, summarized as "Alarm Code Classification and Corresponding Handling Suggestions." This paragraph summary information will serve as a key comparison benchmark for subsequent image-text semantic matching, scientifically determining which part of the context has a stronger logical connection to the image through semantic comparison with the image title.

[0072] In step S203112, after extracting the paragraph summary information, the system initiates a dual-channel semantic matching process. In the first matching channel, the semantic similarity between the paragraph summary information of each group of images and their corresponding image titles is calculated, generating the first matching result. This result quantifies the semantic closeness between the image and its immediate context. In the second matching channel, the system performs cross-regional, global semantic matching between all text data units (including paragraphs, sentences, etc.) obtained from document segmentation and all image titles, generating the second matching result. This result reveals the deep semantic connections between images and potentially non-directly adjacent text content within the document.

[0073] To better illustrate the question-and-answer method provided in the embodiments of this application, a specific embodiment is given below as an example, such as... Figure 13 As shown, it includes: S1301, Input document; S1302, Segment granular segmentation; S1303, Segment to image; S1304, segmentation is divided into: segmentation at the level of a sentence above the graph, summary of a segment above the graph, graph title, summary of a segment below the graph, and segmentation at the level of a sentence below the graph; S1305, Vectorization; S1306. Select a set of paragraphs and sentences with a high degree of similarity to the image title to form an image matching set; S1307, Prompt: Determine the type of knowledge; S13081, Single-point knowledge image collection; S13082, Operational knowledge image collection; S1309, User Question; S13101, Question intent targets a set of single-point knowledge images; S13102, Question intent targets a set of operational knowledge images; S131011. For single-point knowledge image sets, when the user's question is brief, output the segment-level knowledge points and corresponding images; S131012. For single-point knowledge image collections, when the user asks detailed questions, output sentence-level knowledge points and corresponding images. S131021. For operational knowledge image sets, when the user asks detailed questions, output sentence-level knowledge points and corresponding images. S131022. For operational knowledge image sets, if the user's question is brief, output the segment-level knowledge points and corresponding images.

[0074] In one example, the specific steps are as follows: Step 1: Document splitting, the process is the same as the document splitting module above.

[0075] Step 2: Divide the knowledge base into single-point and paragraph-based types, following the same process as the knowledge base construction described above.

[0076] Step 3: User semantic understanding. The user input is vectorized using the Word2Vec method to obtain the user input vector.

[0077] Step 4: Knowledge screening. Based on the intent recognition results of the user input vector, determine whether it is a question about single-point knowledge or an operational knowledge question.

[0078] Step 5: Knowledge Output. Manhattan distance is used for vector matching to search for the knowledge points the user expects to output. If the knowledge point has a related image from Step 1, both the image and text are output simultaneously.

[0079] The above describes a specific implementation of a question-and-answer method provided in this application. Based on the question-and-answer method provided in the above embodiments, this application also provides a specific implementation of a question-and-answer device, as shown in the following embodiments.

[0080] like Figure 14 As shown in the embodiment of this application, a question-and-answer device 1400 is provided, which includes: The data acquisition module 1401 is used to collect the user's original question information; The identification module 1402 is used to identify the intent of the user's original question information and obtain the intent identification result. The intent identification result is used to indicate that the user's original question information belongs to basic facts that do not require proof or knowledge that includes multiple steps and has a clear logical relationship. The matching module 1403 is used to perform vector matching in a preset knowledge base based on the intent recognition result. It searches the preset knowledge base to obtain the knowledge points corresponding to the user's original question information and the images matching the knowledge points. The preset knowledge base is obtained by classifying and summarizing the matching sets obtained by matching multiple sets of image data and multiple sets of text data. The multiple sets of image data and multiple sets of text data are obtained by segmenting the acquired business knowledge documents. Output module 1404 is used to output the knowledge points corresponding to the user's original question information and the images matching the knowledge points.

[0081] Thus, the question-answering device provided in this application embodiment performs intent recognition on the collected original user question information to obtain intent recognition results. Based on the intent recognition results, it performs vector matching in a preset knowledge base to search for the knowledge points corresponding to the original user question information and the images matching the knowledge points in the preset knowledge base. Then, it outputs the knowledge points corresponding to the original user question information and the images matching the knowledge points.

[0082] By using a multimodal document segmentation method, images can be identified during the segmentation process and bound to the preceding or following paragraphs. When the output text involves the preceding or following content, the images can be included as reference images to avoid knowledge loss.

[0083] In another embodiment of this application, the above-described device 1400 may further include: The acquisition module is used to acquire business knowledge documents; The segmentation module is used to segment business knowledge documents to obtain multiple sets of image data and multiple sets of text data. The image data includes images, image titles, paragraphs above and below the images, and text data consists of sentences. The semantic matching module matches the preceding and following paragraphs of images, multiple sets of text data, and image titles in multiple sets of image data to obtain a matching set. The matching set is used to output the corresponding images of the matching image titles as a reference when outputting the target paragraph or target text data. The target paragraph is the paragraph in the preceding and following paragraphs of the image that is semantically closer to the image title. The target text data is the text data in multiple sets of text data whose semantic similarity to the image title is greater than a preset threshold. The generation module is used to classify and summarize the matching set and generate a preset knowledge base. The knowledge points in the preset knowledge base are basic facts that do not require proof or knowledge that includes multiple steps and has a clear logical relationship.

[0084] In another embodiment of this application, the above-mentioned generation module may further include: The inductive unit is used to classify the matching set into single-point knowledge sets and logical knowledge sets. The single-point knowledge set includes multiple single-point knowledge points, which include knowledge points and images matching the knowledge points. The logical knowledge set includes multiple logical knowledge points, which include knowledge points and images matching the knowledge points. Knowledge points are target paragraphs or target text data. Single-point knowledge points are basic facts that do not require proof, while logical knowledge points are knowledge that includes multiple steps and has a clear logical relationship. The generation unit is used to generate a preset knowledge base based on single-point knowledge sets and logical knowledge sets.

[0085] As another embodiment of this application, the semantic matching module may further include: The first matching unit is used to match the preceding paragraphs, following paragraphs, multiple sets of text data, and image titles in multiple sets of image data to obtain matching results. The matching results are used to indicate the semantic similarity between the preceding paragraphs, following paragraphs, multiple sets of text data, and image titles. The second matching unit is used to match the target paragraph or target text data with the image title based on the matching results, and obtain a matching set.

[0086] As another embodiment of this application, the first matching unit may further include: The first matching subunit is used to obtain a first matching result based on the image context paragraphs, image description paragraphs, and image titles in multiple sets of image data, and to obtain a second matching result based on the image titles in multiple sets of text data and multiple sets of image data. The first matching result is used to indicate the degree of semantic similarity between the image context paragraphs and image description paragraphs and the image titles, respectively. The first matching result is used to indicate the degree of semantic similarity between the text data and the image titles. Accordingly, the second matching unit may further include: The second matching subunit is used to match the image corresponding to the image title with the target paragraph based on the first matching result, generating a first image matching set; and to match the image corresponding to the image title with the target text data based on the second matching result, generating a second image matching set. The first image matching set includes multiple sets of image paragraph matching results, which include the image, the image title, and the target paragraph corresponding to the image title. The first image matching set is used to output the image as a reference when outputting the target paragraph, and the second image matching set is used to output the image as a reference when outputting the target text data.

[0087] In another embodiment of this application, the above-mentioned generation module may further include: The knowledge base generation unit is used to classify and summarize the first image matching set and the second image matching set to generate a knowledge base. The knowledge base includes single-point knowledge sets and logical knowledge sets. The single-point knowledge set includes multiple single-point knowledge points, which include knowledge points and images matching those knowledge points. The logical knowledge set includes multiple logical knowledge points, which include knowledge points and images matching those knowledge points. Knowledge points are target paragraphs or target text data. Single-point knowledge points are basic facts that do not require proof, while logical knowledge points are knowledge that includes multiple steps and has a clear logical relationship.

[0088] In another embodiment of this application, the above-described device 1400 may further include: The vectorization module is used to vectorize the user's original query information to obtain the user input vector; Accordingly, the recognition module can be specifically used for: The intent is identified by performing intent recognition on the user input vector.

[0089] As another embodiment of this application, the above-mentioned output module may further include: The determination unit is used to determine the accuracy of the question based on the user's original question information; The output unit is used to output the knowledge points corresponding to the user's original question information and the matching images based on the precision of the question.

[0090] As another embodiment of this application, the knowledge base generation unit described above can also be specifically used for: The knowledge type is determined for the first image matching set and the second image matching set, and a vectorized knowledge set and a knowledge graph are generated. The vectorized knowledge set includes a single-point knowledge set, and the knowledge graph includes a logical knowledge set. The knowledge graph shown is used to indicate the relationship between multiple logical knowledge sets. A knowledge base is generated based on vectorized knowledge sets and knowledge graphs.

[0091] As another embodiment of this application, the first matching subunit described above can also be specifically used for: The preceding and following paragraphs of each image data set are summarized and extracted to obtain the summary information of each paragraph. The summary information is used to indicate the semantics of the preceding and following paragraphs of the image. The first matching result is obtained by semantically matching the paragraph summary information and image titles of multiple sets of image data, and the second matching result is obtained by semantically matching the image titles of multiple sets of text data and multiple sets of image data.

[0092] Based on the question-answering method and apparatus provided in the above embodiments, this application also provides an electronic device 1500, such as... Figure 15 As shown: It includes a processor 1501, a memory 1502, and a computer program stored in the memory 1502 and executable on the processor 1501. When the computer program is executed by the processor 1501, it implements the various processes of the above-described question-and-answer method embodiments and achieves the same technical effect.

[0093] Specifically, the processor 1501 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The memory 1502 may include a mass storage device for data or instructions. For example, and not limitingly, the memory 1502 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 1502 may include removable or non-removable (or fixed) media. Where appropriate, the memory 1502 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, the memory 1502 is a non-volatile solid-state memory.

[0094] In certain embodiments, the memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Thus, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this application.

[0095] The processor 1501 implements any of the question-and-answer methods described in the above embodiments by reading and executing computer program instructions stored in the memory 1502.

[0096] In one example, the electronic device may also include a communication interface 1503 and a bus 1510. As an example, such as... Figure 15 As shown, the processor 1501, memory 1502, and communication interface 1503 are connected through bus 1510 and complete communication with each other.

[0097] The communication interface 1503 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0098] Bus 1510 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 1510 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0099] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described question-and-answer method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0100] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0101] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0102] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0103] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0104] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A question and answer method, characterized by, The method comprises: collecting original question information of a user; performing intention recognition on the original question information of the user to obtain an intention recognition result, the intention recognition result being used to indicate that the original question information of the user belongs to basic facts that do not need to be proved or knowledge that includes multiple steps and has an explicit logical relationship; performing vector matching in a preset knowledge base according to the intention recognition result, searching for a knowledge point corresponding to the original question information of the user and a picture matched with the knowledge point in the preset knowledge base, the preset knowledge base being obtained by classifying and summarizing a matching set obtained by matching a plurality of sets of picture data and a plurality of sets of text data, the plurality of sets of picture data and the plurality of sets of text data being obtained by segmenting obtained business knowledge documents; outputting the knowledge point corresponding to the original question information of the user and the picture matched with the knowledge point.

2. The method of claim 1, wherein, Before the vector matching in the preset knowledge base according to the intention recognition result, the method further comprises: obtaining a business knowledge document; segmenting the business knowledge document to obtain a plurality of sets of picture data and a plurality of sets of text data, the picture data including a picture, a picture caption, a picture preceding paragraph and a picture following paragraph, and the text data being text in a sentence unit; matching the picture preceding paragraph, the picture following paragraph, the plurality of sets of text data and the picture caption in the plurality of sets of picture data to obtain a matching set, the matching set being used to output the picture corresponding to the picture caption as a reference when a target paragraph or a target text data is output, the target paragraph being a paragraph in the picture preceding paragraph and the picture following paragraph that is more semantically close to the picture caption, and the target text data being text data in the plurality of sets of text data that is more similar to the picture caption than a preset threshold; classifying and summarizing the matching set to generate the preset knowledge base, the knowledge point in the preset knowledge base being basic facts that do not need to be proved or knowledge that includes multiple steps and has an explicit logical relationship.

3. The method of claim 2, wherein, The classifying and summarizing the matching set to generate the preset knowledge base comprises: summarizing and classifying the matching set to obtain a single-point knowledge set and a logical knowledge set, the single-point knowledge set including a plurality of single-point knowledge, the single-point knowledge including a knowledge point and a picture matched with the knowledge point, the logical knowledge set including a plurality of logical knowledge, the logical knowledge including a knowledge point and a picture matched with the knowledge point, the knowledge point being the target paragraph or the target text data, the single-point knowledge being basic facts that do not need to be proved, and the logical knowledge being knowledge that includes multiple steps and has an explicit logical relationship; generating a preset knowledge base according to the single-point knowledge set and the logical knowledge set.

4. The method of claim 3, wherein, The matching according to the picture context paragraph, the picture context paragraph, the plurality of groups of text data and the picture title in the plurality of groups of picture data, obtains a matching set, and the matching set includes: The matching according to the picture context paragraph, the picture context paragraph, the plurality of groups of text data and the picture title in the plurality of groups of picture data, obtains a matching result, and the matching result is used to indicate the similarity of the semantic of the picture context paragraph, the picture context paragraph, the plurality of groups of text data and the picture title respectively; According to the matching result, the target paragraph or the target text data is matched with the picture title respectively, and a matching set is obtained.

5. The method of claim 4, wherein, The matching according to the picture context paragraph, the picture context paragraph, the plurality of groups of text data and the picture title in the plurality of groups of picture data, obtains a matching result, and the matching result includes: According to the picture context paragraph, the picture context paragraph and the picture title in the plurality of groups of picture data, a first matching result is obtained, and according to the plurality of groups of text data and the picture title in the plurality of groups of picture data, a second matching result is obtained, the first matching result is used to indicate the similarity of the semantic of the picture context paragraph and the picture context paragraph respectively with the picture title, and the first matching result is used to indicate the similarity of the semantic of the text data with the picture title; The matching according to the matching result, the target paragraph or the target text data is matched with the picture title respectively, and a matching set is obtained. According to the first matching result, the picture corresponding to the picture title is matched with the target paragraph to generate a first picture matching set, and according to the second matching result, the picture corresponding to the picture title is matched with the target text data to generate a second picture matching set, the first picture matching set includes a plurality of picture paragraph matching results, the picture paragraph matching result includes a picture, a picture title and a target paragraph corresponding to the picture title, the first picture matching set is used to output the picture as a reference when the target paragraph is output, and the second picture matching set is used to output the picture as a reference when the target text data is output.

6. The method of claim 5, wherein, The classification and induction of the matching set generates the preset knowledge base, and the classification and induction of the matching set generates the preset knowledge base, including: The classification and induction of the first picture matching set and the second picture matching set generate a knowledge base, the knowledge base includes a single-point knowledge set and a logic knowledge set, the single-point knowledge set includes a plurality of single-point knowledge, the single-point knowledge includes a knowledge point and a picture matched with the knowledge point, the logic knowledge set includes a plurality of logic knowledge, the logic knowledge includes a knowledge point and a picture matched with the knowledge point, the knowledge point is the target paragraph or the target text data, the single-point knowledge is a basic fact without proof, and the logic knowledge is a knowledge including a plurality of steps and having a clear logical relationship.

7. The method of claim 6, wherein, Before the intention recognition of the user's original question information obtains an intention recognition result, it further includes: vectorize the user original question information to obtain a user input vector; the intent recognition of the user original question information obtains an intent recognition result, including: the intent recognition of the user input vector obtains an intent recognition result.

8. The method of claim 7, wherein, the output of the knowledge point corresponding to the user original question information and the picture matched with the knowledge point includes: determine the accuracy of the question according to the user original question information; output the knowledge point corresponding to the user original question information and the picture matched with the knowledge point according to the accuracy of the question.

9. The method of claim 8, wherein, the classification and induction of the first picture matching set and the second picture matching set to generate a knowledge base, including: knowledge type judgment is performed on the first picture matching set and the second picture matching set to generate a vectorized knowledge set and a knowledge graph, the vectorized knowledge set includes a single point type knowledge set, and the knowledge graph includes a logic type knowledge set, the knowledge graph is used to indicate the association relationship between multiple logic type knowledge; generate the knowledge base according to the vectorized knowledge set and the knowledge graph.

10. The method of claim 9, wherein, the first matching result is obtained according to the picture context paragraph, the picture context paragraph and the picture title in the picture data in the multiple groups, and the second matching result is obtained according to the multiple groups of text data and the picture title in the multiple groups of picture data, including: summarize and extract the picture context paragraph and the picture context paragraph in each group of picture data to obtain each group of paragraph summary information, and the paragraph summary information is used to indicate the semantics of the picture context paragraph and the picture context paragraph; the paragraph summary information and the picture title of the multiple groups of picture data are matched to obtain the first matching result, and the multiple groups of text data and the picture title in the multiple groups of picture data are matched to obtain the second matching result.

11. A question answering apparatus characterized by comprising: the device includes: a collection module for collecting user original question information; an identification module for intent recognition of the user original question information to obtain an intent recognition result, the intent recognition result is used to indicate that the user original question information belongs to basic facts without proof or includes multiple steps and has a clear logical relationship; a matching module for vector matching in a preset knowledge base according to the intent recognition result, searching for a knowledge point corresponding to the user original question information and a picture matched with the knowledge point in the preset knowledge base, the preset knowledge base is obtained by classifying and inducing a matching set matched by multiple groups of picture data and multiple groups of text data, and the multiple groups of picture data and multiple groups of text data are obtained by cutting the obtained business knowledge documents; an output module for outputting the knowledge point corresponding to the user original question information and the picture matched with the knowledge point.

12. An electronic device, comprising: the device includes a processor and a memory storing computer program instructions; the processor executes the computer program instructions to realize the question and answer method of any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer program instructions, and the computer program instructions are executed by a processor to implement the question and answer method in any one of claims 1-10.

14. A computer program product, characterised in that, The instructions in the computer program product are executed by a processor of an electronic device to enable the electronic device to perform the question and answer method in any one of claims 1-10.