Multi-model fused vertical domain knowledge question-answering method and system

Through the multi-model fusion architecture, the enhanced small parameter model is used for document analysis and intention understanding, and combined with the large parameter model for knowledge retrieval and style optimization, the problem of insufficient accuracy in professional field knowledge Q&A is solved, and an efficient and economical Q&A system is realized.

CN120336488AActive Publication Date: 2025-07-18POWERCHINA BEIJING ENG CORP

Patent Information

Application Number
CN202510503452.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

When handling knowledge questions and answers in professional fields, the existing single large model architecture is difficult to have the comprehensiveness and accuracy of knowledge, it is easy to generate error information, and it consumes a lot of resources and slow response speed.

Method used

A multi-model fusion architecture is adopted, and the enhanced small parameter model trained for vertical field is used for document analysis, combined with rules to assist in extracting fixed position information, a general language model with lower parameters is used for intent understanding, a model with larger parameters is used for knowledge retrieval and answer generation, and a special small parameter model is used for style optimization, and finally the analysis effect is optimized through the answer reflow mechanism.

Benefits of technology

It significantly improves the accuracy and efficiency of knowledge questions and answers in vertical fields, reduces resource consumption, improves response speed, and realizes the system's self-optimization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336488A_ABST
    Figure CN120336488A_ABST
Patent Text Reader

Abstract

The invention provides a multi-model fusion vertical class domain knowledge question-answering method and system, and belongs to the field of artificial intelligence natural language processing. According to the scheme, a traditional single large model is disassembled into a plurality of specialized models for cooperative work; firstly, an enhanced small parameter model trained for the vertical field is used for performing document analysis, and fixed position information is extracted in combination with rules in an auxiliary manner; then performing intention understanding by using a general language model with relatively low parameters, assisted by a cue word strategy, a rule model or a knowledge graph; then, based on the analyzed document information and the understood user intention, knowledge retrieval and answer generation are carried out through a model with large parameters; and finally, carrying out style optimization by adopting a special small parameter model, and generating an answer conforming to the characteristics of the vertical class field. An answer backflow mechanism is further designed, and continuous optimization of the system is achieved. According to the scheme, the resource consumption is greatly reduced, the accuracy and efficiency of the knowledge questions and answers in the vertical class domain are improved, and the'illusion 'problem existing in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a multi-model fusion vertical domain knowledge question-answering method and system. Background Art

[0002] With the rapid development of artificial intelligence technology, knowledge question-answering systems, as an important application in the field of natural language processing, have been widely used in multiple fields. A knowledge question-answering system can understand questions in the form of natural language from users and retrieve relevant information from a knowledge base to generate accurate and reasonable answers. Such a system has become an important tool for obtaining professional knowledge and solving problems, and is of great significance for improving the efficiency of information acquisition.

[0003] In the existing technical solutions for knowledge question-answering, often a large model is responsible for the entire process from intention understanding to answer generation. This single large model architecture performs well in dealing with general domain problems, but has obvious deficiencies when facing vertical domain knowledge with strong professionalism and complex organizational modalities. For example, knowledge question-answering in the engineering field not only requires understanding multimodal information such as complex design drawings and technical specifications, but also requires a deep professional knowledge background.

[0004] Due to the limitations of the pre-training data of large models, it is difficult for a single model solution to simultaneously possess the comprehensiveness and accuracy of knowledge. Especially for professional fields, when dealing with complex problems that require precise knowledge, it often generates some seemingly reasonable but actually incorrect information, that is, the "hallucination" problem. This phenomenon is particularly obvious when dealing with fields with high-precision requirements such as engineering design and medical diagnosis, and may lead to serious decision-making biases.

[0005] In order to improve the accuracy of the model in professional fields, the current mainstream solution is to increase the scale of model parameters or fine-tune through a large amount of domain data. However, this method means a greater demand for training resources and a longer training time, greatly increasing the development and application costs. For many application scenarios with limited resources, this solution is difficult to implement. At the same time, simply increasing parameters cannot fundamentally solve the problem that the model does not fully understand professional knowledge.

[0006] In addition, the existing single large model architecture has poor parsing effects when dealing with complex format documents (such as engineering design drawings), resulting in inaccurate retrieval results, which in turn affects the quality of question-answering. And for different tasks such as intention understanding, answer generation, and answer style adjustment, a single model lacks the ability to optimize specifically and is difficult to achieve the best effect in each link. Summary of the Invention

[0007] The object of the present invention is to provide a multi-model fusion vertical domain knowledge Q&A method and system, which can not only ensure the accuracy and authority of professional domain knowledge Q&A, but also have advantages in resource consumption and response speed, meeting the actual needs of vertical domain knowledge services.

[0008] To achieve the above object, the present invention is realized through the following technical solutions:

[0009] A multi-model fusion vertical domain knowledge Q&A method includes the following steps:

[0010] S1: Use an enhanced small-parameter model trained for the vertical domain to perform document parsing and extract structured information from the document;

[0011] S2: Use a general language model with lower parameters to perform intent understanding of the user's question;

[0012] S3: Based on the document information parsed in step S1 and the intent understanding result in step S2, use a model with larger parameters to perform knowledge retrieval and answer generation;

[0013] S4: Use a dedicated small-parameter model to optimize the style of the answer generated in step S3 and generate a final answer that conforms to the characteristics of the vertical domain.

[0014] Further: In step S1, it also includes: using rules to assist the enhanced small model in document parsing, and the rules include locating fixed-position information according to the standard format of vertical domain documents.

[0015] Further: In step S2, it also includes: using at least one of a prompt strategy, a rule model, or a knowledge graph as an auxiliary means to supplement the general language model for intent understanding.

[0016] Further: The prompt strategy includes: rewriting the user's question to optimize and expand the question and provide a clear generation direction.

[0017] Further: In step S3, it also includes: for vertical domain questions whose complexity exceeds a preset threshold, first use a small-parameter model to perform rapid knowledge filtering and rough ranking, and then call a large-parameter model to perform in-depth answer generation.

[0018] Further: It also includes: returning the final answer generated in step S4 to the enhanced small model in step S1 for continuously optimizing the document parsing effect and enhancing the generalization of the parsing model.

[0019] A multi-model fusion vertical domain knowledge Q&A system applying the above Q&A method includes:

[0020] The document parsing module is used to parse documents using an enhanced small-parameter model trained for vertical fields to extract structured information from documents;

[0021] An intention understanding module is used to understand the intention of user questions using a general language model with lower parameters; an answer generation module is used to perform knowledge retrieval and answer generation using a model with larger parameters based on the document information parsed by the document parsing module and the intention understanding result of the intention understanding module;

[0022] The style optimization module is used to use a dedicated small parameter model to optimize the style of the answer generated by the answer generation module to generate a final answer that conforms to the characteristics of the vertical field.

[0023] Furthermore: the document parsing module also includes: a rule-assisted unit, which is used to use rules to assist the enhanced small model in document parsing, and the rules include positioning fixed position information according to the standard layout of vertical field documents.

[0024] Furthermore: the intention understanding module also includes: an auxiliary understanding unit, which is used to use at least one of the prompt word strategy, rule model or knowledge graph as an auxiliary means to supplement the general language model for intention understanding.

[0025] Furthermore: it also includes: a model optimization module, which is used to feed the final answer generated by the style optimization module back to the enhanced small model in the document parsing module, so as to continuously optimize the document parsing effect and enhance the generalization of the parsing model.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] 1. The present invention significantly improves the accuracy of knowledge questions and answers in vertical fields. By using an enhanced small model trained for vertical fields to specifically parse documents and combining rules to assist in extracting fixed position information, it is possible to more accurately understand and process highly professional and complex format documents, thereby improving the quality of knowledge extraction from the source. At the same time, answer generation through a model with larger parameters ensures the fluency and logical coherence of the generated content, effectively reducing the occurrence of "hallucination" problems.

[0028] Second, the present invention significantly reduces system resource consumption. The "division of labor and cooperation" model is adopted to enable models of various sizes to play a role in their respective areas of expertise, and the process that originally required hundreds of computing cards for pre-training is transformed into a training process that can be performed by a few cards, which not only saves computing resources but also reduces energy consumption, making the system more economical and efficient, and easier to deploy and apply.

[0029] III. The present invention significantly improves the response speed of the question-answering system. In the intention understanding stage, a general language model with relatively low parameters is used to quickly analyze the user's intention; for complex questions, a strategy of first filtering with a small model and then deeply generating with a large model is adopted, optimizing the processing efficiency while ensuring the quality of the answers; the small-parameter model style optimization stage is also more efficient than using a large model for style adjustment, effectively controlling the response time overall.

[0030] IV. The present invention constructs a mechanism for the system to self-optimize. Through the design of feeding the final answer back to the document parsing small model, the continuous optimization of the system performance and the improvement of the generalization ability are achieved, enabling the system to continuously adapt to new document formats and knowledge structures in practical applications and maintaining long-term effectiveness and applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic flowchart of a multi-model fusion vertical domain knowledge question-answering method of the present invention in one embodiment;

[0032] Figure 2 It is a schematic flowchart of a multi-model fusion vertical domain knowledge question-answering method of the present invention in another embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0035] The technical solutions provided by the present invention can be widely applied to the field of vertical domain knowledge question-answering with strong professionalism and complex organizational modalities, such as, but not limited to, professional fields such as engineering design, medical health, and legal consultation. The following will take the engineering field as an example for detailed description.

[0036] In the prior art, knowledge Q&A systems usually adopt a single large model to be responsible for the whole process from intent understanding to answer generation. However, due to the limitations of the model's pre-training data, when dealing with complex problems in professional fields, it is often difficult to simultaneously possess the comprehensiveness and accuracy of knowledge, and it is easy to generate the "hallucination" problem. To solve these problems, the present invention disassembles the answering process that a large model is responsible for into multiple different models, and each model cooperates to complete the Q&A process, taking into account both the comprehensiveness of general knowledge understanding and ensuring the controllability and accuracy of professional knowledge understanding and generation.

[0037] As Figure 1 shown, the method for vertical domain knowledge Q&A with multi-model fusion of the present invention includes the following steps:

[0038] S1: Use an enhanced small-parameter model trained for the vertical domain to perform document parsing and extract structured information from the document, where the number of parameters of the small-parameter model is in the range of 7M to 10M levels. As a preprocessing process for the application of large model knowledge, the result of document parsing is directly related to the accuracy of Q&A. Especially for documents with strong professionalism and complex formats, the quality of the parsing effect is one of the most important factors affecting knowledge recall. For the engineering field in this embodiment, a model with relatively small parameters can be fine-tuned to understand the format and content of engineering design drawings.

[0039] At the same time, to further improve the parsing effect, the present invention also introduces a rule-assisted parsing mechanism in the document parsing step. For example, when processing engineering design drawings, according to the standard drawing format such as ISO / A3 drawing frame, the title bar and the fixed area of the drawing frame can be located at the lower right corner, and key information can be extracted from these specific positions. This way of combining small models with rules can not only save computing resources but also obtain a vertical model with better effects in less training time.

[0040] S2: Use a general language model to perform intent understanding of user questions, where the number of parameters of the general language model is in the range of 100M to 1B levels. Research shows that as the scale of model parameters increases, the improvement of the accuracy of natural language inference tasks is not proportional. According to the research by Brown et al. in 2020, when the parameters increase by 10 times, the average improvement of the accuracy of natural language inference tasks is only 3-5%. Therefore, a low-parameter model of 7B level can be used for intent understanding. When processing the text input by users, the amount of data to be processed and calculated is small, and the analysis and understanding of the user's intent can be completed in a short time, and a response can be given quickly.

[0041] However, fewer parameters may lead to insufficient understanding ability of the model for complex and diverse user question forms, especially for questions with a low frequency of occurrence or novel and unique expressions in the training data. To solve this problem, the present invention introduces auxiliary means in the intent understanding step, mainly including three categories:

[0042] One is the prompt strategy. By rewriting the user's questions, optimizing and expanding the questions, and providing clear generation directions, the questions are made more targeted and operable. For example, rewrite "What should be noted during the design of a rockfill dam?" as "Please query the precautions for a rockfill dam in terms of dam site selection, material selection, structural design, construction process, etc.".

[0043] The second is the rule model. Utilize its accuracy in processing texts with specific patterns and the advantages of machine learning models in certain local features as a supplement to the general language model to improve the overall performance of intent understanding.

[0044] The third is the knowledge graph, which serves as an external knowledge base to help the model better understand the semantic information in the text. For example, when the user mentions certain professional entities, the model can obtain information such as the attributes and relationships of the relevant entities with the help of the knowledge graph, so as to more accurately understand the user's intent.

[0045] The above three auxiliary means adopt the following principles in specific use: the prompt strategy is used to expand the user's intent; the rule model generally processes scenarios where the document template is relatively fixed and the question and answer formats are relatively fixed, and it is a supplement to the algorithm; the knowledge graph is used to improve the accuracy of recall.

[0046] S3: Based on the document information parsed in step S1 and the intent understanding result in step S2, use a large-parameter model for knowledge retrieval and answer generation, where the number of parameters of the large-parameter model is above the 1B level. The scale of the model parameters is positively correlated with the fluency of answer generation and the logical coherence. Models with more parameters usually generate more fluent and logical content because they can remember more details and knowledge. Therefore, using a model with larger parameters can obtain better-quality generated answers.

[0047] It should be noted that for extremely complex engineering knowledge, the present invention also provides an optimization strategy: first use a small model for rapid filtering and rough ranking to determine the general direction and key points of the answer; then call the large model for in-depth generation to improve the details and logical structure. This hierarchical processing method not only ensures the processing efficiency but also ensures the quality of the final answer.

[0048] S4: Optimize the style of the answer generated in step S3 using a dedicated small-parameter model to generate the final answer that conforms to the characteristics of the vertical domain, where the number of parameters of the dedicated small-parameter model is in the range of 7M to 10M. There are mainly two methods for style optimization. One is to optimize by writing prompts based on existing large language models; the other is to input a training set with the corresponding style for the model to learn and generate. For example, for the engineering field, a specific style optimization model can be trained to make the generated answer more rigorous, accurate, and standardized, conforming to the expression habits and professional standards of engineering professionals.

[0049] In another embodiment, to achieve the continuous optimization of the system, the present invention also designs an answer feedback mechanism in step S5: Feed back the finally optimized answer to the document parsing small model for continuously optimizing the parsing effect and enhancing the generalization of the parsing model. The specific feedback mechanism is to first create a data set, then feed back the question-answer pairs into the data set, and after manual verification, apply them to model training. In this way, the system can learn from actual applications, continuously adapt to new document formats and question types, and maintain long-term effectiveness.

[0050] In another embodiment, the present invention also provides a vertical domain knowledge Q&A system with multi-model fusion to implement the above method. The system includes a document parsing module, an intent understanding module, an answer generation module, and a style optimization module, corresponding to each step in the above method respectively. Among them, the document parsing module includes a rule assistance unit for using rule assistance to enhance the small model for document parsing; the intent understanding module includes an assistance understanding unit for using prompt strategies, rule models, or knowledge graphs as auxiliary means; the system also includes a model optimization module for implementing the answer feedback mechanism.

[0051] In practical applications, taking the engineering design field as an example, when a user asks a question such as "What should be noted in the design of a reinforced concrete frame structure", the system first uses the trained enhanced small model to parse relevant engineering design specifications, design drawings, and other documents; then analyzes the user's intent through the intent understanding module to understand that the user is asking about design considerations; then, the large-parameter model retrieves relevant knowledge and generates an answer based on the parsed document information and the understood user intent, covering the key considerations in the design of a reinforced concrete frame structure; finally, the style optimization module adjusts the expression of the answer to make it conform to the professional expression habits of the engineering field and forms the final answer to be returned to the user.

[0052] Through the above multi-model fusion method, the present invention effectively solves the problems existing in single large models in vertical domain knowledge Q&A, such as insufficient accuracy, high resource consumption, slow response speed, etc., and provides a new technical path for vertical domain intelligent Q&A systems. Compared with traditional methods, the technical solution of the present invention can significantly reduce resource consumption while improving the knowledge accuracy and response efficiency of Q&A systems in professional fields, and has broad application prospects.

[0053] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. All equivalent transformations or modifications made according to the spirit of the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for question answering in a vertical domain with multi-model fusion, characterized in that, Including the following steps: S1: Use an enhanced small-parameter model trained for a vertical domain to perform document parsing and extract structured information from the document, where the number of parameters of the small-parameter model is in the range of 7M to 10M; S2: Use a general language model to perform intent understanding of user questions, where the number of parameters of the general language model is in the range of 100M to 1B; S3: Based on the document information parsed in step S1 and the intent understanding result in step S2, use a large-parameter model to perform knowledge retrieval and answer generation, where the number of parameters of the large-parameter model is above 1B; S4: Use a dedicated small-parameter model to optimize the style of the answer generated in step S3 and generate a final answer that conforms to the characteristics of the vertical domain, where the number of parameters of the dedicated small-parameter model is in the range of 7M to 10M.

2. The multi-model fusion-based vertical domain knowledge Q&A method according to claim 1, wherein In step S1, it further includes: using rules to assist the enhanced small model in document parsing, and the rules include locating fixed-position information according to the standard format of vertical domain documents.

3. A method for vertical domain knowledge Q&A with multi-model fusion according to claim 1, characterized in that, In step S2, it further includes: using at least one of a prompt strategy, a rule model, or a knowledge graph as an auxiliary means to supplement the general language model for intent understanding.

4. A method for vertical domain knowledge Q&A with multi-model fusion according to claim 3, characterized in that, The prompt strategy includes: rewriting the user question to optimize and expand the question and provide a clear generation direction.

5. A method for vertical domain knowledge Q&A with multi-model fusion according to claim 1, characterized in that, In step S3, it further includes: for complex vertical domain questions, first use a small-parameter model to perform rapid knowledge filtering and rough ranking, and then call a large-parameter model to generate in-depth answers.

6. A method for question answering in a vertical domain with multi-model fusion according to claim 1, characterized in that, It further includes step S5: flowing back the final answer generated in step S4 to the enhanced small model in step S1 for continuously optimizing the document parsing effect and enhancing the generalization of the parsing model.

7. A vertical domain knowledge Q&A system with multi-model fusion applying the Q&A method according to any one of claims 1-6 above, characterized in that, Including: A document parsing module for using an enhanced small-parameter model trained for a vertical domain to perform document parsing and extract structured information from the document; An intent understanding module for using a general language model with lower parameters to perform intent understanding of user questions; An answer generation module for performing knowledge retrieval and answer generation using a model with larger parameters based on the document information parsed by the document parsing module and the intent understanding result of the intent understanding module; A style optimization module for using a dedicated small-parameter model to optimize the style of the answer generated by the answer generation module and generate a final answer that conforms to the characteristics of the vertical domain.

8. A vertical domain knowledge Q&A system with multi-model fusion according to claim 7, characterized in that, The document parsing module further includes: a rule assistance unit for using rules to assist the enhanced small model in document parsing, and the rules include locating fixed-position information according to the standard format of vertical domain documents.

9. A vertical domain knowledge Q&A system with multi-model fusion according to claim 7, characterized in that, The intent understanding module further includes: an auxiliary understanding unit for using at least one of a prompt strategy, a rule model, or a knowledge graph as an auxiliary means to supplement the general language model for intent understanding.

10. A vertical domain knowledge Q&A system with multi-model fusion according to claim 7, characterized in that, It further includes: A model optimization module for flowing back the final answer generated by the style optimization module to the enhanced small model in the document parsing module for continuously optimizing the document parsing effect and enhancing the generalization of the parsing model.

Citation Information

Patent Citations

  • Building standard knowledge question and answer model construction method, computer program product, storage medium and electronic equipment

    CN118296118A

  • Model question answering system

    CN118981517A

  • A large language model prompt word engineering method and system based on multi-round RAG technology

    CN119761521A

  • A system for generating answers to multiple questions using rag-based generative artificial intelligence technology

    KR102710159B1

  • Method and system for analyzing natural language data by using domain-specific language models

    US20250013633A1

Cited By

  • A method and system for automated generation of vertical domain knowledge question-answer pairs

    CN122507865A