Enterprise knowledge question-answering system based on large language model

Through the enterprise knowledge question and answer system based on the large language model, the semantic inconsistency problem under the fragmentation of knowledge within the enterprise and the differentiation of permissions is solved, and an accurate and compliant enterprise intelligent question and answer system is realized, which improves the accuracy and security of the question and answer system.

CN120470096AActive Publication Date: 2025-08-12ZHEJIANG THIRDNET TECH

Patent Information

Application Number
CN202510722298.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-12
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In the actual application scenarios of enterprise knowledge questions and answers in multiple departments, different authority levels and inconsistent professional terms, the existing enterprise knowledge question and answer system cannot achieve accurate, compliant, and consistent contextual semantic question and answer, and there are semantic drift and security risks.

Method used

The enterprise knowledge question and answer system based on large language models is adopted, including heterogeneous semantic embedding module, permission mask calculation module, enterprise term alignment module and candidate knowledge query module. Combined with the RAG search enhancement mechanism, multimodal unified embedding, semantic-level permission control and term graph alignment are realized to generate accurate and compliant question and answer.

Benefits of technology

Improve the accuracy and security of the Q&A system, avoid semantic drift and overright leaks, and support cross-format and multi-department enterprise intelligent Q&A.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470096A_ABST
    Figure CN120470096A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large model knowledge questions and answers, and discloses an enterprise knowledge question and answer system based on a large language model, which comprises a heterogeneous semantic embedding module, a permission mask calculation module, an enterprise term alignment module, a candidate knowledge query module and a question and answer generation output module, in the prior art, an enterprise question and answer mode is realized only depending on keyword matching or a fixed FAQ rule, and particularly under the conditions of enterprise internal knowledge distribution fragmentation, authority differentiation and term ambiguity, accurate, compliant and context-consistent semantic question and answer cannot be realized. Due to the fact that multi-mode unified embedding, the semantic level authority control mechanism, the term graph alignment model and the RAG enhanced question and answer mechanism are introduced, cross-format and multi-department enterprise intelligent question and answer is achieved, the problems of semantic drift and unauthorized leakage are avoided, and the accuracy and safety of the question and answer system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large-model knowledge question answering, and in particular to an enterprise knowledge question answering system based on a large language model. Background Art

[0002] Currently, internal enterprise knowledge question-and-answer systems often rely on keyword matching, pre-set FAQ templates, or simple rule-based retrieval algorithms. These systems are unable to handle semantically complex, context-sensitive questions, particularly in real-world enterprise scenarios involving multiple departments, varying levels of authority, and inconsistent terminology. For example, when an employee asks about the payment approval process for last quarter's sales contracts, existing systems are unable to effectively understand the timeframe corresponding to "last quarter," the process nodes involved in "payment approval," or whether the current user has access to such information. This results in answers that are either overly general or exceed user permissions, posing semantic drift or security risks. Furthermore, as enterprise knowledge management becomes increasingly complex, internal data sources are increasingly diverse, including PowerPoint training materials, PDF technical specifications, Word documents, and structured data in databases (such as invoices and approval records). Existing technologies generally lack unified semantic representation and retrieval capabilities for heterogeneous, multi-format documents, making it difficult to support high-quality, multi-scenario, and compliant question-and-answer services. Therefore, there is an urgent need for an enterprise knowledge question-and-answer system that can achieve semantic accuracy, authority compliance, traceability and credibility even when enterprise knowledge is highly fragmented, authority hierarchy is complex, and terminology is inconsistent, so as to improve the level of intelligent office and knowledge utilization efficiency of enterprises. Summary of the Invention

[0003] In response to the above-mentioned technical deficiencies, the purpose of the present invention is to propose an enterprise knowledge question-answering system and method based on a large language model, aiming to solve the technical problem that the enterprise question-answering method in the existing technology only relies on keyword matching or fixed FAQ rules, especially under the conditions of fragmented knowledge distribution, differentiated authority and ambiguous terminology within the enterprise, and cannot achieve accurate, compliant and context-consistent semantic question-answering.

[0004] To solve the above technical problems, the present invention adopts the following technical solutions: The present invention provides an enterprise knowledge question answering system based on a large language model.

[0005] The enterprise knowledge question answering system based on the large language model includes:

[0006] The heterogeneous semantic embedding module is used to obtain heterogeneous knowledge data sources within the enterprise, including unstructured document data sources and structured bill data sources. It performs OCR recognition, structural analysis, and paragraph segmentation on the heterogeneous knowledge data sources to extract information triples. The information triples are input into a preset multimodal embedding function to obtain a unified knowledge semantic vector.

[0007] Permission mask calculation module, used to obtain the current user role identifier , generate the corresponding permission mask vector , the unified knowledge semantic vector is passed through the permission mask vector Perform semantic permission filtering to obtain permission constraint vector; introduce semantic subspace projection matrix Perform vector transformation on the permission constraint vector to obtain a low-dimensional semantic vector ;

[0008] Enterprise term alignment module, used to obtain the current question text vector , for the question text vector Extract professional terms from the source code and construct a term phrase set , the term phrase set Mapping to the preset enterprise terminology ontology , for the term phrase set Each term node in the graph neural network performs a graph neural network embedding calculation to obtain a term semantic vector; the term semantic vector is combined with the question text vector Fusion to build query semantic vector ;

[0009] Candidate knowledge query module is used to query the candidate knowledge based on the low-dimensional semantic vector and query semantic vector The cosine similarity method is used to calculate the similarity score, and K candidate knowledge vectors are selected from high to low scores;

[0010] Question and answer generation output module, used to generate answers based on the RAG retrieval enhancement mechanism combined with the language model DeepSeek And output.

[0011] Preferably, in the heterogeneous semantic embedding module, the unstructured document data source includes slide format documents, Word text documents, PDF text documents and XML text documents; the structured bill data source includes invoice information and bill information; the information triples include text content, image visual embedding features and a structural hierarchy consisting of page numbers and paragraph numbers.

[0012] Preferably, in the permission mask calculation module, the permission mask vector Used to control the dimension-level access rights corresponding to the current user role identifier; low-dimensional semantic vectors are used for content retrieval; semantic subspace projection matrix Used to compress enterprise knowledge content into the current user role identifier in the vector space The corresponding semantic area has access permissions.

[0013] Preferably, in the enterprise term alignment module, the term semantic vector is aligned with the question text vector Fusion to build query semantic vector The steps are as follows: , where n is the total number of term nodes; is the influence weight of the j-th term node, which is used to indicate the semantic contribution of the j-th term node in the current question; is the jth item in the term semantic vector.

[0014] Preferably, in the candidate knowledge query module, the K candidate knowledge vectors are all authorized semantic content within the current user's authority, and carry their original document path, structural location information and authority level identification for subsequent question and answer generation and traceability output.

[0015] Preferably, in the question-answer generation output module, the answer is generated based on the RAG retrieval enhancement mechanism combined with the language model DeepSeek The steps of output include: splicing K candidate knowledge vectors into the prompt template input Prompt Input the prompt template into the DeepSeek language model through the preset API interface and receive the generated answer. And output.

[0016] Preferably, in the question-answer generation output module, ,in, It is a text field in the Json format specified in the preset API interface, used to specify the text range; is the Kth candidate knowledge vector; It is the question field in the Json format specified in the preset API interface, which is used to specify the question content; q is the current question text vector.

[0017] The present invention also provides an enterprise knowledge question answering method based on a large language model, comprising:

[0018] Step S10: Obtain heterogeneous knowledge data sources within the enterprise, including unstructured document data sources and structured bill data sources, perform OCR recognition processing, structural analysis processing, and paragraph segmentation processing on the heterogeneous knowledge data sources in sequence to extract information triples; input the information triples into a preset multimodal embedding function to obtain a unified knowledge semantic vector;

[0019] Step S20: Get the current user role ID , generate the corresponding permission mask vector , the unified knowledge semantic vector is passed through the permission mask vector Perform semantic permission filtering to obtain permission constraint vector; introduce semantic subspace projection matrix Perform vector transformation on the permission constraint vector to obtain a low-dimensional semantic vector ;

[0020] Step S30: Get the current question text vector , for the question text vector Extract professional terms from the source code and construct a term phrase set , the term phrase set Mapping to the preset enterprise terminology ontology , for the term phrase set Each term node in the graph neural network performs a graph neural network embedding calculation to obtain a term semantic vector; the term semantic vector is combined with the question text vector Fusion to build query semantic vector ;

[0021] Step S40: Based on the low-dimensional semantic vector and query semantic vector The cosine similarity method is used to calculate the similarity score, and K candidate knowledge vectors are selected from high to low scores;

[0022] Step S50: Generate answers based on the RAG retrieval enhancement mechanism combined with the language model DeepSeek And output.

[0023] The present invention also provides a computer program product, including an enterprise knowledge question answering program based on a large language model, which implements the enterprise knowledge question answering method based on a large language model when executed by a processor.

[0024] The beneficial effect of the present invention is that compared with the enterprise question-answering method in the prior art that only relies on keyword matching or fixed FAQ rules, especially under the conditions of fragmented knowledge distribution, differentiated permissions and ambiguous terminology within the enterprise, it is impossible to achieve accurate, compliant and context-consistent semantic question-answering. Since this application introduces multimodal unified embedding, semantic-level permission control mechanism, term graph alignment model and RAG enhanced question-answering mechanism, it realizes cross-format and multi-department enterprise intelligent question-answering, thereby avoiding the problems of semantic drift and unauthorized leakage, and improving the accuracy and security of the question-answering system. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1 This is a system diagram of the first embodiment of an enterprise knowledge question-answering system based on a large language model according to the present invention.

[0027] Figure 2 This is a schematic diagram of the equipment of an enterprise knowledge question-answering system based on a large language model according to the present invention. DETAILED DESCRIPTION

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0029] Example 1: Figure 1 1 is a flow chart of the first embodiment of the enterprise knowledge question answering system based on a large language model of the present invention, which proposes the first embodiment of the enterprise knowledge question answering system based on a large language model of the present invention.

[0030] In a first embodiment, the enterprise knowledge question answering system based on a large language model includes:

[0031] The heterogeneous semantic embedding module is used to obtain heterogeneous knowledge data sources within the enterprise, including unstructured document data sources and structured bill data sources. It performs OCR recognition, structural analysis, and paragraph segmentation on the heterogeneous knowledge data sources to extract information triples. The information triples are input into a preset multimodal embedding function to obtain a unified knowledge semantic vector.

[0032] It should be noted that in the heterogeneous semantic embedding module, the unstructured document data sources include slide format documents, Word text documents, PDF text documents and XML text documents; the structured bill data sources include invoice information and bill information; the information triples include text content, image visual embedding features and a structural hierarchy consisting of page numbers and paragraph numbers. OCR recognition processing includes using Tesseract, PaddleOCR or Baidu's general text recognition model to perform optical character recognition on image-based documents, and output the text content and its position information on the image; structural parsing processing includes using a model based on document layout analysis such as LayoutLMv3 to parse the structural relationship between paragraphs, tables, titles and text, and obtain structured document objects; paragraph segmentation processing includes segmenting according to text block boundaries, title levels and semantic pauses such as periods and paragraph markers to obtain logically coherent minimum semantic units for subsequent embedding modeling.

[0033] It is understandable that the goal of the heterogeneous semantic embedding module is to uniformly convert enterprise knowledge fragments in multi-source heterogeneous formats (text, images, structured forms) into vectorized semantic representations that can be used by large language models; the "multimodal embedding function" in this module is based on the combination of the large language model DeepSeek and the multimodal pre-training model BERT-LayoutLM fusion model, supporting joint modeling of text, images and structural levels to generate a unified knowledge semantic vector.

[0034] For example, for a company's "Sales Contract Management Process Manual.pdf", the system can use OCR to identify keywords such as "payment node" and "approval role" in the scanned page, and then extract the chapter "Chapter 3 Approval Process" to which it belongs through structural analysis. Finally, it generates an embedding vector based on "payment node: Chapter 3 content segment" for subsequent knowledge retrieval.

[0035] Permission mask calculation module, used to obtain the current user role identifier , generate the corresponding permission mask vector , the unified knowledge semantic vector is passed through the permission mask vector Perform semantic permission filtering to obtain permission constraint vector; introduce semantic subspace projection matrix Perform vector transformation on the permission constraint vector to obtain a low-dimensional semantic vector ;

[0036] It should be noted that in the permission mask calculation module, the permission mask vector Used to control the dimension-level access rights corresponding to the current user role identifier; the unified knowledge semantic vector is passed through the permission mask vector Perform semantic permission filtering, which includes element-by-element product operations; low-dimensional semantic vectors are used for content retrieval; semantic subspace projection matrix Used to compress enterprise knowledge content into the current user role identifier in the vector space The corresponding semantic area has access permissions.

[0037] It is understandable that the permission mask calculation module not only achieves fine-grained control of user permissions at the vector level, but also provides a computational basis for high-dimensional compression to a low-dimensional semantic subspace for the subsequent knowledge retrieval stage, significantly reducing the security risks caused by permission leakage. The design of the semantic subspace projection matrix can be based on the local structure of the knowledge graph or the permission clustering results, for example, by constructing a low-dimensional projection matrix through Laplace eigendecomposition or principal component extraction based on permission roles.

[0038] For example, if a user is a "financial specialist", their mask vector will block category dimensions such as personnel and management; their semantic projection matrix will only cover financial-related semantic areas such as contract approval and invoice circulation, thereby avoiding unauthorized information when generating questions and answers.

[0039] Enterprise term alignment module, used to obtain the current question text vector , for the question text vector Extract professional terms from the source code and construct a term phrase set , the term phrase set Mapping to the preset enterprise terminology ontology , for the term phrase set Each term node in the graph neural network performs a graph neural network embedding calculation to obtain a term semantic vector; the term semantic vector is combined with the question text vector Fusion to build query semantic vector ;

[0040] It should be noted that in the enterprise term alignment module, the term semantic vector is aligned with the question text vector Fusion to build query semantic vector The steps are as follows: , where n is the total number of term nodes; is the influence weight of the j-th term node, which is used to indicate the semantic contribution of the j-th term node in the current question; is the jth item in the term semantic vector.

[0041] It is understandable that the enterprise term alignment module is used to solve problems such as the different meanings of the same term in different departments or the diverse expressions of synonymous terms, thereby improving the accuracy of the question-answering system's understanding of the internal semantics of the enterprise; the term ontology graph can be pre-constructed by the enterprise knowledge engineering team and stored in standard formats such as RDF / OWL; the parameter training of the graph neural network can be weakly supervised and optimized based on historical question-answering records or term co-occurrence relationships between departmental documents.

[0042] For example, if a user asks, "When will this product be launched?", if their department is "R&D," the term "launch" might correspond to a code deployment process node; for a "Marketing" department, it might refer to the time of the launch conference. The Enterprise Term Alignment module automatically maps the term "launch" to a semantically correct node in the ontology graph based on the question context and user role, ensuring accuracy and consistency between vector retrieval and question-answer generation.

[0043] Candidate knowledge query module is used to query the candidate knowledge based on the low-dimensional semantic vector and query semantic vector The cosine similarity method is used to calculate the similarity score, and K candidate knowledge vectors are selected from high to low scores;

[0044] It should be noted that in the candidate knowledge query module, the K candidate knowledge vectors are all authorized semantic content within the current user's authority scope, and carry their original document path, structural location information and permission level identification for subsequent question and answer generation and traceability output.

[0045] It can be understood that the original document path and structural location information carried in the candidate knowledge vector can be used for source annotation and traceability control in the subsequent answer generation process to ensure that the question and answer results are explainable, auditable, and accountable; at the same time, the permission level identification will be used in the question and answer generation output module to further enhance permission filtering.

[0046] It should be understood that this module is achieved through the fusion of "semantic space retrieval and authority constraint vector" (low-dimensional semantic vector With permission constraints), it realizes a high-precision knowledge recall mechanism that complies with enterprise security regulations; without the need to access the complete original document, it can effectively support document traceability and structure positioning.

[0047] Question and answer generation output module, used to generate answers based on the RAG retrieval enhancement mechanism combined with the language model DeepSeek And output.

[0048] It should be noted that in the question-answer generation output module, the answer is generated based on the RAG retrieval enhancement mechanism combined with the language model DeepSeek The steps of output include: splicing K candidate knowledge vectors into the prompt template input Prompt Input the prompt template into the DeepSeek language model through the preset API interface and receive the generated answer. And output. In the question and answer generation output module, ,in, It is a text field in the Json format specified in the preset API interface, used to specify the text range of the RAG search enhancement mechanism; is the Kth candidate knowledge vector; It is the question field in the Json format specified in the preset API interface, which is used to specify the question content; q is the current question text vector.

[0049] It's understandable that the introduction of the RAG mechanism not only improves the language model's ability to respond to contextual knowledge but also significantly reduces the probability of "hallucination" questions. Compared to traditional retrieval-based or purely generative question answering, this solution is more practical and stable in complex enterprise knowledge scenarios. The advantage of the RAG mechanism lies in its combination of "query relevance" and "language generation capabilities": the retrieval part provides knowledge support, while the generation part enhances expression capabilities, achieving higher-quality, more targeted, and more contextually appropriate intelligent question answering.

[0050] For example, when a user asks, "Did the sales department's office supplies reimbursement exceed the budget in 2023?", the system first retrieves the budget table and reimbursement list paragraphs related to "sales department", "office supplies", and "2023", and inputs them into the DeepSeek model as prompts. The model output will include content such as "The sales department reimbursed a total of 35,000 yuan for office supplies, exceeding the budget limit of 30,000 yuan. It is recommended to submit an approval note", and mark "Information source: / Department Budget / 2023 / Q4 Budget Report.pdf Page 3, Paragraph 5" below to achieve knowledge-verifiable semantic output.

[0051] Embodiment 2: In addition, the enterprise knowledge question-answering method based on a large language model provided by the present invention adopts an enterprise knowledge question-answering system based on a large language model in the above embodiment, which can solve the technical problem of enterprise knowledge question-answering based on a large language model. Compared with the prior art, the beneficial effects of the enterprise knowledge question-answering method based on a large language model provided by the present invention are the same as the beneficial effects of the enterprise knowledge question-answering system based on a large language model provided by the above embodiment, and the other technical features of the enterprise knowledge question-answering method based on a large language model are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0052] Example 3: The present invention provides an enterprise knowledge question answering device based on a large language model, please refer to Figure 2 A large language model-based enterprise knowledge question-answering device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the large language model-based enterprise knowledge question-answering method of the first embodiment described above. An enterprise knowledge question-answering device based on a large language model in an embodiment of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. An enterprise knowledge question-answering device based on a large language model is merely an example and should not limit the functionality and scope of use of the embodiments of the present invention. A large language model-based enterprise knowledge question-answering device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the large language model-based enterprise knowledge question-answering device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 can allow a large language model-based enterprise knowledge question-and-answer device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a large language model-based enterprise knowledge question-and-answer device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have alternatively.

[0053] Example 4: The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned enterprise knowledge question-answering method based on a large language model. The computer program product provided by the present invention can solve the technical problem of enterprise knowledge question-answering based on a large language model. Compared with the prior art, the beneficial effects of the computer program product provided by the present invention are the same as the beneficial effects of the enterprise knowledge question-answering method based on a large language model provided in the above-mentioned embodiment, and will not be repeated here.

[0054] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present invention are performed.

[0055] It should be understood that the various parts disclosed in the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.

[0056] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. An enterprise knowledge question answering system based on a large language model, characterized by: The system includes: The heterogeneous semantic embedding module is used to obtain heterogeneous knowledge data sources within the enterprise, including unstructured document data sources and structured bill data sources. It performs OCR recognition, structural analysis, and paragraph segmentation on the heterogeneous knowledge data sources to extract information triples. The information triples are input into a preset multimodal embedding function to obtain a unified knowledge semantic vector. Permission mask calculation module, used to obtain the current user role identifier , generate the corresponding permission mask vector , the unified knowledge semantic vector is passed through the permission mask vector Perform semantic permission filtering to obtain permission constraint vector; introduce semantic subspace projection matrix Perform vector transformation on the permission constraint vector to obtain a low-dimensional semantic vector ; Enterprise term alignment module, used to obtain the current question text vector , for the question text vector Extract professional terms from the source code and construct a term phrase set , the term phrase set Mapping to the preset enterprise terminology ontology , for the term phrase set Each term node in the graph neural network performs a graph neural network embedding calculation to obtain a term semantic vector; the term semantic vector is combined with the question text vector Fusion to build query semantic vector ; Candidate knowledge query module is used to query the candidate knowledge based on the low-dimensional semantic vector and query semantic vector The cosine similarity method is used to calculate the similarity score, and K candidate knowledge vectors are selected from high to low scores; Question and answer generation output module, used to generate answers based on the RAG retrieval enhancement mechanism combined with the language model DeepSeek And output.

2. The enterprise knowledge question answering system based on a large language model according to claim 1, characterized in that: In the heterogeneous semantic embedding module, the unstructured document data sources include slide format documents, Word text documents, PDF text documents, and XML text documents; the structured bill data source includes invoice information and bill information; the information triples include text content, image visual embedding features, and a structural hierarchy consisting of page numbers and paragraph numbers.

3. The enterprise knowledge question answering system based on a large language model according to claim 1, characterized in that: In the permission mask calculation module, the permission mask vector Used to control the dimension-level access rights corresponding to the current user role identifier; low-dimensional semantic vectors are used for content retrieval; semantic subspace projection matrix Used to compress enterprise knowledge content into the current user role identifier in the vector space The corresponding semantic area has access permissions.

4. The enterprise knowledge question answering system based on a large language model according to claim 1, characterized in that: In the enterprise term alignment module, the term semantic vector is aligned with the question text vector Fusion to build query semantic vector The steps are as follows: , where n is the total number of term nodes; is the influence weight of the j-th term node, which is used to indicate the semantic contribution of the j-th term node in the current question; is the jth item in the term semantic vector.

5. The enterprise knowledge question answering system based on a large language model according to claim 1, characterized in that: In the candidate knowledge query module, the K candidate knowledge vectors are all authorized semantic content within the current user's authority scope, and carry their original document path, structural location information and permission level identification for subsequent question and answer generation and traceability output.

6. The enterprise knowledge question answering system based on a large language model according to claim 1, characterized in that: In the question-answer generation output module, the answer is generated based on the RAG retrieval enhancement mechanism combined with the language model DeepSeek The steps of output include: splicing K candidate knowledge vectors into the prompt template input Prompt Input the prompt template into the DeepSeek language model through the preset API interface and receive the generated answer. And output.

7. The enterprise knowledge question answering system based on a large language model according to claim 6, characterized in that: In the question-answer generation output module, prompt template input ,in, It is a text field in the Json format specified in the preset API interface, used to specify the text range; is the Kth candidate knowledge vector; It is the question field in the Json format specified in the preset API interface, which is used to specify the question content; q is the current question text vector.

8. A method for enterprise knowledge question answering based on a large language model, applied to an enterprise knowledge question answering system based on a large language model according to any one of claims 1 to 7, characterized in that: Methods include: Step S10: Obtain heterogeneous knowledge data sources within the enterprise, including unstructured document data sources and structured bill data sources, perform OCR recognition processing, structural analysis processing, and paragraph segmentation processing on the heterogeneous knowledge data sources in sequence to extract information triples; input the information triples into a preset multimodal embedding function to obtain a unified knowledge semantic vector; Step S20: Get the current user role ID , generate the corresponding permission mask vector , the unified knowledge semantic vector is passed through the permission mask vector Perform semantic permission filtering to obtain permission constraint vector; introduce semantic subspace projection matrix Perform vector transformation on the permission constraint vector to obtain a low-dimensional semantic vector ; Step S30: Get the current question text vector , for the question text vector Extract professional terms from the source code and construct a term phrase set , the term phrase set Mapping to the preset enterprise terminology ontology , for the term phrase set Each term node in the graph neural network performs a graph neural network embedding calculation to obtain a term semantic vector; the term semantic vector is combined with the question text vector Fusion to build query semantic vector ; Step S40: Based on the low-dimensional semantic vector and query semantic vector The cosine similarity method is used to calculate the similarity score, and K candidate knowledge vectors are selected from high to low scores; Step S50: Generate answers based on the RAG retrieval enhancement mechanism combined with the language model DeepSeek And output.

9. An enterprise knowledge question-answering device based on a large language model, characterized in that: The enterprise knowledge question and answer device based on a large language model includes: a memory, a processor, and an enterprise knowledge question and answer program based on a large language model stored in the memory and executable on the processor. When the enterprise knowledge question and answer program based on a large language model is executed by the processor, an enterprise knowledge question and answer system based on a large language model according to any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that The computer program product includes an enterprise knowledge question answering program based on a large language model, and when the enterprise knowledge question answering program based on a large language model is executed by a processor, an enterprise knowledge question answering system based on a large language model according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Enterprise-level knowledge management system based on large language model

    CN119476460A

  • Digital base fusion system and electronic equipment

    CN119760007A

Cited By

  • Enterprise process intelligent analysis system based on large language model

    CN121119659A

  • Cross-modal heterogeneous data retrieval method and system based on semantic information

    CN121167001A

  • A cross-modal heterogeneous data retrieval method and system based on semantic information

    CN121167001B

  • Method, device and equipment for retrieving and enhancing enterprise multi-source knowledge traceability and storage medium

    CN121256067A

  • Retrieval enhancement method based on domain term association mining and term closure expansion

    CN121352043A