Private AI question and answer method, system and device and medium
By building a local digital knowledge base and integrating an AI Q&A model, the data security and personalized needs of traditional cloud AI Q&A systems are solved, and an efficient Q&A system is realized, which improves data privacy protection and intelligent Q&A capabilities.
Patent Information
- Application Number
- CN202510167028.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional cloud AI Q&A systems have problems that cannot meet data security and personalized needs, and cannot be deeply customized locally, resulting in insufficient data privacy protection and efficient performance.
Collect file archive data sets for data cleaning and classification indexing, build a local digital knowledge base, select a suitable AI semantic big model architecture for annotation and enhancement, build a knowledge base AI Q&A model, and establish a search mechanism based on user access control rules, integrate the knowledge base search mechanism with the AI Q&A model, and generate a search-enhanced AI Q&A model.
It improves the intelligence and data access control of the Q&A system, outputs accurate Q&A results, and solves the shortcomings in data security and information retrieval efficiency.
Smart Images

Figure CN120277180A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to a privatized AI question-answering method, system, device and medium. Background Art
[0002] With the rapid development of artificial intelligence technology, traditional question-answering systems are facing many challenges, especially in terms of data security, personalized needs and high efficiency. Although traditional cloud-based AI question-answering systems can provide efficient question-answering services, due to the dependence of data storage and processing on external cloud servers, data privacy protection and security issues have arisen. In addition, traditional AI question-answering systems usually cannot be deeply customized according to the specific needs of enterprises or institutions, and cannot fully meet the requirements of different business scenarios. Summary of the Invention
[0003] The present disclosure provides a privatized AI question-answering method, system, device and medium for solving the technical problems existing in the prior art in terms of data security, intelligent question-answering and information retrieval efficiency.
[0004] In view of the above problems, the present disclosure provides a privatized AI question-answering method, system, device and medium.
[0005] In the first aspect of the present disclosure, a privatized AI question-answering method is provided, and the method includes: Collecting a file archive data set, performing data cleaning and classification indexing on the file archive data set to construct a local digital knowledge base; selecting an AI semantic large model architecture according to the application requirements of the privatized scenario; using the AI semantic large model architecture to perform annotation enhancement on the local digital knowledge base to construct a knowledge base AI question-answering model; obtaining user access control rules, performing permission marking on the local digital knowledge base based on the user access control rules, and establishing a knowledge base retrieval mechanism; integrating and fusing the knowledge base retrieval mechanism and the knowledge base AI question-answering model to generate a retrieval-enhanced AI question-answering model; obtaining request question information of a target user, and performing semantic question-answering retrieval on the request question information based on the retrieval-enhanced AI question-answering model to output a user request question-answering result.
[0006] In the second aspect of the present disclosure, a privatized AI question-answering system is provided, and the system includes: A knowledge base construction module for collecting file and archive data sets, cleaning and classifying and indexing the file and archive data sets to construct a local digital knowledge base; a large model architecture selection module for selecting an AI semantic large model architecture according to the application requirements of the privatization scenario; an annotation enhancement module for using the AI semantic large model architecture to enhance the annotation of the local digital knowledge base to construct a knowledge base AI Q&A model; a permission marking module for obtaining user access control rules and marking permissions for the local digital knowledge base based on the user access control rules to establish a knowledge base retrieval mechanism; an integration and fusion module for integrating and fusing the knowledge base retrieval mechanism and the knowledge base AI Q&A model to generate a retrieval-enhanced AI Q&A model; a semantic Q&A retrieval module for obtaining request question information of a target user and performing semantic Q&A retrieval on the request question information based on the retrieval-enhanced AI Q&A model to output a user request Q&A result.
[0007] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a memory for storing executable instructions; a processor for implementing a privatized AI Q&A method provided by the present disclosure when executing the executable instructions stored in the memory.
[0008] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor, implements a privatized AI Q&A method provided by the present disclosure.
[0009] One or more technical solutions provided in the present disclosure have at least the following technical effects or advantages: The present disclosure collects a file and archive data set, performs data cleaning and classification indexing on the file and archive data set, and constructs a local digital knowledge base; according to the application requirements of the privatization scenario, an AI semantic large model architecture is selected; the local digital knowledge base is enhanced by annotation using the AI semantic large model architecture to construct a knowledge base AI question-answering model; user access control rules are obtained, and based on the user access control rules, permission marking is performed on the local digital knowledge base to establish a knowledge base retrieval mechanism; the knowledge base retrieval mechanism and the knowledge base AI question-answering model are integrated and fused to generate a retrieval-enhanced AI question-answering model; request question information of a target user is obtained, and semantic question-answering retrieval is performed on the request question information based on the retrieval-enhanced AI question-answering model to output a user request question-answering result. The present disclosure solves the technical problems existing in the prior art in terms of data security, intelligent question-answering, and information retrieval efficiency. By collecting file and archive data and performing cleaning, classification, and indexing, a local digital knowledge base is constructed, a suitable AI semantic large model architecture is selected for annotation enhancement to construct a knowledge base AI question-answering model, a retrieval mechanism is established based on user access control rules, and it is integrated with the AI question-answering model to generate a retrieval-enhanced question-answering model. Finally, through this model, semantic question-answering retrieval is performed on user questions to output accurate answer results, achieving the technical effects of improving the intelligence of the question-answering system and data access control. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 Schematic flowchart of a privatized AI question-answering method provided by an embodiment of the present disclosure.
[0012] Figure 2 Schematic structural diagram of a privatized AI question-answering system provided by an embodiment of the present disclosure.
[0013] Figure 3 Schematic structural diagram of an exemplary electronic device of the present disclosure.
[0014] REFERENCE SIGNS Knowledge base construction module 11, large model architecture selection module 12, annotation enhancement module 13, permission marking module 14, integration and fusion module 15, semantic question-answering retrieval module 16, bus 300, receiver 301, processor 302, transmitter 303, memory 304, bus interface 305. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The present disclosure provides a privatized AI question-answering method, system, device, and medium, aiming to solve the technical problems of the existing technology in terms of data security, intelligent question-answering, and information retrieval efficiency. By collecting file and archive data, cleaning, classifying, and indexing it, a local digital knowledge base is constructed. A suitable AI semantic large model architecture is selected for annotation enhancement to build a knowledge base AI question-answering model. A retrieval mechanism is established based on user access control rules and integrated with the AI question-answering model to generate a retrieval-enhanced question-answering model. Finally, semantic question-answering retrieval is performed on user questions through this model to output accurate answer results, achieving the technical effects of improving the intelligence of the question-answering system and data access control.
[0016] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0017] It should be noted that any variations of the terms "include" and "have" are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0018] Embodiment 1, as Figure 1 shown, the present disclosure provides a privatized AI question-answering method, and the method includes: Step S100: Collect a file and archive data set, perform data cleaning and classification indexing on the file and archive data set, and construct a local digital knowledge base.
[0019] In the embodiments of the present disclosure, in the data collection stage, a file and archive data set is collected from various data sources such as various internal systems, document management platforms, databases, emails, reports, manuals, and meeting records. These data sets include structured data (such as tables, database records), semi-structured data (such as XML, JSON-formatted log files), and unstructured data (such as PDF documents, Word files, pictures, etc.).
[0020] Next, format specifications, duplicate removal, and correction are performed on different types of data to ensure data quality; finally, classification indexing is performed, and data is labeled according to the file content, and the data is organized into a structure that is easy to retrieve, thereby establishing an efficient digital knowledge base.
[0021] Further, in the method provided by the application embodiment, the construction of the local digital knowledge base further includes: Classify the file and archive dataset structurally to obtain structured archive data, semi-structured archive data, and unstructured archive data; construct a data structure type cleaning program set according to the structural type of the archive data; respectively perform data cleaning processing on the structured archive data, semi-structured archive data, and unstructured archive data based on the data structure type cleaning program set to obtain an available file and archive dataset; build an archive classification system, and perform tagged indexing on the available file and archive dataset based on the archive classification system to construct a local digital knowledge base.
[0022] In the embodiment of the present disclosure, first, the file and archive dataset is structurally classified, and the collected file and archive data is divided into structured data, semi-structured data, and unstructured data. Specifically, for structured data, such as data from relational databases (such as customer information, financial statements, etc.), this data already has clear fields and data types, and is directly extracted and sorted through SQL queries. For example, table data is extracted from the database through SQL statements and classified as structured data. For semi-structured data, such as XML and JSON format files, although this data has a certain structure, it is not as standardized as relational databases. For this type of data, an XML parser or JSON parser is used to parse and convert it, and its content is converted into a unified standard format for subsequent processing. For unstructured data, such as PDF documents, scanned images, and Word files, OCR (Optical Character Recognition) technology is adopted. For example, the Tesseract OCR tool is used to extract the text information in the scanned image or PDF and convert it into a machine-readable text format. Through this process, structured archive data, semi-structured archive data, and unstructured archive data are obtained.
[0023] Next, a data structure type cleaning program set is constructed according to the structural type of the archive data. This program set defines corresponding cleaning rules and methods according to the characteristics of different data types. For structured archive data, the cleaning program set will include operations such as removing duplicate data, repairing inconsistent values, and filling in missing data to ensure the integrity, accuracy, and consistency of the information in the data table; for semi-structured archive data, the cleaning program set will focus on formatting the data structure, removing invalid information, and standardizing tags to ensure that its content conforms to a unified specification after parsing; for unstructured archive data, the cleaning program set needs to use OCR technology to extract text from images or PDF documents, and perform operations such as denoising and correcting the text format to ensure that the data extracted from scanned files and images can be used correctly. Through this process, the construction of the data structure type cleaning program set is completed.
[0024] After the construction of the cleaning program set is completed, the cleaning program set cleans the structured archive data, semi-structured archive data, and unstructured archive data respectively according to the data structure type. In this process, corresponding technologies are adopted for each data type. The structured archive data removes duplicate values, fills in missing items, and ensures consistent data formats through database queries and script tools; the semi-structured archive data is normalized through parsers (such as XML parsers or JSON parsers) to convert the data into a unified standard format; for unstructured archive data, after the text is extracted through OCR technology, denoising processing is performed, and format and grammar problems are corrected to ensure the high accuracy and usability of the extracted data. Through this process, an available file archive data set is obtained.
[0025] Finally, a file classification system is built. Specifically, the file archive data is divided into different categories and sub-categories. First, the main categories of the archives are defined according to the characteristics such as the type, use, and attributes of the files. For example, the files are divided into major categories such as technical archives, administrative archives, and financial archives. Then, under each major category, the sub-categories are further refined according to the specific content, usage scenarios, and business requirements. For example, under the "technical archives" category, there can be "design archives", "construction archives", "equipment maintenance archives", etc. Next, according to the content and category of the data, it is tagged and classified for indexing. Through this classification system, all available file archive data will be efficiently indexed according to the tags, facilitating subsequent query and use. Finally, based on this tagged indexing system, a local digital knowledge base is constructed.
[0026] Furthermore, in the method provided by the application embodiment, the construction of the local digital knowledge base further includes: Performing element extraction on the file classification system to determine a file archive classification element set; classifying knowledge nodes of the available file archive data set based on the file archive classification element set to obtain a file archive knowledge graph; creating an archive knowledge tag library, and performing tag assignment coding on the file archive knowledge graph based on the archive knowledge tag library to obtain an archive knowledge tag coding set; performing tagged indexing on the available file archive data set based on the archive knowledge tag coding set to construct the local digital knowledge base.
[0027] In the embodiments of the present disclosure, first, element extraction is performed on the file classification system to extract the classification element set of the document files. This step uses text mining and natural language processing (NLP) techniques to first extract key classification elements from the metadata of the document files. These elements may include the type of the file (such as technical file, financial file, etc.), creation date, relevant department, keywords, etc. At the same time, the actual content of the file is analyzed to extract information such as the theme, keywords, and abstract of the document. Through this process, the finally determined classification element set contains all the key information that may affect file classification, such as the type of the file, creation time, responsible department, relevant theme, etc.
[0028] Next, based on the determined classification element set of the document files, knowledge node classification is performed on the available document file data set, and a document file knowledge graph is constructed. In this step, various classification elements of the file (such as document type, keywords, time information, etc.) are transformed into knowledge nodes, which represent the key features and elements in the file data. Through a graph database or a dedicated knowledge graph construction tool (such as Neo4j), the relationships between these nodes are further organized and visualized to form a document file knowledge graph. The knowledge graph reveals the internal relationships between different file elements through the relationships of nodes (such as "technical document", "financial statement") and edges (such as "belongs to", "created in").
[0029] On this basis, a file knowledge label library is created, and label assignment coding is performed on the document file knowledge graph based on this library to obtain a file knowledge label coding set. In this step, all knowledge nodes are assigned specific labels according to their characteristics and content. These labels are defined in the file knowledge label library and can clarify the key attributes of the document files. The labels may include "technical document", "financial statement", "XXXX year", etc., and corresponding labels are assigned to each knowledge node according to the actual needs of the file content. By assigning labels to each node in the graph, a complete file knowledge label coding set is formed. This set includes the label and coding information of all document files, providing a standardized identifier for subsequent indexing and retrieval.
[0030] Finally, based on the file knowledge label coding set, label-based indexing is performed on the available document file data set to build a local digital knowledge base. In this process, the file data is associated with the classification elements and the knowledge graph through labeling and coding, thereby forming a digital knowledge base that can support efficient query and accurate retrieval. The label-based indexing provides a standardized identifier for each file, enabling users to quickly locate and access relevant files through labels, greatly improving the efficiency of file management and data retrieval. Finally, a local digital knowledge base is built, containing all classified, labeled, and coded document files, ensuring that file information can be quickly stored, managed, and retrieved.
[0031] Step S200: Select an AI semantic large model architecture according to the application requirements of the privatization scenario.
[0032] In the embodiments of the present disclosure, in the application of the privatization scenario, the process of selecting an AI semantic large model architecture first needs to clarify the specific requirements of the application, such as data security, privacy protection, and the customization ability of the model. Since privatization deployment requires the model and data to run in a local environment and cannot rely on external cloud services, the selected AI semantic large model architecture needs to meet the requirements of data privacy protection and be able to run efficiently on a local server. According to the application requirements of the privatization scenario, different types of AI semantic large model architectures are selected according to specific requirements. For example, if the application scenario needs to process a large amount of text data and has high requirements for context understanding and information extraction, a model based on the BERT architecture is selected. BERT has strong context understanding ability and is suitable for tasks such as question answering systems, text classification, and information extraction, and can provide efficient semantic understanding after local deployment.
[0033] Step S300: Use the AI semantic large model architecture to enhance the annotation of the local digital knowledge base and construct a knowledge base AI question answering model.
[0034] In the embodiments of the present disclosure, in the process of using the AI semantic large model architecture to enhance the annotation of the local digital knowledge base, first, relevant question and answer data are obtained from the knowledge base through question and answer data collection, and natural language processing technology is used for semantic understanding to extract key information in the questions and answers. Then, a question and answer semantic tag library is constructed, and semantic tags such as question type, answer type, and domain tag are added to each data item for semantic analysis and annotation to form a target question and answer data set. Subsequently, the question and answer data set is expanded through data enhancement technology to increase the training data volume and diversity of the model. Finally, the enhanced question and answer data set is trained using the AI semantic large model to generate an initial AI question answering model. Finally, based on the generative question answering technology, the initial AI question answering model is optimized through interactive tuning to enable it to better understand and respond to user questions, and finally a knowledge base AI question answering model is obtained.
[0035] Furthermore, in the method provided by the application embodiments, the construction of the knowledge base AI question answering model further includes: Performing question and answer data collection and semantic analysis annotation based on the local digital knowledge base to obtain a target question and answer data set; performing data enhancement on the target question and answer data set to obtain a question and answer enhanced data set; using the AI semantic large model architecture to perform question answering model training on the question and answer enhanced data set to generate an initial AI question answering model; and performing interactive tuning on the initial AI question answering model based on the generative question answering technology to obtain the knowledge base AI question answering model.
[0036] In the embodiments of the present disclosure, first, question-and-answer data collection and semantic analysis annotation are performed based on a local digital knowledge base. Specifically, first, question-and-answer data collection is carried out on the information in the knowledge base. This step collects relevant question-and-answer pairs by extracting documents, entries, FAQs, etc. from the knowledge base to form a local knowledge base question-and-answer data set. Subsequently, natural language processing techniques (such as text analysis, entity recognition, etc.) are used to semantically understand these question-and-answer data, extract the core semantic information of the questions and answers therefrom, and form a question-and-answer semantic entity data set. This data set contains semantic entities in each question-and-answer pair, such as the intention of the question, the category of the answer, domain information, etc. By further constructing a question-and-answer semantic tag library (including question types, answer types, domain tags, and semantic relationship tags), semantic analysis annotation is performed on the semantic entity data set to obtain the target question-and-answer data set.
[0037] Next, data augmentation is performed on the target question-and-answer data set. The purpose of data augmentation is to expand the original data set by generating new data samples to enhance the training effect of the model. Specific methods include synonym replacement, word order adjustment, text generation, etc., aiming to increase the diversity of the data set and improve the robustness and generalization ability of the model. Through these methods, a question-and-answer augmented data set is finally obtained.
[0038] Then, an AI semantic large model architecture is used to train the question-and-answer augmented data set. In this step, a deep learning model, usually based on the Transformer architecture (such as BERT), is used to train the augmented data set to generate an initial AI question-and-answer model. This model can process and understand various user questions by learning the language features, context relationships, and semantic information in the question-and-answer data set, and initially realizes intelligent answering of user questions.
[0039] Finally, based on generative question-and-answer technology, interactive tuning is performed on the initial AI question-and-answer model. In this process, through interaction and feedback with users, the answer quality of the model is gradually adjusted and optimized. Generative question-and-answer technology continuously improves the performance of the model in actual applications by generating new question-and-answer pairs. Finally, after interactive tuning, a fully trained and optimized knowledge base AI question-and-answer model is generated.
[0040] Furthermore, in the method provided by the application embodiments, obtaining the target question-and-answer data set further includes: Collect Q&A data from the local digital knowledge base according to the Q&A data collection objective to obtain a local knowledge base Q&A data set; use natural language processing technology to perform semantic understanding on the local knowledge base Q&A data set to obtain a Q&A semantic entity data set; construct a Q&A semantic tag library, which includes question types, answer types, domain tags, and semantic relationship tags; perform semantic analysis and annotation on the Q&A semantic entity data set based on the Q&A semantic tag library to obtain the target Q&A data set.
[0041] In the embodiments of the present disclosure, first, Q&A data is collected from the local digital knowledge base according to the Q&A data collection objective. For example, in the field of power generation, this step can collect relevant Q&A data from sources such as power equipment maintenance manuals, internal documents of power companies, frequently asked questions (FAQs), and equipment operation guides. This data may include content such as operation procedures, fault diagnosis, and preventive measures of power equipment. Through database queries or API interfaces, etc., the relevant Q&A data is extracted into a local knowledge base Q&A data set, including typical Q&A pairs such as operation questions and fault troubleshooting.
[0042] Next, use natural language processing technology to perform semantic understanding on the local knowledge base Q&A data set. For example, in the field of power generation, natural language processing technologies such as word segmentation and named entity recognition (NER) are used to identify professional terms and power equipment names in the document. For example, terms such as "transformer failure" or "power load analysis" are identified in the Q&A pair, and the relationships between words in the question are analyzed through methods such as dependency syntactic analysis, such as judging the causal relationship between "power equipment with overload" and "transformer failure". Finally, a Q&A semantic entity data set is obtained, which contains key information such as power equipment, operation instructions, and fault handling procedures.
[0043] Then, construct a Q&A semantic tag library, which is used to further classify and annotate the Q&A semantic entity data in the field of power generation. The tag library includes question types (such as equipment fault diagnosis, operating parameter setting, maintenance operation, etc.), answer types (such as step lists, phrase answers, graphic explanations, etc.), domain tags (such as power equipment, grid maintenance, dispatching control, etc.), and semantic relationship tags (such as causal relationship, conditional relationship, etc.).
[0044] Finally, semantic analysis and annotation are performed on the Q&A semantic entity dataset based on the Q&A semantic tag library. In this step, the tags in the tag library are used to perform detailed annotation on the Q&A data. For example, for a question like "What should be done when an overload fault occurs in a power transformer", "transformer" is marked as the device type, "overload" is marked as the fault type, "solution method" is marked as the answer type, and "operation steps" will be further subdivided into step tags. Through this annotation, the target Q&A dataset is obtained.
[0045] Step S400: Obtain the user access control rules, perform permission marking on the local digital knowledge base based on the user access control rules, and establish a knowledge base retrieval mechanism.
[0046] In the embodiment of the present disclosure, during the process of constructing the knowledge base retrieval mechanism, first, a knowledge base usage role library is defined according to the privatization scenario requirements, and corresponding permissions are assigned to each role to form role access permission information. Then, based on this information, access content marking is performed on the archival knowledge tag encoding set to obtain the role access permission encoding content set, ensuring that different users can only access specific content. Subsequently, according to the user access control rules, access permission marking is performed on the local digital knowledge base to form the user access permission knowledge content set, and based on this content, the user's access permission shielding content set is determined. Finally, a retrieval encryption mechanism is introduced to perform permission-level encryption on the shielding content to complete the establishment of the knowledge base retrieval mechanism.
[0047] Furthermore, in the method provided by the application embodiment, the obtaining of the user access control rules further includes: Defining a knowledge base usage role library according to the privatization scenario application requirements; respectively assigning permissions to each usage role in the knowledge base usage role library to obtain usage role access permission information; performing access content marking on the archival knowledge tag encoding set based on the usage role access permission information to obtain the role access permission encoding content set; and determining the user access control rules according to the role access permission encoding content set.
[0048] In the embodiment of the present disclosure, according to the privatization scenario application requirements, a knowledge base usage role library is first defined. This step requires clarifying the roles and responsibilities of different users within the enterprise according to the actual business requirements and organizational structure. For example, in a power generation enterprise, it may include equipment maintenance personnel, operators, technical support personnel, and management personnel, etc. Each role has different work tasks and data access requirements, so it is necessary to determine the specific data access requirements of each role through interviews, job responsibility analysis, and work process analysis. Through this process, a knowledge base usage role library is finally obtained, including each role and its corresponding responsibilities and data requirements.
[0049] Next, permissions are assigned to the knowledge base for each role in the role library. At this stage, according to the scope of responsibilities of the roles, the data and information that each role can access are clearly divided. For example, equipment maintenance personnel can access data related to equipment status and faults, operators can access daily operation data, and technical support personnel can access the detailed maintenance history and technical parameters of the equipment. Through this method of permission assignment, the access permissions of each role will be clearly defined and divided, so as to ensure that each role can only access relevant information within its scope of responsibilities during the work process. Through the above process, the access permission information of the usage roles is obtained.
[0050] After the permission assignment is completed, based on the access permission information of the usage roles, the access content of the archival knowledge tag encoding set is marked. At this time, according to the defined role permissions, each archival data or document needs to be matched with the corresponding access permission tag. For example, a certain equipment failure report may only be visible to equipment maintenance personnel, so this report will be marked as "only visible to maintenance personnel"; while another power generation operation report may only be allowed to be viewed by management, so this report will be marked as "only visible to management". This marking process ensures that each piece of information in the knowledge base can be matched with the access requirements of specific roles. After the marking is completed, the role access permission encoding content set is obtained.
[0051] Finally, by obtaining the role access permission encoding content set, it is clear which users can access which information, avoiding the risk of information leakage or misuse, and determining the user access control rules.
[0052] Furthermore, in the method provided by the application embodiment, the establishment of the knowledge base retrieval mechanism further includes: Based on the user access control rules, access permissions are marked for the local digital knowledge base to obtain a user access permission knowledge content set; according to the user access permission knowledge content set, a user access permission shielding content set is determined; a retrieval encryption mechanism is introduced to perform hierarchical encryption of permissions on the user access permission shielding content set, and the knowledge base retrieval mechanism is established.
[0053] In the embodiment of the present disclosure, based on the user access control rules, first, access permissions are marked for the local digital knowledge base. This process involves dividing and marking the access permissions of each user according to their roles and authorization scopes. For example, engineers in a power generation enterprise may only be able to access power equipment operation data related to their work, while management can access global data. Finally, through the application of the access control rules, a user access permission knowledge content set is obtained, which marks the specific data content that each user can access.
[0054] Next, based on the user access privilege knowledge content set, determine the user access privilege masked content set. At this stage, according to the access privileges of each user, mask the data that they cannot access. For data outside the scope of the user's privileges, it is masked through an automated system to ensure that users cannot access sensitive or irrelevant information. For example, if an operator is only authorized to access the equipment data of a certain power station, the data of other power stations will be masked. Through this process, the user access privilege masked content set is obtained.
[0055] Finally, for the user access privilege masked content set, introduce the privilege hierarchical encryption technology. Specifically, use a combination of symmetric encryption algorithms (such as AES-256) and asymmetric encryption algorithms (such as RSA) to encrypt the masked content. The symmetric encryption algorithm is used to encrypt a large amount of content that needs to be protected, while the asymmetric encryption is used to protect the security of the encryption key itself. Through this encryption mechanism, even if sensitive data is illegally accessed or stolen, unauthorized users still cannot decrypt and read the data. For the content that users can access, provide the corresponding decryption key to ensure that legitimate users can smoothly read the required data. Through this encryption step, establish a knowledge base retrieval mechanism.
[0056] Step S500: Integrate and fuse the knowledge base retrieval mechanism and the knowledge base AI Q&A model to generate a retrieval enhanced AI Q&A model.
[0057] In the embodiment of the present disclosure, first integrate and fuse the knowledge base retrieval mechanism and the knowledge base AI Q&A model to generate an initial retrieval enhanced AI Q&A model. In this integration process, the retrieval mechanism filters out relevant knowledge content, and the AI Q&A model processes and answers it, thereby improving the accuracy and relevance of the Q&A. Then, through performance evaluation of the model, obtain the key parameters of its performance (such as accuracy rate, response time, etc.). Based on the evaluation results, further iteratively optimize the initial model to improve its effect in specific application scenarios. Finally, generate an optimized retrieval enhanced AI Q&A model.
[0058] Furthermore, in the method provided by the application embodiment, the generation of the retrieval enhanced AI Q&A model further includes: Integrate and fuse the knowledge base retrieval mechanism and the knowledge base AI Q&A model to obtain an initial retrieval enhanced Q&A model; perform performance evaluation on the initial retrieval enhanced Q&A model to obtain the Q&A model performance evaluation parameter information; based on the Q&A model performance evaluation parameter information, iteratively optimize the initial retrieval enhanced Q&A model to generate a retrieval enhanced AI Q&A model.
[0059] In the embodiments of the present disclosure, first, the knowledge base retrieval mechanism is integrated with the AI question answering model. The core objective of this step is to combine the advantages of both, provide relevant background information for the question answering model through the retrieval mechanism, and generate the final natural language answer through the question answering model. Specifically, the retrieval mechanism uses vector retrieval techniques (such as BERT-based vector search) to quickly screen the documents in the knowledge base and extract the content related to the user's question. The AI question answering model further reasons based on these retrieval results to generate accurate answers. Finally, the integrated system forms an initial retrieval-enhanced question answering model.
[0060] Next, the performance of the initial retrieval-enhanced question answering model is evaluated. In this link, accuracy is used as the evaluation metric. Accuracy quantifies the performance of the model by calculating the matching degree between the generated answer and the true answer. The specific operation of the evaluation step is to compare the model's answers to a set of preset test questions with the standard answers to obtain the accuracy information of the system. The obtained accuracy information is used as the performance evaluation parameter information of the question answering model.
[0061] Finally, the initial retrieval-enhanced question answering model is iteratively optimized according to the performance evaluation parameter information of the question answering model. This step adjusts the model according to the performance evaluation parameter information of the question answering model. For example, by increasing the amount of training data, adjusting the algorithm structure, etc., to further improve the performance of the model. Through this process, a retrieval-enhanced AI question answering model is finally generated.
[0062] Furthermore, in the method provided by the application embodiment, the generation of the retrieval-enhanced AI question answering model further includes: Performing an optimization strategy analysis on the performance evaluation parameter information of the question answering model to obtain question answering model optimization strategy information; iteratively optimizing the initial retrieval-enhanced question answering model based on the question answering model optimization strategy information to obtain the retrieval-enhanced AI question answering model.
[0063] In the embodiments of the present disclosure, in the process of performing an optimization strategy analysis on the performance evaluation parameter information of the question answering model, first, a detailed analysis is performed on the accuracy information in the performance evaluation parameter information of the question answering model to identify the deficiencies of the initial retrieval-enhanced question answering model in specific question types or tasks. For example, if the model has a low accuracy in knowledge answering in certain fields or specific types of questions, through comparative analysis, identify which question types, contexts, or answer patterns lead to deviations in the model's answers. At this time, the optimization strategy analysis mainly proposes specific improvement directions based on the evaluation results, such as increasing the training samples of a certain type of data or adjusting the algorithm weights of the model. Through this process, the question answering model optimization strategy information is determined.
[0064] Next, based on the Q&A model optimization strategy information, the initial retrieval-enhanced Q&A model is iteratively optimized. This optimization process includes adjusting the hyperparameters of the model, expanding the training set, and fine-tuning the model architecture, etc., to enhance its ability to handle different types of questions. Through repeated training and optimization, a retrieval-enhanced AI Q&A model with better performance and better meeting the actual needs is generated.
[0065] Step S600: Obtain the request question information of the target user, and perform semantic Q&A retrieval on the request question information based on the retrieval-enhanced AI Q&A model, and output the user request Q&A result.
[0066] In the embodiment of the present disclosure, first, the request question information of the target user is obtained. This process involves the specific questions raised by the user, which includes various forms of queries, including text or voice input. Next, based on the retrieval-enhanced AI Q&A model, semantic Q&A retrieval is performed on the request question information. This step uses the model to perform semantic understanding of the question, that is, through natural language processing techniques (such as word embedding, semantic parsing, etc.) to extract the deep semantics of the question, beyond the surface text matching, to understand the user's intention. After completing the semantic understanding, Q&A retrieval is performed according to the knowledge base in the retrieval-enhanced AI Q&A model to find the answer most relevant to the user's question. Finally, the user request Q&A result is output, that is, the most accurate answer obtained based on the semantic analysis and retrieval process.
[0067] In the embodiment of the present disclosure, in summary, the embodiment of the present disclosure has at least the following technical effects: The present disclosure collects a file and archive data set, performs data cleaning and classification indexing on the file and archive data set, and constructs a local digital knowledge base; according to the application requirements of the privatization scenario, selects an AI semantic large model architecture; uses the AI semantic large model architecture to perform annotation enhancement on the local digital knowledge base, and constructs a knowledge base AI question-and-answer model; obtains user access control rules, performs permission marking on the local digital knowledge base based on the user access control rules, and establishes a knowledge base retrieval mechanism; integrates and fuses the knowledge base retrieval mechanism and the knowledge base AI question-and-answer model to generate a retrieval-enhanced AI question-and-answer model; obtains the request question information of the target user, and performs semantic question-and-answer retrieval on the request question information based on the retrieval-enhanced AI question-and-answer model, and outputs the user request question-and-answer result. The present disclosure solves the technical problems existing in the prior art in terms of data security, intelligent question-and-answer, and information retrieval efficiency. By collecting file and archive data and performing cleaning, classification, and indexing, a local digital knowledge base is constructed, a suitable AI semantic large model architecture is selected for annotation enhancement, a knowledge base AI question-and-answer model is constructed, a retrieval mechanism is established based on user access control rules, and it is integrated with the AI question-and-answer model to generate a retrieval-enhanced question-and-answer model. Finally, through this model, semantic question-and-answer retrieval is performed on user questions, and accurate answer results are output, achieving the technical effects of improving the intelligence of the question-and-answer system and data access control.
[0068] Embodiment 2, based on the same inventive concept as the privatized AI question-and-answer method in the foregoing embodiment, as Figure 2 shown, the present disclosure provides a privatized AI question-and-answer system. The system in the embodiments of the present disclosure and the method embodiments are based on the same inventive concept. Among them, the system includes: A knowledge base construction module 11, which collects a file and archive data set, performs data cleaning and classification indexing on the file and archive data set, and constructs a local digital knowledge base; a large model architecture selection module 12, which selects an AI semantic large model architecture according to the application requirements of the privatization scenario; an annotation enhancement module 13, which uses the AI semantic large model architecture to perform annotation enhancement on the local digital knowledge base and constructs a knowledge base AI question-and-answer model; a permission marking module 14, which obtains user access control rules, performs permission marking on the local digital knowledge base based on the user access control rules, and establishes a knowledge base retrieval mechanism; an integration and fusion module 15, which integrates and fuses the knowledge base retrieval mechanism and the knowledge base AI question-and-answer model to generate a retrieval-enhanced AI question-and-answer model; a semantic question-and-answer retrieval module 16, which obtains the request question information of the target user, performs semantic question-and-answer retrieval on the request question information based on the retrieval-enhanced AI question-and-answer model, and outputs the user request question-and-answer result.
[0069] Furthermore, the system is also used to implement the following functions: Classify the file archive dataset structurally to obtain structured archive data, semi-structured archive data, and unstructured archive data; construct a data structure type cleaning program set according to the structural type of the archive data; respectively perform data cleaning processing on the structured archive data, semi-structured archive data, and unstructured archive data based on the data structure type cleaning program set to obtain an available file archive dataset; build an archive classification system, and perform labeled indexing on the available file archive dataset based on the archive classification system to construct a local digital knowledge base.
[0070] Furthermore, the system is also used to implement the following functions: Extract elements from the archive classification system to determine the file archive classification element set; classify knowledge nodes of the available file archive dataset based on the file archive classification element set to obtain a file archive knowledge graph; create an archive knowledge label library, and perform label assignment coding on the file archive knowledge graph based on the archive knowledge label library to obtain an archive knowledge label coding set; perform labeled indexing on the available file archive dataset based on the archive knowledge label coding set to construct the local digital knowledge base.
[0071] Furthermore, the system is also used to implement the following functions: Collect question-and-answer data and perform semantic analysis and annotation based on the local digital knowledge base to obtain a target question-and-answer dataset; perform data augmentation on the target question-and-answer dataset to obtain a question-and-answer augmented dataset; use the AI semantic large model architecture to train a question-and-answer model on the question-and-answer augmented dataset to generate an initial AI question-and-answer model; perform interactive optimization on the initial AI question-and-answer model based on generative question-and-answer technology to obtain the knowledge base AI question-and-answer model.
[0072] Furthermore, the system is also used to implement the following functions: Collect question-and-answer data from the local digital knowledge base according to the question-and-answer data collection target to obtain a local knowledge base question-and-answer dataset; use natural language processing technology to perform semantic understanding on the local knowledge base question-and-answer dataset to obtain a question-and-answer semantic entity dataset; construct a question-and-answer semantic label library, which includes question type, answer type, domain label, and semantic relationship label; perform semantic analysis and annotation on the question-and-answer semantic entity dataset based on the question-and-answer semantic label library to obtain the target question-and-answer dataset.
[0073] Furthermore, the system is also used to implement the following functions: According to the application requirements of the privatization scenario, define a knowledge base usage role library; assign permissions to each usage role in the knowledge base usage role library to obtain usage role access permission information; based on the usage role access permission information, mark the access content of the file knowledge tag encoding set to obtain a role access permission encoding content set; determine the user access control rules according to the role access permission encoding content set.
[0074] Further, the system is also used to implement the following functions: Based on the user access control rules, mark the access permissions of the local digital knowledge base to obtain a user access permission knowledge content set; determine a user access permission shielding content set according to the user access permission knowledge content set; introduce a retrieval encryption mechanism to perform permission-level encryption on the user access permission shielding content set, and establish the knowledge base retrieval mechanism.
[0075] Further, the system is also used to implement the following functions: Integrate and fuse the knowledge base retrieval mechanism and the knowledge base AI Q&A model to obtain an initial retrieval-enhanced Q&A model; perform performance evaluation on the initial retrieval-enhanced Q&A model to obtain Q&A model performance evaluation parameter information; iteratively optimize the initial retrieval-enhanced Q&A model based on the Q&A model performance evaluation parameter information to generate a retrieval-enhanced AI Q&A model.
[0076] Further, the system is also used to implement the following functions: Analyze the optimization strategy of the Q&A model performance evaluation parameter information to obtain Q&A model optimization strategy information; iteratively optimize the initial retrieval-enhanced Q&A model based on the Q&A model optimization strategy information to obtain the retrieval-enhanced AI Q&A model.
[0077] Embodiment 3. Based on the inventive concept of a privatized AI Q&A method in the foregoing embodiments, the present disclosure also provides an electronic device, including: at least one processor; a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of any one of the methods in the foregoing Embodiment 1.
[0078] Figure 3 This is a schematic structural diagram of an exemplary electronic device of the present disclosure. In Figure 3Among them, the bus architecture is represented by bus 300. Bus 300 may include any number of interconnected buses and bridges. Bus 300 connects various circuits of one or more processors represented by processor 302 and a memory represented by memory 304 together. Bus 300 may also connect various other circuits together, such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art and thus will not be further described herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 302 is responsible for managing bus 300 and general processing, while memory 304 may be used to store data used by processor 302 when performing operations.
[0079] Embodiment 4. Based on a privatized AI Q&A method in the foregoing embodiments and with the same inventive concept, the present disclosure also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed, it implements the steps of any one of the methods described in Embodiment 1 above.
[0080] It should be noted that the above sequence of the embodiments of the present disclosure is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. The processes depicted in the drawings do not necessarily require the specific order and continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0081] The above are only the preferred embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
[0082] This specification and the drawings are only exemplary descriptions of the present disclosure and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present disclosure. Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the present disclosure and its equivalent technologies, the present disclosure is intended to include these changes and modifications.
Claims
1. A privatized AI question-answering method, characterized in that, The method includes: Collecting a file archive data set, cleaning and classifying and indexing the file archive data set, and constructing a local digital knowledge base; Selecting an AI semantic large model architecture according to the application requirements of the privatization scenario; Using the AI semantic large model architecture to enhance the annotation of the local digital knowledge base and constructing a knowledge base AI question and answer model; Obtaining user access control rules, marking permissions for the local digital knowledge base based on the user access control rules, and establishing a knowledge base retrieval mechanism; Integrating and fusing the knowledge base retrieval mechanism and the knowledge base AI question and answer model to generate a retrieval-enhanced AI question and answer model; Obtaining the request question information of the target user, and performing semantic question and answer retrieval on the request question information based on the retrieval-enhanced AI question and answer model to output the user request question and answer result.
2. The privatized AI question-answering method according to claim 1, characterized in that, The construction of the local digital knowledge base includes: Structurally classifying the file archive data set to obtain structured archive data, semi-structured archive data, and unstructured archive data; Constructing a data structure type cleaning program set according to the structural type of the archive data; Based on the data structure type cleaning program set, respectively performing data cleaning processing on the structured archive data, semi-structured archive data, and unstructured archive data to obtain an available file archive data set; Building an archive classification system, and performing tagged indexing on the available file archive data set based on the archive classification system to construct a local digital knowledge base.
3. The privatized AI question-and-answer method according to claim 2, wherein, The construction of the local digital knowledge base includes: Extracting elements from the archive classification system to determine a file archive classification element set; Based on the file archive classification element set, performing knowledge node classification on the available file archive data set to obtain a file archive knowledge graph; Creating an archive knowledge tag library, and performing tag assignment encoding on the file archive knowledge graph based on the archive knowledge tag library to obtain an archive knowledge tag encoding set; Based on the archive knowledge tag encoding set, performing tagged indexing on the available file archive data set to construct the local digital knowledge base.
4. The privatized AI question-answering method according to claim 1, characterized in that, The construction of the knowledge base AI question and answer model includes: Performing question and answer data collection and semantic analysis annotation based on the local digital knowledge base to obtain a target question and answer data set; Performing data enhancement on the target question and answer data set to obtain a question and answer enhanced data set; Using the AI semantic large model architecture to train a question and answer model on the question and answer enhanced data set to generate an initial AI question and answer model; Performing interactive optimization on the initial AI question and answer model based on generative question and answer technology to obtain the knowledge base AI question and answer model.
5. The privatized AI Q&A method according to claim 4, wherein The obtaining of the target question and answer data set includes: Collecting question and answer data for the local digital knowledge base according to the question and answer data collection target to obtain a local knowledge base question and answer data set; Using natural language processing technology to perform semantic understanding on the local knowledge base question and answer data set to obtain a question and answer semantic entity data set; Constructing a question and answer semantic tag library, where the question and answer semantic tag library includes question types, answer types, domain tags, and semantic relationship tags; Semantically analyze and annotate the Q&A semantic entity dataset based on the Q&A semantic tag library to obtain the target Q&A dataset.
6. The privatized AI question and answer method according to claim 3, wherein, The obtaining of the user access control rules includes: Define a knowledge base usage role library according to the application requirements of the privatization scenario; Grant permissions to each usage role in the knowledge base usage role library respectively to obtain usage role access permission information; Mark the access content of the file knowledge tag encoding set based on the usage role access permission information to obtain a role access permission encoding content set; Determine the user access control rules according to the role access permission encoding content set.
7. The privatized AI Q&A method according to claim 1, characterized in that, The establishment of the knowledge base retrieval mechanism includes: Mark the access permissions of the local digital knowledge base based on the user access control rules to obtain a user access permission knowledge content set; Determine the user access permission shielding content set according to the user access permission knowledge content set; Introduce a retrieval encryption mechanism to perform permission-level encryption on the user access permission shielding content set to establish the knowledge base retrieval mechanism.
8. The privatized AI question-answering method according to claim 1, characterized in that, The generation of the retrieval enhanced AI Q&A model includes: Integrate and fuse the knowledge base retrieval mechanism and the knowledge base AI Q&A model to obtain an initial retrieval enhanced Q&A model; Evaluate the performance of the initial retrieval enhanced Q&A model to obtain Q&A model performance evaluation parameter information; Iteratively optimize the initial retrieval enhanced Q&A model based on the Q&A model performance evaluation parameter information to generate a retrieval enhanced AI Q&A model.
9. The privatized AI Q&A method according to claim 8, characterized in that, The generation of the retrieval enhanced AI Q&A model includes: Analyze the optimization strategy of the Q&A model performance evaluation parameter information to obtain Q&A model optimization strategy information; Iteratively optimize the initial retrieval enhanced Q&A model based on the Q&A model optimization strategy information to obtain the retrieval enhanced AI Q&A model.
10. A privatized AI Q&A system, characterized in that, For implementing a privatized AI Q&A method according to any one of claims 1-9, the system includes: A knowledge base construction module for collecting a file archive dataset, performing data cleaning and classification indexing on the file archive dataset, and constructing a local digital knowledge base; A large model architecture selection module for selecting an AI semantic large model architecture according to the application requirements of the privatization scenario; A annotation enhancement module for using the AI semantic large model architecture to perform annotation enhancement on the local digital knowledge base to construct a knowledge base AI Q&A model; A permission marking module for obtaining user access control rules and marking the permissions of the local digital knowledge base based on the user access control rules to establish a knowledge base retrieval mechanism; An integration and fusion module for integrating and fusing the knowledge base retrieval mechanism and the knowledge base AI Q&A model to generate a retrieval enhanced AI Q&A model; A semantic Q&A retrieval module for obtaining the request question information of the target user, performing semantic Q&A retrieval on the request question information based on the retrieval enhanced AI Q&A model, and outputting the user request Q&A result.
11. An electronic device, characterized in that, The electronic device includes: A memory for storing executable instructions; A processor, when executing the executable instructions stored in the memory, implements a privatized AI Q&A method according to any one of claims 1-9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements a privatized AI Q&A method according to any one of claims 1-9.
Citation Information
Cited By
Computer equipment fault monitoring system and method based on artificial intelligence
CN120508477A
Whole-process engineering consultation project information interaction method and system
CN120578747A
Enhanced retrieval generation optimization method for private domain knowledge base
CN121029802A
Information consultation method and system based on artificial intelligence
CN121031796A
Enterprise knowledge base construction method and device suitable for question and answer large model, equipment and medium
CN121072699A