LLM-based professional knowledge question and answer method and system
By building a professional knowledge question and answer system based on LLM, the problems of low knowledge acquisition efficiency, data dispersion and knowledge lag in the laboratory are solved, and fast and accurate knowledge retrieval and dynamic update are achieved, which enhances operational security and data security.
Patent Information
- Application Number
- CN202510222650.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
The existing question and answer system cannot access sensitive data inside the laboratory, resulting in low efficiency in knowledge acquisition; the laboratory data storage and management are scattered, resulting in knowledge silos; the traditional knowledge base update method cannot keep up with changes in experimental data in real time, resulting in knowledge lag.
Using professional knowledge questions and answer methods and systems based on LLM, we use privatized knowledge bases, localized LLM fine-tuning, realize multimodal interaction and device linkage, and carry out security and privacy protection.
It improves the speed and accuracy of knowledge retrieval, realizes dynamic updates of the knowledge base, enhances operational security, and ensures data security and compliance.
Smart Images

Figure CN120144710A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge question answering systems, and in particular to a method and system for question answering professional knowledge based on LLM. Background Art
[0002] In today's digital age, the acquisition and application of knowledge in professional fields are crucial to the efficient performance of work. Especially in laboratory scenarios, accurate and timely acquisition of professional knowledge and effective operational guidance directly affect the output of scientific research results and the safety of experiments.
[0003] At present, relevant technologies have many shortcomings in meeting laboratory needs, as follows:
[0004] 1) Limitations of general models. Existing question-and-answer systems represented by ChatGPT widely rely on cloud service architecture. This architecture makes it have obvious shortcomings when facing laboratory scenarios, and it is impossible to access sensitive data stored inside the laboratory, such as undisclosed compound formulas, detailed equipment operation logs, etc. These sensitive data are often the core part of laboratory scientific research results, containing a large amount of professional knowledge and key information, and play a decisive role in the advancement of experiments, the deepening of research, and the protection of results. However, due to the reliance of general models on cloud services and the limitations of security policies, laboratory staff cannot use such question-and-answer systems to quickly obtain the knowledge contained in these internal sensitive data, which seriously restricts its application efficiency in professional laboratory scenarios.
[0005] 2) Knowledge island problem. The storage and management methods of laboratory data are relatively scattered. On the one hand, a large amount of data is stored in different computer devices in the form of local files, lacking unified organization and indexing. Finding specific data is like looking for a needle in a haystack, which is time-consuming and laborious. On the other hand, data in independent devices, such as the storage systems that come with various experimental instruments, are difficult to integrate and interact with other data due to differences in device interfaces and data formats. In addition, a considerable amount of important data is recorded in paper documents, which not only makes it difficult to quickly retrieve and share data, but also easily leads to permanent loss of data due to damage or loss of paper documents. This scattered data situation has led to the formation of knowledge islands. The lack of a unified knowledge integration and retrieval mechanism has forced experimenters to spend a lot of time switching between different data sources when they need to obtain relevant knowledge, greatly reducing experimental efficiency and hindering the smooth progress of scientific research.
[0006] 3) Lack of real-time performance. The traditional knowledge base update method mainly relies on manual input. In a laboratory environment, experimental data is generated at all times, such as data collected by various real-time sensors, experimental results output by instrument equipment, etc. These data reflect the latest progress and real-time status of the experiment, and have extremely high timeliness and value. However, the update method of manual input cannot keep up with the speed of experimental data generation, resulting in the knowledge in the knowledge base often lagging behind the actual experimental situation. When experimental personnel operate or make decisions based on the knowledge base, they may make wrong judgments because they use outdated data and knowledge, which will affect the accuracy and reliability of the experiment, and may even lead to the failure of the experiment, causing waste of time and resources.
[0007] In response to the problems in the related technologies, no effective solution has been proposed yet. Summary of the Invention
[0008] In response to the problems in the related technologies, the present invention proposes a professional knowledge Q&A method and system based on LLM to overcome the above-mentioned technical problems existing in the existing related technologies.
[0009] The technical solution of the present invention is implemented as follows:
[0010] On the one hand of the present invention:
[0011] A professional knowledge Q&A method based on LLM, comprising the following steps:
[0012] Pre-build a private knowledge base, which includes collecting data from laboratory equipment, structured documents and unstructured data, integrating, structuring, and cleaning the collected data, converting it into a high-dimensional vector representation, using the approximate nearest neighbor algorithm, selecting one or more of the graph index HNSW, locality-sensitive hashing LSH, and product quantization PQ techniques to balance accuracy and performance, and establishing an index dynamic maintenance mechanism to build a laboratory knowledge graph and store it in the graph database Neo4j of the present invention;
[0013] Perform LLM local fine-tuning, which includes injecting laboratory domain corpora based on LLM, including equipment manuals and safety regulations, using the parameter-efficient fine-tuning technology PEFT, updating 0.5% of the model parameters through the LORA adapter, and establishing a real-time learning mechanism to automatically generate QA pairs from experimental logs for model training;
[0014] Perform multimodal interaction and device linkage, which includes taking text, voice, and images as inputs, outputting step-by-step operation guides and embedding safety warning icons, realizing AR visualization guidance through Hololens, and sending real-time control instructions to devices; among them, the output operation guides and AR visualization guidance content are associated with the input information, and the sent real-time control instructions are generated based on the Q&A results;
[0015] Implement security and privacy protection, including localizing the system deployment on the internal server of the laboratory, prohibiting external network access, encrypting and storing sensitive data using AES-256, using the TLS1.3 protocol for data transmission, and setting different access permissions for different roles based on the RBAC model.
[0016] Among them, for the construction of the privatized knowledge base in the steps, the data cleaning includes deduplication, noise filtering, and format standardization operations to ensure the accuracy and consistency of the knowledge base data; sparse coding or quantization compression technology is used to reduce the vector storage cost, enabling the knowledge base to support dynamic updates while improving storage efficiency.
[0017] Among them, for the LLM local fine-tuning in the steps, when automatically generating QA pairs from experiment logs, records are made for user questions and corresponding answers, which are organized into a standard format and then added to the training set to continuously optimize the model's ability to answer questions in the laboratory field.
[0018] Among them, for the multi-modal interaction and device linkage in the steps, when the image is input, for the image of the device failure interface taken, CV technology is used to identify the error codes therein, and corresponding fault handling guidance content is generated based on the recognition results; in the AR visual guidance, the device disassembly and assembly animations shown through Hololens are precisely matched with the actual operation steps of the device, providing intuitive and accurate operation guidance for experimenters.
[0019] Among them, for the security and privacy protection in the steps, when setting permissions based on the RBAC model, permissions are set according to the user identity. Among them, interns are only given the permission to query the basic operation guide, while the laboratory supervisor has higher permissions including querying, modifying, auditing, etc. The permission settings for different roles are clear and do not interfere with each other.
[0020] On the other hand, the present invention:
[0021] A professional knowledge Q&A system based on LLM, including:
[0022] Knowledge base engine module: used to dynamically fuse device real-time data and document knowledge, where the device real-time data is collected through the Modbus / OPCUA protocol; supports version backtracking and audit tracking, automatically records document revision records, can update the knowledge graph in real time according to data changes, and provides data support for generating trustworthy answers;
[0023] Device interaction interface module: Connects to laboratory devices through industrial protocols OPCUA and Modbus-TCP to achieve two-way communication for reading device status and issuing commands; when the device has an abnormal situation, automatically pushes the corresponding operation process and locks the device to ensure the safe operation of the device.
[0024] Trusted Answer Generation Module: Combine the output of the LLM with the knowledge graph verification to generate the answer content and label the answer source; Set the confidence threshold to 85%. When the confidence of the generated answer is lower than this threshold, trigger the manual review process and transfer it to the laboratory supervisor for review and processing.
[0025] In addition, when the knowledge base engine module fuses data, for the device exception log data, it automatically associates the device failure nodes in the knowledge graph, improves the content of the knowledge graph, and enhances the system's ability to answer device failure problems.
[0026] In addition, during the two-way communication process, the device interaction interface module verifies the validity of the issued control instructions to ensure that the instructions comply with the device operation specifications and safety requirements, preventing incorrect instructions from damaging the device.
[0027] In addition, when the trusted answer generation module labels the answer source, it clarifies the specific reference document name, chapter, and knowledge graph node information, enhancing the credibility and traceability of the answer.
[0028] Advantages of the present invention:
[0029] By constructing a professional knowledge Q&A method and system based on the LLM, the present invention effectively solves many problems in the application of the prior art in the laboratory scenario, and has achieved remarkable beneficial effects in improving knowledge acquisition efficiency, ensuring data security, enhancing operation safety, etc., providing strong support for the efficient operation of the laboratory and the smooth progress of scientific research work. Among them, by constructing a private knowledge base that integrates multi-source data and applying the advanced approximate nearest neighbor algorithm, the speed and accuracy of knowledge retrieval have been greatly improved. After entering the keyword, relevant information can be quickly located, and the provided content is detailed and accurate, greatly shortening the search time, enabling experimental personnel to devote more energy to the core experimental work, and significantly improving the overall experimental efficiency. At the same time, the private knowledge base of the system has a dynamic update mechanism, which can collect laboratory device data in real time and automatically integrate new experimental records, research results and other information. This enables the knowledge in the knowledge base to always be synchronized with the experimental progress, can meet the rapidly changing experimental needs, and effectively guarantees the accuracy and reliability of the experiment.
[0030] In addition, with the help of the multi-modal interaction and device linkage function, the system provides more intuitive and accurate operation guidance for experimenters, and at the same time has the ability of device anomaly monitoring and automatic intervention, significantly reducing the device misoperation rate. In addition, the present invention can deeply analyze the user's needs based on the user's historical questions and usage habits, and provide personalized knowledge recommendations and services for different users. And it adopts a local deployment method, running entirely on the laboratory internal server, prohibiting external network access, isolating potential external security threats at the physical level, effectively preventing the leakage of sensitive data, and strictly meeting the confidentiality specification requirements such as laboratory GLP / GMP, providing a solid guarantee for the data security and compliant operation of the laboratory. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0032] Figure 1 is a flowchart of a professional knowledge Q&A method based on LLM according to an embodiment of the present invention;
[0033] Figure 2 is a schematic block diagram of the principle of a professional knowledge Q&A system based on LLM according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present invention.
[0035] According to an embodiment of the present invention, a professional knowledge Q&A method based on LLM is provided.
[0036] As Figure 1 shown, the professional knowledge Q&A method based on LLM according to an embodiment of the present invention includes the following steps:
[0037] Build a privatized knowledge base in advance, specifically as follows:
[0038] Among them, for data collection, in a laboratory environment, for various types of equipment, such as centrifuges, temperature control boxes, mass spectrometers, etc., data collection channels are established through the Modbus / OPCUA protocol. Taking the temperature control box as an example, the system automatically collects its temperature data at regular intervals (such as every 5 seconds) to ensure obtaining the real-time operating status of the equipment. For structured documents, SOP operation manuals, MSDS chemical safety sheets, etc. are regularly (such as weekly) collected from the laboratory document management system and stored in the designated data storage area. For unstructured data, automated scripts are used to regularly scan the scientific research paper storage folder to obtain new PDF files; at the same time, with the help of special OCR devices, the handwritten notes submitted by experimenters are scanned and recognized at a fixed time every day, and the recognized text data is preliminarily sorted out.
[0039] Among them, for data processing and storage, the collected data is first subjected to data cleaning operations. Using a deduplication algorithm, duplicate data records are deleted; through noise filtering rules, data containing error flags or abnormal formats is removed; format standardization tools are used to uniformly convert data in different formats into a standard format recognizable by the system. When converting data into a high-dimensional vector representation, a feature extraction model based on Transformer is adopted, combined with sparse coding technology, to effectively reduce the vector storage cost while ensuring the integrity of data features. To improve the retrieval efficiency, the system selects the HNSW algorithm to construct a graph index according to the data scale and characteristics. In practical applications, if a batch of experimental equipment data is newly added, the system automatically adds the new data to the index structure through the dynamic maintenance mechanism of the index and adjusts the relevant index parameters to ensure the accuracy and efficiency of data retrieval. In terms of knowledge graph construction, using the Neo4j graph database, data from different sources is modeled according to entities and relationships. For example, equipment, experimental operations, chemicals, etc. are used as entities, and their associations (such as chemicals used by equipment, equipment involved in operations, etc.) are stored as relationships for subsequent Gremlin queries and dynamic updates.
[0040] For LLM local fine-tuning, among them, model selection and corpus injection, Llama2 is selected as the basic general LLM, and corpus in fields such as equipment manuals and safety regulations is extracted from the laboratory internal information management system. These corpora are preprocessed, including removing irrelevant format symbols, unifying term expressions, etc. The preprocessed corpora are divided according to a certain ratio (such as 80% for training and 20% for validation), and then through a special fine-tuning tool, these corpora are gradually injected into the Llama2 model.
[0041] Among them, for parameter fine-tuning and real-time learning, the LORA adapter in the PEFT technology is used to fine-tune the model parameters, and only 0.5% of the key model parameters are updated. During the actual fine-tuning process, by setting reasonable learning rates (such as 1e-4) and the number of training epochs (such as 5 epochs), the model is trained using the computing servers within the laboratory. At the same time, the system establishes a real-time learning mechanism. Whenever an experimenter asks a question and gets an answer in the system, the system automatically organizes the question content and the answer content into a QA pair and stores it in the training dataset in a specific format (such as {"question": "question content", "answer": "answer content"}). When the system is idle at night, it automatically uses the newly added QA pairs for incremental training of the model, continuously improving the model's ability to answer professional questions in the laboratory.
[0042] For multi-modal interaction and device linkage, the details are as follows:
[0043] Among them, for multi-modal input processing, when the experimenter asks a question in text form, the system directly receives the input content and performs preprocessing operations such as word segmentation and part-of-speech tagging using natural language processing technology to extract the key information of the question. If voice input is used, the experimenter asks questions through the professional recording equipment equipped in the laboratory, and the system uses ASR technology to convert the voice into text in real time and then performs subsequent processing. In terms of image input, the experimenter uses the dedicated high-definition camera in the laboratory to take pictures of the equipment failure interface, and the system uses CV technology to extract the features of the image and identify the error codes or key failure features therein. For example, when taking pictures of the failure indicator light interface of a PCR instrument, the system can accurately identify the color and blinking frequency of the indicator light and convert them into corresponding failure code information.
[0044] Among them, for multi-modal output presentation, for text questions, if the system gives a step-by-step operation guide, corresponding safety warning icons will be embedded in each step. For example, in the steps involving chemical operations, the chemical hazard level identification icon will be embedded. In terms of AR visual guidance, through the Hololens device, the experimenter can see the disassembly and assembly animations closely combined with the actual equipment. For example, when repairing complex instrument equipment, the experimenter wears the Hololens, and the device will superimpose the corresponding disassembly and assembly animations on the real equipment according to the current operation steps to guide the experimenter to operate correctly. In terms of real-time control instructions, when the system determines that equipment control is required, such as when the temperature in the temperature control box is too high, the system will automatically generate control instructions and send them to the temperature control box through the device interaction interface to adjust its temperature settings.
[0045] Security and privacy protection is carried out, including local deployment. A dedicated server cluster is built inside the laboratory, and all components of the question-answering system (including knowledge base, models, applications, etc.) are deployed on the cluster. Through network configuration, external networks are prohibited from accessing the server cluster, and only devices inside the laboratory are allowed to access through specific LAN IP segments. At the same time, firewall devices are deployed at the front end of the server cluster, and strict access rules are set to further ensure the security of the system.
[0046] Data encryption and permission control are performed. For sensitive data, such as compound formulas, the AES-256 encryption algorithm is used for encryption during storage, and the encrypted data is stored in a dedicated encrypted database. During data transmission, the TLS1.3 protocol is used to encrypt data transmission to ensure the security of data during network transmission. Permission control is performed based on the RBAC model. When the system is initialized, corresponding permissions are assigned to different roles (such as interns, laboratory technicians, laboratory supervisors, etc.). For example, interns can only query basic operating guides and cannot access sensitive experimental data and advanced operating instructions; laboratory technicians can query data and operating guides related to their own experiments, but cannot perform system management operations; laboratory supervisors have the highest permissions, including system management, data review, and access to advanced experimental data.
[0047] Conduct system operation and effect evaluation. Specifically, in terms of daily system operation, in daily laboratory work, experimenters can use the system for knowledge query and operation guidance at any time. When experimenters encounter equipment failure, they can ask questions to the system through multimodal interaction, and the system will quickly give answers and solutions. For example, the experimenter found that the mass spectrometer had an abnormal signal. He took a picture of the mass spectrometer fault interface and uploaded it to the system. After the system identified the fault code, it combined the knowledge base and LLM model to give possible causes of failure (such as ion source contamination) and detailed maintenance steps (such as how to disassemble and clean the ion source), and embedded safety warning information in the maintenance steps (such as wearing protective gloves and goggles during operation).
[0048] In terms of effect evaluation, the system effect is evaluated by regularly collecting feedback from experimenters and system operation data. In terms of efficiency improvement, the knowledge query time of experimenters before and after using this system is compared, and it is found that the query time is reduced by an average of 75% after using the system. In terms of operational safety, the number of equipment misoperations within 12 months is counted, and the equipment misoperation rate is reduced by 68% after using this system. At the same time, the results of the experimenter satisfaction survey on the system show that more than 90% of the experimenters believe that the system is of great help to their work and improves the accuracy and efficiency of the experiment.
[0049] According to an embodiment of the present invention, a professional knowledge question and answer system based on LLM is provided.
[0050] As shown Figure 2 in the figure, the LLM-based professional knowledge Q&A system according to an embodiment of the present invention includes:
[0051] Knowledge base engine module 1: used to dynamically fuse device real-time data and document knowledge, where the device real-time data is collected through the Modbus / OPCUA protocol; supports version backtracking and audit tracking, automatically records document revision records, can update the knowledge graph in real time according to data changes, and provides data support for generating credible answers;
[0052] Device interaction interface module 2: Connects to laboratory devices through industrial protocols OPCUA and Modbus-TCP to achieve two-way communication for reading device status and issuing instructions; when the device has an abnormal situation, automatically pushes the corresponding operation process and locks the device to ensure the safe operation of the device.
[0053] Credible answer generation module 3: Combines the LLM output and knowledge graph verification to generate answer content and marks the answer source; sets the confidence threshold to 85%, and when the confidence of the generated answer is lower than this threshold, triggers the manual review process and transfers it to the laboratory supervisor for review and processing.
[0054] In addition, when fusing data, the knowledge base engine module 1 automatically associates the device abnormal log data with the device failure nodes in the knowledge graph, improves the content of the knowledge graph, and enhances the system's ability to answer device failure problems.
[0055] In addition, during the two-way communication process, the device interaction interface module 2 validates the effectiveness of the issued control instructions to ensure that the instructions comply with the device operation specifications and safety requirements, and prevent incorrect instructions from damaging the device.
[0056] In addition, when marking the answer source, the credible answer generation module 3 specifies the specific reference document name, chapter, and knowledge graph node information to enhance the credibility and traceability of the answer.
[0057] In summary, by means of the above technical solutions of the present invention, the following effects can be achieved:
[0058] By constructing a professional knowledge Q&A method and system based on LLM, the present invention effectively solves many problems in the application of existing technologies in laboratory scenarios, and has achieved remarkable beneficial effects in improving knowledge acquisition efficiency, ensuring data security, enhancing operation safety, etc., providing strong support for the efficient operation of laboratories and the smooth progress of scientific research work. Among them, by constructing a private knowledge base that integrates multi-source data and applying an advanced approximate nearest neighbor algorithm, the speed and accuracy of knowledge retrieval have been greatly improved. After entering keywords, relevant information can be quickly located, and the provided content is detailed and accurate, greatly shortening the search time, enabling experimental personnel to devote more energy to core experimental work, and significantly improving the overall experimental efficiency. At the same time, the private knowledge base of the system has a dynamic update mechanism, which can collect laboratory equipment data in real time and automatically integrate new experimental records, research results and other information. This enables the knowledge in the knowledge base to always be synchronized with the experimental progress, can meet the rapidly changing experimental needs, and effectively guarantees the accuracy and reliability of the experiment.
[0059] In addition, with the help of the multi-modal interaction and device linkage function, the system provides more intuitive and accurate operation guidance for experimental personnel, and at the same time has the ability of device anomaly monitoring and automatic intervention, significantly reducing the device misoperation rate. The test data for 12 months shows that the device misoperation rate has been reduced by 68%. For example, during the operation of a centrifuge, if abnormal situations such as overspeed occur, the system will immediately issue an alarm, automatically push the emergency stop operation process, and lock the device at the same time to prevent safety accidents caused by the operator's failure to detect the anomaly in time or improper operation, effectively guaranteeing the personal safety of experimental personnel and the normal operation of the device.
[0060] Furthermore, the present invention can deeply analyze the user's needs based on the user's historical questions and usage habits, and provide personalized knowledge recommendations and services for different users. And it adopts a local deployment method, running entirely on the laboratory internal server, prohibiting external network access, isolating potential external security threats at the physical level, effectively preventing the leakage of sensitive data, and strictly meeting the confidentiality specification requirements such as laboratory GLP / GMP, providing a solid guarantee for the data security and compliant operation of the laboratory.
[0061] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will easily think of other implementation schemes of the present disclosure after considering the disclosure in the specification and embodiments. This application aims to cover any variations, uses or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0062] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A professional knowledge question-answering method based on LLM, characterized in that: The following steps are involved: Preliminary construction of a private knowledge base includes data collection from laboratory equipment, structured documents, and unstructured data, integration, structural processing, and data cleaning of the collected data, and conversion into high-dimensional vector representation. Approximate nearest neighbor algorithms are used to select one or more of the graph index HNSW, locality sensitive hashing LSH, and product quantization PQ technologies to balance accuracy and performance, and establish a dynamic index maintenance mechanism to build a laboratory knowledge graph stored in the graph database Neo4j. Perform LLM localization fine-tuning, which includes injecting laboratory domain corpus based on LLM, including equipment manuals and safety regulations, using parameter efficient fine-tuning technology PEFT, updating 0.5% model parameters through LORA adapter, establishing a real-time learning mechanism, and automatically generating QA pairs from experimental logs for model training; Conduct multimodal interaction and device linkage, including taking text, voice and images as input, outputting step-by-step operation instructions with embedded safety warning icons, implementing AR visual guidance through Hololens, and issuing real-time control instructions to the device; wherein the output operation instructions and AR visual guidance content are associated with the input information, and the issued real-time control instructions are generated based on the question and answer results; Security and privacy protection are carried out, including local deployment of the system on the laboratory's internal server, prohibiting external network access, using AES-256 encryption to store sensitive data, using TLS1.3 protocol for data transmission, and setting differentiated access rights for different roles based on the RBAC model.
2. The LLM-based professional knowledge question-answering method according to claim 1, characterized in that: The steps of constructing a privatized knowledge base and cleaning the data include deduplication, noise filtering, and format standardization operations to ensure the accuracy and consistency of the knowledge base data.
3. The LLM-based professional knowledge question-answering method according to claim 1, characterized in that: The LLM localization fine-tuning described in the step automatically generates QA pairs from the experimental log, records user questions and corresponding answers, organizes them into a standard format and adds them to the training set to continuously optimize the model's ability to answer laboratory field questions.
4. The LLM-based professional knowledge question-answering method and system according to claim 1, characterized in that: The multimodal interaction described in the step is linked with the device. When the image is input, the CV technology is used to identify the error code in the captured image of the device fault interface, and the corresponding fault handling guidance content is generated based on the recognition result; wherein, in the AR visual guidance, the device disassembly and assembly animation displayed by Hololens is accurately matched with the actual operation steps of the device, so as to provide intuitive and accurate operation guidance for the experimenter.
5. The LLM-based professional knowledge question-answering method according to claim 1, characterized in that: The security and privacy protection described in the step, when setting permissions based on the RBAC model, sets permissions based on user identity.
6. A professional knowledge question answering system based on LLM, used for the professional knowledge question answering method based on LLM according to any one of claims 1 to 5, characterized in that: include: Knowledge base engine module (1): used to dynamically integrate device real-time data and document knowledge, where the device real-time data is collected through Modbus / OPCUA protocol; It supports version backtracking and audit tracking, automatically records document revision records, can update the knowledge graph in real time according to data changes, and provide data support for the generation of trusted answers; Equipment interaction interface module (2): connects to laboratory equipment through industrial protocols OPC UA and Modbus-TCP to achieve two-way communication for equipment status reading and command issuance; When an abnormal situation occurs in the device, the corresponding operation process is automatically pushed and the device is locked to ensure the safe operation of the device. Trusted answer generation module (3): Combines LLM output with knowledge graph verification to generate answer content and annotates the source of the answer; sets the confidence threshold to 85%. When the confidence of the generated answer is lower than the threshold, the manual review process is triggered and the answer is transferred to the laboratory supervisor for review and processing.
7. The LLM-based professional knowledge question-answering system according to claim 6, characterized in that: When fusing data, the knowledge base engine module (1) automatically associates the device failure nodes in the knowledge graph with the device abnormality log data, so as to improve the content of the knowledge graph and enhance the system's ability to answer device failure questions.
8. The LLM-based professional knowledge question-answering system according to claim 6, characterized in that: The device interaction interface module (2) verifies the validity of the control instructions issued during the two-way communication process to ensure that the instructions comply with the device operation specifications and safety requirements, thereby preventing erroneous instructions from causing damage to the device.
9. The LLM-based professional knowledge question-answering system according to claim 6, characterized in that: The credible answer generation module (3) specifies the specific reference document name, chapter and knowledge graph node information when marking the source of the answer, so as to enhance the credibility and traceability of the answer.