Industrial internet equipment operation and maintenance knowledge question answering method, system, equipment and medium
By extracting and structuring the operation and maintenance documents of industrial Internet equipment, building an operation and maintenance knowledge base, and using a mixed search strategy of sparse and dense vectors, combining a large language model to generate intelligent question-and-answer questions, the efficiency and accuracy problems of traditional operation and maintenance methods in dealing with massive documents and data are solved, and efficient operation and maintenance knowledge questions and answers are achieved.
Patent Information
- Application Number
- CN202510267289.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
AI Technical Summary
When traditional operation and maintenance methods process massive documents and data from industrial Internet equipment, they lack intelligent data analysis and processing capabilities, and it is difficult to quickly and accurately extract key content, and lack a systematic knowledge integration and sharing mechanism, which makes it difficult to meet actual needs for operation and maintenance efficiency and accuracy.
The text data in the operation and maintenance document is extracted through the MinerU tool to generate structured text files; the langchain library is used for text chunking and vectorization processing to build an operation and maintenance knowledge base; a hybrid search strategy combining sparse vectors and dense vectors is adopted to retrieve target text data from the knowledge base, and a large language model is used to generate intelligent question-and-answer questions and answers.
It improves operation and maintenance efficiency, reduces the operation and maintenance costs of massive equipment, realizes efficient operation and maintenance of complex industrial Internet scenarios, and improves the response quality and user experience of the Q&A system.
Smart Images

Figure CN120069091A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent operation and maintenance, and more specifically, relates to a method, system, device and medium for answering questions about the operation and maintenance knowledge of industrial Internet devices. Background Art
[0002] At present, with the development of the industrial Internet, a large number of devices are connected to the Internet of Things and organically combined with the Internet to build an innovative system architecture with cloud-edge-end collaboration. The formation of this architecture has greatly promoted the deep integration of industrial systems and the Internet, effectively improving the intelligent level of industrial production and significantly enhancing the overall efficiency.
[0003] However, this rapid development has also brought a series of thorny problems. On the one hand, the devices in the industrial Internet are of various types and large in number, and have the distributed characteristics of spanning regions and territories. On the other hand, the wide popularization of virtualization, cloudification and containerization technologies has brought many conveniences to the industrial Internet, but also made the industrial Internet system architecture more complex.
[0004] These problems have directly led to a sharp increase in the complexity of the operation and maintenance work of industrial Internet devices. The operation and maintenance work involves a large amount of documents and data. Traditional operation and maintenance methods rely on simple document retrieval tools in terms of technical means and lack intelligent data analysis and processing capabilities, making it difficult to quickly and accurately extract key content from a large amount of information. From the perspective of the management mode, traditional operation and maintenance lacks a systematic knowledge integration and sharing mechanism, and the operation and maintenance information of each link is isolated from each other, unable to form an efficient synergy effect. In the face of large-scale, complex and dynamic industrial Internet scenarios, traditional operation and maintenance methods are difficult to meet the actual needs in terms of efficiency and accuracy. Summary of the Invention
[0005] In view of the above problems, the purpose of the present invention is to provide a method, system, device and medium for answering questions about the operation and maintenance knowledge of industrial Internet devices. By automatically retrieving relevant knowledge of operation and maintenance questions according to device operation and maintenance documents and historical knowledge bases, and using large models for automated intelligent answering, the operation and maintenance efficiency is improved, and the operation and maintenance costs of a large number of devices in the industrial Internet are reduced.
[0006] To achieve the above object, the present invention is realized through the following technical solutions: In a first aspect, an embodiment of the present application provides a method for answering questions about the operation and maintenance knowledge of industrial Internet devices, including: Obtain the operation and maintenance documents of industrial Internet devices, extract the text data in the operation and maintenance documents through the MinerU tool, and generate a structured text file; Based on the structured text file, use the langchain library to perform text chunking and vectorization processing to generate an operation and maintenance knowledge base; Obtain the user's query question. According to the query question, adopt a hybrid retrieval strategy combining sparse vectors and dense vectors to retrieve target text data from the operation and maintenance knowledge base; Generate prompt information based on the query question and the target text data, and input the prompt information into the large language model to generate an answer.
[0007] In an optional implementation manner, to obtain the operation and maintenance documents of industrial Internet devices, use the MinerU tool to extract the text data in the operation and maintenance documents to generate a structured text file, including: Collect the operation and maintenance documents of industrial Internet devices, where the operation and maintenance documents include text documents and non-text documents; Use the text extraction algorithm of the MinerU tool to extract the text data in the non-text document as the first text data; Use the picture extraction algorithm to extract the pictures in the non-text document, use the OCR recognition algorithm to recognize the text in the pictures, and convert it into the second text data; Use the document structure recognition algorithm to recognize the hierarchical structures in the text documents, the first text data, and the second text data, and mark them to generate a structured text file.
[0008] In an optional implementation manner, based on the structured text file, use the langchain library to perform text chunking and vectorization processing to generate an operation and maintenance knowledge base, including: Based on the structured text file, use the RecursiveCharacterTextSplitter splitting method of the langchain library to split the structured text into multiple text chunks; Use the BM25 algorithm to convert the text chunks into sparse vectors, and use the BAAI General Embedding model to map the text chunks into dense vectors; Store the sparse vectors, dense vectors, and the corresponding text chunks into the Milvus vector database, configure the index type and similarity measurement method, and generate an operation and maintenance knowledge base.
[0009] In an optional implementation manner, based on the structured text file, using the RecursiveCharacterTextSplitter splitting method of the langchain library to split the structured text into multiple text chunks, including: Based on the structured text file, use the RecursiveCharacterTextSplitter splitting method of the langchain library, with full stops, commas, and semicolons as delimiters, and split the structured text into multiple text chunks according to each text chunk including at most 256 characters.
[0010] In an optional embodiment, converting the text block into a sparse vector using the BM25 algorithm and mapping the text block into a dense vector using the BAAI General Embedding model includes: Converting the text block into a term-frequency-based sparse vector using the BM25 algorithm; Mapping the text block into a 768-dimensional semantic vector using the BAAI General Embedding model.
[0011] In an optional embodiment, storing the sparse vector and the dense vector, as well as the corresponding text block, into the Milvus vector database, configuring the index type and the similarity measurement method, and generating an operation and maintenance knowledge base includes: Storing the sparse vector and the dense vector, as well as the corresponding text block, into the Milvus vector database; setting the index type to IVF_FLAT and the similarity measurement method to inner product; Establishing a mapping relationship where the text block ID corresponds to the text block, the sparse vector, or the dense vector, and generating an operation and maintenance knowledge base.
[0012] In an optional embodiment, obtaining the user's query question, and according to the query question, retrieving target text data from the operation and maintenance knowledge base using a hybrid retrieval strategy that combines sparse vectors and dense vectors includes: Obtaining the user's query question, generating a sparse vector of the query question based on the BM25 algorithm, and retrieving the sparse vectors in the operation and maintenance knowledge base, and finding text data with keyword matching by calculating the similarity of the sparse vectors; Generating a dense vector of the query question based on the BAAI General Embedding model, and retrieving the dense vectors in the operation and maintenance knowledge base, and finding text data with semantic similarity by calculating the similarity of the dense vectors; Merging the text data with keyword matching and the text data with semantic similarity to generate process text data; For each piece of process text data, calculating a similarity score based on the similarity of the sparse vector and the similarity of the dense vector according to a preset similarity weight; Sorting the similarity scores, and screening out the target text data according to the sorting result.
[0013] In a second aspect, an industrial Internet device operation and maintenance knowledge Q&A system provided by an embodiment of the present application further includes: An operation and maintenance document preprocessing module, configured to obtain the operation and maintenance documents of industrial Internet devices, extract the text data in the operation and maintenance documents through the MinerU tool, and generate a structured text file; An operation and maintenance knowledge base construction module, which is used to perform text chunking on a structured text file using the langchain library, perform vectorization processing, and generate an operation and maintenance knowledge base; An operation and maintenance knowledge retrieval module, which is used to obtain a user's query file, and retrieve target text data from the operation and maintenance knowledge base according to the query problem by adopting a hybrid retrieval strategy that combines sparse vectors and dense vectors; An intelligent question answering module, which is used to generate prompt information according to the query problem and the target text data using a prompt template, input the prompt information into a large language model, and generate an answer.
[0014] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the industrial Internet device operation and maintenance knowledge question answering method described in any one of the above are implemented.
[0015] In a fourth aspect, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the industrial Internet device operation and maintenance knowledge question answering method described in any one of the above are implemented.
[0016] From the above technical solutions, it can be seen that the present invention has the following advantages: In the industrial Internet device operation and maintenance knowledge question answering method provided by the present application, the document content is extracted by the MinerU tool, the knowledge base is constructed using the langchain library, the sparse vector and the dense vector are combined for hybrid retrieval, and a specific prompt template is designed to ensure that the model output is based on the document content. Finally, a detailed and logical answer is generated using the large speech model. This method automatically retrieves relevant knowledge of operation and maintenance problems and uses the large model to perform automatic summary and answer to reduce labor costs and effectively solve the operation and maintenance problems in large, complex, and dynamic industrial Internet scenarios.
[0017] The present application extracts text data from operation and maintenance documents in various formats through the MinerU tool, and combines the OCR technology and the document structure recognition algorithm to generate a structured text file. This process not only improves the efficiency of data processing, but also ensures the integrity and accuracy of information, laying a solid foundation for subsequent knowledge base construction.
[0018] The present application adopts a hybrid retrieval strategy that combines sparse vectors and dense vectors, which can consider both keyword matching and semantic similarity, so as to more accurately retrieve target text data related to the user's query problem from the operation and maintenance knowledge base. This strategy effectively improves the response quality of the question answering system and the user experience.
[0019] Through the combination of prompt templates and large language models, the system can generate accurate and natural answers based on query questions and retrieved target text data. This method not only improves the intelligence level of question answering, but also enhances the practicality and scalability of the system, and is applicable to complex industrial Internet device operation and maintenance scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic flowchart of the method for answering questions about the operation and maintenance knowledge of industrial Internet devices provided by this application.
[0022] Figure 2 It is a schematic flowchart of the method for constructing an operation and maintenance knowledge base provided by this application.
[0023] Figure 3 It is a schematic flowchart of the method for retrieving operation and maintenance knowledge provided by this application.
[0024] Figure 4 It is a schematic structural diagram of the industrial Internet device operation and maintenance knowledge question answering system provided by this application.
[0025] Figure 5 It is a schematic structural diagram of the electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In the following, the specific steps of the method for answering questions about the operation and maintenance knowledge of industrial Internet devices will be described in detail, and various embodiments of the present disclosure will be described more comprehensively. The present disclosure can have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0027] In the following, the term "comprising" or "may comprise" that can be used in various embodiments of the present disclosure indicates the presence of the disclosed functions, operations, or elements, and does not limit the addition of one or more functions, operations, or elements. Further, as used in various embodiments of the present disclosure, the terms "comprising", "having" and their cognates are only intended to indicate a specific feature, number, step, operation, element, component, or combination of the foregoing items, and should not be construed as precluding the existence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items.
[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] Please refer to Figure 1 The following is a flowchart of a method for answering questions about the operation and maintenance of industrial Internet devices in a specific embodiment. The method includes: S1: Obtain the operation and maintenance documents of industrial Internet devices, extract the text data in the operation and maintenance documents through the MinerU tool, and generate a structured text file.
[0030] In the specific implementation, first collect the operation and maintenance documents of industrial Internet devices. The operation and maintenance documents include text documents and non-text documents. Among them, the operation and maintenance files are operation and maintenance documents in different formats, including TXT, PDF, Word documents, PPT slides, Excel spreadsheets, etc. The operation and maintenance files in PDF, Word, PPT, and Excel formats are non-text documents, and the operation and maintenance files in TXT format are text documents.
[0031] Then, use the text extraction algorithm of the MinerU tool to extract the text data in the non-text document as the first text data. At the same time, use the image extraction algorithm to extract the images in the non-text document, use the OCR recognition algorithm to recognize the text in the images, and convert it into the second text data.
[0032] Finally, use the document structure recognition algorithm to recognize the hierarchical structures in the text document, the first text data, and the second text data, and mark them to generate a structured text file. Use the document structure recognition algorithm to recognize the structures such as titles, paragraphs, and lists in the document to better organize and understand the document content.
[0033] Among them, the document structure recognition algorithm can adopt a recognition algorithm based on regular expressions or directly use a layout analysis model to recognize structures such as titles, paragraphs, and lists; after recognition, it outputs text with hierarchical tags (such as <Title 1>Fault Diagnosis Process< / Title 1>). The structured text file adopts the JSON format and contains information such as content, picture links, and hierarchical structures.
[0034] S2: Based on the structured text file, use the langchain library to perform text chunking and vectorization processing to generate an operation and maintenance knowledge base.
[0035] S3: Obtain the user's query question, and according to the query question, adopt a hybrid retrieval strategy combining sparse vectors and dense vectors to retrieve target text data from the operation and maintenance knowledge base.
[0036] S4: Generate prompt information based on the query question and the target text data, and input the prompt information into a large language model to generate an answer.
[0037] The purpose of this step is to achieve intelligent question answering through a large model. When the user's query triggers the large model to generate an answer, the design of the prompt is crucial.
[0038] As an example, the prompt template constructed in this step is as follows: “ You are a knowledgeable AI assistant. According to the following provided document content, answer the user's question. Please ensure that the answer is based on the content of these documents, rather than generating new, unsupported information. Relevant document content: {context} User question: {question} Your answer should explain in detail and logically, and clearly cite the information in the document. ” This prompt can ensure that the output of the model is always based on the provided document content, avoiding generating unfounded content. In the model inference stage, the qwen-14b large language model is used to generate answers. With its powerful language generation ability, this model can provide users with detailed answers based on the context, ensuring that the content is accurate and logical.
[0039] In this embodiment, through structured text extraction, hybrid retrieval strategy, and intelligent question answering generation, the performance of the industrial Internet device operation and maintenance knowledge question answering system is significantly improved. Combining keyword matching and semantic similarity retrieval, accurately obtain target data, and use the large language model to generate accurate answers, improving the response speed and intelligence level, enhancing the practicability and scalability of the system, and providing an efficient solution for device operation and maintenance.
[0040] In one embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.
[0041] Refer to Figure 2 As shown, this embodiment discloses an operation and maintenance knowledge base construction method, which specifically includes the following steps: S201: Based on the structured text file, use the RecursiveCharacterTextSplitter splitting method of the langchain library to split the structured text into multiple text blocks.
[0042] Exemplarily, use the RecursiveCharacterTextSplitter splitting method in the langchain library, specify multiple delimiters (such as full stops, commas, semicolons, etc.), and split the text into smaller text blocks. Set the chunk_size of the maximum text block to 256 characters to ensure that the text blocks are easy to process. Among them, after splitting, the context association between text blocks needs to be retained (such as marking consecutive serial numbers for adjacent blocks).
[0043] S202: Use the BM25 algorithm to convert the text blocks into sparse vectors, and use the BAAI General Embedding model to map the text blocks into dense vectors.
[0044] Exemplarily, generate sparse vectors based on word frequency through the BM25 algorithm (suitable for keyword matching); at the same time, use the BGE model to map the text into 768-dimensional semantic vectors.
[0045] Among them, the BM25 algorithm is a probability-based ranking function used for information retrieval and text mining, especially suitable for keyword matching. The BGE model is a semantic-based vector model that can capture the deep semantic information of the text.
[0046] S203: Store the sparse vectors, dense vectors, and the corresponding text blocks into the Milvus vector database, configure the index type and similarity measurement method, and generate an operation and maintenance knowledge base.
[0047] Exemplarily, first store the sparse vectors, dense vectors, and the corresponding text blocks into the Milvus vector database; set the index type to IVF_FLAT and the similarity measurement method to inner product. Then, establish a mapping relationship between the text block ID and the corresponding text block, sparse vector, or dense vector, and generate an operation and maintenance knowledge base.
[0048] In this step, the Milvus database is used to store semantic vectors and their corresponding text data. To achieve efficient retrieval, the index uses the vector inner product calculation as the similarity metric to ensure that the most semantically similar documents can be retrieved quickly.
[0049] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.
[0050] Reference Figure 3 As shown, this embodiment discloses an operation and maintenance knowledge retrieval method, which specifically includes the following steps: S301: Obtain the user's query question, generate a sparse vector of the query question based on the BM25 algorithm, and retrieve the sparse vectors in the operation and maintenance knowledge base. Find the text data with keyword matching by calculating the similarity of the sparse vectors.
[0051] For example, first obtain the query question q input by the user, and convert the tokenized query q into a BM25 sparse vector based on the BM25 algorithm, representing the weight of each word in the query. Then, in the operation and maintenance knowledge base, calculate the similarity between the sparse vector of q and all document sparse vectors in the operation and maintenance knowledge base, sort them in descending order according to the similarity value, and take the document data corresponding to the top 20 sparse vectors as the text data with keyword matching. S302: Generate a dense vector of the query question based on the BAAI General Embedding model, and retrieve the dense vectors in the operation and maintenance knowledge base. Find the text data with semantic similarity by calculating the similarity of the dense vectors.
[0052] For example, first use the BGE model to encode q into a 768-dimensional dense vector to capture the semantic information of the query. Then, in the operation and maintenance knowledge base, calculate the inner product similarity between the dense vector of q and all document dense vectors in the operation and maintenance knowledge base, sort them in descending order according to the inner product similarity value, and take the document data corresponding to the top 20 dense vectors as the text data with semantic similarity. S303: Merge the text data with keyword matching and the text data with semantic similarity to generate process text data.
[0053] For example, through comparison, extract the text data with keyword matching and the text data with semantic similarity with the same content as the process text data. The purpose of this step is to ensure that the retrieved text data has both the similarity value of the dense vector and the similarity value of the dense vector for subsequent steps to comprehensively screen the two similarity values.
[0054] S304: For each piece of process text data, calculate a similarity score based on the similarity of sparse vectors and the similarity of dense vectors according to a preset similarity weight.
[0055] Exemplarily, for each piece of process text data, first normalize the similarity of its sparse vector and the similarity of its dense vector, and convert them into BM25 scores and BGE scores respectively. Then, use the following formula to calculate the similarity score X: X = α × A + (1 - α) × B; where A is the BM25 score, B is the BGE score, and α = 0.5, that is, the weights of the sparse vector and the dense vector are equal.
[0056] S305: Sort the similarity scores, and filter out the target text data according to the sorting result.
[0057] Exemplarily, sort all the similarity scores in descending order, and take the process text data corresponding to the top 5 similarity scores as the target text data.
[0058] It can be seen that the present invention adopts a hybrid retrieval strategy that combines sparse vectors and dense vectors. Among them, sparse vectors are suitable for exact matching (such as keyword search), while dense vectors are suitable for capturing broader semantic similarities. First, set the similarity weights of sparse vectors and dense vectors to be the same to ensure that the retrieval results can balance keyword matching and semantic relevance. After the user inputs a query request, the system first uses the BM25 algorithm to retrieve the sparse vectors and find the documents that match the keywords. Then, use the BGE model to retrieve the dense vectors converted from the user input query and find the semantically similar documents. Combine the two retrieval results and sort them according to the similarity weight, and finally return them to the user. Finally, according to the user feedback and the accuracy of the retrieval results, continuously adjust and optimize the similarity weights of sparse vectors and dense vectors, as well as the parameters of the BM25 and BGE models, to improve the accuracy and efficiency of the retrieval.
[0059] As Figure 4 shown, the following is an embodiment of the industrial Internet device operation and maintenance knowledge Q&A system provided by the embodiments of the present disclosure. This system belongs to the same inventive concept as the industrial Internet device operation and maintenance knowledge Q&A method in the above embodiments. For the details not described in detail in the embodiment of the industrial Internet device operation and maintenance knowledge Q&A system, reference can be made to the embodiments of the above industrial Internet device operation and maintenance knowledge Q&A method.
[0060] An industrial Internet device operation and maintenance knowledge Q&A system includes: an operation and maintenance document preprocessing module, an operation and maintenance knowledge base construction module, an operation and maintenance knowledge retrieval module, and an intelligent Q&A module.
[0061] The operation and maintenance document preprocessing module is used to obtain the operation and maintenance documents of industrial Internet devices, extract the text data in the operation and maintenance documents through the MinerU tool, and generate a structured text file; The operation and maintenance knowledge base construction module is used to perform text chunking and vectorization processing on the structured text file using the langchain library to generate an operation and maintenance knowledge base; The operation and maintenance knowledge retrieval module is used to obtain the user's query file, and according to the query problem, retrieve the target text data from the operation and maintenance knowledge base by adopting a hybrid retrieval strategy that combines sparse vectors and dense vectors; The intelligent question and answer module is used to generate prompt information according to the query problem and the target text data using a prompt template, and input the prompt information into a large language model to generate an answer.
[0062] The industrial Internet device operation and maintenance knowledge question and answer system provided in this embodiment efficiently constructs an operation and maintenance knowledge base through structured text extraction, a hybrid retrieval strategy, and intelligent question and answer generation, accurately matches user queries, generates accurate answers, and improves the operation and maintenance efficiency and intelligent level of industrial equipment.
[0063] Figure 5 The hardware structure diagram of an electronic device for implementing various embodiments of the present invention.
[0064] The industrial Internet device operation and maintenance knowledge question and answer method provided in the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the figure, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.
[0065] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a key, a camera, a display screen, and a SIM card interface, etc.
[0066] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0067] Among them, the processor may be the nerve center and command center of the electronic device. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0068] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.
[0069] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.
[0070] The internal memory can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. The internal memory may include a high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0071] The wireless communication function of the electronic device can be implemented through an antenna, a wireless communication module, a modem processor, a baseband processor, etc.
[0072] The wireless communication module can provide solutions for wireless communications applied to electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0073] The electronic device can implement audio functions through an audio module, speaker, receiver, microphone, headphone jack, application processor, etc.
[0074] The electronic device can implement shooting functions through an ISP, camera, video codec, GPU, display screen, application processor, etc.
[0075] The electronic device can implement display functions through a GPU, display screen, application processor, etc.
[0076] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs, which execute program instructions to generate or change display information.
[0077] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0078] The above-mentioned electronic device realizes the industrial Internet device operation and maintenance knowledge Q&A method of generating structured text by extracting operation and maintenance documents, combining a hybrid retrieval strategy of sparse and dense vectors, accurately matching user queries, and using a large language model to generate intelligent answers. It achieves the beneficial effect of improving the efficiency and intelligent level of industrial device operation and maintenance knowledge Q&A.
[0079] In the storage medium provided by this application, there is a program product capable of implementing the industrial Internet device operation and maintenance knowledge Q&A method.
[0080] The method for answering questions about the operation and maintenance of industrial Internet devices includes: obtaining the operation and maintenance documents of industrial Internet devices, extracting text data from the operation and maintenance documents through the MinerU tool, and generating a structured text file; based on the structured text file, using the langchain library to perform text chunking and vectorization processing to generate an operation and maintenance knowledge base; obtaining the query questions of the user, and according to the query questions, adopting a hybrid retrieval strategy combining sparse vectors and dense vectors to retrieve target text data from the operation and maintenance knowledge base; generating prompt information according to the query questions and the target text data using a prompt template, and inputting the prompt information into a large language model to generate an answer.
[0081] In some possible implementation manners, the method for answering questions about the operation and maintenance of industrial Internet devices of the present disclosure may be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0082] The storage medium of the present disclosure may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0083] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for question-answering knowledge about operation and maintenance of industrial Internet equipment, characterized in that: include: Obtain the operation and maintenance documents of industrial Internet equipment, extract text data in the operation and maintenance documents through the MinerU tool, and generate a structured text file; Based on structured text files, the langchain library is used to segment text into blocks and perform vectorization processing to generate an operation and maintenance knowledge base; Obtaining a user's query question, and based on the query question, using a hybrid retrieval strategy combining sparse vectors and dense vectors to retrieve target text data from an operation and maintenance knowledge base; Prompt word information is generated using a prompt word template according to the query question and target text data, and the prompt word information is input into a large language model to generate an answer.
2. The industrial Internet equipment operation and maintenance knowledge question and answer method according to claim 1 is characterized in that: The obtaining of the operation and maintenance documents of the industrial Internet equipment, extracting text data in the operation and maintenance documents through the MinerU tool, and generating a structured text file includes: Collecting operation and maintenance documents of industrial Internet devices, wherein the operation and maintenance documents include text documents and non-text documents; Use the text extraction algorithm of the MinerU tool to extract text data from the non-text document as the first text data; Extracting images from non-text documents using an image extraction algorithm, recognizing text in the images using an OCR recognition algorithm, and converting the text into second text data; A document structure recognition algorithm is used to recognize hierarchical structures in the text document, the first text data, and the second text data, and mark them to generate a structured text file.
3. The industrial Internet equipment operation and maintenance knowledge question and answer method according to claim 2 is characterized in that: Based on the structured text file, the langchain library is used to segment the text and perform vectorization processing to generate an operation and maintenance knowledge base, including: Based on the structured text file, the RecursiveCharacterTextSplitter segmentation method of the langchain library is used to split the structured text into multiple text blocks; Use the BM25 algorithm to convert text blocks into sparse vectors, and use the BAAI General Embedding model to map text blocks into dense vectors; Store the sparse vectors and dense vectors, as well as the corresponding text blocks, into the Milvus vector database, configure the index type and similarity measurement method, and generate an operation and maintenance knowledge base.
4. The industrial Internet equipment operation and maintenance knowledge question and answer method according to claim 3 is characterized in that: Based on the structured text file, the RecursiveCharacterTextSplitter segmentation method of the langchain library is used to segment the structured text into multiple text blocks, including: Based on the structured text file, the RecursiveCharacterTextSplitter segmentation method of the langchain library is used to split the structured text into multiple text blocks with periods, commas and semicolons as delimiters, and each text block includes a maximum of 256 characters.
5. The industrial Internet equipment operation and maintenance knowledge question and answer method according to claim 3 is characterized in that: The BM25 algorithm is used to convert the text block into a sparse vector, and the BAAI General Embedding model is used to map the text block into a dense vector, including: Use the BM25 algorithm to convert text blocks into sparse vectors based on word frequency; The BAAI General Embedding model is used to map text blocks into 768-dimensional semantic vectors.
6. The industrial Internet equipment operation and maintenance knowledge question and answer method according to claim 3 is characterized in that: The sparse vectors and dense vectors, as well as the corresponding text blocks, are stored in the Milvus vector database, the index type and similarity measurement method are configured, and the operation and maintenance knowledge base is generated, including: Store the sparse vectors and dense vectors, as well as the corresponding text blocks, into the Milvus vector database; set the index type to IVF_FLAT and the similarity measure to inner product; A mapping relationship between a text block ID and a corresponding text block, a sparse vector or a dense vector is established, and an operation and maintenance knowledge base is generated.
7. The industrial Internet equipment operation and maintenance knowledge question and answer method according to claim 3 is characterized in that: The step of obtaining a user's query question and retrieving target text data from the operation and maintenance knowledge base using a hybrid retrieval strategy combining sparse vectors and dense vectors according to the query question includes: Obtain the user's query question, generate a sparse vector of the query question based on the BM25 algorithm, search the sparse vector in the operation and maintenance knowledge base, and find text data matching the keywords by calculating the similarity of the sparse vectors; Generate a dense vector of the query question based on the BAAI General Embedding model, search the dense vector in the operation and maintenance knowledge base, and find text data with similar semantics by calculating the similarity of the dense vector; Merge keyword-matched text data and semantically similar text data to generate process text data; For each process text data, a similarity score is calculated based on the similarity of the sparse vector and the similarity of the dense vector according to the preset similarity weight; Sort the similarity scores and filter out the target text data based on the sorting results.
8. An industrial Internet equipment operation and maintenance knowledge question and answer system, characterized in that: The system adopts the industrial Internet equipment operation and maintenance knowledge question and answer method as described in any one of claims 1 to 7; The system comprises: The operation and maintenance document preprocessing module is used to obtain the operation and maintenance documents of the industrial Internet equipment, extract the text data in the operation and maintenance documents through the MinerU tool, and generate a structured text file; The operation and maintenance knowledge base construction module is used to generate the operation and maintenance knowledge base based on structured text files, using the langchain library to segment text and perform vectorization processing; The operation and maintenance knowledge retrieval module is used to obtain the user's query file and retrieve the target text data from the operation and maintenance knowledge base using a hybrid retrieval strategy combining sparse vectors and dense vectors according to the query question; The intelligent question-answering module is used to generate prompt word information using a prompt word template according to the query question and target text data, input the prompt word information into a large language model, and generate an answer.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the industrial Internet equipment operation and maintenance knowledge question and answer method as described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the industrial Internet equipment operation and maintenance knowledge question and answer method as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Industrial knowledge query method and device, equipment and medium
CN120849554A
Large-model-driven emergency disposal process generation method and system
CN121010104A
Medical record text processing method and device, equipment and medium
CN121096687A
Question and answer dialogue method, system and device
CN121255977A
Question and answer method, device and equipment based on RAG operation and maintenance library and medium
CN121524274A